Electronic device and control method thereof
By combining a microphone, communication interface, and trained neural network model, the system identifies and verifies user voice input, solving the problem of users repeatedly speaking and improving the quality of wake-up voice input and the usability of electronic devices.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- SAMSUNG ELECTRONICS CO LTD
- Filing Date
- 2021-09-17
- Publication Date
- 2026-04-14
AI Technical Summary
In existing technologies, users need to repeatedly utter the same words to register and activate voice input, resulting in a poor user experience and the risk of accidental activation of voice input.
By using a microphone, communication interface, and processor, combined with a trained neural network model and an external server, user voice input is identified and verified. Based on feature vector similarity and text information, the number of times the user speaks is reduced to register high-quality wake-up voice input.
This approach improves the registration quality of wake-up voice input and the usability of electronic devices while minimizing the number of times users need to speak, and reduces the possibility of accidental activation.
Smart Images

Figure CN116686046B_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to electronic devices and methods for controlling electronic devices. More specifically, this disclosure relates to electronic devices configured to register user voice input for waking up electronic devices and methods for controlling such devices. Background Technology
[0002] Recently, technologies have been developed for controlling electronic devices via user voice input. In particular, electronic devices can receive wake-up voice input to activate the device, or to activate specific applications of the device (e.g., artificial intelligence applications).
[0003] To register traditional wake-up voice input, or to clearly identify text included in the user's voice input to be registered as wake-up voice input, an electronic device can register wake-up voice input by simply uttering the same word multiple times (e.g., five or more times). In this case, usability is limited because users often feel embarrassed or uncomfortable uttering the same word multiple times. If wake-up voice input is registered by uttering the user's voice input to be registered as wake-up voice input only once, there may be problems where the electronic device is activated by user voice input including registration text and similar text.
[0004] In addition, if a user's voice input that they want to register as wake-up voice input is not suitable as wake-up voice input, the user may need to be notified of this. Summary of the Invention
[0005] Technical issues
[0006] This disclosure provides an electronic device and its control method capable of registering high-quality wake-up voice input while minimizing the number of times the user speaks to register the wake-up voice input.
[0007] Technical solution
[0008] According to one aspect of an exemplary embodiment, an electronic device may include a microphone, a communication interface, a memory, and a processor, wherein the memory is configured to store at least one instruction, and the processor is configured to execute at least one instruction to: obtain user voice input via the microphone for registering wake-up voice input; input the user voice input into a trained neural network model to obtain a first feature vector corresponding to text included in the user voice input; receive, via the communication interface, a verification dataset determined based on information relating to the text included in the user voice input from an external server; input verification voice input included in the verification dataset into the trained neural network model to obtain a second feature vector corresponding to the verification voice input; and identify whether to register the user voice input as wake-up voice input based on the similarity between the first feature vector and the second feature vector.
[0009] The processor can recognize user voice input to obtain information related to the text included in the user voice input; and send the information related to the text included in the user voice input to an external server via a communication interface. The external server can use the text-related information to obtain a verification voice input based on a first phoneme sequence of the verification voice text, wherein the first phoneme sequence of the verification voice text is a number of common phonemes comprised of a second phoneme sequence of the text stored on the external server and a phoneme sequence of the information related to the text.
[0010] The verification voice input is the voice data corresponding to the verification voice text whose number of common phonemes is equal to or greater than the threshold.
[0011] The processor may, based on the similarity between the first feature vector and the second feature vector being less than a threshold, input another verification voice input included in the verification dataset into a trained neural network model to obtain a third feature vector corresponding to the other verification voice input; compare another similarity between the first feature vector and the third feature vector; and, based on the other similarity between the first feature vector and the third feature vector being equal to or greater than a threshold, provide a guiding message requesting additional user voice input for registering wake-up voice input.
[0012] The processor can register a user's voice input as a wake-up voice input based on multiple similarities less than a threshold between the feature vectors corresponding to all verified voice inputs included in the verification dataset and a first feature vector.
[0013] The processor can input user voice input into a speech recognition model to obtain text included in the user's voice; and based on at least one of the length and repetition of phonemes in the text included in the user's voice input, identify whether to register the text included in the user's voice input as wake-up voice input.
[0014] The processor may provide a prompting message that requests the phonation of another text attached to the user's voice input for registration of the wake-up voice input, based on the fact that the number of phonemes in the text included in the user's voice input is less than a first threshold or the number of repeating phonemes in the text included in the user's voice input is greater than a second threshold.
[0015] The guidance message is configured to include a message that recommends another text, determined based on the electronic device’s usage history, as the wake-up voice input.
[0016] The processor can feed user voice input into a trained speech recognition model to obtain feature values indicating whether the user voice input is a user voice input that utters a specific text; and based on the feature values, determine whether to register the user voice input as a wake-up voice input.
[0017] The processor can provide a guiding message for requesting additional user voice input to register wake-up voice input based on the feature value being less than a threshold.
[0018] According to one aspect of an exemplary embodiment, a method for controlling an electronic device may include: obtaining user voice input for registering wake-up voice input; inputting the user voice input into a trained neural network model to obtain a first feature vector corresponding to text included in the user voice input; receiving a verification dataset determined from an external server based on information relating to the text included in the user voice input; inputting verification voice input included in the verification dataset into the trained neural network model to obtain a second feature vector corresponding to the verification voice input; and identifying whether to register the user voice input as wake-up voice input based on the similarity between the first feature vector and the second feature vector.
[0019] The method may include: recognizing user voice input to obtain information related to text included in the user voice input; and sending the information related to the text included in the user voice input to an external server. The external server may obtain the verification voice input based on a first phoneme sequence of the verification voice text, the first phoneme sequence of the verification voice text being a number of common phonemes comprised of a second phoneme sequence of the text stored on the external server and a phoneme sequence of the information related to the text.
[0020] The verification voice input can be voice data corresponding to verification voice text with a number of common phonemes equal to or greater than a threshold.
[0021] The method may include: inputting another verification voice input, included in the verification dataset, into a trained neural network model to obtain a third feature vector corresponding to the other verification voice input, based on the similarity between a first feature vector and a second feature vector being less than a threshold; comparing another similarity between the first feature vector and the third feature vector; and providing a guiding message requesting additional user voice input for registering wake-up voice input based on the other similarity between the first feature vector and the third feature vector being equal to or greater than a threshold.
[0022] The method may include: registering user voice input as wake-up voice input based on multiple similarities less than a threshold between the feature vectors corresponding to all verified voice inputs included in the verification dataset and a first feature vector.
[0023] Beneficial effects
[0024] According to the embodiments of the present disclosure as described above, the convenience or usability of the wake-up voice input registration process can be increased by minimizing the number of voices uttered by the user for registering the wake-up voice, and a technology that enables control of electronic devices via voice recognition can be realized by registering high-quality wake-up voices. Attached Figure Description
[0025] The above and other aspects, features and advantages of certain embodiments of this disclosure will become more apparent from the following description taken in conjunction with the accompanying drawings.
[0026] Figure 1a This is a diagram illustrating the process of registering and waking up voice input.
[0027] Figure 1b This is a diagram illustrating the process of waking up an electronic device using a registered wake-up voice input.
[0028] Figure 2 This is a block diagram schematically illustrating the configuration of an electronic device according to an embodiment.
[0029] Figure 3 This is a block diagram illustrating the configuration of the text evaluation module according to an embodiment.
[0030] Figure 4 This is a block diagram illustrating the configuration of a non-voice input evaluation module according to an embodiment.
[0031] Figure 5 This is a sequence diagram illustrating a method for obtaining a verification dataset according to an embodiment.
[0032] Figure 6 This is a flowchart illustrating a method for verifying user voice input using a verification dataset according to an embodiment.
[0033] Figures 7a to 8This is a diagram illustrating the boot message according to an embodiment.
[0034] Figure 9 This is a flowchart illustrating a method for controlling an electronic device according to an embodiment.
[0035] Figure 10 This is a detailed block diagram illustrating the configuration of an electronic device according to an embodiment. Detailed Implementation
[0036] This disclosure may have several embodiments, and various modifications can be made to the embodiments. Specific embodiments are provided in the following description with reference to the accompanying drawings and their detailed description. However, it should be understood that this disclosure is not limited to the specific embodiments described below, but includes various modifications, equivalents, and / or substitutions of the embodiments of this disclosure. Regarding the interpretation of the drawings, similar reference numerals may be used for similar constituent elements.
[0037] In describing exemplary embodiments, if a detailed description of a related known function or component makes the description of the subject matter unclear, that detailed description may be omitted.
[0038] Furthermore, exemplary embodiments may be modified in various forms, and therefore the scope of the technology is not limited to the following exemplary embodiments. Rather, these exemplary embodiments are provided to make this disclosure thorough and complete.
[0039] The terminology used herein is intended only to explain certain exemplary embodiments and does not limit the scope of this disclosure. Unless the context clearly indicates otherwise, the singular form is intended to include the plural form.
[0040] The terms “having,” “may have,” “including,” and “may include” used in embodiments of this disclosure indicate the presence of a corresponding feature (e.g., an element such as a numerical value, function, operation, or part thereof) and do not exclude the presence of additional features.
[0041] In this specification, the terms “A or B”, “at least one of A and / or B”, or “one or more of A and / or B” can include all possible combinations of the items listed together. For example, the terms “A or B” or “at least one of A and / or B” can specify (1) at least one A; (2) at least one B; or (3) both at least one A and at least one B.
[0042] As used herein, the terms “1,” “2,” “first,” or “second” can modify a wide variety of elements, regardless of their order and / or importance, and are used only to distinguish one element from another. Therefore, the corresponding elements are not limited.
[0043] When a component (e.g., a first component) is operatively or communicatively coupled to or connected to another component (e.g., a second component), the component may be directly coupled to the other component or may be coupled through another component (e.g., a third component). When a component (e.g., a first component) is directly coupled to or connected to another component (e.g., a second component), there may be no component (e.g., a third component) between the component and the other component.
[0044] When an element (e.g., a first element) is directly coupled to or directly connected to another element (e.g., a second element), there may be no element (e.g., a third element) between the element and the other element.
[0045] In this specification, the term "configured as" may be changed in certain circumstances to, for example, "suitable," "capable," "designed to," "adapted to," "made of," or "able to." The term "configured as (set up)" does not necessarily mean "specifically designed as" at the hardware level.
[0046] In some cases, the term "device configured to" can mean "the device is able to" do certain things together with another device or component. For example, a processor configured to perform "A, B, and C" can be implemented as a dedicated processor (e.g., an embedded processor) or a general-purpose processor for performing functions by running one or more software programs stored in a memory device (e.g., a central processing unit (CPU) or an application processor (AP)).
[0047] In the embodiments disclosed herein, the terms "module" or "unit" refer to an element that performs at least one function or operation. A "module" or "unit" can be implemented as hardware, software, or a combination thereof. Furthermore, in addition to "modules" or "units" that should be implemented in specific hardware, multiple "modules" or "units" can be integrated into at least one module and can be implemented in an integrated manner as at least one processor.
[0048] Furthermore, the various elements and areas shown in the accompanying drawings are schematically illustrated. Therefore, the technical ideas are not limited by the relative dimensions or spacing shown in the accompanying drawings.
[0049] Electronic devices according to various embodiments of this disclosure may include at least one of, for example, smartphones, tablet PCs, desktop PCs, laptop PCs, and wearable devices. Wearable devices may include at least one of accessories (e.g., watches, rings, bracelets, anklets, necklaces, glasses, contact lenses, or head-mounted devices (HMDs)), fabrics or clothing (e.g., electronic clothing), body attachment types (e.g., pads or tattoos), or bio-implantable circuitry.
[0050] According to another embodiment, the electronic device can be a household appliance. Household appliances include, for example, televisions, digital video disc (DVD) players, audio equipment, refrigerators, air conditioners, vacuum cleaners, ovens, microwave ovens, washing machines, set-top boxes, home automation control panels, security control panels, and television boxes (e.g., Samsung HomeSync). TM Apple TV TM or Google TV TM ), game consoles (e.g., Xbox) TM PlayStation TM (etc.), electronic dictionary, electronic key, portable camera or electronic photo frame.
[0051] According to embodiments of this disclosure, wake-up voice input can be user voice input that causes electronic device 100 to perform a wake-up operation. In this case, a wake-up operation can refer to an operation used to activate electronic device 100, activate a specific application, or activate a specific function of electronic device 100. Furthermore, activation can refer to the state where the power, application, or function of electronic device 100 is turned off or switched from standby mode to an on state. Additionally, wake-up voice input according to embodiments of this disclosure can be used as another term such as triggering voice input.
[0052] Exemplary embodiments of this disclosure will now be described in more detail in a manner that will be understood by those skilled in the art.
[0053] Figure 1a This is a diagram illustrating the process of registering and waking up voice input.
[0054] When entering the wake-up voice input registration mode, the electronic device 100 can receive first user voice input to be registered as wake-up voice input via a microphone. In this case, the first user voice input may include specific text as keywords for the wake-up voice input to be registered.
[0055] The electronic device 100 can perform preprocessing on the received first user voice input during operation 10. Specifically, the electronic device 100 can perform preprocessing operations such as noise removal and sound quality enhancement.
[0056] The electronic device 100 can extract speech features from the preprocessed first user speech input during operation 20. Specifically, the electronic device 100 can extract speech features by converting the preprocessed first user speech input from a time dimension to a frequency dimension.
[0057] The electronic device 100 can input user speech input, converted to a frequency dimension, into each of the first neural network model 30-1 and the second neural network model 30-2. In this case, the first neural network model (e.g., a keyword recognition model, a speech recognition model, etc.) can be a neural network model trained to obtain feature vectors corresponding to text included in the user speech input, and the second neural network model (e.g., a speaker recognition model) can be a neural network model trained to obtain feature vectors corresponding to unique features (e.g., glottis) of the speech of the speaker who issued the user speech input.
[0058] In addition, the electronic device 100 can register the feature vector output from the first neural network model 30-1 that corresponds to the text included in the first user's voice input as the keyword feature vector 40-1, and the feature vector output from the second neural network model 30-2 that corresponds to the voice features of the speaker who made the first user's voice input can be registered as the speaker feature vector 40-2, so as to register the first user's voice input as a wake-up voice input.
[0059] Specifically, the electronic device 100 can guide the user to issue user voice input containing the same text approximately five times, and can also communicate with... Figure 1a The same process is used to register the first user's voice input as wake-up voice input multiple times.
[0060] Figure 1b This is a diagram illustrating the process of waking up an electronic device using a registered wake-up voice input.
[0061] The electronic device 100 can receive second user voice input for waking up the electronic device 100 via a microphone.
[0062] The electronic device 100 can perform preprocessing on the received second user voice input in operation 10, and can extract the voice features of the preprocessed second user voice input in operation 20.
[0063] The electronic device 100 can input user voice input into each of the first neural network model 30-1 and the second neural network model 30-2.
[0064] The electronic device 100 can compare the similarity between the feature vector output from the first neural network model 30-1 corresponding to the text included in the second user's voice input and the pre-registered keyword feature vector 40-1 (in operation 50-1), and can identify the similarity between the feature vector output from the second neural network model 30-2 corresponding to the speech features of the speaker who gave the second user's voice input and the previously registered speaker feature vector 40-2 (in operation 50-2). In this case, the similarity of the feature vectors can be identified by the distance between the feature vectors. More specifically, when the distance between the feature vectors is short, the similarity of the feature vectors can be identified as high, and when the distance between the feature vectors is large, the similarity of the feature vectors can be identified as low.
[0065] When the similarity between the feature vector corresponding to the text included in the second user's voice input and the pre-registered keyword feature vector 40-1, and the similarity between the feature vector corresponding to the voice feature of the speaker who made the second user's voice input and the pre-registered speaker feature vector 40-2, is equal to or greater than the threshold (operation 60), the electronic device 100 can be woken up based on the second user's voice input.
[0066] However, when the similarity between the feature vector corresponding to the text included in the second user's voice input and the pre-registered keyword feature vector 40-1, and the similarity between the feature vector corresponding to the voice feature of the speaker who issued the second user's voice input and the pre-registered speaker feature vector 40-2, is less than the threshold (operation 60), the electronic device 100 may ignore the second user's voice input and not wake up the electronic device 100 based on the acquired second user's voice input.
[0067] The embodiments of this disclosure will be referred to as follows. Figure 1a and Figure 1b Register wake-up voice input for performing operations to wake up electronic device 100, as described above. More specifically, this disclosure relates to a method for registering higher quality wake-up voice input while minimizing the process of registering the wake-up voice input.
[0068] In the following, exemplary embodiments will be described in detail with reference to the accompanying drawings. Figure 2 This is a block diagram illustrating the configuration of a control device according to an exemplary embodiment.
[0069] Electronic device 100 may include microphone 110, communication interface 120, memory 130, and processor 140. Here, electronic device 100 may be a smartphone. However, electronic device 100 according to this disclosure is not limited to a specific type of device and can be implemented as various types of electronic devices 100 such as tablet PCs, laptop PCs, and digital TVs.
[0070] Microphone 110 can receive user voice input. In this case, microphone 110 can convert the received user voice input into an electrical signal representing voltage changes over time.
[0071] In this case, the microphone 110 may be located inside the electronic device 100, but this is only an exemplary embodiment, and it may be located outside the device and electrically connected to the device.
[0072] The communication interface 120 includes circuitry and can communicate with external devices. Specifically, the processor 140 can receive various data or information from external devices connected via the communication interface 120, and can send various data or information to external devices.
[0073] The communication interface 120 may include at least one of a Wi-Fi module, a Bluetooth module, a wireless communication module, and a Near Field Communication (NFC) module. Specifically, the Wi-Fi module and the Bluetooth module may perform communication using Wi-Fi and Bluetooth methods, respectively. If a Wi-Fi module or a Bluetooth module is used, various types of connection information, such as a Service Set Identifier (SSID) and a session key, are first sent and received, and various types of information can be sent and received after communication is established.
[0074] Wireless communication chips can perform communication according to various communication standards such as IEEE, ZigBee, 3G, 3GPP, and LTE. An NFC module refers to a module that operates using the NFC method in the 13.56MHz band among various radio frequency identification (RFID) frequency bands such as 135kHz, 13.56MHz, 433MHz, 860–960MHz, and 2.45GHz.
[0075] Specifically, according to various embodiments of this disclosure, the communication interface 120 can communicate with an external server 200 (see...). Figure 5 The system sends information related to the text corresponding to the user's voice input to be registered as a wake-up voice input, and may receive a verification dataset obtained from the external server 200 based on the user's voice input to be registered as a wake-up voice input. In this case, the verification dataset may include at least one verification voice input containing a phoneme sequence in the user's voice input to be recorded as a wake-up voice input and a common phoneme sequence having a threshold or greater.
[0076] The memory 130 can store instructions for controlling the electronic device 100. An instruction refers to an action statement that can be directly executed by the processor 140 in a programming language, and is the smallest unit for program execution or action.
[0077] Specifically, memory 130 may store data for modules used to register wake-up voice input to perform various operations. Modules for registering wake-up voice input may include a preprocessing module 141, a voice feature extraction module 142, a text evaluation module 143, a non-voice input evaluation module 144, a validation set evaluation module 145, a wake-up voice input registration module 146, and a message providing module 147. Furthermore, memory 130 may store a keyword recognition model, a speech recognition model, a speaker recognition model, and a voice activity detection model. The keyword recognition model is trained to recognize specific keywords included in the user's voice input for registering wake-up voice input; the speech recognition model is trained to acquire the text corresponding to the user's voice input; the speaker recognition model is trained to acquire the voice features of the speaker who issued the user's voice input; and the voice activity detection model is trained to detect speech portions in the audio containing the user's voice. Additionally, memory 130 may store a usage history database (DB) including records of users of the electronic device 100 (e.g., search records, execution records, purchase records, etc.).
[0078] The memory 130 may include non-volatile memory and volatile memory. The non-volatile memory retains stored information even when the power supply is interrupted, while the volatile memory requires a continuous power supply to retain the stored information. Data for modules used to register wake-up voice input to perform various operations can be stored in the non-volatile memory. Furthermore, various neural network models, such as keyword recognition models (or speech recognition models), speaker recognition models, and speech segment detection models, can also be stored in the non-volatile memory.
[0079] The processor 140 can be electrically connected to the memory 130 to control the overall functions and operation of the electronic device 100.
[0080] When a user command for registering wake-up voice input is input, processor 140 can load data from modules stored in non-volatile memory for registering wake-up voice input to perform various operations into volatile memory. Furthermore, processor 140 can load neural network models such as keyword recognition models, speaker recognition models, and voice activity detection models into volatile memory. Processor 140 can then perform various operations based on the data loaded into volatile memory using these modules and neural network models. Here, "loading" refers to the operation of loading data stored in non-volatile memory and storing it in volatile memory so that processor 140 can access the data.
[0081] When a user command is input to register wake-up voice input, processor 140 can enter a mode for registering wake-up voice input. Specifically, processor 140 can provide a user interface (UI) for guiding the registration of wake-up voice input. The UI may include messages guiding the user's voice input.
[0082] The processor 140 can acquire user voice input to be registered as wake-up voice input via the microphone 110. The user voice input to be registered as wake-up voice input may include keywords, such as text like a password for waking up the electronic device 100.
[0083] When user voice input is acquired via microphone 110, processor 140 can perform preprocessing operations on the acquired user voice input via preprocessing module 141. Specifically, preprocessing module 141 can remove noise included in the acquired user voice input and can perform operations such as sound quality enhancement to clarify the user voice input included in the audio signal.
[0084] The processor 140 can extract speech features from the preprocessed user speech input using the speech feature extraction module 142. In this case, extracting speech features may mean converting the preprocessed user speech input from a time dimension to a frequency dimension. In this case, the speech feature extraction module 142 can transform the user speech input in the time dimension to the frequency dimension using methods such as Fourier transform.
[0085] The processor 140 can use the text evaluation module 143 to evaluate the text included in the user's voice input after frequency dimension transformation to verify whether the user's voice input is registered as a wake-up voice input.
[0086] like Figure 3 As shown, the text evaluation module 143 may include a phoneme length evaluation module 310, a phoneme repetition evaluation module 320, and a pronoun evaluation module 330.
[0087] Specifically, the text evaluation module 143 can obtain information related to the text included in the user's voice input by performing speech recognition on the user's voice input via a speech recognition model. In this case, the information related to the text can be information related to the phonemes included in the text.
[0088] The phoneme length evaluation module 310 can verify whether a user's voice input can be registered as a wake-up voice input based on the length of the phonemes included in the text. In this case, the phoneme length evaluation module 310 can verify whether the user's voice input can be registered as a wake-up voice input by identifying whether the length of the phonemes included in the text is equal to or less than a threshold. For example, when the user's voice input is "cha", the text evaluation module 143 can obtain "JA" as the phoneme for the user's voice input, and the phoneme length evaluation module 310 can identify that the number of phonemes is equal to or less than a threshold (e.g., 2), and can identify that the user's voice input is not suitable for registration as a wake-up voice input.
[0089] The phoneme repetition evaluation module 320 can verify whether a user's voice input can be registered as a wake-up voice input based on whether phonemes included in the text are repeated. In this case, the phoneme repetition evaluation module 320 can verify whether the user's voice input can be registered as a wake-up voice input by identifying whether the number of repeated phonemes included in the text is equal to or greater than a threshold. For example, if the user's voice input is "yayaya", the text evaluation module 143 can obtain "JApapaya JA" as the phoneme for the user's voice input, and the phoneme repetition evaluation module 320 can identify that the phoneme is repeated as a threshold (e.g., three times) and identify that the user's voice input is not suitable for registration as a wake-up voice input.
[0090] The pronoun evaluation module 330 can verify whether a user's voice input can be registered as a wake-up voice input based on whether the text contains pronouns. For example, if the user's voice is "that," the pronoun evaluation module 330 can identify that the user's voice input contains a pronoun and that the user's voice input is not suitable for registration as a wake-up voice input.
[0091] In the above embodiments, the text evaluation module 143 may identify whether the user's voice input is suitable for registration as a wake-up voice input based on the length of the phonemes, whether they are repeated, and whether they include pronouns. However, this is only an embodiment, and the text evaluation module 143 may identify whether the user's voice input is suitable for registration as a wake-up voice input based on other features of the text (e.g., when the length of the text is equal to or greater than a threshold).
[0092] Return to reference Figure 2 The processor 140 can identify whether the user's voice input is intended to produce specific text through the non-voice input evaluation module 144. In other words, the non-voice input evaluation module 144 can identify whether the user's voice input acquired through the microphone 110 is not human voice input, or whether the voice input is not intended to produce text (such as snoring, resting, etc.), in order to verify whether the user's voice input should be registered as wake-up voice input.
[0093] Specifically, such as Figure 4 As shown, the non-voice input evaluation module 144 may include a voice activity detection module 410 and a voice evaluation module 420. In this case, the voice activity detection module 410 can identify whether non-voice input is included in the user's voice input using a learned voice activity detection model. Specifically, the voice activity detection model is a neural network model learned using training data for voice and non-voice input, and can obtain feature values indicating whether voice activity is included in the user's voice input. The voice evaluation module 420 can identify whether non-voice input is included in the user's voice input based on the feature values obtained by the voice activity detection module 410. Specifically, when the feature value obtained by the voice activity detection module 410 is less than a threshold, the voice evaluation module 420 can identify that the user's voice input includes non-voice input in addition to voice activity, and when the feature value obtained by the voice activity detection module 410 is equal to or greater than the threshold, the voice evaluation module 420 can identify that non-voice input is not included in the user's voice input.
[0094] In the above embodiment, the non-voice input evaluation module 144 uses a voice activity detection model to verify whether non-voice input is included in the user's voice input. However, this is only one embodiment, and the non-voice input evaluation module 144 can use another method to verify whether the user's voice input includes non-voice input. For example, the non-voice input evaluation module 144 can model features related to the characteristics of the voice input (e.g., zero-crossing rate, spectral entropy, etc.) and identify whether voice activity is detected in the user's voice input based on whether the modeled features appear in the user's voice input. In other words, when non-voice activity instead of voice activity is detected in the user's voice, the non-voice input evaluation module 144 can identify that the user's voice input includes non-voice input and determine not to register the user's voice input as wake-up voice input.
[0095] Return to reference Figure 2 The processor 140 can verify, through the verification set evaluation module 145, whether text similar to the text included in the user's voice input wakes up the electronic device 100.
[0096] Specifically, the validation set evaluation module 145 can obtain a validation dataset for validating the user's voice input based on the text included in the user's voice input. This will refer to... Figure 5 Describe it.
[0097] The electronic device 100 can acquire text included in the user's voice input (operation S510). In this case, the electronic device 100 can acquire text included in the user's voice input by inputting the user's voice input into a speech recognition model.
[0098] The electronic device 100 can convert the acquired text into a phoneme sequence (operation S520). In this case, a phoneme is the smallest sound unit that distinguishes the meaning of a word, and a phoneme sequence means that the phonemes included in a word are arranged sequentially.
[0099] Electronic device 100 can send phoneme sequence information, which is text-related information, to server 200 (operation S530). In this case, electronic device 100 can send information related to other texts besides phoneme sequence information.
[0100] Server 200 can obtain a verification dataset based on the maximum common phoneme sequence (operation S540). Specifically, server 200 can obtain the verification dataset based on the phoneme sequences of multiple texts stored in server 200 and the length of the phoneme sequences included in the received phoneme sequence information. In other words, server 200 can identify texts among multiple texts stored in server 200 whose length, together with the phoneme sequences included in the received phoneme sequence information, is equal to or greater than a threshold. Server 200 can recognize the speech data corresponding to the recognized text as verification speech input and obtain a verification dataset including at least one recognized verification speech input.
[0101] The length of the phoneme sequence included together with the phoneme sequence included in the received phoneme sequence information can refer to the number of phoneme sequences that are included both in the phoneme sequence of the text and in the phoneme sequence contained in the received phoneme sequence information.
[0102] For example, when the user's voice is "Halli Galli", the phoneme sequence corresponding to the user's voice could be (HH AAL R IY K AA LR IY). Furthermore, server 200 can obtain voice data for Halli Geondam (HH AA LR IY K AAL R IY) (which is text whose length is equal to or greater than a threshold (e.g., 5) of the phoneme sequence commonly included in the phoneme sequence corresponding to the user's voice), voice data for Harleys (HH AA LR IY SS), etc., as verification voice.
[0103] Server 200 can send the obtained verification dataset to electronic device 100 (operation S550).
[0104] Electronic device 100 can use a verification dataset to verify user voice input (operation S560). (Refer to...) Figure 6 The method described herein is as follows: the verification set evaluation module 145 of the electronic device 100 uses a verification dataset to verify user voice input.
[0105] In the above embodiments, the use of the maximum common phoneme sequence to obtain text similar to the text included in the user's voice input has been described, but this is only an embodiment, and other methods can be used to obtain text similar to the text included in the user's voice input.
[0106] Figure 6 This is a flowchart illustrating a method by which an electronic device 100, according to an embodiment of the present disclosure, verifies user voice input using a verification dataset through a verification set evaluation module 145.
[0107] The validation set evaluation module 145 can select validation speech input from the validation dataset received from the server 200 (operation S610). In this case, the validation set evaluation module 145 can select the validation speech input from the validation dataset that corresponds to the text with the highest similarity to the user's speech input (e.g., including the most common phoneme sequences).
[0108] At this point, the validation set evaluation module 145 can perform preprocessing on the obtained validation speech input through the preprocessing module 141, and extract speech features from the preprocessed validation speech input through the speech feature extraction module 142. However, when the validation speech input received from the server 200 is speech data that has already undergone the preprocessing and speech extraction processes, the preprocessing and speech extraction processes can be omitted.
[0109] The validation set evaluation module 145 can obtain a feature vector corresponding to the selected validation speech input by inputting the selected validation speech input into a neural network model (operation S620). In this case, the neural network model can be one of a keyword recognition model trained to detect whether the user's speech input includes specific text or a speech recognition model trained to obtain the text included in the user's speech input.
[0110] The validation set evaluation module 145 can compare the similarity between the feature vector corresponding to the selected validation speech input and the feature vector corresponding to the user's speech input (operation S630). Specifically, the validation set evaluation module 145 can compare the similarity between the feature vector obtained in operation S620 and the feature vector obtained by inputting the user's speech input into the neural network model. In this case, the validation set evaluation module 145 can calculate the similarity based on the cosine distance between the two feature vectors.
[0111] The validation set evaluation module 145 can identify whether the similarity is equal to or greater than a threshold (operation S640). Specifically, when the similarity between the feature vector of the selected validation voice input and the feature vector of the user's voice is low, the probability of misidentification of the keyword used to perform the wake-up operation is low when the user's voice input including similar keywords is emitted, and therefore the electronic device 100 may overestimate the possibility of registering the user's voice input; and when the similarity between the feature vector of the selected validation voice input and the feature vector of the user's voice input is high, the probability of misidentification of the keyword used to perform the wake-up operation is high when the user's voice input including similar keywords is emitted, and the validation set evaluation module 145 may underestimate the registration probability of the user's voice input.
[0112] When the similarity is less than the threshold (operation S640 - No), the validation set evaluation module 145 can identify whether all validation speech inputs have been evaluated (operation S650).
[0113] When not all verification voice inputs are evaluated (operation S650 - No), the verification set evaluation module 145 can select the next verification voice input and repeat operations S610 to S640. When all verification voice inputs are evaluated (operation S650 - Yes), the verification set evaluation module 145 can register the corresponding user voice input as a wake-up voice input (operation S660).
[0114] However, when the similarity is equal to or greater than the threshold (operation S640 - Yes), the validation set evaluation module 145 can provide a guidance message (operation S670). In this case, the guidance message can include a message to guide the user to additionally emit user voice input containing the same text. In other words, when the pronunciation of the user's voice input is inaccurate or the user's voice input is distorted due to external factors, the possibility of misidentification through similar keywords is high, and therefore a guidance message requesting the user to make additional sounds can be provided.
[0115] In other words, such as Figure 6 As shown, user voice input is validated via internal validation using a validation dataset. When the first user voice input is clear, the user voice input is registered as a single utterance, and a minimum number of utterances can be made to register the user voice input as a wake-up voice input when it is registered.
[0116] Return to reference Figure 2The processor 140 can register the user's voice input as wake-up voice input through the wake-up voice input registration module 146. In this case, the wake-up voice input registration module 146 can register the user's voice input as wake-up voice input based on the verification results of the text evaluation module 143, the non-voice input evaluation module 144, and the verification set evaluation module 145. In other words, when the verification results of the text evaluation module 143, the non-voice input evaluation module 144, and the verification results of the verification set evaluation module 145 all verify that the user's voice input should be registered as wake-up voice input, the wake-up voice input registration module 146 can register the user's voice input as wake-up voice input and store the feature vector obtained by the first neural network model 30-1 (e.g., a keyword recognition model or a speech recognition model) and the feature vector obtained by the second neural network model 30-2 (speaker input model) in the memory 130.
[0117] In this case, the wake-up voice input registration module 146 can sequentially obtain the verification results of the text evaluation module 143, the non-voice input evaluation module 144, and the verification set evaluation module 145. However, this is only an embodiment, and the verification results of the text evaluation module 143, the non-voice input evaluation module 144, and the verification set evaluation module 145 can be obtained in parallel regardless of their order.
[0118] The message providing module 147 can provide a guidance message based on the verification results of the text evaluation module 143, the non-voice input evaluation module 144, and the validation set evaluation module 145. In other words, if one of the verification results of the text evaluation module 143, the non-voice input evaluation module 144, and the validation set evaluation module 145 is determined to be inappropriate, the message providing module 147 can provide a guidance message corresponding to the module identified as inappropriate.
[0119] Specifically, if the verification result of the text evaluation module 143 is deemed inappropriate, the message providing module 147 can provide a guidance message explaining the reason for the inappropriateness and alternative text. In this case, the alternative text can be determined based on the usage history database. In other words, based on the usage history database, the alternative text can be identified as text frequently used by the user, text in areas of interest to the user, etc. For example, when a user wants to register "yayaya" as wake-up voice input, the message providing module 147 can notify the user that the voice input has been repeated and provide a guidance message explaining the reason for the inappropriateness and alternative text. Figure 7aThe example shown is a guidance message 710 for guiding alternative text. In this case, "Cheolsooya," determined as the alternative text, could be text frequently used by the user, while some phonemes are repeated with "yayaya" (which is text included in the user's voice input). As another example, when the user wants to register "cha" as wake-up voice input, the message providing module 147 can provide a guidance message 720 for guiding the user's voice input and guiding alternative text, such as... Figure 7b As shown in Figure 7C. In this case, "Jadongcha," determined as the alternative text, could be text related to a field of interest to the user, while some phonemes would be repeated with "cha" (which is text included in the user's voice input). As another example, when a user wants to register "Geugue" as wake-up voice input, the message providing module 147 can notify the user that the voice input includes a pronoun, as shown in Figure 7C, and provide a guidance message 730 to guide the alternative text. In this case, "Galaxy," determined as the alternative text, could be text frequently used by the user.
[0120] When the verification result of either the non-voice input evaluation module 144 or the verification set evaluation module 145 is deemed inappropriate, the message providing module 147 can provide a guiding message requesting additional user voice input containing the same text included in the user's voice input. For example, the message providing module 147 can provide a guiding message 810 requesting additional user voice input, such as... Figure 8 As shown in the diagram. However, when the non-voice input evaluation module 144 and the validation set evaluation module 145 identify the number of inappropriate actions as exceeding a threshold number (e.g., 5 times), the message providing module 147 can provide a guiding message requesting user voice input including another text.
[0121] Figure 9 This is a flowchart illustrating a method for controlling an electronic device according to an embodiment of the present disclosure. Figure 9 This diagram illustrates the operation after the electronic device 100 enters the mode for registering wake-up voice input.
[0122] Electronic device 100 can receive user voice input (operation S910). In this case, user voice input can be received via microphone 110, but this is only an example and can be received from an external source. The user voice input may include text that the user wants to register as wake-up voice input.
[0123] The electronic device 100 can preprocess the acquired user voice input (operation S920). Specifically, the electronic device 100 can perform preprocessing on the user voice input, such as noise removal and sound quality enhancement.
[0124] The electronic device 100 can extract speech features from the preprocessed user speech input (operation S930). Specifically, the electronic device 100 can extract speech features by converting the time-dimensional user speech data into the frequency-dimensional user speech data.
[0125] The electronic device 100 can verify text included in the user's voice input (operation S940). Specifically, as shown in the reference... Figure 2 and Figure 3 The electronic device 100 can verify, through the text evaluation module 143, whether the text included in the user's voice input can be registered as text for wake-up voice input.
[0126] If the text verification result is deemed appropriate (operation S950 - No), the electronic device 100 can verify whether the non-voice input is included in the user's voice input (operation S960). Specifically, as referenced... Figure 2 and Figure 4 The electronic device 100 can verify whether non-voice input is included in the user's voice input through the non-voice input evaluation module 144.
[0127] If the non-voice input verification result is deemed appropriate (operation S970 - No), the electronic device 100 can use the verification dataset to verify the user's voice input (operation S980). Specifically, as referenced... Figure 2 , Figure 5 and Figure 6 The electronic device 100 can obtain a verification dataset through the verification set evaluation module 145 and use the verification speech included in the verification dataset to verify the user's voice input.
[0128] If the verification result is deemed appropriate using the verification dataset (Operation S990 - No), the electronic device 100 can register the user's voice input as a wake-up voice input (Operation S991).
[0129] However, if the text verification result is inappropriate (operation S950 - Yes), the electronic device 100 can provide a guidance message (operation S993). In this case, as... Figures 7a to 7c As shown, electronic device 100 can provide a guidance message that includes the reason for the inappropriateness and alternative text.
[0130] If the non-voice input verification result is identified as inappropriate (operation S970 - Yes) or if the result of verification using the verification dataset is identified as inappropriate (operation S990 - Yes), then the electronic device 100 can provide a guidance message (operation S993). In this case, such as Figure 8As shown, electronic device 100 can provide guidance messages for guiding additional phonics input to a user's voice.
[0131] exist Figure 9 The text verification (operation S940), non-voice input verification (operation S960), and verification using the verification dataset (operation S980) have been described as being performed sequentially, but this is only an example, and the text verification (operation S940), non-voice input verification (operation S960), and verification using the verification dataset (operation S980) can be performed in parallel.
[0132] Figure 10 This is a detailed block diagram illustrating the configuration of an electronic device according to embodiments of the present disclosure. Figure 10 As shown, the electronic device 1000 according to this disclosure may include a display 1010, a speaker 1020, a camera 1030, a memory 1040, a communication interface 1050, an input interface 1060, a sensor 1070, and a processor 1080. However, this configuration is exemplary, and new configurations may be added or some configurations may be omitted in addition to implementing this configuration of the present disclosure. The communication interface 1050, memory 1040, and processor 1080 may have the same configuration as the communication interface 120, memory 130, and processor 140 described with reference to FIG1, and therefore redundant descriptions will be omitted.
[0133] The display 1010 can display images acquired from an external source or images captured by the camera 1030. Furthermore, the display 1010 can display a UI screen for registering wake-up voice input, and can display inappropriate guidance messages used to guide input as a result of verification of the user's voice input.
[0134] The display 1010 can be implemented as a liquid crystal display panel (LCD), an organic light-emitting diode (OLED), etc., and in some cases, the display 1010 can be implemented as a flexible display, a transparent display, etc. However, the display 1010 according to this disclosure is not limited to a particular type.
[0135] The speaker 1020 can output voice messages. Specifically, the speaker 1020 may be included in the electronic device 1000, but this is merely an example, and may be electrically connected to the electronic device 1000 and located externally. In this case, the speaker 1020 can output voice messages guiding the user through voice verification results.
[0136] Camera 1030 can capture images. Specifically, camera 1030 can capture images including the user. In this case, the image can be a still image or a moving image. Furthermore, camera 1030 can include multiple lenses that are different from each other. Here, multiple lenses that are different from each other can include situations where each of the multiple lenses has a different field of view (FOV) and situations where each of the multiple lenses is positioned differently, etc.
[0137] Input interface 1060 may include circuitry, and processor 1080 may receive user instructions for controlling the operation of electronic device 1000 via input interface 1060. Specifically, input interface 1060 may include display 1010 as a touch screen, but this is only one example, and may include components such as buttons, microphone 110, and remote control signal receiver.
[0138] Sensor 1070 can acquire various information related to electronic device 1000. In particular, sensor 1070 may include a Global Positioning System (GPS) capable of acquiring location information of electronic device 1000, as well as various sensors, such as biometric sensors (e.g., heart rate sensors, photoplethysmography (PPG) sensors, etc.) for acquiring biometric information of users of electronic device 1000, and motion sensors for detecting movement of electronic device 1000.
[0139] The processor 1080 can be electrically connected as Figure 10 The components shown include a display 1010, a speaker 1020, a camera 1030, a memory 1040, a communication interface 1050, an input interface 1060, and a sensor 1070, which control the overall function and operation of the electronic device 1000.
[0140] Specifically, processor 1080 can use a verification dataset to verify user voice input to be registered as wake-up voice input. Specifically, processor 1080 can obtain user voice input for registration via microphone 110, input the user voice input into a learned neural network model (e.g., a keyword recognition model) to obtain a first feature vector corresponding to the text included in the user voice input, receive a verification dataset determined based on information related to the text included in the user voice input from external server 200 via communication interface 1050, obtain a second feature vector corresponding to the verification voice input by inputting the verification voice input included in the received verification dataset into the neural network model, and verify whether to register the user voice input as wake-up voice input based on the similarity between the first and second feature vectors.
[0141] Specifically, processor 1080 can recognize user voice input to obtain information about text included in the user voice input, and send the information about the text included in the user voice input to external server 200 via communication interface 1050. In this case, external server 200 can obtain verification voice input based on the length of the phoneme sequences of multiple texts stored in external server 200 and the length of the phoneme sequence commonly included in the phoneme sequence included in the information about the text. Verification voice input can be voice data corresponding to text among multiple texts stored in external server 200, the length of which is equal to or greater than a threshold in the phoneme sequence commonly included in the phoneme sequence in the information about the text.
[0142] When the similarity between the first and second feature vectors is less than a threshold, the processor 1080 can input another verification voice input included in the verification dataset into the neural network model to obtain a third feature vector corresponding to the other verification voice inputs, and verify the user voice input by comparing the similarity between the first and third feature vectors. However, when the similarity between the first and second feature vectors is equal to or greater than the threshold, the processor 1080 can provide a guiding message requesting additional user voice input for registering the wake-up voice input, such as... Figure 8 As shown in the image.
[0143] When the similarity between the first feature vector and the feature vector corresponding to all the verification voices included in the verification dataset is less than a threshold, the processor 1080 can register the user's voice input as a wake-up voice input.
[0144] Furthermore, the processor 1080 can verify the text included in the user's voice input. Specifically, the processor 1080 can input the user's voice input into a speech recognition model to obtain text included in the user's voice input, and verify whether to register the text included in the user's voice input as the text for waking up the voice input based on at least one of the length and repetition of the phonemes in the text included in the user's voice input. In other words, when the number of phonemes in the text included in the user's voice input is less than a first threshold or the number of phonemes in the text included in the user's voice input repeats more than a second threshold, such as... Figures 7a to 7c As shown, the processor 1080 can provide a guiding message that requests the utterance of additional user voice input, including other text for registering wake-up voice input. In this case, the guiding message may include a message for recommending text determined based on the usage history information of the electronic device as the text for wake-up voice input.
[0145] Furthermore, the processor 1080 can verify whether non-voice input is included in the user's voice input. Specifically, the processor 1080 can input the user's voice input into a learned speech determination model to obtain feature values indicating whether the user's voice input is the user's voice input uttering a specific text, and verify whether the user's voice input is registered via wake-up voice input based on the feature values. In this case, if the feature value is less than a threshold, the processor 1080 can provide a guiding message requesting additional user voice input for registration of wake-up voice input, such as... Figure 8 As shown in the image.
[0146] According to the embodiments of this disclosure as described above, the convenience or usability of the wake-up voice input registration process can be increased by minimizing the number of times a user speaks to register the wake-up voice input, and the performance of the technology capable of controlling electronic devices via voice recognition can be improved by registering high-quality wake-up voice input.
[0147] The artificial intelligence-related functions according to this disclosure are operated via processor 1080 and memory 1040. Processor 1080 may include one or more processors. In this case, the one or more processors may include general-purpose processors such as CPUs, APs, digital signal processors (DSPs), etc., and single graphics processors such as graphics processing units (GPUs) and vision processing units (VPUs). Alternatively, it may be a processor dedicated to artificial intelligence such as neural processing units (NPUs).
[0148] One or more processors 1080 can be controlled to process input data according to predetermined operating rules or artificial intelligence models stored in memory 1040. Alternatively, when one or more processors are a single AI processor, the single AI processor can be designed with a hardware architecture dedicated to processing a specific AI model.
[0149] The functions related to the neural network model described above can be executed via memory and a processor. The processor may include one or more processors. In this case, the one or more processors may be general-purpose processors (such as CPUs and APs), GPUs, single graphics processing units (such as VPUs), or single artificial intelligence processors (such as NPUs). The one or more processors are controlled to process input data according to predetermined operating rules or artificial intelligence models stored in non-volatile memory and volatile memory. The predetermined operating rules or artificial intelligence models are configured to be generated through learning.
[0150] Here, learning refers to generating predetermined operating rules or artificial intelligence models by applying a learning algorithm to multiple learning data to produce desired characteristics. This learning can be performed within the device itself, on which the artificial intelligence according to this disclosure is executed, or via a separate server / system.
[0151] Artificial intelligence models can consist of multiple neural network layers. Each layer has multiple weight values, and layer operations are performed by combining the operations of the previous layer with the operations of the multiple weight values. Examples of neural networks include convolutional neural networks (CNNs), deep neural networks (DNNs), recurrent neural networks (RNNs), restricted Boltzmann machines (RBMs), deep belief networks (DBNs), bidirectional recurrent deep neural networks (BRDNNs), generative adversarial networks (GANs), and deep Q-networks, and the neural networks in this disclosure are not limited to the examples mentioned above, unless otherwise stated.
[0152] A learning algorithm is a method of training a predetermined target device (e.g., a robot) using multiple training data sets, enabling the predetermined target device to make decisions or predictions on its own. Examples of learning algorithms may include supervised learning, unsupervised learning, semi-supervised learning, or reinforcement learning, and the learning algorithms in this disclosure are not limited to the examples described above.
[0153] Machine-readable storage media may be provided in the form of non-transitory storage media. "Non-transitory storage media" means that the storage medium does not contain signals (e.g., electromagnetic waves) and is tangible, but does not distinguish whether data is stored semi-permanently or temporarily in the storage medium. For example, the term "non-transitory" may include a cache for temporarily storing data.
[0154] Furthermore, according to embodiments, the methods according to the various embodiments described above can be provided as part of a computer program product. The computer program product can be traded between a seller and a buyer. The computer program product can be distributed in the form of a machine-readable storage medium (e.g., an optical disc read-only memory (CD-ROM)) or through an app store (e.g., the Play Store). TM Online distribution. In the case of online distribution, at least a portion of a computer program product (e.g., a downloadable app) may be temporarily stored or generated on a storage medium (such as memory in a manufacturer's server, an app store's server, or a relay server).
[0155] Furthermore, each of the components (e.g., modules or programs) according to the various embodiments described above may consist of a single entity or multiple entities, and some of the sub-components described above may be omitted, or other sub-components may be included in the various embodiments. Typically, or additionally, some components (e.g., modules or programs) may be integrated into a single entity to perform the same or similar functions performed by each respective component prior to integration. According to the various embodiments, the operations performed by modules, programs, or other components may be performed sequentially, in parallel, or both iteratively or tentatively, or at least some operations may be performed in a different order, omitted, or additional operations may be added.
[0156] According to various exemplary embodiments, the operations performed by a module, program module or other component may be performed sequentially, in parallel, or both iteratively or tentatively, or at least some operations may be performed in a different order, omitted, or additional operations may be added.
[0157] As used herein, the term "module" includes units comprised of hardware, software, or firmware, and is used interchangeably with terms such as logic, logic block, component, or circuit. A "module" can be an integral building block or the smallest unit or part thereof that performs one or more functions. For example, a module can be configured as an application-specific integrated circuit (ASIC).
[0158] According to embodiments, the various embodiments described above can be implemented as software including instructions stored in a machine-readable storage medium that can be read by a machine (e.g., a computer). The device may include an electronic device according to the disclosed embodiments, as a device that invokes stored instructions from the storage medium and is capable of operating according to the invoked instructions.
[0159] When an instruction is executed by a processor, the processor may instruct other components to perform the function corresponding to the instruction, or the function may be performed under the control of the processor. Instructions may include code generated or executed by a compiler or interpreter.
[0160] The foregoing exemplary embodiments and advantages are merely illustrative and should not be construed as limiting this disclosure. This teaching can be readily applied to other types of devices. Furthermore, the description of exemplary embodiments of this disclosure is intended to be illustrative and not to limit the scope of the claims, and many substitutions, modifications, and variations will be apparent to those skilled in the art.
Claims
1. Electronic devices, including: microphone; Communication interface; The memory is configured to store at least one instruction. as well as Processor, configured to execute at least one of the instructions to: User voice input for registering wake-up voice input is obtained via the microphone; The user's voice input is fed into a trained neural network model to obtain a first feature vector corresponding to the text included in the user's voice input; The verification dataset is received from an external server via the communication interface. The verification dataset is determined based on information relating to the text included in the user's voice input. The verification speech input included in the verification dataset is input into the trained neural network model to obtain a second feature vector corresponding to the verification speech input; as well as Based on the similarity between the first feature vector and the second feature vector, it is determined whether the user's voice input should be registered as the wake-up voice input.
2. The electronic device according to claim 1, wherein, The processor is also configured to: Recognize the user's voice input to obtain the information relating to the text included in the user's voice input; and The information relating to the text included in the user's voice input is sent to the external server via the communication interface. The external server is configured to: use the information related to the text to obtain the verification voice input based on a first phoneme sequence of the verification voice text, wherein the first phoneme sequence of the verification voice text is a number of common phonemes jointly included by a second phoneme sequence of the text stored in the external server and a phoneme sequence of the information related to the text.
3. The electronic device according to claim 2, wherein, The verification voice input is voice data corresponding to the verification voice text whose number of common phonemes is equal to or greater than a threshold.
4. The electronic device according to claim 2, wherein, The processor is also configured to: Based on the fact that the similarity between the first feature vector and the second feature vector is less than a threshold, another verification speech input included in the verification dataset is input into the trained neural network model to obtain a third feature vector corresponding to the other verification speech input; Compare another similarity between the first feature vector and the third feature vector; and Based on the fact that the other similarity between the first feature vector and the third feature vector is equal to or greater than the threshold, a guiding message is provided to request additional user voice input for registering wake-up voice input.
5. The electronic device according to claim 4, wherein, The processor is also configured to: The user's voice input is registered as the wake-up voice input based on multiple similarities between the feature vectors corresponding to all verified voice inputs included in the verification dataset and the first feature vector, which are less than the threshold.
6. The electronic device according to claim 1, wherein, The processor is also configured to: The user's voice input is fed into a speech recognition model to obtain text included in the user's voice input; and Based on at least one of the length and repetition of phonemes included in the user's voice input, it is determined whether the text included in the user's voice input is registered as the wake-up voice input.
7. The electronic device according to claim 6, wherein, The processor is also configured to: Based on the fact that the number of phonemes in the text included in the user's voice input is less than a first threshold or the number of repeating phonemes in the text included in the user's voice input is greater than a second threshold, a guiding message is provided requesting the pronunciation of additional user voice input, including another text for registering the wake-up voice input.
8. The electronic device according to claim 7, wherein, The guidance message is configured to include a message recommending another text, determined based on the usage history information of the electronic device, as the wake-up voice input.
9. The electronic device according to claim 1, wherein, The processor is also configured to: The user's voice input is fed into a trained speech recognition model to obtain feature values that indicate whether the user's voice input is a user voice input that utters a specific text; as well as Based on the feature value, it is determined whether the user's voice input should be registered as the wake-up voice input.
10. The electronic device according to claim 9, wherein, The processor is also configured to: Based on the fact that the feature value is less than the threshold, a guiding message is provided to request additional user voice input for registering the wake-up voice input.
11. A method for controlling an electronic device, the method comprising: Obtain user voice input for registering wake-up voice input; The user's voice input is fed into a trained neural network model to obtain a first feature vector corresponding to the text included in the user's voice input; Receive a verification dataset from an external server, the verification dataset being determined based on information relating to the text included in the user's voice input; The verification speech input included in the verification dataset is input into the trained neural network model to obtain a second feature vector corresponding to the verification speech input; as well as Based on the similarity between the first feature vector and the second feature vector, it is determined whether the user's voice input should be registered as the wake-up voice input.
12. The method of claim 11, further comprising: Recognize the user's voice input to obtain the information relating to the text included in the user's voice input; as well as Send the information relating to the text included in the user's voice input to the external server. The external server is configured to obtain the verification voice input based on a first phoneme sequence of the verification voice text, wherein the first phoneme sequence of the verification voice text is a number of common phonemes comprised of a second phoneme sequence of the text stored in the external server and a phoneme sequence of the information related to the text.
13. The method according to claim 12, wherein, The verification voice input is voice data corresponding to the verification voice text whose number of common phonemes is equal to or greater than a threshold.
14. The method of claim 12, further comprising: Based on the fact that the similarity between the first feature vector and the second feature vector is less than a threshold, another verification speech input included in the verification dataset is input into the trained neural network model to obtain a third feature vector corresponding to the other verification speech input; Compare another similarity between the first feature vector and the third feature vector; and Based on the fact that the other similarity between the first feature vector and the third feature vector is equal to or greater than the threshold, a guiding message is provided to request additional user voice input for registering wake-up voice input.
15. The method of claim 14, further comprising: The user's voice input is registered as the wake-up voice input based on multiple similarities between the feature vectors corresponding to all verified voice inputs included in the verification dataset and the first feature vector, which are less than the threshold.
Citation Information
Patent Citations
Method and device for awakening intelligent equipment
CN110515449A
Speech-controlled apparatus for preventing false detections of keyword and method of operating the same
KR1020180127065A