Artificial intelligence apparatus and method for recognizing plurality of wake-up word

KR103022265B1Active Publication Date: 2026-09-22LG ELECTRONICS INC
View PDF 12 Cites 0 Cited by

Patent Information

Application Number
KR1020200120742
Authority / Receiving Office
KR · KR
Patent Type
Patents
Current Assignee / Owner
Filing Date
2020-09-18
Publication Date
2026-09-22
Estimated Expiration
2040-09-18

Smart Images

  • Figure 112020099524094-PAT00006_ABST
    Figure 112020099524094-PAT00006_ABST
Patent Text Reader

Abstract

One embodiment of the present disclosure provides an artificial intelligence device for recognizing a plurality of command words, comprising: a microphone; a memory storing a first command word recognition engine; a communication unit communicating with an artificial intelligence server storing a second command word recognition engine; and a processor that receives an input audio signal through the microphone, generates a preprocessed audio signal from the input audio signal, extracts a voice segment from the preprocessed audio signal, sets a command word recognition segment including the voice segment and a buffer segment corresponding to the voice segment among the preprocessed audio signals, and transmits the command word recognition segment among the preprocessed audio signals to the first command word recognition engine and the second command word recognition engine.
Need to check novelty before this filing date? Find Prior Art

Description

Technology Field

[0001] The present disclosure relates to an artificial intelligence device and a method for recognizing a plurality of command words. Background Technology

[0002] Recently, there has been an increase in artificial intelligence devices equipped with voice recognition functions that recognize the user's spoken voice. These voice recognition functions are typically configured to be activated by predetermined button inputs, touch inputs, or voice inputs, and voice input may refer to the recognition of a predetermined trigger word (or voice recognition trigger word). Since the voice recognition trigger word must be recognized to determine whether the voice recognition function is activated, the trigger word recognition model for recognizing the voice recognition trigger word is almost always active, and consequently, a considerable amount of resources are required for trigger word recognition.

[0003] Since different speech recognition platforms recognize different commands using different command recognition engines, multiple command recognition engines must be installed in a single AI device to support multiple speech recognition platforms. Furthermore, since multiple command recognition engines operate continuously to recognize their respective commands, this requires significant resources, leading to a problem where the utilization of the processor or central processing unit (CPU) increases significantly. When CPU utilization by command recognition engines is high in this manner, the AI ​​device may slow down the execution of other high-load tasks, or conversely, command recognition may fail to function properly during the execution of other high-load tasks. The problem to be solved

[0004] The present disclosure aims to provide an artificial intelligence device and a method for recognizing a plurality of voice recognition trigger words.

[0005] In addition, the present disclosure aims to provide an artificial intelligence device and a method for recognizing a plurality of speech recognition start words only in a portion of an input audio signal. means of solving the problem

[0006] One embodiment of the present disclosure provides an artificial intelligence device for recognizing a plurality of command words, comprising: a microphone; a memory storing a first command word recognition engine; a communication unit communicating with an artificial intelligence server storing a second command word recognition engine; and a processor that receives an input audio signal through the microphone, generates a preprocessed audio signal from the input audio signal, extracts a voice segment from the preprocessed audio signal, sets a command word recognition segment including the voice segment and a buffer segment corresponding to the voice segment among the preprocessed audio signals, and transmits the command word recognition segment among the preprocessed audio signals to the first command word recognition engine and the second command word recognition engine.

[0007] The processor can extract the voice segment from the preprocessed audio signal through a voice activity detection (VAD) function.

[0008] The processor may set a preceding section of a first length from the voice section as a first buffer section, set a subsequent section of a second length from the voice section as a second buffer section, and set a starting word recognition section including the voice section, the first buffer section, and the second buffer section.

[0009] The above processor can obtain a command word recognition result for a first command word through the first command word recognition engine and obtain a command word recognition result for a second command word through the second command word recognition engine.

[0010] The above processor may disable the voice activity detection function when the first trigger word or the second trigger word is recognized, obtain a voice recognition result for a command recognition section after the trigger word section for the recognized trigger word among the preprocessed audio signals, perform an operation based on the voice recognition result, and enable the voice activity detection function.

[0011] The processor obtains the speech recognition result for the command recognition section using speech engines of a speech recognition platform corresponding to the recognized command word, and the speech engines may include a Speech-To-Text (STT) engine, a Natural Language Processing (NLP) engine, and a speech synthesis engine.

[0012] The processor can transmit the command word recognition section to the artificial intelligence server through an API (Application Programming Interface) for the second command word recognition engine and obtain a command word recognition result for the second command word.

[0013] The processor can obtain a probability of voice presence from the preprocessed audio signal using the voice activity detection function and extract the voice segment using the probability of voice presence.

[0014] The processor can extract a section where the probability of the voice existence is greater than the first reference value as the voice section.

[0015] The processor can extract a section as the voice section where the value obtained by multiplying the amplitude of the preprocessed audio signal and the probability of voice existence is greater than the second reference value.

[0016] The above processor can disable the voice activity detection function when the operating mode is voice registration mode, and enable the voice activity detection function after the voice registration function is terminated.

[0017] Additionally, one embodiment of the present disclosure provides a method for recognizing a plurality of command words or a recording medium recording the same, comprising the steps of: receiving an input audio signal through a microphone; generating a preprocessed audio signal from the input audio signal; extracting a voice segment from the preprocessed audio signal; setting a command word recognition segment among the preprocessed audio signal, the voice segment and a buffer segment corresponding to the voice segment; and transmitting the command word recognition segment among the preprocessed audio signal to a first command word recognition engine stored in memory and a second command word recognition engine stored in an artificial intelligence server. Effects of the invention

[0018] According to various embodiments of the present disclosure, a plurality of speech recognition models can be incorporated to support multiple speech recognition platforms in a single artificial intelligence device.

[0019] In addition, according to various embodiments of the present disclosure, even if a plurality of command word recognition models are installed, the resources consumed by the plurality of command word recognition models in an idle state can be effectively reduced. Brief explanation of the drawing

[0020] FIG. 1 is a block diagram showing an artificial intelligence device (100) according to one embodiment of the present disclosure. FIG. 2 is a block diagram showing a remote control device (200) according to one embodiment of the present disclosure. FIG. 3 is a drawing showing a remote control device (200) according to one embodiment of the present disclosure. FIG. 4 illustrates an example of interacting with an artificial intelligence device (100) through a remote control device (200) in one embodiment of the present disclosure. FIG. 5 is a block diagram showing an artificial intelligence server (400) according to one embodiment of the present disclosure. FIG. 6 is a flowchart illustrating a method for recognizing a plurality of voice recognition trigger words according to one embodiment of the present disclosure. FIG. 7 is a diagram showing voice servers according to one embodiment of the present disclosure. Figure 8 is a diagram showing an example of a preprocessed audio signal and a corresponding command word recognition section. FIG. 9 is a flowchart illustrating an example of the step (S613) of providing a voice recognition service illustrated in FIG. 6. Figure 10 is a diagram showing an example of controlling a voice activity detection function based on activation word recognition. FIG. 11 is a flowchart illustrating a method for recognizing a plurality of voice recognition trigger words according to one embodiment of the present disclosure. FIG. 12 is a drawing showing a voice registration interface according to one embodiment of the present disclosure. FIG. 13 is a diagram showing an example of controlling the voice activity detection function in voice registration mode. FIG. 14 is a ladder diagram illustrating a method for recognizing a plurality of activators according to one embodiment of the present disclosure. FIG. 15 is a flowchart illustrating an example of the step (S605) of extracting a voice segment illustrated in FIG. 6. FIG. 16 is a flowchart showing an example of the step (S605) of extracting a voice segment illustrated in FIG. 6. Figure 17 is a diagram illustrating a method for extracting speech segments from a preprocessed audio signal. Specific details for implementing the invention

[0021] Hereinafter, embodiments disclosed in this specification will be described in detail with reference to the attached drawings. Identical or similar components, regardless of drawing symbols, are assigned the same reference number, and redundant descriptions thereof will be omitted. The suffixes 'module' and 'part' for components used in the following description are assigned or used interchangeably solely for the ease of drafting the specification and do not inherently possess distinct meanings or roles. Furthermore, in describing embodiments disclosed in this specification, if it is determined that a detailed description of related prior art could obscure the essence of the embodiments disclosed in this specification, such detailed description will be omitted. Additionally, the attached drawings are intended only to facilitate understanding of the embodiments disclosed in this specification; the technical concept disclosed in this specification is not limited by the attached drawings, and it should be understood that they include all modifications, equivalents, and substitutions that fall within the concept and technical scope of this disclosure.

[0022] Terms including ordinal numbers, such as first, second, etc., may be used to describe various components, but said components are not limited by said terms. These terms are used solely for the purpose of distinguishing one component from another.

[0023] When it is stated that one component is 'connected' or 'connected' to another component, it should be understood that while it may be directly connected or connected to that other component, there may also be other components in between. On the other hand, when it is stated that one component is 'directly connected' or 'directly connected' to another component, it should be understood that there are no other components in between.

[0025] FIG. 1 is a block diagram showing an artificial intelligence device (100) according to one embodiment of the present disclosure.

[0026] Referring to FIG. 1, the artificial intelligence device (100) can be connected to at least one of a remote control device (200), a user terminal (300), an artificial intelligence server (400), or a content provider (500) to transmit or receive data or signals.

[0027] The artificial intelligence device (100) may be a display device capable of outputting images, including a display unit (180, or a display panel). For example, the artificial intelligence device (100) may be implemented as a fixed device or a mobile device, such as a TV, projector, mobile phone, smartphone, desktop computer, laptop, digital broadcasting terminal, PDA (personal digital assistants), PMP (portable multimedia player), navigation, tablet PC, wearable device, set-top box (STB), DMB receiver, radio, speaker, washing machine, refrigerator, digital signage, robot, vehicle, etc.

[0028] The user terminal (300) can be implemented as a mobile phone, smartphone, tablet PC, laptop, wearable device, PDA, etc. The user terminal (300) may also simply be referred to as a terminal (300).

[0029] The content provider (500) refers to a device that provides content data corresponding to the content to be output by the artificial intelligence device (100), and the artificial intelligence device (100) can receive content data from the content provider (500) and output the content.

[0030] The artificial intelligence device (100) may include a communication unit (110), a broadcast receiving unit (130), an external device interface unit (135), a memory (140), an input unit (150), a processor (170), a display unit (180), an audio output unit (185), and a power supply unit (190).

[0031] The communication unit (110) can communicate with an external device via wired or wireless communication. For example, the communication unit (110) can transmit and receive sensor information, user input, learning models, control signals, etc., with external devices such as other artificial intelligence devices. Here, the other artificial intelligence device (100) may be a mobile terminal such as a wearable device (e.g., a smartwatch, smart glass, HMD (head mounted display)) or a smartphone that is capable of exchanging data with (or interoperable with) the artificial intelligence device (100) according to the present invention.

[0032] The communication unit (110) can detect (or recognize) a communicable wearable device around the artificial intelligence device (100). Furthermore, if the detected wearable device is a device authenticated to communicate with the artificial intelligence device (100), the processor (170) can transmit at least a portion of the data processed by the artificial intelligence device (100) to the wearable device through the communication unit (110). Thus, the user of the wearable device can use the data processed by the artificial intelligence device (100) through the wearable device.

[0033] The communication technologies used by the communication department (110) include GSM (Global System for Mobile communication), CDMA (Code Division Multi Access), LTE (Long Term Evolution), 5G, WLAN (Wireless LAN), Wi-Fi (Wireless-Fidelity), Bluetooth (Bluetooth), RFID (Radio Frequency Identification), Infrared Data Association (IrDA), ZigBee, NFC (Near Field Communication), etc.

[0034] The communication unit (110) may also be called a communication modem or a communication interface.

[0035] The broadcast receiving unit (130) may include a tuner (131), a demodulating unit (132), and a network interface unit (133).

[0036] The tuner (131) can tune to a specific broadcast channel according to a channel tuning command. The tuner (131) can receive a broadcast signal for the tuned specific broadcast channel.

[0037] The demodulator (132) can separate the received broadcast signal into a video signal, an audio signal, and a data signal related to the broadcast program, and can restore the separated video signal, audio signal, and data signal into a form that can be output.

[0038] The external device interface unit (135) can receive an application or a list of applications within an adjacent external device and transmit it to a processor (170) or memory (140).

[0039] The external device interface section (135) can provide a connection path between the artificial intelligence device (100) and an external device. The external device interface section (135) can receive one or more of video and audio output from an external device connected to the artificial intelligence device (100) wirelessly or via a wired connection and transmit them to the processor (170). The external device interface section (135) may include a plurality of external input terminals. The plurality of external input terminals may include an RGB terminal, one or more HDMI (High Definition Multimedia Interface) terminals, and a component terminal.

[0040] The video signal of an external device input through the external device interface unit (135) can be output through the display unit (180). The voice signal of an external device input through the external device interface unit (135) can be output through the audio output unit (185).

[0041] The external device that can be connected to the external device interface section (135) may be any one of a set-top box, Blu-ray player, DVD player, game console, soundbar, smartphone, PC, USB memory, or home theater, but this is merely an example.

[0042] The network interface unit (133) may provide an interface for connecting the artificial intelligence device (100) to a wired / wireless network including the Internet. The network interface unit (133) may transmit or receive data to or from other users or other electronic devices through the connected network or another network linked to the connected network.

[0043] Some content data stored in the artificial intelligence device (100) can be transmitted to another user or other electronic device selected among other users or other electronic devices that are previously registered in the artificial intelligence device (100).

[0044] The network interface unit (133) can access a specific web page through a connected network or another network linked to the connected network. That is, it can access a specific web page through a network and transmit or receive data with the corresponding server.

[0045] The network interface unit (133) can receive content or data provided by a content provider or network operator. That is, the network interface unit (133) can receive content such as movies, advertisements, games, VOD, broadcast signals, and related information provided by a content provider or network provider through a network.

[0046] The network interface unit (133) can receive firmware update information and update files provided by the network operator and can transmit data to the internet or a content provider or network operator.

[0047] The network interface unit (133) can select and receive a desired application among the applications that are open to the public through the network.

[0048] The memory (140) can store programs for each signal processing and control within the processor (170), and can store signal-processed video, audio, or data signals. For example, the memory (140) can store input data, learning data, learning models, learning history, etc., obtained from the input unit (150).

[0049] The memory (140) may perform the function of temporarily storing video, audio, or data signals input from the external device interface unit (135) or the network interface unit (133), and may also store information regarding a predetermined image through a channel memory function.

[0050] The memory (140) can store an application or a list of applications input from an external device interface unit (135) or a network interface unit (133).

[0051] The artificial intelligence device (100) can play content files (video files, still image files, music files, document files, application files, etc.) stored in memory (140) and provide them to the user.

[0052] The input unit (150) can acquire various types of data. The input unit (150) may include a camera for inputting video signals, a microphone for receiving audio signals, a user input unit for receiving information from a user, etc.

[0053] The user input unit can transmit a signal input by the user to the processor (170) or transmit a signal from the processor (170) to the user. For example, the user input unit can receive and process control signals such as power on / off, channel selection, and screen setting from the remote control device (200) according to various communication methods such as Bluetooth, Ultra Wideband (WB), ZigBee, Radio Frequency (RF) communication, or Infrared (IR) communication, or process control signals from the processor (170) to be transmitted to the remote control device (200).

[0054] The user input unit can transmit control signals input from local keys (not shown), such as power key, channel key, volume key, and setting value, to the processor (170).

[0055] The learning processor (160) can train a model composed of an artificial neural network using training data. Here, the trained artificial neural network may be referred to as a learning model. The learning model can be used to infer a result value for new input data other than the training data, and the inferred value can be used as a basis for judgment to perform an action.

[0056] The learning processor (160) can perform artificial intelligence processing together with the learning processor (440) of the artificial intelligence server (400).

[0057] The learning processor (160) may include memory integrated into or implemented in the artificial intelligence device (100). Alternatively, the learning processor (160) may be implemented using memory (170), external memory directly coupled to the artificial intelligence device (100), or memory maintained in an external device.

[0058] The image signal processed by the processor (170) can be input to the display unit (180) and displayed as an image corresponding to the image signal. Additionally, the image signal processed by the processor (170) can be input to an external output device through the external device interface unit (135).

[0059] The voice signal processed by the processor (170) can be output as audio to the audio output unit (185). Additionally, the voice signal processed by the processor (170) can be input to an external output device through the external device interface unit (135).

[0060] The processor (170) can control the overall operation of the artificial intelligence device (100).

[0061] The processor (170) can control the artificial intelligence device (100) by means of a user command or internal program entered through the user input unit, and can connect to a network to allow the user to download an application or a list of applications desired by the user into the artificial intelligence device (100).

[0062] The processor (170) enables the processed video or audio signal, such as channel information selected by the user, to be output through the display unit (180) or audio output unit (185).

[0063] The processor (170) enables a video signal or audio signal from an external device, such as a camera or camcorder, which is input through the external device interface unit (135), to be output through the display unit (180) or audio output unit (185) in accordance with an external device video playback command received through the user input unit.

[0064] The processor (170) can control the display unit (180) to display an image, for example, a broadcast image input through the tuner (131), an external input image input through the external device interface unit (135), an image input through the network interface unit (133), or an image stored in the memory (140) can be controlled to be displayed on the display unit (180). In this case, the image displayed on the display unit (180) may be a still image or a video, and may be a 2D image or a 3D image.

[0065] The processor (170) can control the playback of content stored in the artificial intelligence device (100), received broadcast content, or external input content input from the outside, and the content may be in various forms such as broadcast video, external input video, audio file, still image, connected web screen, and document file.

[0066] The processor (170) can determine at least one executable operation of the artificial intelligence device (100) based on information determined or generated using a data analysis algorithm or a machine learning algorithm. The processor (170) can control the components of the artificial intelligence device (100) to perform the determined operation.

[0067] To this end, the processor (170) can request, search, receive, or utilize data from the learning processor (160) or memory (140), and can control the components of the artificial intelligence device (100) to execute a predicted operation or a preferred operation among the at least one executable operation.

[0068] The processor (170) can obtain intent information regarding user input and determine the user's requirements based on the obtained intent information.

[0069] The processor (170) can obtain intent information corresponding to user input by using at least one of a Speech To Text (STT) engine for converting voice input into a string or a Natural Language Processing (NLP) engine for obtaining intent information of natural language.

[0070] At least one of the STT engine or NLP engine may be composed of an artificial neural network, at least a portion of which is trained according to a machine learning algorithm. Additionally, at least one of the STT engine or NLP engine may be trained by a learning processor (160), trained by a learning processor (440) of an artificial intelligence server (400), or trained by distributed processing thereof.

[0071] The processor (170) can collect history information, including the operation details of the artificial intelligence device (100) or user feedback regarding the operation, and store it in memory (140) or a learning processor (160), or transmit it to an external device such as an artificial intelligence server (400). The collected history information can be used to update the learning model.

[0072] The display unit (180) can output an image by converting the video signal, data signal, OSD signal processed by the processor (170) or the video signal, data signal, etc. received from the external device interface unit (135) into R, G, and B signals, respectively.

[0073] Meanwhile, the artificial intelligence device (100) illustrated in FIG. 1 is merely an example of the present disclosure, and some of the illustrated components may be integrated, added, or omitted depending on the specifications of the artificial intelligence device (100) actually implemented.

[0074] In one embodiment, two or more components of the artificial intelligence device (100) may be combined into a single component, or a single component may be subdivided into two or more components. Additionally, the functions performed in each block are intended to explain the embodiments of the present disclosure, and the specific operations or devices thereof do not limit the scope of the rights of the present disclosure.

[0075] According to one embodiment of the present disclosure, the artificial intelligence device (100) may receive and play video through a network interface unit (133) or an external device interface unit (135) without having a tuner (131) and a demodulator (132), unlike as shown in FIG. 1. For example, the artificial intelligence device (100) may be implemented by being separated into a video processing device, such as a set-top box, for receiving broadcast signals or content according to various network services, and a content playback device for playing content input from said video processing device. In this case, the method of operation of the artificial intelligence device according to one embodiment of the present disclosure described below may be performed not only by the artificial intelligence device (100) as described with reference to FIG. 1, but also by any one of the separated video processing device, such as a set-top box, or a content playback device having a display unit (180) and an audio output unit (185).

[0077] FIG. 2 is a block diagram showing a remote control device (200) according to one embodiment of the present disclosure.

[0078] Referring to FIG. 2, the remote control device (200) may include a fingerprint recognition unit (210), a communication unit (220), a user input unit (230), a sensor unit (240), an output unit (250), a power supply unit (260), a memory (270), a processor (280), and a voice acquisition unit (290).

[0079] The communication unit (220) can transmit and receive signals with any one of the artificial intelligence devices (100) according to the embodiments of the present disclosure described above.

[0080] The remote control device (200) is equipped with an RF module (221) capable of transmitting and receiving signals to and from the artificial intelligence device (100) according to RF communication standards, and may be equipped with an IR module (223) capable of transmitting and receiving signals to and from the artificial intelligence device (100) according to IR communication standards. Additionally, the remote control device (200) may be equipped with a Bluetooth module (225) capable of transmitting and receiving signals to and from the artificial intelligence device (100) according to Bluetooth communication standards. Furthermore, the remote control device (200) may be equipped with an NFC module (227) capable of transmitting and receiving signals to and from the artificial intelligence device (100) according to NFC (Near Field Communication) communication standards, and may be equipped with a WLAN module (229) capable of transmitting and receiving signals to and from the artificial intelligence device (100) according to WLAN (Wireless LAN) communication standards.

[0081] A signal containing information regarding the movement of the remote control device (200), etc., can be transmitted to the artificial intelligence device (100) through the communication unit (220).

[0082] The remote control device (200) can receive a signal transmitted by the artificial intelligence device (100) through the RF module (221), and, if necessary, can transmit commands regarding power on / off, channel change, volume change, etc. to the artificial intelligence device (100) through the IR module (223).

[0083] The user input unit (230) may be composed of a keypad, buttons, a touchpad, or a touch screen. The user can input commands related to the artificial intelligence device (100) to the remote control device (200) by operating the user input unit (230). If the user input unit (230) is equipped with a hard key button, the user can input commands related to the artificial intelligence device (100) to the remote control device (200) through a push operation of the hard key button.

[0084] If a touch screen is included in the user input unit (230), the user can input commands related to the artificial intelligence device (100) to the remote control device (200) by touching the soft keys on the touch screen. Additionally, the user input unit (230) may be equipped with various types of input means that the user can operate, such as a scroll key or a jog key.

[0085] The sensor unit (240) may be equipped with a gyroscope sensor (241) or an accelerometer sensor (243), and the gyroscope sensor (241) may sense information regarding the movement of the remote control device (200). For example, the gyroscope sensor (241) may sense information regarding the operation of the remote control device (200) based on the x, y, and z axes, and the accelerometer sensor (243) may sense information regarding the movement speed of the remote control device (200). Meanwhile, the remote control device (200) may further be equipped with a distance measuring sensor to sense the distance to the display unit (180) of the artificial intelligence device (100).

[0086] The output unit (250) can output a video or audio signal corresponding to the operation of the user input unit (230) or a signal transmitted from the artificial intelligence device (100). The user can recognize whether the user input unit (230) is operated or whether the artificial intelligence device (100) is controlled through the output unit (250). For example, the output unit (250) may be equipped with an LED module (251) that lights up when the user input unit (230) is operated or when a signal is transmitted or received with the artificial intelligence device (100) through the communication unit (225), a vibration module (253) that generates vibration, a sound output module (255) that outputs sound, or a display module (257) that outputs video.

[0087] The power supply unit (260) can supply power to the remote control device (200). The power supply unit (260) can reduce power waste by stopping the power supply when the remote control device (200) does not move for a predetermined period of time. The power supply unit (260) can resume the power supply when a predetermined key provided in the remote control device (200) is operated.

[0088] The memory (270) can store various types of programs, application data, etc. required for the control or operation of the remote control device (200).

[0089] When the remote control device (200) wirelessly transmits and receives signals to and from the artificial intelligence device (100) through the RF module (221), the remote control device (200) and the artificial intelligence device (100) can transmit and receive signals through a predetermined frequency band. To this end, the processor (280) of the remote control device (200) can store and refer to information regarding the frequency band, etc., for which signals can be wirelessly transmitted and received to and from the artificial intelligence device (100) paired with the remote control device (200) in the memory (270).

[0090] The processor (280) can control all matters related to the control of the remote control device (200). The processor (280) can transmit a signal corresponding to a predetermined key operation of the user input unit (230) or a signal corresponding to the movement of the remote control device (200) sensed by the sensor unit (240) to the artificial intelligence device (100) through the communication unit (225).

[0091] The voice acquisition unit (290) can acquire voice. The voice acquisition unit (290) may include at least one microphone (291) and can acquire voice through the microphone (291).

[0093] FIG. 3 is a drawing showing a remote control device (200) according to one embodiment of the present disclosure.

[0094] Referring to FIG. 3, the remote control device (200) may include a plurality of buttons. The plurality of buttons included in the remote control device (200) may include a fingerprint recognition button (212), a power button (231), a home button (232), a live button (233), an external input button (234), a volume control button (235), a voice recognition button (236), a channel change button (237), a confirmation button (238), and a back button (239).

[0095] The fingerprint recognition button (212) may be a button for recognizing a user's fingerprint. In one embodiment, the fingerprint recognition button (212) may be capable of a push operation and may receive a push operation and a fingerprint recognition operation. The power button (231) may be a button for turning the power of the artificial intelligence device (100) on / off. The home button (232) may be a button for moving to the home screen of the artificial intelligence device (100). The live button (233) may be a button for displaying a real-time broadcast program. The external input button (234) may be a button for receiving an external input connected to the artificial intelligence device (100). The volume control button (235) may be a button for adjusting the volume output by the artificial intelligence device (100). The voice recognition button (236) may be a button for receiving the user's voice and recognizing the received voice. The channel change button (237) may be a button for receiving a broadcast signal of a specific broadcast channel. The confirmation button (238) may be a button for selecting a specific function, and the back button (239) may be a button for returning to the previous screen.

[0097] FIG. 4 illustrates an example of interacting with an artificial intelligence device (100) through a remote control device (200) in one embodiment of the present disclosure.

[0098] Referring to FIG. 4, a pointer (205) corresponding to a remote control device (200) can be displayed on the display unit (180).

[0099] Referring to FIG. 4 (a), the user can move or rotate the remote control device (200) up / down and left / right.

[0100] The pointer (205) displayed on the display unit (180) of the artificial intelligence device (100) can move in response to the movement of the remote control device (200). Since the pointer (205) moves and is displayed according to the movement of the remote control device (200) in 3D space, it can be named a spatial remote control.

[0101] Referring to FIG. 4(b), when the user moves the remote control device (200) to the left, the pointer (205) displayed on the display unit (180) of the artificial intelligence device (100) can also move to the left in response.

[0102] Information regarding the movement of the remote control device (200) detected through the sensor of the remote control device (200) can be transmitted to the artificial intelligence device (100). The artificial intelligence device (100) can calculate the coordinates of the pointer (205) from the information regarding the movement of the remote control device (200) and display the pointer (205) to correspond to the calculated coordinates.

[0103] Referring to FIG. 4(c), when a user moves the remote control device (200) away from the display unit (180) while pressing a specific button within the remote control device (200), the selected area within the display unit (180) corresponding to the pointer (205) can be zoomed in and enlarged. Conversely, when a user moves the remote control device (200) closer to the display unit (180) while pressing a specific button within the remote control device (200), the selected area within the display unit (180) corresponding to the pointer (205) can be zoomed out and reduced.

[0104] Meanwhile, when the remote control device (200) moves away from the display unit (180), the selection area may be zoomed out, and when the remote control device (200) moves closer to the display unit (180), the selection area may be zoomed in.

[0105] Additionally, when a specific button within the remote control device (200) is pressed, recognition of up / down and left / right movement may be excluded. That is, when the remote control device (200) moves away from or closer to the display unit (180), up / down / left / right movement of the remote control device (200) is not recognized, and only forward / backward movement may be recognized. In this case, when the specific button within the remote control device (200) is not pressed, only the pointer (205) can move according to the up / down / left / right movement of the remote control device (200).

[0106] The movement speed or direction of movement of the pointer (205) can correspond to the movement speed or direction of movement of the remote control device (200).

[0107] In the present disclosure, a pointer (205) may refer to an object displayed on a display unit (180) in response to the operation of a remote control device (200). Accordingly, the pointer (205) may be an object of various shapes in addition to the arrow shape shown in FIG. 4. For example, the pointer (205) may include a point, a cursor, a prompt, a thick outline, etc. Furthermore, the pointer (205) may be displayed corresponding to either a point on the horizontal axis or a vertical axis on the display unit (180), as well as to multiple points such as a line or a surface.

[0109] FIG. 5 is a block diagram showing an artificial intelligence server (400) according to one embodiment of the present disclosure.

[0110] Referring to FIG. 5, the artificial intelligence server (400) may refer to a device that trains an artificial neural network using a machine learning algorithm or uses a trained artificial neural network. Here, the artificial intelligence server (400) may be composed of multiple servers to perform distributed processing and may be defined as a 5G network.

[0111] The artificial intelligence server (400) may also perform at least some of the artificial intelligence processing of the artificial intelligence device (100). Artificial intelligence processing may refer to operations required for learning an artificial intelligence model.

[0112] The artificial intelligence server (400) may include a communication unit (410), memory (430), a learning processor (440), and a processor (460), etc.

[0113] The communication unit (410) can transmit and receive data with an external device such as an artificial intelligence device (100).

[0114] The memory (430) may include a model storage unit (431). The model storage unit (431) may store a model (431a, or artificial neural network) that is being learned or has been learned through the learning processor (440).

[0115] The learning processor (440) can train the artificial neural network (431a) using training data. The training model may be used while mounted on the artificial intelligence server (400) of the artificial neural network, or it may be used while mounted on an external device such as an artificial intelligence device (100).

[0116] The learning model may be implemented in hardware, software, or a combination of hardware and software. If part or all of the learning model is implemented in software, one or more instructions constituting the learning model may be stored in memory (430).

[0117] The processor (460) can use a learning model to infer a result value for new input data and generate a response or control command based on the inferred result value.

[0119] The voice recognition process largely consists of a step of recognizing a voice recognition trigger word to activate the voice recognition function, and a step of recognizing speech spoken while the voice recognition function is activated. The voice recognition trigger word may be a pre-set word (specifically set by the manufacturer or developer).

[0120] Typically, since speech spoken while the speech recognition function is activated must recognize various sentences composed of multiple words rather than recognizing pre-set words, a speech engine that recognizes general spoken speech (e.g., STT engine, NLP engine, NLU engine, etc.) requires much more complex and extensive computation than a speech recognition engine that recognizes speech recognition trigger words. Accordingly, the artificial intelligence device (100) can recognize general spoken speech by directly using a speech engine when the processor (170) has sufficient computational power, and can recognize general spoken speech through an external artificial intelligence server (400) when the computational power is insufficient.

[0121] In contrast, the voice recognition engine for recognizing voice recognition commands only needs to recognize preset commands, so it is less complex and requires fewer operations compared to a voice engine that recognizes general speech voice. Accordingly, the artificial intelligence device (100) can recognize voice recognition commands using the internally mounted voice recognition engine without the help of an external artificial intelligence server (400).

[0123] FIG. 6 is a flowchart illustrating a method for recognizing a plurality of voice recognition trigger words according to one embodiment of the present disclosure.

[0124] Referring to FIG. 6, the processor (170) of the artificial intelligence device (100) receives an input audio signal through a microphone (S601).

[0125] The artificial intelligence device (100) can operate a microphone (not shown) at all times to provide a voice recognition function, and the processor (170) can receive an input audio signal at all times through the microphone (not shown) to provide a voice recognition function. In that the input audio signal is received at all times, the input audio signal may also be referred to as an input audio stream.

[0126] The input audio signal may or may not include the user's voice. Additionally, even if the input audio signal includes the user's voice, it may or may not include a voice recognition trigger word.

[0127] And, the processor (170) of the artificial intelligence device (100) generates a pre-processed audio signal from the input audio signal (S603).

[0128] Preprocessing of the input audio signal may include noise removal, voice enhancement, etc. All sounds excluding the user's voice may be considered as noise, and the noise may include not only ambient noise but also sound output from the audio output unit (185) of the artificial intelligence device (100).

[0129] The processor (170) can remove sound (or audio signal) corresponding to the output audio signal from the input audio signal by considering the output audio signal output through the audio output unit (185). Additionally, the processor (170) can remove noise included in the input audio signal by using a noise removal engine composed of a band-pass filter, etc.

[0130] In the following, the audio signal may refer to a preprocessed audio signal.

[0131] Then, the processor (170) of the artificial intelligence device (100) extracts a voice segment from the preprocessed audio signal (S605).

[0132] The processor (170) can extract a voice segment, which is a segment containing voice, from the preprocessed audio signal.

[0133] In one embodiment, the processor (170) can extract voice segments from a preprocessed audio signal through Voice Activation Detection (VAD). Voice Activation Detection may refer to the ability or function to distinguish between voice segments containing voice and non-voice segments not containing voice in a preprocessed audio signal. A non-voice segment may refer only to a segment in the preprocessed audio signal that does not contain the user's voice at all, or it may refer to a segment that includes a segment where the amplitude of the user's voice is smaller than a reference value.

[0134] And, the processor (170) of the artificial intelligence device (100) sets a voice recognition section including a voice section and a buffer section corresponding to the voice section among the preprocessed audio signals (S607).

[0135] A command word recognition section may include a speech section and a buffer section corresponding to the speech section, and may refer to a section transmitted to command word recognition engines to recognize command words within the preprocessed audio signal. That is, a command word recognition section may refer to a section in the preprocessed audio signal where there is a possibility that a command word is included.

[0136] The buffer section may include a first buffer section consisting of a previous section of a first length from the extracted voice section and a second buffer section consisting of a subsequent section of a second length from the extracted voice section. The first buffer section may be referred to as the previous buffer section, and the second buffer section may be referred to as the subsequent buffer section. For example, the first buffer section may be set to 4 seconds, etc., and the second buffer section may be set to 2 to 3 seconds, etc.

[0137] Since the voice word recognition engines determine whether a voice word is uttered in relation to surrounding sounds, there is a problem that when a voice word is recognized using only a voice segment, the accuracy is lower than when a voice word is recognized using a voice segment and its surrounding segments. To prevent this problem, the processor (170) can set a voice word recognition segment by including not only the voice segment but also a first buffer segment and a second buffer segment corresponding to the voice segment.

[0138] The processor (170) sets up a starter word recognition section by including all voice segments, and furthermore, can set up a starter word recognition section by additionally including a first buffer section and a second buffer section corresponding to each voice segment. That is, if another voice segment or the first buffer section of another voice segment overlaps with the second buffer section corresponding to a specific voice segment, both voice segments and the section between them can all be included in the starter word recognition section.

[0139] And, the processor (170) of the artificial intelligence device (100) transmits the command word recognition section among the preprocessed audio signals to each of the plurality of command word recognition engines (S609).

[0140] The command word recognition engines may be stored in the memory (140) of the artificial intelligence device (100) or in an external server. According to an embodiment, a plurality of command word recognition engines may all be stored in the memory (140), some may be stored in the memory (140), or all may be stored in an external server. When a command word recognition engine is stored in an external server, each command word recognition engine may be stored in a different external server.

[0141] A plurality of command word recognition engines are called by different command words individually configured for each speech recognition platform, and the artificial intelligence device (100) can recognize command words using a command word recognition engine stored in memory (140) or an external server. To this end, the processor (170) can transmit a command word recognition section from a preprocessed audio signal to each of the plurality of command word recognition engines.

[0142] The processor (170) transmits a command word recognition section of a preprocessed audio signal to each of a plurality of command word recognition engines, and through each command word recognition engine, can recognize a command word corresponding to each command word recognition engine in the command word recognition section. For example, a first command word recognition engine can recognize one or more first command words corresponding to the first command word recognition engine in the command word recognition section, and a second command word recognition engine can recognize one or more second command words corresponding to the second command word recognition engine in the command word recognition section.

[0143] When the processor (170) uses a command word recognition engine stored in memory (140), it recognizes a command word in the command word recognition section using the command word recognition engine directly, and when it uses a command word recognition engine stored in an external server, it transmits only the command word recognition section to the external server through the communication unit (110) and receives the command word recognition result from the external server.

[0144] Since the processor (170) transmits only the command word recognition section from the preprocessed audio signal to the command word recognition engines, it is sufficient for the command word recognition engines to operate only when the command word recognition section is transmitted, rather than operating constantly. Accordingly, the resources required by the command word recognition engines are reduced, thereby lowering the utilization of the processor (170) of the command word recognition engines.

[0145] And, the processor (170) determines whether the start word is recognized (S611).

[0146] If the processor (170) determines that the command word has been recognized, even if it determines that at least one of the multiple command word recognition engines has recognized the command word, the command word has been recognized.

[0147] Since it is undesirable for multiple voice recognition platforms to operate simultaneously, it would be desirable for only one command word recognition engine to recognize the command word, but this may not be the case. If it is determined that two or more command word recognition engines have recognized their respective command words, the processor (170) may select one command word recognition engine among the two or more command word recognition engines that recognized the command word, and determine that the command word was recognized only by the selected command word recognition engine.

[0148] If two or more command word recognition engines determine that they have recognized their respective command words, the processor (170) may select only one recognized command word based on a predetermined priority among the speech recognition platforms (or command word recognition engines) or a command word recognition score (or command word recognition accuracy) in the command word recognition engine, and activate only the speech recognition platform corresponding to the selected command word. For example, if the first command word recognition engine has a higher priority than the second command word recognition engine, and both the first command word recognition engine and the second command word recognition engine have recognized their respective command words, the processor (170) may determine that only the first command word recognition engine with the higher priority has recognized the command word, and activate only the first speech recognition platform corresponding to the first command word recognition engine. In another example, when both the first and second voice recognition engines recognize their respective voice words, and the voice recognition score of the first voice recognition engine is 0.8 and the voice recognition score of the second voice recognition engine is 0.9, the processor (170) determines that only the second voice recognition engine with the higher voice recognition score recognized the voice word, and can activate only the second voice recognition platform corresponding to the second voice recognition engine.

[0149] If, as a result of the judgment in step (S611), the starter word is not recognized, the process returns to step (S601) and the processor (170) receives an input audio signal through the microphone.

[0150] If, as a result of the judgment in step (S611), a trigger word is recognized, the processor (170) provides a voice recognition service of a voice recognition platform corresponding to the trigger word recognition engine in which the trigger word was recognized (S613).

[0151] Providing a voice recognition service of a specific voice recognition platform may mean recognizing a user's voice based on the voice recognition platform, performing control appropriate to the recognized voice, or providing a response. Providing a voice recognition service of a specific voice recognition platform may mean activating a specific voice recognition platform or a voice recognition service of a specific voice recognition platform. To this end, the processor (170) may recognize a user's voice included in a preprocessed audio signal using voice engines of the voice recognition platform, and may perform appropriate control or provide a response based thereon.

[0152] Speech engines that recognize a user's voice in a speech recognition platform may include a speech-to-text (STT) engine that converts spoken speech contained in a preprocessed audio signal into text, a natural language processing (NLP) engine that determines the intent of the utterance converted into text, and a speech synthesis engine or text-to-speech (TTS) engine that synthesizes a response generated based on the determined intent into speech. The speech synthesis engines may be stored in memory (140) or may be stored on an external server (e.g., an artificial intelligence server (400)).

[0153] The processor (170) can provide a voice recognition service corresponding to a specific voice recognition platform by using voice engines stored in memory (140) or voice engines stored in an external server.

[0154] The order of the steps illustrated in FIG. 6 is merely an example and the present disclosure is not limited thereto. That is, in one embodiment, the order of some of the steps illustrated in FIG. 6 may be reversed. Also, in one embodiment, some of the steps illustrated in FIG. 6 may be performed in parallel. Additionally, only some of the steps illustrated in FIG. 6 may be performed.

[0155] FIG. 6 illustrates a method for recognizing multiple voice recognition commands in only one cycle, but the method for recognizing multiple voice recognition commands illustrated in FIG. 6 can be performed repeatedly. That is, after performing the step of providing a voice recognition service (S613), the step of receiving an input audio signal (S601) can be performed again.

[0157] FIG. 7 is a diagram showing voice servers according to one embodiment of the present disclosure.

[0158] Referring to FIG. 7, an artificial intelligence device (100) may communicate with one or more voice servers to provide voice recognition services. The voice servers may include a command word recognition server (710) that recognizes a command word included in an audio signal using a command word recognition engine, an STT server (720) that converts a spoken voice included in an audio signal into text using an STT engine, an NLP server (730) that determines the intent of the utterance converted into text using an NLP engine, and a voice synthesis server (740) that synthesizes a response generated based on the determined intent into speech using a TTS engine. The control or response corresponding to the intent of the utterance may be performed in the NLP server (730) or in the voice synthesis server (740).

[0159] These voice servers may exist separately for each voice recognition platform, and the artificial intelligence device (100) may provide voice recognition services by communicating with voice servers corresponding to the activated voice recognition platform. For example, when the first voice recognition platform is activated, the artificial intelligence device (100) may provide voice recognition services by communicating with the first STT server, the first NLP server, and the first speech synthesis server corresponding to the first voice recognition platform.

[0160] The speech recognition server (710), STT server (720), NLP server (730), and speech synthesis server (740) may be configured as separate servers distinct from one another, but two or more may be configured as a single server. For example, the speech recognition server (710), STT server (720), NLP server (730), and speech synthesis server (740) may be configured as a single artificial intelligence server (400), in which case the speech recognition server (710), STT server (720), NLP server (730), and speech synthesis server (740) may represent individual functions of the artificial intelligence server (400).

[0162] Figure 8 is a diagram showing an example of a preprocessed audio signal and a corresponding command word recognition section.

[0163] Referring to FIG. 8, when the processor (170) acquires a preprocessed audio signal (810), it can extract a voice segment (820) through voice activity detection (VAD).

[0164] In the example illustrated in FIG. 8, the processor (170) can extract the t2-t3, t4-t5, t6-t7 and t8-t9 intervals from the preprocessed audio signal (810) as voice intervals (820) as a result of voice activity detection.

[0165] Additionally, the processor (170) may set a starting word recognition section (830) including an extracted voice section (820) and a buffer section corresponding to the voice section (820). Specifically, the processor (170) may set a first buffer section for a first length (T1) preceding each extracted voice section (820), set a second buffer section for a second length (T2) following each extracted voice section (820), and set a starting word recognition section (830) including the extracted voice section (820) and the corresponding buffer sections.

[0166] In FIG. 8, the size of the first buffer section (T1) is shown as being smaller than the size of the second buffer section (T2), but the present disclosure is not limited thereto. According to various embodiments, the size of the first buffer section and the size of the second buffer section may be the same, or the size of the first buffer section may be larger than the size of the second buffer section. For example, the size of the first buffer section may be set to 4 seconds, and the size of the second buffer section may be set to 3 seconds.

[0167] In the example illustrated in FIG. 8, the processor (170) processes the t1–t2, t3–t4, t5–t6, t7–t8 and t9–t among the preprocessed audio signals (810). 10 Set the section as a buffer section, and t including the voice section (820) and the buffer section. 1~ t 10 The interval can be set as a starting word recognition interval (830). The t3~t4 buffer interval is a second buffer interval for the t2~t3 voice interval and at the same time a first buffer interval for the t4~t5 voice interval. That is, the buffer interval between two adjacent voice intervals can be a first buffer interval and at the same time a second buffer interval.

[0168] Among the preprocessed audio signals (810), the sections that are not the trigger word recognition sections (830) can be referred to as idle sections, and in the example illustrated in FIG. 8, the t0~t1 section and t 10 ~t 11 The section is an idle section.

[0169] The processor (170) can set a first buffer section corresponding to a voice section using a circular queue. For example, the processor (170) can sequentially fill a circular queue of 5 seconds with preprocessed audio signals, and when voice activity is detected within the circular queue and a voice section is extracted, the section preceding the extracted voice section by a predetermined length (e.g., 4 seconds) can be set as the first buffer section. The size (or length) of the circular queue is larger than the size (or length) of the first buffer section.

[0170] The processor (170) can set a second buffer period corresponding to a voice segment using a timer. For example, when an extracted voice segment ends, the processor (170) can start a timer for a predetermined length (e.g., 3 seconds) to determine whether a new voice segment is extracted within the timer period, and if a new voice segment is extracted within the timer period, it can start a timer again for a predetermined length from the end point of the new voice segment to determine whether another new voice segment is extracted within the timer period. A timer period of a predetermined length from the end point of the voice segment can be set as the second buffer period.

[0171] And, the processor (170) can transmit only the command word recognition section (830) among the preprocessed audio signal (810) to each command word recognition engine.

[0172] In the example illustrated in FIG. 8, the processor (170) recognizes the starter word recognition interval (830) t1~t among the preprocessed audio signals (810). 10Only the segment can be transmitted to each starting word recognition engine. Conventionally, the processor (170) transmits the entire segment t0~t of the preprocessed audio signal (810). 11 Although the entire section had to be transmitted to each starting word recognition engine, in the present disclosure, the processor (170) transmits some starting word recognition sections t1~t 10 Since only the segments are transmitted to each command word recognition engine, the amount of CPU computation can be effectively reduced. In other words, compared to prior art, the present disclosure can prevent the waste of unnecessary resources in idle segments.

[0174] FIG. 9 is a flowchart illustrating an example of the step (S613) of providing a voice recognition service illustrated in FIG. 6.

[0175] Referring to FIG. 9, the processor (170) disables the voice activity detection (VAD) function after the starter word interval (S901).

[0176] The processor (170) extracts a voice segment from a preprocessed audio signal through voice activity detection for the recognition of a trigger word, and sets a trigger word recognition segment based on the extracted voice segment. However, after the trigger word is recognized, the intention of the spoken voice included in the preprocessed audio signal is determined using voice engines of a voice recognition platform corresponding to the recognized trigger word, so the setting of a trigger word recognition segment is unnecessary. Accordingly, the processor (170) can disable the voice activity detection function after the trigger word segment in which the trigger word is recognized.

[0177] One reason the processor (170) disables the voice activity detection function is that there is no longer a need to set the voice recognition interval because the voice activity has been detected, but also to ensure performance in the voice engines because the entire interval of the original audio signal, which includes ambient sounds, is used for the learning of the voice engines, not just the voice intervals extracted according to voice activity detection.

[0178] And, the processor (170) preprocesses after the starter word section The audio signal is transmitted to the voice engines of the voice recognition platform corresponding to the recognized command word (S903).

[0179] As described above, voice engines may be stored in memory (140) or in an external server (voice server). The processor (170) can provide a voice recognition service based on a specific voice recognition platform in which the command word is recognized by transmitting a preprocessed audio signal after the command word section to the voice engines of the voice recognition platform corresponding to the recognized command word among the voice engines of the various voice recognition platforms.

[0180] The command word section may refer to a section of recognized command words. In the preprocessed audio signal, the section following the command word section contains commands that are the subject of speech recognition, and this can be referred to as the command recognition section.

[0181] Then, the processor (170) obtains the voice recognition result (S905).

[0182] The processor (170) can transmit a preprocessed audio signal after the starter word section to voice engines (e.g., STT engine, NLP engine, speech synthesis engine, etc.) stored in memory (140) to identify the intent of the spoken voice included in the preprocessed audio signal and determine a corresponding voice recognition result (e.g., control or response). Alternatively, the processor (170) can transmit the preprocessed audio signal after the starter word section to an external server (or voice server) via the communication unit (110) and receive a voice recognition result (e.g., control or response) corresponding to the intent of the spoken voice included in the preprocessed audio signal transmitted from the external server (or voice server).

[0183] And, the processor (170) performs an action based on the voice recognition result (S907).

[0184] The processor (170) can perform control corresponding to the input speech voice based on the speech recognition result, output a response corresponding to the input speech voice, or both.

[0185] And, the processor (170) enables the voice activity detection (VAD) function (S909).

[0186] Since the voice recognition function is performed after the utterance of the utterance, the processor (170) can activate the voice activity detection function to recognize the utterance.

[0187] The order of the steps illustrated in FIG. 9 is merely an example and the present disclosure is not limited thereto. That is, in one embodiment, the order of some of the steps illustrated in FIG. 9 may be reversed. Also, in one embodiment, some of the steps illustrated in FIG. 9 may be performed in parallel. Additionally, only some of the steps illustrated in FIG. 9 may be performed.

[0189] Figure 10 is a diagram showing an example of controlling a voice activity detection function based on activation word recognition.

[0190] Referring to FIG. 10, when the processor (170) acquires a preprocessed audio signal (1010), it can activate a voice activity detection function (1020) to recognize a start word (1011).

[0191] In the example illustrated in FIG. 10, a wake-up word (1011) is included in the t2 to t3 interval of the preprocessed audio signal (1010). The processor (170) recognizes the wake-up word (1011) in the t2 to t3 interval and, accordingly, can disable the voice activity detection function (1020) to recognize a command (1012) included in the spoken voice at time t3 or at time t4, which is a certain interval after time t3. For example, time t4 may be a time 1 second after time t3, when the wake-up word ends.

[0192] And, the processor (170) can set the section of the preprocessed audio signal (1010) in which the voice activity detection function (1020) is disabled as the command recognition section (1030).

[0193] In the example illustrated in FIG. 10, the processor (170) can obtain a voice recognition result for a command (1012) by setting the t4 to t5 interval, in which the voice activity detection function (1020) is disabled, as a command recognition interval (1030) and transmitting the command recognition interval (1030) to voice engines of a voice recognition platform corresponding to the recognized command word (1011).

[0194] And, the processor (170) can activate the voice activity detection function (1020) when the recognition of the instruction (1012) is finished.

[0195] In the example illustrated in FIG. 10, the processor (170) can activate the voice activity detection function (102) after time t5, when the recognition of the instruction (1012) is terminated.

[0196] In this way, when recognizing commands, only the preprocessed audio signal is transmitted to the voice engines with the voice activity detection function disabled, thereby enabling more accurate recognition of the commands contained in the preprocessed audio signal.

[0198] FIG. 11 is a flowchart illustrating a method for recognizing a plurality of voice recognition trigger words according to one embodiment of the present disclosure.

[0199] Referring to FIG. 11, the processor (170) determines whether the current operation mode is voice registration mode (S1101).

[0200] The voice registration mode refers to a mode for registering a specific user's voice when providing a voice recognition service, and may be provided to improve the accuracy of voice recognition for individual users or to set different voice recognition settings for each individual user.

[0201] If, as a result of the judgment in step (S1101), the operation mode is the voice registration mode, the processor (170) disables the voice activity detection function (S1103), provides the voice registration function (S1105), and when the voice registration function is terminated, enables the voice activity detection function (S1107).

[0202] Given that the voice registration mode is a mode for registering a specific user's voice, it is desirable to register the user's voice by using the audio signal (or pre-processed audio signal) from which only noise or echo has been removed. Accordingly, the processor (170) can disable the voice activity detection function and provide a function to register the user's voice. For example, the processor (170) can provide a function to register the user's voice by providing a voice registration interface.

[0203] If, as a result of the determination of step (S1101), the operation mode is not a voice registration mode, the processor (170) performs the step (S601) of acquiring an input audio signal.

[0204] The processor (170) can perform the steps (S601 to S613) illustrated in FIG. 6 for recognizing a start word when the artificial intelligence device (100) is not operating in voice registration mode.

[0205] The order of the steps illustrated in FIG. 11 is merely an example and the present disclosure is not limited thereto. That is, in one embodiment, the order of some of the steps illustrated in FIG. 11 may be reversed. Also, in one embodiment, some of the steps illustrated in FIG. 11 may be performed in parallel. Additionally, only some of the steps illustrated in FIG. 11 may be performed.

[0206] FIG. 11 illustrates a method for recognizing multiple voice recognition trigger words in only one cycle, but the method for recognizing multiple voice recognition trigger words illustrated in FIG. 11 can be performed repeatedly. That is, after performing the step of activating the voice activity detection function (S1107), the step of determining whether the operation mode is a voice registration mode (S1101) can be performed again.

[0208] FIG. 12 is a drawing showing a voice registration interface according to one embodiment of the present disclosure.

[0209] Referring to FIG. 12, the voice registration interface (1210) may include a notice (1211) requesting the utterance of a trigger word until the trigger word is recognized a predetermined number of times, and information (1212) indicating the number of trigger words successfully recognized. Furthermore, the voice registration interface (1210) may further include a sound visualization image (1213) that changes color or shape according to the input sound. The user can check whether sound is being properly input to the current display device (100) through the sound visualization image (1213).

[0210] The processor (170) may provide a voice registration interface (1210) such as (a) of FIG. 12 when registering a new voice, and a voice registration interface (1210) such as (b) of FIG. 12 when re-registering an existing registered voice.

[0211] Although not illustrated in FIG. 12, if the user's voice is successfully acquired through the voice registration interface (1210) illustrated in FIG. 12, the processor (170) may provide an interface for setting the name or designation of the voice.

[0213] FIG. 13 is a diagram showing an example of controlling the voice activity detection function in voice registration mode.

[0214] Referring to FIG. 13, when the operating mode of the artificial intelligence device (100) is voice registration mode, the processor (170) can disable the voice activity detection function (1320) and provide the voice registration function while the voice registration mode is active.

[0215] In the example illustrated in FIG. 13, the voice registration mode is activated at time t1 and the voice registration mode is deactivated at time t2, and the processor (170) can disable the voice activity detection function (1320) during the t1 to t2 interval. Then, the processor (170) can register the user's voice (or voice) using the t1 to t2 interval among the preprocessed audio signals (1310).

[0217] FIG. 14 is a ladder diagram illustrating a method for recognizing a plurality of activators according to one embodiment of the present disclosure.

[0218] Referring to FIG. 14, the artificial intelligence device (100) supports a first voice recognition platform and a second voice recognition platform, and may have a first voice recognition engine that recognizes a first voice recognition word corresponding to the first voice recognition platform, and may not have a second voice recognition engine that recognizes a second voice recognition word corresponding to the second voice recognition platform. For example, the second voice recognition platform may refer to an external voice recognition platform based on the artificial intelligence device (100), and the artificial intelligence device (100) may provide voice recognition services of the second voice recognition platform using an API (Application Programming Interface). The first voice recognition word may be composed of multiple voice recognition words, and likewise, the second voice recognition word may also be composed of multiple voice recognition words.

[0219] The first artificial intelligence server (400_1) refers to an artificial intelligence server (400) that provides a voice recognition service for the first voice recognition platform and can store at least one of the first command word recognition engine or the first voice recognition engines for the first voice recognition platform. The second artificial intelligence server (400_2) refers to an artificial intelligence server (400) that provides a voice recognition service for the second voice recognition platform and can store at least one of the second command word recognition engine or the second voice recognition engines for the second voice recognition platform.

[0220] The processor (170) of the artificial intelligence device (100) receives an input audio signal (S1401), generates a preprocessed audio signal from the input audio signal (S1403), extracts a voice segment from the preprocessed audio signal through voice activity detection (VAD) (S1405), and sets a speech recognition segment based on the voice segment (S1407).

[0221] Then, the processor (170) of the artificial intelligence device (100) transmits a command word recognition section to the first command word recognition engine stored in memory (140) (S1409), and recognizes the first command word through the first command word recognition engine (S1413).

[0222] Then, the processor (170) of the artificial intelligence device (100) transmits the command word recognition section to the second artificial intelligence server (400_2) that stores the second command word recognition engine through the communication unit (110) (S1411), the processor (460) of the second artificial intelligence server (400_2) recognizes the second command word through the second command word recognition engine (S1415), and transmits the command word recognition result to the artificial intelligence device (100) through the communication unit (410) (S1417).

[0223] The processor (170) can transmit a command word recognition section to the second artificial intelligence server (400_2) using an API for the second command word recognition engine provided by the second artificial intelligence server (400_2), and obtain a command word recognition result for the second command word.

[0224] Since the artificial intelligence device (100) can recognize both a first command word corresponding to a first voice recognition platform and a second command word corresponding to a second voice recognition platform, the steps of recognizing the first command word (S1409 and S1413) and the steps of recognizing the second command word (S1411, S1415 and S1417) can be performed in parallel with each other.

[0225] And, the processor (170) of the artificial intelligence device (100) determines which start word is recognized (S1421).

[0226] If, as a result of the judgment in step (S1421), the start word is not recognized, the process proceeds to the step of receiving the input audio signal (S1401).

[0227] If, as a result of the judgment in step (S1421), the recognized command word is the first command word, the processor (170) of the artificial intelligence device (100) transmits the command recognition section to the first artificial intelligence server (400_1) that stores the first voice engines through the communication unit (110) (S1423), the processor (460) of the first artificial intelligence server (400_1) recognizes the command through the first voice engines (S1425), and transmits the voice recognition result to the artificial intelligence device (100) through the communication unit (410) (S1427).

[0228] The processor (170) can transmit a command recognition section to the first artificial intelligence server (400_1) using an API for the first voice engines provided by the first artificial intelligence server (400_1) and obtain a voice recognition result.

[0229] If, as a result of the judgment in step (S1421), the recognized command word is the second command word, the processor (170) of the artificial intelligence device (100) transmits the command recognition section to the second artificial intelligence server (400_1) that stores the second voice engines through the communication unit (110) (S1429), the processor (460) of the second artificial intelligence server (400_2) recognizes the command through the second voice engines (S1431), and transmits the voice recognition result to the artificial intelligence device (100) through the communication unit (410) (S1433).

[0230] The processor (170) can transmit a command recognition section to the second artificial intelligence server (400_2) using an API for the second voice engines provided by the second artificial intelligence server (400_2) and obtain a voice recognition result.

[0231] And, the processor (170) of the artificial intelligence device (100) performs an operation based on the acquired voice recognition result (S1435).

[0232] The order of the steps illustrated in FIG. 14 is merely an example and the present disclosure is not limited thereto. That is, in one embodiment, the order of some of the steps illustrated in FIG. 14 may be reversed. Also, in one embodiment, some of the steps illustrated in FIG. 14 may be performed in parallel. Additionally, only some of the steps illustrated in FIG. 14 may be performed.

[0233] FIG. 14 illustrates a method for recognizing multiple voice recognition commands in only one cycle, but the method for recognizing multiple voice recognition commands illustrated in FIG. 14 can be performed repeatedly. That is, after performing the step of performing an operation based on the voice recognition result (S1435), the step of receiving an input audio signal (S1401) can be performed again.

[0234] FIG. 14 illustrates an embodiment in which the first voice engines are stored in the first artificial intelligence server (400_1), but the present disclosure is not limited thereto. That is, in one embodiment, the artificial intelligence device (100) stores not only the first voice engines but also the first voice engines, and the artificial intelligence device (100) can provide the voice recognition service of the first voice recognition platform directly without going through the first artificial intelligence server (400_1).

[0236] FIG. 15 is a flowchart illustrating an example of the step (S605) of extracting a voice segment illustrated in FIG. 6.

[0237] Referring to FIG. 15, the processor (170) of the artificial intelligence device (100) obtains a probability of voice presence through voice activity detection in a preprocessed audio signal (S1501).

[0238] The processor (170) can obtain the probability or possibility that voice is present at each point in time of the preprocessed audio signal through voice activity detection.

[0239] And, the processor (170) of the artificial intelligence device (100) determines the interval in which the probability of voice existence is greater than the first reference value as the voice interval (S1503).

[0240] That is, the processor (170) can extract a voice segment from a preprocessed audio signal based solely on the probability of voice existence.

[0242] FIG. 16 is a flowchart showing an example of the step (S605) of extracting a voice segment illustrated in FIG. 6.

[0243] Referring to FIG. 16, the processor (170) of the artificial intelligence device (100) obtains a probability of voice presence through voice activity detection in a preprocessed audio signal (S1601).

[0244] Then, the processor (170) of the artificial intelligence device (100) multiplies the amplitude of the preprocessed audio signal and the corresponding probability of voice existence (S1603).

[0245] The preprocessed audio signal can be expressed as an amplitude over time, and the processor (170) can multiply the amplitude of the preprocessed audio signal at each time and the corresponding probability of voice presence.

[0246] And, the processor (170) of the artificial intelligence device (100) determines the section where the product of the amplitude of the preprocessed audio signal and the probability of voice existence is greater than the second reference value as the voice section (S1605).

[0247] That is, the processor (170) can extract a voice segment from a preprocessed audio signal by considering the probability of voice existence as well as the preprocessed audio signal.

[0248] As a result of the actual experiment, the speech segment was determined based on the product of the amplitude of the preprocessed audio signal and the probability of speech existence (example in Fig. 16), and the speech recognition performance was better than the case where the speech segment was determined based solely on the probability of speech existence (example in Fig. 15).

[0250] Figure 17 is a diagram illustrating a method for extracting speech segments from a preprocessed audio signal.

[0251] Referring to FIG. 17, when the processor (170) acquires a preprocessed audio signal (1710), it can acquire a voice presence probability (1720) corresponding to the preprocessed audio signal (1710) through voice activity detection (VAD).

[0252] And, the processor (170) can determine and extract a section as a voice section (1740) in which the value (1730) obtained by multiplying the amplitude of the preprocessed audio signal (1710) and the probability of voice existence (1730) is greater than a predetermined reference value (1731).

[0253] In the example illustrated in FIG. 17, the processor (170) has the intervals t1–t2, t3–t4, t5–t6, t7–t8, and t9–t, in which the value (1730) obtained by multiplying the amplitude of the preprocessed audio signal (1710) by the probability of speech presence (1730) is greater than a predetermined reference value (1731). 10 The section can be extracted as a voice section (1740).

[0255] According to one embodiment of the present disclosure, the above-described method can be implemented as computer-readable code on a medium on which a program is recorded. A computer-readable medium includes all types of recording devices in which data that can be read by a computer system is stored. Examples of computer-readable media include HDD (Hard Disk Drive), SSD (Solid State Disk), SSD (Silicon Disk Drive), ROM, RAM, CD-ROM, magnetic tape, floppy disk, optical data storage device, etc.

Claims

Claim 1 An artificial intelligence device for recognizing multiple starter words, comprising: a microphone; a memory for storing a first starter word recognition engine; and a communication unit for communicating with an artificial intelligence server for storing a second starter word recognition engine. An artificial intelligence device comprising a processor that receives an input audio signal through the microphone, generates a preprocessed audio signal from the input audio signal, extracts a voice segment from the preprocessed audio signal through a Voice Activation Detection (VAD) function, sets a command word recognition segment including the voice segment and a buffer segment corresponding to the voice segment among the preprocessed audio signals, transmits the command word recognition segment among the preprocessed audio signals to the first command word recognition engine and the second command word recognition engine, obtains a command word recognition result for the first command word through the first command word recognition engine and obtains a command word recognition result for the second command word through the second command word recognition engine, disables the Voice Activation Detection function when the first command word or the second command word is recognized, obtains a voice recognition result for a command recognition segment after the command word segment for the recognized command word among the preprocessed audio signals, performs an operation based on the voice recognition result, and activates the Voice Activation Detection function. Claim 2 delete Claim 3 An artificial intelligence device according to claim 1, wherein the processor sets a preceding section of a first length from the voice section as a first buffer section, sets a subsequent section of a second length from the voice section as a second buffer section, and sets the starting word recognition section including the voice section, the first buffer section, and the second buffer section. Claim 4 delete Claim 5 delete Claim 6 An artificial intelligence device according to claim 1, wherein the processor obtains the voice recognition result for the command recognition section using voice engines of a voice recognition platform corresponding to the recognized command word. Claim 7 An artificial intelligence device according to claim 1, wherein the processor obtains a probability of voice presence in the preprocessed audio signal using the voice activity detection function and extracts the voice segment using the probability of voice presence. Claim 8 An artificial intelligence device according to claim 7, wherein the processor extracts a section in which the probability of voice existence is greater than a first reference value as the voice section, or extracts a section in which the value obtained by multiplying the amplitude of the preprocessed audio signal by the probability of voice existence is greater than a second reference value as the voice section. Claim 9 A plurality of command word recognition methods performed in an artificial intelligence device comprising a microphone and a processor that executes instructions stored in memory, the method comprising: receiving an input audio signal through the microphone; generating a preprocessed audio signal from the input audio signal; extracting a voice segment from the preprocessed audio signal through a Voice Activation Detection (VAD) function; setting a command word recognition segment among the preprocessed audio signal, the segment including the voice segment and a buffer segment corresponding to the voice segment; transmitting the command word recognition segment among the preprocessed audio signal to a first command word recognition engine stored in memory and a second command word recognition engine stored in an artificial intelligence server; obtaining a command word recognition result for a first command word through the first command word recognition engine and obtaining a command word recognition result for a second command word through the second command word recognition engine. A method comprising the steps of: deactivating the voice activity detection function when the first or second trigger word is recognized; obtaining a voice recognition result for a command recognition section after the trigger word section for the recognized trigger word among the preprocessed audio signals; performing an operation based on the voice recognition result; and activating the voice activity detection function. Claim 10 A recording medium recording a plurality of command word recognition methods performed in an artificial intelligence device comprising a microphone and a processor that executes instructions stored in memory, wherein the method comprises: receiving an input audio signal through a microphone; generating a preprocessed audio signal from the input audio signal; extracting a voice segment from the preprocessed audio signal through a Voice Activation Detection (VAD) function; setting a command word recognition segment among the preprocessed audio signals, the voice segment and a buffer segment corresponding to the voice segment; transmitting the command word recognition segment among the preprocessed audio signals to a first command word recognition engine stored in memory and a second command word recognition engine stored in an artificial intelligence server; obtaining a command word recognition result for a first command word through the first command word recognition engine and obtaining a command word recognition result for a second command word through the second command word recognition engine. A recording medium comprising the steps of: deactivating the voice activity detection function when the first or second trigger word is recognized; obtaining a voice recognition result for a command recognition section after the trigger word section for the recognized trigger word among the preprocessed audio signals; performing an operation based on the voice recognition result; and activating the voice activity detection function. Claim 11 delete Claim 12 delete Claim 13 delete

Citation Information

Patent Citations

  • Apparatus and method for recognizing speech, vehicle system

    KR1020190050225A

  • Method of increasing speech recognition and device of implementing thereof

    KR1020190065094A

  • Speech recognition method and apparatus therefor

    KR1020190089128A

  • Artificial intelligence device and operating method thereof

    KR1020190094301A

  • Low-power, always-on, voice command detection and capture

    KR1020190100270A