Apparatus and method for identifying wake-up word
By identifying a preset wake-up word and detecting a wake-up word in an outputable sound source, the problem that the voice recognition device is recognized as a user's intention because the wake-up word in the audio output other than the user's voice is improved, and the accuracy of the wake-up operation and user trust are improved.
Patent Information
- Application Number
- CN202411664365.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2023-12-22
- Filing Date
- 2024-11-20
- Publication Date
- 2025-06-24
AI Technical Summary
When the existing voice recognition device is based on the voice wake-up method, it is easy for the wake-up words in audio output other than the user's voice (such as broadcast, radio, song, etc.) to be recognized as user intention, resulting in a wake-up operation error, which reduces the success rate of voice recognition and user trust.
The device that starts the service by identifying a preset wake-up word, receives an audio signal using one of the server and the voice recognition device, recognizes whether the wake-up word is included in the audio signal, and detects the wake-up word in the output source. Only when the wake-up word is recognized in the audio signal and the wake-up word is not detected in the output source, a wake-up signal is generated to start the service.
It improves the accuracy of wake-up operations, prevents user inconvenience caused by unintentional wake-up, ensures the correct start of voice recognition services, and enhances users' trust in voice recognition functions.
Smart Images

Figure CN120199248A_ABST
Abstract
Description
[0001] Cross - reference to related applications
[0002] This application claims the priority of Korean Patent Application No. 10 - 2023 - 0189213, filed on December 22, 2023, the entire contents of which are incorporated herein for all purposes by this reference. Technical field
[0003] The present invention relates to an apparatus and method for recognizing a wake - up word. More specifically, the present invention relates to a wake - up word recognition apparatus and method capable of improving the recognition of a wake - up instruction. Background art
[0004] The content described in this section only provides background information for this embodiment and does not constitute related art.
[0005] Speech recognition is a series of processes that extract phonemes or language information from acoustic information included in speech and enable a machine to recognize the extracted information and respond to it.
[0006] Voice dialogue is considered the most natural and simplest method among many information exchange media between humans and machines. However, for voice communication with a machine, there is a limitation that human speech must be converted into code that the machine can process. The process of converting into code is speech recognition.
[0007] Recently, advanced speech recognition technology has been applied to automobiles, so that simple convenience devices, such as raising and lowering windows, turning on and off windshield wipers, operating an air conditioner, and turning on and off vehicle lights, can be driven only by a driver's voice command.
[0008] A speech recognition device can start a speech recognition service based on a voice wake - up method. For example, when a voice command signal including a wake - up word is input, the speech recognition device can prepare for speech recognition based on the wake - up word and provide a speech recognition service according to the voice command signal input through a microphone. Summary of the invention
[0009] In view of the above - mentioned situation, the present invention provides a voice interface with improved wake - up operation performance, such that the operation is not started by an audio output other than the user's voice (e.g., broadcast, radio, and song, etc.).
[0010] In addition, the present invention provides a wake - up word recognition method that can get rid of the limitation of selecting a wake - up word in a device that starts a service based on a voice wake - up method, that is, forcing the wake - up instruction (i.e., wake - up word (WuW)) to be an in - house term that is not commonly used in daily life.
[0011] The problems to be solved by the present invention are not limited to the above problems, and those skilled in the art can clearly understand other problems not mentioned from the following description.
[0012] According to one aspect, the present invention provides a method for recognizing a wake-up word, which is used for a device to start a service by recognizing a preset wake-up word, and is implemented by at least one of a server and a voice recognition device. The method includes: receiving an audio signal from an audio input device; recognizing whether the audio signal includes a wake-up word; detecting the wake-up word in an output sound source output by at least one audio output device; and generating a wake-up signal to start the service in response to recognizing that the audio signal includes a wake-up word and not detecting the wake-up word in the output sound source.
[0013] According to one aspect of the present invention, a voice interface with improved wake-up operation performance is provided, so that in a device that starts a service based on a voice wake-up method, the operation is not started by an audio output other than the user's voice.
[0014] According to another aspect of the present invention, various wake-up words can be selected, thus getting rid of the limitation of forcibly selecting a wake-up word as an inherent term not commonly used in daily life in a device that starts a service based on a voice wake-up method.
[0015] The effects provided by the technology of the present invention are not limited to the above effects, and those skilled in the art can clearly understand other effects not mentioned from the following description. BRIEF DESCRIPTION OF THE DRAWINGS
[0016] Figure 1 is a flowchart of a wake-up word recognition method according to an embodiment of the present invention.
[0017] Figure 2 is a flowchart of a wake-up word recognition method according to a first embodiment of the present invention.
[0018] Figure 3 is a flowchart of a wake-up word recognition method according to a second embodiment of the present invention.
[0019] Figure 4 is a flowchart of a wake-up word recognition method according to a third embodiment of the present invention.
[0020] Figure 5 is a flowchart of a wake-up word recognition method according to a fourth embodiment of the present invention.
[0021] Figure 6 is a block diagram of a voice recognition device and a voice recognition system operating according to a wake-up word recognition method according to an embodiment of the present invention. DETAILED DESCRIPTION
[0022] In the following, some embodiments of the present invention will be described in detail with reference to the accompanying drawings. In the following description, the same reference numerals preferably denote the same elements, although these elements are shown in different drawings. In addition, for the sake of clarity and brevity, when the detailed description of related known components and functions is considered to obscure the subject matter of the present invention, the detailed description of related known components and functions will be omitted in the description of some of the following embodiments.
[0023] Various ordinal numbers or alphabetical codes such as first, second, i), ii), a), b), etc. are only used as prefixes to distinguish one component from another, rather than teaching or implying the substance, order or sequence of the components. Throughout the specification, when a component "comprises" or "includes" a component, the component is intended to further include other components, rather than excluding other components, unless specifically stated to the contrary. In addition, terms such as "unit" and "module" in the specification refer to a unit that processes at least one function or operation, which can be implemented by hardware, software, or a combination of hardware and software.
[0024] The description of the present invention presented below in conjunction with the accompanying drawings is intended to describe exemplary embodiments of the present invention and is not intended to represent the only embodiments in which the technical idea of the present invention can be practiced.
[0025] Figure 1 A flowchart of a wake-up word recognition method according to an embodiment of the present invention is shown.
[0026] Referring to Figure 1 , the wake-up word recognition method (S100) according to an embodiment of the present invention includes: a step of receiving an audio signal (S110), a step of recognizing whether the audio signal includes a wake-up word (S120), a step of obtaining an outputtable sound source (S130), a step of detecting a wake-up word in the outputtable sound source (S140), and a step of generating a wake-up signal (S150).
[0027] The wake-up word recognition method according to an embodiment of the present invention is applied to a device that starts a service by recognizing a preset wake-up word (wake-up word, WuW). For example, devices that start a service through wake-up word recognition include smart speakers, mobile phones, household appliances, or voice recognition devices installed in vehicles and performing voice recognition functions.
[0028] A wake-up instruction or wake-up word is a start instruction for starting speech instruction recognition and should be caused by a user's utterance. However, according to the audio signal played by the media, this wake-up word may be included. The speech recognition device may recognize the wake-up word in the audio signal generated by the media playback (instead of the user's utterance) around the device as the wake-up word generated by the user's utterance and perform operations, resulting in an incorrect wake-up operation. This ultimately reduces the success rate of the speech recognition device in recognizing the wake-up word issued by the user and reduces the user's trust in the speech recognition function of the device.
[0029] Embodiments of the present invention solve the above problems through the following process: obtaining the sound source of the audio signal of the media that can be played around the device, and detecting whether the playable (or outputtable) sound source includes a wake-up word.
[0030] The step of receiving the audio signal (S110) includes a device (e.g., a speech recognition device) that starts a service through preset wake-up word recognition, receiving the audio signal from an audio input device. The audio input device may be a microphone that converts sound waves in the air into an electrical audio signal.
[0031] The step of identifying whether the audio signal includes a wake-up word (S120) includes detecting a preset wake-up word in the audio signal received in step S110. In other words, step S120 is a process of identifying a wake-up word in the audio signal. In step S120, the speech part is detected from the audio signal, the signal of the speech part is analyzed to detect the characteristic pattern of the speech signal, and the detected characteristic pattern is compared with the speech signal of the preset wake-up word spoken to detect the wake-up word. Alternatively, in step S120, the speech signal is converted into text data, and it is identified whether the text data includes a wake-up word to detect the wake-up word.
[0032] In step S120, the wake-up word can be specified as a basic wake-up instruction and pre-stored, or the wake-up word can be pre-stored by directly setting the desired instruction by the user. In the latter case, the wake-up word recognition method according to the embodiment of the present invention further includes the process of inputting the desired instruction of the user as the wake-up word and setting and storing the instruction. Here, inputting the wake-up word specified by the user can be performed using the above-mentioned audio input device and / or text input device.
[0033] If it is recognized in step S120 that the audio signal includes a wake-up word, then the step S140 of detecting whether the sound source (which can be output around the user) includes a wake-up word can be executed, or the step S130 of obtaining the outputtable sound source can be sequentially executed to execute step S140, as Figure 1 shown.
[0034] If it is determined in step S120 that the audio signal does not include a wake word, the process returns to step S110 to receive an audio signal. Thereafter, if no speech signal is detected from the received audio signal for a predetermined time period, the speech recognition device may enter a standby mode for speech recognition.
[0035] Step S130 of obtaining an outputtable sound source is to obtain a sound source that can be output by at least one audio output device related to the speech recognition device. The audio input device receives not only the sound from the user's speech but also the sound from the audio output device around the speech recognition device. The wake word recognition method according to the present invention can prevent a response to a wake word originating from the surrounding audio output device to only respond to a wake word spoken by the user (i.e., activate a wake-up or service).
[0036] The at least one audio output device is a speaker and can be electrically connected to the speech recognition device and a device providing a sound source. The sound sources that can be output by the audio output device include: broadcast data, streaming media data, and media data. The broadcast data is from a broadcast output device such as a radio, digital multimedia broadcasting (DMB), etc. The streaming media data is from a streaming media device connected to a user terminal through Bluetooth communication. The media data is recorded on a storage medium such as a universal serial bus (USB), a compact disc (CD), and a digital versatile disc (DVD), and from a storage medium playback device that plays the recorded data.
[0037] In step S130, when the outputtable sound source is broadcast data, it is possible to monitor the broadcast data from a broadcast channel output from the audio output device, or identify that the broadcast data from multiple broadcast channels includes a wake word, and record the broadcast channel (which is the source of the broadcast data including the wake word) and identification information including the identification time.
[0038] Step S130 may include a process of recording a sound source played by at least one audio output device. In this case, the sound source may be recorded in a buffer. When a streaming media device, a storage medium playback device, or a broadcast output device sends a sound source corresponding to media data or broadcast data to the audio output device for playback or output, it may be recorded in the buffer before the sound source is output from the audio output device. In this case, a part of the sound source may be continuously stored in the buffer for a predetermined time period. The predetermined time period may be specified in advance with respect to the speech recognition part of the speech recognition device.
[0039] Although Figure 1 step S130 is shown as being executed as a determination result in step S120, in the following Figures 2 to 5In the description, step S130 can be executed separately from step S120. For a sound source from a widely known broadcast channel, step S130 of obtaining the sound source that can be output by the available audio output device can always be executed regardless of the determination result in step S120. Additionally, under the condition of sending the sound source to the audio output device and playing the voice signal, step S130 of obtaining the sound source that can be output by the available audio output device can be executed.
[0040] The step of detecting a wake word in the outputtable sound source (S140) includes a process of identifying whether the sound source obtained in step S130 includes a wake word. The step S140 of detecting a wake word from the outputtable sound source includes a process of identifying whether the sound source obtained in step S130 includes a wake word. Step S140 includes the step of detecting a wake word from the sound source from the broadcast channel. Step S140 may include a process of comparing the time when the wake word is broadcast from the sound source from the broadcast channel with the time when the wake word is input to the audio input device, and a process of determining whether the wake word is detected based on the comparison result. Step S140 may include detecting a wake word from the sound source recorded in step S130.
[0041] If a wake word is detected from the outputtable sound source in step S140, the step returns to step S110 to receive an audio signal. After that, if no voice signal is detected from the received audio signal for more than a predetermined time, the voice recognition device may enter a standby mode for voice recognition.
[0042] If no wake word is detected in the outputtable sound source in step S140, step S150 of generating a wake signal is executed. In addition to the wake word detection process through audio input, the wake word detection process is also performed based on the playability of the surrounding sound sources. The wake word recognition method according to the present invention improves the accuracy of the voice recognition service initiated by the user's intention and prevents user inconvenience caused by unintentional wake-up.
[0043] When it is recognized in step S120 that the audio signal from the audio input device includes a wake word and no wake word is detected from the sound source output by the available audio output device in step S140, step S150 of generating a wake signal is executed. The device can be switched from the power-saving mode or the sleep mode to the operating mode through the wake signal generated in step S150. If the device is a voice recognition device, the operating mode can be a voice command recognition mode.
[0044] Figure 2 is a flowchart of a wake word recognition method according to the first embodiment of the present invention.
[0045] Reference Figure 2, the wake-up word recognition method S200 according to the first embodiment of the present invention includes: a step S210 of receiving an audio signal, a step S220 of recognizing whether the audio signal includes a wake-up word, a step (S232) of receiving information of a device that recognizes a word and starts a service, a step (S234) of monitoring a sound source of a currently output broadcast channel, a step S240 of detecting a wake-up word from the sound source, and a step S250 of generating a wake-up signal.
[0046] Hereinafter, parts of the description of method S200 that are the same as the foregoing content of method S100 will be omitted.
[0047] Method S200 includes selecting a broadcast channel using information related to the device and detecting a wake-up word in the sound source from the selected broadcast channel. Therefore, method S200 includes a step S232 of receiving device information and a step S234 of monitoring a sound source of a currently output broadcast channel.
[0048] In step S232, information about a device that recognizes a wake-up word and starts a service (for example, a voice recognition device) is received. Information about the device may include: the location of the device-equipped device, the broadcast channel played around the device or in the device-equipped device, and the time information when the wake-up word is input to the audio input device when the device recognizes that the audio signal includes a wake-up word in step S220, etc.
[0049] Step S234 includes a process of monitoring a sound source from a broadcast channel that is being output from at least one audio output device using the information about the device received in step S232. Step S234 is independently executed without depending on the determination result in step S220, and the monitoring of the sound source from the corresponding broadcast channel is always performed.
[0050] Step S240 includes a process of detecting a wake-up word in the sound source of the broadcast channel monitored in step S234. At this time, when the information about the device is received in step S232, especially if the device recognizes that the audio signal includes a wake-up word in step S220, in step S240, the time information when the wake-up word is input to the audio input device is used to determine whether a wake-up word is detected in the sound source from the corresponding broadcast channel.
[0051] In step S240, if the sound source from the corresponding broadcast channel includes a wake word and the broadcast time of the wake word is within the range where a margin is applied to the time when the wake word is input to the audio input device, it can be determined that the wake word is detected in the sound source from the corresponding broadcast channel. On the contrary, in step S240, if the sound source from the corresponding broadcast channel does not include a wake word, or even if it includes a wake word, but the broadcast time of the wake word is outside the range where a margin is applied to the input time, it can be determined that the wake word is not detected in the sound source from the corresponding broadcast channel. The range where a margin is applied to the time can be pre-specified to be related to the voice part for identifying whether the audio signal includes a wake word or the voice recognition part of the voice recognition device.
[0052] If it is recognized in step S220 that the audio signal does not include a wake word, the process returns to step S210. Or if it is recognized in step S220 that the audio signal includes a wake word and it is determined in step S240 that the wake word is detected in the sound source from the corresponding broadcast channel, it is determined that the wake word is caused by the sound source from the corresponding broadcast channel played through the audio output device, and even if the wake word is recognized, the process returns to step S210.
[0053] If it is determined in step S220 that the audio signal includes a wake word, and if it is determined in step S240 that the wake word is not detected in the sound source from the corresponding broadcast channel, step S250 is executed.
[0054] Figure 3 The flowchart showing the wake word recognition method according to the second embodiment of the present invention is presented.
[0055] Reference Figure 3 , the wake word recognition method S300 according to the second embodiment of the present invention includes: step S310 of receiving an audio signal, step S320 of identifying whether the audio signal includes a wake word, step S332 of storing the recognition information if the sound source of the broadcast channel includes a wake word, step S334 of comparing with the recognition information, step S340 of detecting the wake word in the sound source, and step S350 of generating a wake signal.
[0056] Hereinafter, the parts of the description of method S300 that are the same as the foregoing information about method S100 will be omitted.
[0057] Method S300 includes a process of detecting a wake word in the sound sources from multiple broadcast channels. Therefore, method S300 includes: step (S332) of storing the recognition information when the sound source of the broadcast channel includes a wake word and step (S334) of comparing with the recognition information.
[0058] Step S332 includes: when it is recognized that a wake-up word is included in sound sources from multiple broadcast channels, a process of storing recognition information, where the recognition information includes information about the time of the broadcast wake-up word and information about the broadcast channel of the broadcast wake-up word.
[0059] Step S334 includes: when it is recognized in step S320 that a wake-up word is included in the audio signal received from the audio input device, a process of comparing the time when the wake-up word is input into the audio input device with the recognition information stored in step S332. In step S334, the time of the broadcast wake-up word is compared with the time when the wake-up word is input into the audio input device in the recognition information.
[0060] Step S340 includes: a process of determining whether a wake-up word is detected based on the comparison result of step S334. In step S340, if the broadcast time of the wake-up word is within the range where a margin is applied to the time when the wake-up word is input into the audio input device, it can be determined that a wake-up word is detected in the sound source. On the contrary, in step S340, if the broadcast time of the wake-up word is outside the range where a margin is applied to the input time of the audio input device, it can be determined that a wake-up word is not detected in the sound source. The range to which the margin is applied to the time can be pre-specified to be related to the voice part for recognizing whether the audio signal includes a wake-up word or the voice recognition part of the voice recognition device.
[0061] In method S300, on the premise that it is recognized in step S320 that the audio signal includes a wake-up word, steps S334 and S340 are executed. If it is determined in step S340 that a wake-up word is not detected in the sound source, step S350 is executed. In step S340, if it is determined that a wake-up word is detected in the sound source, it is determined that it corresponds to the case where the wake-up word is caused by the sound source from the corresponding broadcast channel played by the audio output device, and the step returns to step S310 without generating a wake-up signal.
[0062] Figure 4 is a flowchart of a wake-up word recognition method according to the third embodiment of the present invention.
[0063] Reference Figure 4 According to
[0064] Hereinafter, parts of the description of method S400 that are the same as the foregoing content of method S100 will be omitted.
[0065] In method S400, the sound sources being played by the audio output devices around the recording device are recorded, and a wake word is detected among the recorded sound sources. Therefore, method S400 includes step S430 of recording the sound sources being played by the audio output devices.
[0066] Step S430 is a process of recording the sound sources being played by at least one audio output device. In step S430, the sound sources can be recorded in a buffer. When a streaming media device, a storage medium playback device, or a broadcast output device sends the sound sources corresponding to media data or broadcast data to the audio output device for playback or output, the sound sources can be recorded in the buffer before being output from the audio output device. In this case, a part of the sound sources can be continuously stored in the buffer for a predetermined period. The predetermined period can be specified in advance with respect to the voice part for identifying whether the audio signal includes a wake word or the voice recognition part of the voice recognition device.
[0067] Step S440 includes a process of detecting a wake word among the sound sources recorded in step S430. Step S440 can further include: when a wake word is detected, additionally recording and storing the detection time to determine whether a wake word is detected among the recorded sound sources by periodically executing. Alternatively, step S440 can be executed on the premise that it is determined in step S420 that the audio signal includes a wake word.
[0068] If it is recognized in step S420 that the audio signal does not include a wake word, the process returns to step S410. Or if it is recognized in step S420 that the audio signal includes a wake word and it is determined in step S440 that a wake word is detected among the sound sources from the corresponding broadcast channel, it is determined that the wake word is caused by the sound sources from the corresponding broadcast channel played by the audio output device, and even if a wake word is recognized in the audio signal, the process returns to step S410 without generating a wake signal.
[0069] If it is determined in step S420 that the audio signal from the audio input device includes a wake word, and if it is determined in step S440 that a wake word is not detected among the sound sources from the audio output device, then step S450 is executed.
[0070] Figure 5 is a flowchart of a wake word recognition method according to the fourth embodiment of the present invention.
[0071] Reference Figure 5, the wake word recognition method S500 according to the fourth embodiment of the present invention includes: a step S510 of receiving an audio signal, a step S520 of identifying whether the audio signal includes a wake word, a step S525 of identifying whether the utterance of the wake word in the audio signal is from a registered speaker, a step S530 of obtaining an outputtable sound source, a step S540 of detecting the wake word in the outputtable sound source, and a step S550 of generating a wake signal.
[0072] Hereinafter, parts of the description of method S500 that are the same as the foregoing content of method S100 will be omitted.
[0073] Compared with method S100, method S500 further includes a process of checking whether the wake word in the audio signal from the audio input device is spoken by a pre-registered speaker. Therefore, method S500 includes a step S525 of identifying whether the utterance of the wake word in the audio signal is the utterance of a registered speaker. In addition, method S500 may further include: inputting the user's utterance of the wake word into the audio input device, and setting and storing the user's utterance of the wake word to register the user to the speech recognition device.
[0074] Step S525 includes: comparing the wake word generated by the utterance of the wake word of the registered speaker with the audio signal of the utterance of the wake word identified in step S520 to identify whether the utterance of the wake word in the audio signal is the utterance of a registered speaker. If it is identified in step S520 that the audio signal from the audio input device includes a wake word, then step S525 is executed, and if it is identified in step S525 that the utterance of the wake word in the audio signal is spoken by a registered speaker, then step S550 of generating a wake signal is executed. If it is identified in step S525 that the utterance of the wake word in the audio signal is not the utterance of a registered speaker, then step S530 of obtaining a sound source that can be output by the audio output device and step S540 of detecting the wake word in the sound source obtained in step S530 are executed.
[0075] When step S525 is executed when it is identified in step S520 that the audio signal from the audio input device includes a wake word and when it is identified in step S525 that the utterance of the wake word in the audio signal is not the utterance of a registered speaker, if the wake word is not detected in the sound source that can be output by the audio output device in step S540, then the process of generating a wake signal in step S550 is executed.
[0076] Figure 6 The block diagrams of a speech recognition device and a speech recognition system operating by a wake word recognition method according to an embodiment of the present invention are shown.
[0077] Reference Figure 6, the speech recognition system 10 including a speech recognition device (which operates by a wake word recognition method) according to an embodiment of the present invention includes: a media playback device 100, an audio output device 200, an audio input device 300, a speech recognition device 400, and a communication module 500.
[0078] The media playback device 100 is a device for playing media data including a sound source, and may include various types of media playback devices. For example, the media playback device 100 includes: a streaming media device 120, a storage medium playback device 130, and a broadcast output device 140. The streaming media device 120 is connected to a user terminal (not shown) through Bluetooth communication and streams media data. The storage medium playback device 130 plays media data recorded on a storage medium (e.g., a universal serial bus (USB), a compact disc (CD), a digital versatile disc (DVD)). The broadcast output device 140 receives and plays broadcast data (e.g., a radio and a digital multimedia broadcast (DMB)). In addition, the media playback device 100 includes a sound source buffer 110, which temporarily records and stores the sound source when sending the sound source from the media playback device to the audio output device for playing in the air.
[0079] The audio output device 200 is a device for outputting an audio signal, and includes a speaker, an amplifier, etc. When playing media data including an audio file through the media playback device, the audio output device 200 can receive and output the audio signal from the media playback device.
[0080] The audio input device 300 is a device for receiving an audio signal (which includes a voice signal), and includes a microphone.
[0081] The speech recognition device 400 can perform speech recognition on the audio signal input through the audio input device 300 and output a speech recognition result (e.g., a voice command). The speech recognition device 400 may include a speech recognition module 410, a wake-up determination module 420, and a speech processing module 430.
[0082] When receiving an audio signal through the audio input device 300, the speech recognition module 410 can perform preprocessing (e.g., noise removal) and detect the speech part from the preprocessed audio signal. When detecting the speech part from the preprocessed audio signal, the speech recognition module 410 analyzes the signal of the speech part to detect the characteristic pattern of the speech signal, and compares the detected characteristic pattern with a preset reference speech signal to recognize the speech. Alternatively, the speech recognition module 410 converts the speech signal into text data to recognize the speech.
[0083] When no voice signal is detected from the received audio signal for more than a predetermined period, the voice recognition module 410 may enter a standby mode for voice recognition. If a wake-up instruction (i.e., a voice signal corresponding to a wake-up word) is recognized from the audio signal while operating in the standby mode, the voice recognition module 410 may output the recognition result to the wake-up determination module or the server. After that, when a wake-up signal is generated and the service is started, the voice recognition module 410 enters the voice command recognition mode and waits for a voice command input.
[0084] When a voice command is recognized from the audio signal in the voice command recognition mode, the voice recognition module 410 outputs a voice recognition result including the recognized voice command to the voice processing module 430. The voice processing module 430 that receives the voice recognition result generates output information based on the voice recognition result and outputs the generated output information to a controller (not shown).
[0085] The controller that receives the output information may perform a corresponding function in response to the voice command recognized by the voice recognition device. If the voice command recognition is successfully terminated in the voice command recognition mode, or if no voice command is recognized from the audio signal within a predetermined period after entering the voice command recognition mode, the voice recognition module 410 may enter the standby mode again and wait to receive a wake-up instruction.
[0086] The wake-up instruction or wake-up word is a start instruction for starting voice command recognition. If a voice command is recognized within a predetermined time after the wake-up word is recognized, the controller may perform a specific function in response to the recognized voice command. In other words, using the wake-up word, the voice recognition module and the controller can recognize that a voice command will be input within a predetermined time and perform the function of switching to the voice command recognition mode. The wake-up word should have a high recognition success rate in any environment, especially in a noisy situation where audio signals from media playback are mixed in addition to the voice signals from the user's speech.
[0087] Server 20 includes: an automatic speech recognition (ASR) server, a natural language processing (NLP) server, and a text-to-speech (TTS) server 1113. The automatic speech recognition server receives voice data from a voice recognition device and converts the received voice data. The natural language processing server receives text data from the ASR server, analyzes the received text data to determine a voice command, and sends a response signal based on the determined voice command to the voice recognition device. The text-to-speech server 1113 receives a signal including text corresponding to the response signal from the voice recognition device, converts the text included in the received signal into voice data, and sends the voice data to the voice recognition device. Server 20 is connected to memory 30.
[0088] The wake word recognition methods S100, S200, S300, S400, and S500 according to the embodiments of the present invention can be executed by the voice recognition device 400 and / or the server 20. That is, some of the steps included in the wake word recognition methods S100, S200, S300, S400, and S500 can be executed by the voice recognition device 400, and other steps can be executed by the server 20.
[0089] For example, steps S110, S210, S310, S410, and S510 can be executed by the voice recognition device 400, and other steps can be executed by the server 20. In this case, the voice recognition device 400 sends the received audio signal to the server 20 through the communication module 500. Alternatively, steps S232, S234, S240, and S332 can be executed by the server 20, and other steps can be executed by the voice recognition device 400. In this case, the server 20 can send the wake word detection result or the recognition information in the sound source to the voice recognition device 400 through the communication module 500.
[0090] The embodiments of the present invention can be summarized as follows.
[0091] A method for recognizing a wake word, which is used for a device that starts a service by recognizing a preset wake word and is implemented by at least one of a server and a voice recognition device. The method includes: receiving an audio signal from an audio input device; recognizing whether the audio signal includes a wake word; detecting a wake word in an outputtable sound source output by at least one audio output device; and generating a wake signal to start a service in response to recognizing that the audio signal includes a wake word and not detecting a wake word in the outputtable sound source.
[0092] In an embodiment, the method further comprises: receiving information about a device; using the information about the device to monitor a sound source from a broadcast channel being output from at least one audio output device; wherein detecting a wake word comprises detecting the wake word in the sound source from the broadcast channel.
[0093] In an embodiment, the method further comprises: identifying whether a wake word is included in sound sources from a plurality of broadcast channels; in response to identifying that a wake word is included in the sound sources from the plurality of broadcast channels, storing identification information, the identification information including information about the time of the broadcast wake word and information about the broadcast channel on which the wake word is identified; wherein detecting the wake word comprises: comparing the time of the broadcast wake word in the identification information with the time when the wake word is input to the audio input device; and determining whether the wake word is detected based on the comparison result.
[0094] In an embodiment, wherein detecting the wake word comprises: detecting the wake word from a sound source recorded by a media playback device, the media playback device recording the sound source being played by at least one audio output device.
[0095] In an embodiment, the wake word recognition method further comprises: identifying whether the wake word in the audio signal is spoken by a registered speaker.
[0096] The various illustrative embodiments of the systems and methods described herein can be implemented by digital electronic circuits, integrated circuits, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include those implemented in one or more computer programs executable on a programmable system. The programmable system includes at least one programmable processor, the at least one programmable processor being coupled to receive data and instructions from, and to transmit data and instructions to, a storage system, at least one input device, and at least one output device, wherein the programmable processor can be a dedicated processor or a general-purpose processor. A computer program (which is also referred to as a program, software, software application, or code) contains instructions for the programmable processor and is stored in a "computer-readable recording medium".
[0097] A computer-readable recording medium includes any type of recording device on which data readable by a computer system can be recorded. Examples of computer-readable recording media include non-volatile or non-transitory media such as ROM, CD-ROM, magnetic tape, floppy disk, memory card, hard disk, optical disk / disk, storage device, etc. The computer-readable recording medium may further include transitory media such as data transmission media. In addition, the computer-readable recording medium may be distributed in computer systems connected via a network, where the computer-readable code may be stored and executed in a distributed manner.
[0098] Various embodiments of the systems and techniques described herein can be implemented by a programmable computer. The computer includes a programmable processor, a data storage system (including volatile memory, non-volatile memory, or other types of storage systems or a combination thereof), and at least one communication interface. For example, the programmable computer can be one of a server, a network device, a set-top box, an embedded device, a computer expansion module, a personal computer, a laptop computer, a personal data assistant (PDA), a cloud computing system, or a mobile device.
[0099] Although exemplary embodiments of the present invention have been described for illustrative purposes, those skilled in the art should understand that various modified embodiments, additional embodiments, and alternative embodiments are possible without departing from the spirit and scope of the claimed invention. Therefore, the exemplary embodiments of the present invention have been described for simplicity and clarity. The scope of the technical idea of the embodiments of the present invention is not limited by the drawings. Accordingly, those of ordinary skill in the art should understand that the scope of the claimed invention is not limited by the embodiments described explicitly above, but is limited by the claims and their equivalents.
Claims
1. A method for identifying a wake-up word, the method is used to start a service device by identifying a preset wake-up word, and is implemented by at least one of a server and a speech recognition device, the method comprising: receiving an audio signal from an audio input device; Identify whether the audio signal includes a wake-up word; detecting a wake-up word in an outputtable sound source outputted using at least one audio output device; In response to identifying that the audio signal includes a wake-up word and the wake-up word is not detected in the outputtable sound source, a wake-up signal is generated to start the service.
2. The method according to claim 1, further comprising: receiving information about the device; monitoring a source of sound from a broadcast channel being output from at least one audio output device using information about the device; The detecting of the wake-up word includes detecting, by the server, the wake-up word in a sound source from a broadcast channel.
3. The method according to claim 1, further comprising: Identify whether the sound source from multiple broadcast channels includes a wake-up word; In response to recognizing that the wake-up word is included in the sound sources from the plurality of broadcast channels, storing recognition information including information about the time when the wake-up word is broadcast and information about the broadcast channel in which the wake-up word is recognized; Among them, detecting the wake-up word includes: comparing the time when the wake-up word is broadcast in the identification information with the time when the wake-up word is input into the audio input device, and determining whether the wake-up word is detected based on the comparison result.
4. The method according to claim 1, wherein: Detecting the wake-up word includes detecting the wake-up word from a sound source recorded by a media playback device, wherein the media playback device records the sound source being played using at least one audio output device.
5. The method according to claim 1, further comprising: Identify whether the wake word in the audio signal is spoken by a registered speaker.