Differentiate between voice commands

By communicating with the voice command device, receiving and analyzing data from the blocked direction, and distinguishing voice commands from human users and non-human devices, the problem of voice command devices being misresponsive or ignoring legitimate commands is solved, and more accurate voice command processing is achieved.

CN114097030BActive Publication Date: 2025-06-06INTERNATIONAL BUSINESS MACHINE CORPORATION
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202080049611.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2019-08-20
Filing Date
2020-08-13
Publication Date
2025-06-06
Estimated Expiration
2040-08-13

AI Technical Summary

Technical Problem

When a voice command device receives voice commands from other devices, it is difficult for a voice command from a human user to a non-human device to be distinguished, resulting in the problem of false response or ignoring legitimate commands.

Method used

By establishing communication with the voice command device, data indicating the blocked direction is received, a voice command is determined whether the voice command comes from the blocked direction, and the command is ignored when determining that it is a command of a non-human device.

Benefits of technology

Effectively distinguish between voice commands of human users and commands of non-human devices, avoid misresponses, and ensure that legal commands are processed.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114097030B_ABST
    Figure CN114097030B_ABST
Patent Text Reader

Abstract

Aspects of the present disclosure relate to voice command filtering. Establishing communication with a voice control device located at a location. Receive data indicating a blocked direction from the voice control device. Receive a voice command. Determine that the voice command is received from a blocked direction indicated in the data. Then ignore the received voice command.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to voice command devices and, more particularly, to voice command filtering. Background Art

[0002] A voice command device (VCD) is controlled by human voice commands. The device is controlled by human voice commands, eliminating the need to operate the device using manual controls such as buttons, dials, switches, user interfaces, etc. This allows the user to operate the device while both hands are busy with other tasks or when the user is not close enough to touch the device.

[0003] VCDs can take various forms, such as dedicated devices such as household appliances, controllers for other devices, or personal assistants. VCDs in the form of virtual personal assistants can be integrated into computing devices such as mobile phones. Virtual personal assistants may include voice-activated instructions for performing tasks or services in response to voice commands and inputs.

[0004] A VCD may be activated by a voice command in the form of one or more trigger words. A VCD may use voice recognition and be programmed to respond only to the voice of a registered individual or a group of registered individuals. This prevents non-registered users from issuing commands. Other types of VCDs are not tuned for registered users, allowing any user to issue commands in the form of specified command words and instructions.

[0005] A voice command device (VCD) is controlled by human voice commands. The device is controlled by human voice commands, eliminating the need to operate the device using manual controls such as buttons, dials, switches, user interfaces, etc. This allows the user to operate the device while both hands are busy with other tasks or when the user is not close enough to touch the device.

[0006] Complications arise when the VCD is triggered by voice commands from a television, radio, computer, or other non-human device that emits speech in the vicinity of the VCD.

[0007] For example, a VCD in the form of a smart speaker combined with a voice-controlled smart personal assistant may be provided in the living room. The smart speaker may mistakenly respond to audio from the television. Sometimes, this may be a benign command that the smart speaker does not understand; however, occasionally, the audio is a valid command or trigger word that may cause an action by the smart personal assistant.

[0008] Therefore, there is a need in the art to solve the above problems. Summary of the invention

[0009] Viewed from a first aspect, the present invention provides a computer-implemented method for voice command filtering, the method comprising: establishing communication with a voice command device at a location; receiving data indicating a blocked direction from the voice command device; receiving a voice command; determining that the voice command was received from a blocked direction indicated in the data; and ignoring the received voice command.

[0010] From another aspect, the present invention provides a system for voice command filtering, the system comprising: a memory storing program instructions; and a processor configured to execute the program instructions to perform a method comprising: establishing communication with a voice command device located at a location; receiving data indicating a blocked direction from the voice command device; receiving a voice command; determining that the voice command is received from a blocked direction indicated in the data; and ignoring the received voice command.

[0011] Viewed from another aspect, the present invention provides a computer program product for voice command filtering, the computer program product comprising a computer-readable storage medium readable by a processing circuit and storing instructions for execution by the processing circuit to perform a method for performing the steps of the present invention.

[0012] Viewed from another aspect, the invention provides a computer program stored on a computer readable medium and loadable into the internal memory of a digital computer, comprising software code portions for performing the steps of the invention when said program is run on a computer.

[0013] From another aspect, the present invention provides a computer program product, comprising a computer-readable storage medium having program instructions, wherein the computer-readable storage medium itself is not a temporary signal, and the program instructions can be executed by a processor to cause the processor to perform a method, the method comprising: establishing communication with a voice command device located at a location; receiving data indicating a blocked direction from the voice command device; receiving a voice command; determining that the voice command is received from the blocked direction indicated in the data; and ignoring the received voice command.

[0014] Embodiments of the present disclosure include methods, computer program products, and systems for voice command filtering. Communication may be established with a voice command device located at a location. Data indicating a blocked direction may be received from the voice command device. A voice command may be received. It may be determined that the voice command was received from a blocked direction indicated in the data. The received voice command may then be ignored.

[0015] The above summary is not intended to describe each illustrated embodiment or every implementation of the present disclosure. BRIEF DESCRIPTION OF THE DRAWINGS

[0016] The accompanying drawings included in the present disclosure are incorporated into the specification and form a part of the specification. They illustrate embodiments of the present disclosure and, together with the specification, are used to explain the principles of the present disclosure. The accompanying drawings are only illustrations of typical embodiments and do not limit the present disclosure.

[0017] Figure 1 is a schematic diagram illustrating an environment in which embodiments of the present disclosure may be implemented.

[0018] Figure 2A is a flow chart illustrating an example method for voice command filtering based on blocked directions according to an embodiment of the present disclosure.

[0019] Figure 2B is a flow chart illustrating an example method for determining whether to execute a voice command according to an embodiment of the present disclosure.

[0020] Figure 2C is a flow chart illustrating an example method for querying an audio output device to determine whether a voice command should be ignored according to an embodiment of the present disclosure.

[0021] Figure 3 is a flow chart illustrating an example method for transferring an audio file to a voice command device according to an embodiment of the present disclosure.

[0022] Figure 4 is a block diagram of a voice command device according to an embodiment of the present disclosure.

[0023] Figure 5 is a block diagram of an audio output device according to an embodiment of the present disclosure.

[0024] Figure 6 is a flow chart illustrating an example method for filtering voice commands at a mobile device according to an embodiment of the present disclosure.

[0025] Figure 7 is a flow chart illustrating an example method for updating a voice command device with blocked directions corresponding to a location of a mobile device according to an embodiment of the present disclosure.

[0026] Figure 8 is a high-level block diagram illustrating an example computer system that may be used to implement one or more of the methods, tools, and modules described herein, and any related functionality, according to an embodiment of the present disclosure.

[0027] Fig. 9 is a diagram illustrating a cloud computing environment according to an embodiment of the present disclosure.

[0028] Fig.10 is a block diagram illustrating an abstract model layer according to an embodiment of the present disclosure.

[0029] Although the embodiments described herein may have various modifications and alternative forms, the details thereof have been shown by way of example in the accompanying drawings and will be described in detail. However, it should be understood that the specific embodiments described should not be construed as limiting. On the contrary, the present invention will encompass all modifications, equivalents, and alternatives that fall within the scope of the present disclosure. DETAILED DESCRIPTION

[0030] Aspects of the present disclosure generally relate to live voice command devices, and specifically to voice command filtering. Although the present disclosure is not necessarily limited to such applications, various aspects of the present disclosure may be understood through discussion of various examples using this context.

[0031] Aspects of the present disclosure relate to distinguishing voice commands at a voice command device. If, over time, an audio output device (e.g., a television) is emitting background noise, the audio output device can be added to the blocked direction so that the VCD is not triggered in response to the audio output of the background noise. The voice command device can be configured to identify a registered user from the blocked direction so that the registered user's command is executed from the blocked direction. However, if an unregistered user attempts to issue a command from a blocked direction (e.g., through a television), the unregistered user may be mistakenly ignored. Therefore, aspects of the present disclosure overcome the above-mentioned complex situation by querying the audio output device to determine whether it is emitting audio. If the audio output device does not output audio when issuing a command, the command of the unregistered user can be processed. If the audio output device is outputting audio when issuing a command, an audio file can be collected from the audio output device and compared with the received voice command. If the voice command and the audio file are substantially matched, the command can be ignored (e.g., because it may be received from a television). If the voice command and the audio file do not substantially match, the voice command can be processed because it does not originate from the audio output device.

[0032] Aspects also recognize that programming blocked directions into the VCD in a static manner may cause problems when a new mobile device enters the vicinity. The mobile device itself may become a source of background noise, which may also cause voice commands to be erroneously triggered at the VCD. Therefore, aspects of the present disclosure are directed to updating the VCD when a new background source (e.g., a mobile device) enters the vicinity of the VCD.

[0033] In addition, the mobile device itself may also include voice command functionality. However, the mobile device may not be updated with the currently blocked direction, authenticated user voice, and other important considerations. Therefore, various aspects enable synchronization of VCD data (e.g., blocked direction, authenticated user voice, and other considerations such as audio output device volume status) between VCDs (e.g., dedicated VCDs and mobile devices with VCD functionality).

[0034] Reference now Figure 1 , box 100 shows an environment 110 (e.g., a room) in which a voice command device (VCD) 120 may be conventionally placed. For example, VCD 120 may be in the form of a smart speaker, including a voice-controlled smart personal assistant located on a table next to sofa 117 in environment 110.

[0035] The environment 110 may include a television 114 from which audio may emanate from two speakers 115, 116 associated with the television 114. The environment 110 may also include a radio 112 having speakers.

[0036] The VCD 120 may receive audio inputs at various times from the two television speakers 115, 116 and the radio 112. These audio inputs may include speech including command words that inadvertently trigger the VCD 120 or provide input to the VCD 120.

[0037] Aspects of the present disclosure provide VCD 120 with additional functionality to learn the directions (eg, relative angles) of audio inputs that should be ignored for a given position of VCD 120 .

[0038] Over time, the VCD 120 may learn to identify background noise sources in the environment 110 by the direction of their audio input to the VCD 120. In this example, the radio 112 is located at approximately 0 degrees relative to the VCD 120, and the dashed triangle 131 shows how the audio output of the radio 112 may be received at the VCD 120. The two speakers 115, 116 of the television 114 may be detected as being located at directions of approximately 15-20 degrees and 40-45 degrees, respectively, and the dashed triangles 132, 133 show how the audio output of the speakers 115, 116 may be received at the VCD 120. Over time, the VCD 120 may learn these directions as blocked directions, ignoring audio commands from these blocked directions.

[0039] In another example, VCD 120 may be a stationary appliance (eg, a washing machine), and may learn blocked directions for an audio input source (eg, a radio) that is in the same room as the washing machine.

[0040] If VCD 120 receives a command from these blocked directions, it may ignore the command unless it is configured to accept commands from these directions from the voice of a known registered user.

[0041] In an embodiment, the VCD 120 may be configured to distinguish commands from an unknown person 140 speaking from a blocked direction (e.g., the direction of one of the audio speakers 115, 116 of the television 114 or the direction of the radio 112). As further described below, the VCD 120 may be configured to interrogate an audio output device (e.g., the television 114 or the radio 112) to determine the state of the audio device (e.g., whether the audio device is outputting audio). If the state of the audio device indicates that the audio device is not outputting audio, the command from the unknown person 140 may be executed even if the unknown person is speaking from a blocked direction.

[0042] Embodiments enable a VCD 120 located at a given position (eg, relative to the orientation of the VCD 120) to ignore undesired command sources at that position without missing commands from unknown users.

[0043] In addition, aspects depict a mobile device (MD) 150 that may have entered environment 110. Mobile device 150 may be a smart phone with local (such as Bluetooth and Wi-Fi) and wide area (such as 4G) wireless capabilities as well as processing and storage capabilities. Mobile device 150 may also include audio output (one or more speakers) and audio input (one or more microphones) technology. Mobile device 150 may also have voice control / command software functionality provided thereon that allows mobile device 150 to be controlled by the user's voice.

[0044] The dynamic introduction of mobile device 150 into environment 110 has the potential to confuse the operation of existing static VCD 120, as mobile device 150 may be considered an additional source of (background) audio output that may cause VCD 120 to inadvertently receive commands from mobile device 150 and act upon them. Similarly, audio output from existing background audio sources (radio 112 and television 114) may confuse the operation of voice control functionality present on mobile device 150 (assuming mobile device 150 has such functionality). The introduction of new devices such as mobile device 150 may be handled by VCD 120 to minimize the potential for disruption of voice control operation of VCD 120 and operation of mobile device 150.

[0045] For further discussion below, the VCD 120 may store details of one or more blocked sources of background voice noise. The details stored by the VCD 120 may be structured in one or more different ways, such as using directional information and / or further information about the identity of the device providing the background voice noise and / or characteristics (e.g., tone and pitch) of the background voice noise that may be emitted by a particular device that may output an unexpected voice command. The VCD 120 retains this information for its own purpose, namely filtering out unwanted voice commands that may originate from a background voice noise source rather than from the user.

[0046] A mobile device 150 entering an environment 110 where a static VCD 120 is located can cause the VCD 120 to take a series of actions to achieve better operation of both the VCD 120 and the mobile device 150. The VCD 120 can be considered static (e.g., in a fixed position within the environment 110), and the mobile device 150 can be considered dynamic (e.g., moving within the environment 110). Ultimately, both devices are capable of functioning as VCDs, but one is static and one is dynamic. The devices 120 and 150 are operated so that whenever the mobile device 150 is triggered by an audio command, the mobile device 150 communicates with the static VCD 120 to utilize information about background noise sources that the VCD 120 has obtained. This then allows the mobile device 150 to prevent background noise sources from executing unwanted commands when the mobile device 150 moves to a new environment. Additionally, devices 120 and 150 are operated to allow VCD 120 to gain knowledge of mobile device 150 as a temporary background noise source so that, for example, if mobile device 150 is to play a video, stream a TV show, movie, audio track, etc., then static VCD 120 can block unwanted commands from the mobile device's speakers. This collaboration between VCD 120 and mobile device 150 improves the operation of both devices.

[0047] This improved operation of VCD 120 and mobile device 150 can be implemented as a software update to current VCD and personal assistant devices such as mobile phones. When mobile device 150 enters a new environment where VCD 120 exists, this process can be triggered. Mobile device 150 can communicate with VCD 120 located in environment 110 to establish pairing via Bluetooth, Wi-Fi or another localized communication technology. As part of this initial handshake, mobile device 150 can send appropriate trigger words and / or phrases for activating the personal assistant of the mobile device. VCD 120 then places these trigger words in memory. VCD 120 will also store the details of previously paired mobile devices 150 to eliminate the need to send trigger words each time the same mobile device 150 enters environment 110.

[0048] VCD 120 will then listen for all occurrences of the stored mobile device trigger word / phrase, and if VCD 120 detects any instance of the trigger word / phrase, VCD 120 will store the instance in temporary memory, detailing the phrase used, the time / date the phrase was received, and whether the phrase came from the direction of a known background noise source. Likewise, when the mobile device 150 detects a trigger word / phrase, before processing the command, the mobile device 150 sends a query to VCD 120 to see if the command came from one of the known background noise sources of VCD 120. If the trigger word / phrase originates from a known background noise source in the environment (according to VCD 120), the mobile device 150 may ignore the command. Otherwise, the mobile device 150 may execute the command normally.

[0049] VCD 120 can periodically determine whether mobile device 150 is still in environment 110 by emitting a signal (e.g., a high frequency audio signal emitted by mobile device 150). When mobile device 150 leaves environment 110, VCD 120 can clear its temporary storage of observed trigger words / phrases and stop storing any future commands. In an embodiment, mobile device 150 updates VCD 120 with its location using the signal, so that VCD 120 can update the direction it blocks with the location of mobile device 150.

[0050] Figure 2A is a flow chart illustrating an example method 200 for voice command filtering based on blocked directions according to an embodiment of the present disclosure.

[0051] Method 200 begins by determining one or more directions of background speech noise. This is shown in step 201. VCD 120 can learn the direction of the location by receiving background audio input from these directions and analyzing the relative directions from which the audio input is received (which may include speech and non-speech background noise). In some embodiments, determining one or more directions of background noise is accomplished by triangulation (e.g., by measuring the audio input at two or more known locations (e.g., microphones installed in VCD 120) and determining the direction and / or location of the audio input by measuring the angle with the known points). Triangulation can be accomplished by including two or more microphones in VCD 120 and cross-referencing the audio data received at the two or more microphones. In some embodiments, one or more blocked directions can be determined by time difference of arrival (TDOA). The method can similarly utilize two or more microphones. Data received at two or more microphones can be analyzed to determine the location of the received audio input based on the difference in arrival of the audio input. In some embodiments, the location of the received audio input can be determined by combining a sensor (e.g., an optical sensor, a GPS sensor, an RFID tag, etc.) with a speaker in the room (e.g., Figure 1 The one or more blocked directions can be determined by associating the VCD 120 with the speakers 112, 115, and 116 of the room and using sensors to determine the blocked directions. For example, an optical sensor can be associated with the VCD 120 and one or more speakers in the room. The blocked directions can then be determined by the optical sensor associated with the VCD 120. Alternatively, the direction of the background voice noise can be configured by the user. In an embodiment, the direction of the background voice noise can be determined based on one or more mobile devices (e.g., Figure 1 The blocked direction is determined by using the position of the mobile device 150).

[0052] The one or more blocked directions are then stored. This is shown in step 202. The one or more blocked directions may be stored in any suitable memory (e.g., flash memory, RAM, hard disk memory, etc.). In some embodiments, the one or more blocked directions are stored in a local memory on VCD 120. In some embodiments, the one or more blocked directions may be stored on another machine and transmitted over a network.

[0053] In addition, the VCD 120 determines the recognized voice biometric data. This is shown in step 203. Determining the recognized voice biometric data can be accomplished by, for example, applying voice recognition to one or more registered voices. Voice recognition can utilize characteristics of voice, such as pitch and tone. The recognized voice biometric data can be stored on the VCD 120 to determine whether the incoming voice is registered with the VCD 120. This can be achieved by comparing the incoming voice with the recognized voice biometric data to determine whether the incoming voice is the recognized voice.

[0054] Then voice input is received. This is shown in step 204. Voice input may be received by a human or non-human entity. Thus, the term "voice input" does not necessarily have to be speech, but may include background noise (e.g., such as a running washing machine, music from a speaker, etc.).

[0055] It is then determined whether the voice input is from a blocked direction. This is shown in step 205. Determining whether the voice input is from a blocked direction can be accomplished by comparing the stored blocked direction with the direction from which the voice input was received to determine whether the received voice input is associated with the stored blocked direction. If the voice input is from a direction different from the stored blocked direction, it can be determined that the voice input is not from a blocked direction.

[0056] If it is determined that the voice input is not from a blocked direction, the voice input is processed. This is shown in step 206. In some embodiments, processing includes recognizing a command and executing the received command. In some embodiments, processing may include comparing the received voice input with stored command data (e.g., data specifying a command word and a command initiation protocol) to determine whether the received voice input corresponds to (e.g., matches) a command of the stored command data. For example, if the voice input includes the phrase "power off" and "power off" is specified as a command initiation phrase in the stored command data, it can be determined that the voice input is a command and the command can be executed (e.g., the power can be turned off).

[0057] If it is determined that voice input is received from a blocked direction, the audio device associated with the blocked direction may be queried to verify the voice input. This is shown in step 207. For example, referring to Figure 1 If audio input is received from speaker 116, television 114 may be queried to determine if the television is off or muted. If it is determined that the television is off or muted (not emitting audio), the voice input may be processed (e.g., a voice command may be executed) because the audio input is received from unknown person 140.

[0058] The above operations may be completed in any order and are not limited to those described. In addition, some, all, or none of the above operations may be completed and still remain within the scope of the present disclosure.

[0059] Figure 2B-2C 2 are flow charts that collectively illustrate a method 250 for filtering audio voice data received by a voice command device (eg, VCD 120 ) according to an embodiment of the present disclosure. Figure 2B is a flow chart illustrating an example method for determining whether to execute a voice command based on direction and recognized voice biometric data according to an embodiment of the present disclosure. Figure 2C is a flow chart illustrating an example method for querying an audio output device to determine whether a voice command should be ignored according to an embodiment of the present disclosure.

[0060] Method 250 begins with registering one or more recognized voices on the VCD. This is shown in step 252. For example, the VCD can be configured to register the primary user of the VCD to ensure that the voice biometric data of the voice is easily recognized. Voice registration can be completed by analyzing various voice inputs from the primary user. The analyzed voice input can then be used to distinguish the tone and pitch of the primary user. Alternatively, regularly received voices can be automatically learned and registered by recording the pitch and tone data received during the use of the device. Having a voice recognition function may not exclude commands or inputs from other unregistered voices from being accepted by the VCD. In some embodiments, the voice may not be registered, and the method may block the direction of all voice inputs.

[0061] Then receive voice input. This is shown in step 253. In some embodiments, non-voice input is automatically filtered, which can be an inherent function of the VCD. In some embodiments, the voice input can include background noise voice or string (e.g., it includes non-human voice input).

[0062] It is then determined whether the voice input is associated with the recognized voice. This is shown in step 254. Determining whether the voice input is associated with the recognized voice can be accomplished by analyzing voice biometric data of the received voice input (e.g., by performing a tone and pitch analysis of the received voice input) and comparing the analyzed voice input data to the registered voice biometric data.

[0063] If the voice input is associated with the recognized voice, and the voice input is a command, the command is executed. This is shown in step 255. This enables the VCD to respond to voice input from a registered user regardless of the direction from which the voice input is received, including whether the voice input is received from an obstructed direction. The method can store data points of valid commands for learning purposes, including optionally storing the time and direction from which the voice command was received. This can be used, for example, to learn common directions of valid commands, such as from a favorite position relative to the VCD, which can be promoted for more sensitive command recognition.

[0064] If the voice input is not from the recognized voice, it can be determined whether the voice input originates from a blocked direction. This is shown in step 256. Determining the direction of the voice input can include measuring the angle of the incoming voice input. The VCD may have a known function for evaluating the direction of the incoming sound. For example, multiple microphones can be located on the device, and sound detection across multiple microphones can enable the determination of the position (e.g., using triangulation or arrival time difference). The blocked direction can be stored as a range of the angle of incidence of the incoming sound to the receiver. In the case of multiple microphones, the blocked direction can correspond to a voice input that is primarily or more strongly received at one or more of the multiple microphones. The direction of the incoming sound can be determined in a three-dimensional arrangement, where the input direction is determined from above or below and in a lateral direction around the VCD.

[0065] If the voice input is not from a blocked direction, it is determined whether the voice input is a command. This is shown in step 257. If it is determined that the voice input is a command, the command is executed in step 255. This enables commands to be executed from unregistered voices, i.e., from new or guest users, and does not limit users of the VCD to registered users. In some embodiments, commands may be stored as command data points including the direction in which the command was received. This can be analyzed for further voice registration, or to determine whether the command is covered by subsequent user input. The method can then end and wait for further voice input (e.g., in step 253).

[0066] If it is determined that the voice input is not from a blocked direction and is not a command, the audio input is determined to be a background noise source. The background noise data is then stored along with the time, date, and direction (e.g., angle of incidence) from which the background noise was received. This is shown in step 259. The direction can then be added as a blocked direction so that if voice input is repeatedly received from that direction, the voice input can be blocked. A threshold can be implemented to determine when voice input that does not specify a command received from an unblocked direction can be determined to be background noise.

[0067] For example, multiple voice inputs may be received from a specific direction. Each of the multiple voice inputs may be received at different times. Multiple received voice inputs may be compared with stored command data to determine whether each of the multiple received voice inputs corresponds to stored command data. The number of voice inputs that do not correspond to the stored command data may be determined. The number of voice inputs that do not correspond to the stored command data may be compared with a non-command voice input threshold (e.g., a threshold that specifies the number of non-command voice inputs that can be received for a given direction before the given direction is stored as one or more blocked directions). In response to the number of voice inputs that do not correspond to the stored command data exceeding the non-command voice input threshold, a specific direction may be stored as one or more blocked directions. In some embodiments, the background noise threshold includes sound characteristic information, such as frequency and amplitude.

[0068] If it is determined that the voice input is from a blocked direction, it is determined whether the voice input is a command. This is shown in step 258. If the voice input is from a blocked direction and is not a command, the voice input is stored as a background noise source. This is shown in step 259.

[0069] If it is determined that the voice input from the blocked direction is a command, method 250 proceeds to Figure 2C , wherein one or more audio output devices in the blocked direction are identified. This is shown in step 271. Data storage can be maintained by the VCD in the blocked direction and the audio output devices in each blocked direction together with the communication channel or address for querying each audio output device. In some embodiments, the method can identify all audio output devices near the VCD, and can use overlay communication to communicate with all audio output devices.

[0070] Query one or more identified audio output devices. This is shown in step 272. In an embodiment where there are multiple audio output devices in the blocked direction, multiple queries can be sent to two or more audio output devices simultaneously. Querying the audio output device can include sending a request signal to the device to collect audio data associated with the device. The audio output device can be queried for volume status (e.g., whether the device is silent or the current volume level) and power status (e.g., whether the device is powered on). Any appropriate connection (e.g., intranet, Internet, Bluetooth, etc.) can be used to send a status request.

[0071] Determine whether the queried audio device is emitting audio. This is shown in step 273. Determine whether the audio output device is emitting audio based on the response of the audio output device to the query request. For example, if the response to the query request indicates that the audio output device is muted, then determine in step 273 that the audio output device is not emitting audio. Similarly, if the response to the query request (or no response) indicates that the audio output device is turned off, then determine that the audio output device is not emitting audio. In some embodiments, the determination can be completed based on a volume level threshold. For example, if the volume level threshold is set to 50% volume, and the response to the query request indicates that the volume of the audio output device is 40%, then it can be determined that the audio output device does not emit audio that is loud enough to trigger the VCD.

[0072] If it is determined that the audio output device is not emitting audio (or the audio is too quiet to trigger the VCD based on the volume threshold), the command is executed. This is shown in step 274. If it is determined that the audio output device is outputting audio (or the audio is loud enough to trigger the VCD based on the volume threshold), an audio sample is requested from the audio output device. This is shown at step 275. In some embodiments, the audio sample (e.g., an audio file, an audio clip, etc.) corresponds to the time point at which the VCD is triggered. For example, a 2-10 second audio sample containing the time point at which the VCD is triggered can be requested. However, any suitable audio sample length (e.g., past hours, days, or VCD power sessions) can be requested.

[0073] Then, VCD receives the requested audio file. This is shown in step 276. In addition, the audio file detected by the microphone obtained by VCD is also obtained. This is shown in step 277. Then in step 278, the audio file is processed for comparison. In some embodiments, processing can include cleaning up the audio file (e.g., removing static and background noise from the audio file). In some embodiments, processing can include cutting the audio file (e.g., the audio file from the audio output device and the audio file obtained from the VCD microphone) into a uniform length. In some embodiments, processing can include amplifying the audio file. In some embodiments, processing can include dynamically adjusting the timing of the audio file so that the words present in the audio file are properly aligned for comparison. However, in step 278, the audio file can be processed in any other suitable manner. For example, in some embodiments, the audio file is converted into text (e.g., using conventional speech to text conversion) so that the text (e.g., transcript) between the audio files can be compared.

[0074] Then, compare the audio file obtained from the audio output device and the audio file obtained at the VCD microphone. This is shown in step 279. In some embodiments, a Fast Fourier Transform (FFT) is used to compare the audio files. In embodiments where the audio files are converted to text, the transcripts of each audio file can be compared to determine whether the strings in each transcript match. In some embodiments, the text transcript can be converted into a phonetic representation to avoid false positives, where the words sound like commands but are actually different. Using known string similarity and text comparison methods, slight differences in detection can be explained.

[0075] It is then determined whether there is a match between the audio file received from the audio output device and the audio file obtained at the VCD microphone. This is shown in step 280. In an embodiment, determining whether there is a match can be done based on one or more thresholds. For example, if the audio files are compared via FFT, there can be a match certainty threshold to determine whether the audio files substantially match. For example, if the match certainty threshold is set to 70%, if the FFT comparison indicates 60% similarity, it can be determined that the audio files do not substantially match. As another example, if the comparison indicates 75% similarity, it can be determined that the audio files do substantially match.

[0076] In embodiments where audio files are converted to text and compared, the match certainty threshold can be based on the number of matching characters / words in the audio file transcripts. For example, the match certainty threshold can specify a first number of characters within each transcript that need to match in order for the audio files to substantially match (e.g., 20 characters must match in order to determine that the voice command and the audio file of the audio output device substantially match). As another example, the match certainty threshold can specify a number of words within each transcript that need to match in order for the audio files to substantially match (e.g., 5 words must match in order to determine that the voice command and the audio file of the audio output device substantially match).

[0077] If there is a match, the command has originated from the audio output device and can be ignored. This is shown in step 281. If there is no match, the command can be processed and executed because the sound is emitted from a source other than the audio output device. This is shown in step 274.

[0078] If the command is ignored at step 281, the command may be stored as an ignored command data point with a timestamp of the time and date and the direction of the received voice input. This data may be used to analyze blocked directions. Additionally, by referencing the time and date of the stored background noise data point, the data may be used to analyze whether the device is still in the blocked direction. A threshold number of unidentified voice commands from a given direction may be stored before adding that direction to the blocked directions.

[0079] Analysis of stored data points of valid voice commands recorded along with incoming directions, background noise input, and ignored or invalid commands can be performed to learn blocked directions and, optionally, common directions of valid commands. Periodic data cleaning (e.g., formatting and filtering) can be performed as background processing of the stored data points.

[0080] Data points may be stored to allow the method and system to more accurately identify directions in which noise should be ignored.Noise from known blocking directions that is not from recognized speech may be ignored, regardless of whether it is a command or not.

[0081] Storing data points for different types of noise allows for further analysis of the background noise and therefore gives a finer discrimination as to which directions to block. For example, a VCD may receive an oven beep from the direction of the oven. Over time, the background noise data points may be analyzed to identify that the background noise data points in that direction have never included a command, or that the background noise data points from a given direction are very similar in audio content. In this case, the command from that direction may be allowed.

[0082] The VCD may include a user input mechanism to override blocked directional input or override the execution of a command. The method may also learn from such override input from the user to improve performance.

[0083] This approach assumes that the VCD stays in the same location in the same room, which is often the case. If the VCD is moved to a new location, it can relearn its environment to identify blocked directions for non-human sources relative to the VCD at the new location. The approach can store blocked directions relative to a given location of the VCD so that the VCD can be returned to a previous location and reconfigure blocked directions without relearning its environment.

[0084] In some embodiments, the method may allow for configuration of known directions to be blocked. This may eliminate the need to learn the blocked directions over time or in addition to learning the blocked directions over time, may enable users to pre-set the blocked directions based on their knowledge of the directions from which interfering audio may be received.

[0085] The user can position the VCD in a location, and the VCD can allow input to configure the blocked location. This can be via a graphical user interface, via a remote programming service, etc. In one embodiment, the configuration of the blocked direction can be performed using voice commands by the user standing at the angle to be blocked (e.g., in front of the TV) and commanding that direction to be blocked. In another embodiment, a pre-configuration of a room can be loaded and stored for use if the VCD is moved.

[0086] use Figure 1 In the example shown, VCD 120 can start picking up commands from TV 114 and can store the directions of these incoming commands. The method can then filter out commands and background noise that are always issued from the same direction. Commands from that direction that are not consistent with voices that usually give commands from other directions will be ignored.

[0087] The advantage of the method described is that unlike devices that block all commands not from known users, the VCD can include an additional level of verification in the form of direction. When a new user arrives near the VCD and issues a command, the VCD can still execute the command because it is coming from a direction that is not blocked by association with a static audio emitting object.

[0088] The technical problem solved is to enable a device to recognize that the sound is from another device rather than a human user. In addition, the present disclosure provides a technical benefit of determining whether a command from a blocked direction is the result of an audio output device or a human by querying the device for the most recent audio sample.

[0089] The above operations may be completed in any order and are not limited to those described. In addition, some, all, or none of the above operations may be completed and still remain within the scope of the present disclosure.

[0090] Figure 3 is a flow chart illustrating an example method for transferring an audio file to a voice command device according to an embodiment of the present disclosure. An audio output device that outputs sound, such as a television, radio, or mobile device, may include software components that allow them to communicate with the VCD upon request (e.g., a smart speaker or smart TV with a network interface controller (NIC)). The audio output device may require a software update to include this functionality, and the audio output device may also need to be connected to the same network as the VCD, such as the user's home WiFi network.

[0091] Method 300 begins with an audio device performing audio output monitoring. This is shown in step 301. The audio output may be buffered for a predetermined or configured duration. This is shown in step 302. For example, the audio output device may store the last 10 seconds, 1 minute, 5 minutes, etc. of the audio output.

[0092] The audio output device may receive a status request from the VCD. This is shown in step 303. The request may be associated with the reference Figure 2CThe request described in step 275 of step 304 may be the same or substantially similar. In an embodiment, the request may query the volume status and / or power status. The audio output device may confirm the request with a status reply. This is shown in step 304. The status reply may include volume status and / or power status data. If the VCD determines that the audio output device is outputting audio based on the status reply, a request for an audio output sample may be received from the VCD. This is shown in step 305. In some embodiments, if the audio output device is emitting audio, it may automatically send the most recently buffered audio output to the VCD in the form of an audio file. This is shown in step 306. In some embodiments, the audio output device may wait to receive an audio sample request from the VCD and then transfer the buffered audio file to the VCD.

[0093] The above operations may be completed in any order and are not limited to those described. In addition, some, all, or none of the above operations may be completed and still remain within the scope of the present disclosure.

[0094] Reference Figure 4 , shows a block diagram of a VCD 420 according to an embodiment of the present disclosure. The VCD 420 may be used with Figure 1 In an embodiment, the components depicted within VCD 420 may be the same or substantially similar to VCD 120 described in , and in an embodiment, the components depicted within VCD 420 may be processor-executable instructions configured to be executed by a processor.

[0095] VCD 420 may be part of a dedicated device or a multi-purpose computing device including at least one processor 401, which may include hardware modules or circuits for performing the functions of the described components, which may be software units executed on at least one processor 401. Multiple processors running parallel processing threads may be provided to enable parallel processing of some or all functions of the components. Memory 402 may be configured to provide computer instructions 403 to at least one processor 401 to perform the functions of the components.

[0096] VCD 120 may include components for known functions of a VCD, which depend on the type of device and known speech processing. In an embodiment, VCD 420 includes a speech input receiver 404, which includes multiple (e.g., two or more) microphones configured as an array to receive speech input from different directions relative to VCD 420. This audio received at multiple microphones of speech input receiver 404 can be used to determine the position (e.g., direction) of incoming noise. This can be accomplished by triangulation or time difference of arrival.

[0097] The VCD 420 may include a command processing system 406 in the form of existing software of the VCD for receiving and processing voice commands. In addition, a voice command recognition system 410 may be provided for determining the direction to be blocked and identifying commands from the blocked direction. The VCD 420 may further include a voice command differentiation system 440 configured to distinguish voice input from a known audio output device in the blocked direction from real voice input commands that may be from an unknown or unregistered user.

[0098] The VCD software including the voice command recognition processing may be provided locally to the VCD 420 or computing device, or may be provided as a remote service over a network, such as a cloud-based service. The voice command recognition system 410 and the voice command differentiation system 440 may be provided as a downloadable update at the VCD software, or may be provided as a separate additional remote service over a network, such as a cloud-based service. The remote service may also provide an application or application update for the audio output device to provide the described functionality at the audio output device.

[0099] The voice command differentiation system 440 may include a blocked direction component 421 for accessing one or more stored blocked directions of background voice noise for the location of the VCD 420 provided by the voice command recognition system 410. The voice command recognition system 410 may maintain a data store 430 of blocked directions with associated audio output devices, including stored communication channels for the audio output devices.

[0100] The voice input receiver 404 may be configured to receive voice input at the VCD 420 at the location and determine that the voice input is received from a blocked direction.

[0101] Voice command differentiation system 440 may include identification component 423 for identifying audio output devices associated with blocked directions by referencing data store 430. Identification component 423 may determine devices located in blocked directions via triangulation or time difference of arrival.

[0102] Voice command differentiation system 440 may include query component 424, which is configured to query the state of the identified audio output device to determine whether it is currently emitting audio output, and if it is currently emitting audio output, the audio file of the latest audio output is obtained from the audio output device. Query component 424 includes a state request component 425 configured to request one or more states (e.g., volume or power state) from one or more devices. Query component 424 also includes audio determination component 426 for determining whether the queried audio device is emitting audio (e.g., based on volume / power state). In some embodiments, query component 424 includes a threshold component 427 that implements one or more thresholds to determine whether the audio output device is emitting audio. For example, threshold component 427 can be configured to set a volume threshold for determining whether the device is emitting audio. The query component may include a state receiving component 431 for receiving a state from an audio output device.

[0103] The query component 424 can query the status of all audio output devices within range of the voice command device to determine which of them is currently emitting audio output. The query component 424 can obtain the audio files buffered at the audio output device.

[0104] The voice command differentiation system 440 may include an audio file acquisition component 432 for receiving an audio file from the queried audio output device. The obtained audio file may be compared with the voice input received at the voice input receiver 404.

[0105] The voice command differentiation system 440 may include a comparison component 428 for comparing the obtained audio file with the received voice input. The comparison component 428 may include a processing component 429 for processing the obtained audio file and the received voice input to convert the audio file and the voice input into text and represent the text as a voice string for comparison. However, the processing component 429 may process the audio file in any other suitable manner (e.g., modifying the amplitude, length, etc. of the audio file).

[0106] The voice command differentiation system 440 may include a voice input ignoring component 433 for ignoring received voice input if there is a substantial match to the obtained audio file.

[0107] The voice command differentiation system 440 may also include a mobile device interaction component 460. The mobile device interaction component 460 may be configured to synchronize the blocked directions stored in the data store 430 with nearby mobile devices. In addition, the mobile device interaction component 460 may also be configured to receive a signal from the mobile device so that the VCD 420 can store the updated position of the corresponding mobile device over time. In an embodiment, the mobile device interaction component 460 may transmit other VCD data stored in the data store 430 to nearby mobile devices, such as voice recognition data (e.g., registered user voiceprints), background noise data (e.g., characteristics of the background noise), and audio state data (e.g., received by the query component 424).

[0108] refer to Figure 5 , a block diagram of an example audio output device 550 according to an embodiment of the present disclosure is shown. In an embodiment, the audio output device 550 can be any suitable audio output device. For example, the audio output device can be Figure 1 14, radio 112, or mobile device 150 described above, however, audio output device 550 may be a smart watch, mobile device, speaker, voice command device, computer system (e.g., laptop computer, desktop computer, etc.), or any other suitable audio output device.

[0109] The audio output device 550 can be any form of device having an audio output and having at least one processor 551, which can be a hardware module or a circuit for performing the functions of the described components, which can be a software unit executed on at least one processor 551. Multiple processors running parallel processing threads can be provided to enable some or all functions of the components to be processed in parallel. The memory 552 can be configured to provide computer instructions 553 to the at least one processor 551 to perform the functions of the components.

[0110] Audio output device 550 includes a software component in the form of audio output providing system 560 that allows it to communicate with a VCD (e.g., VCD 120 or VCD 420) upon request from the VCD. Audio output device 550 may require a software update to include this functionality, and audio output device 550 may also need to be connected to the same network as the VCD, such as a user's home WiFi network.

[0111] The audio output providing system 560 may include a monitoring component 561 for monitoring the audio output of the audio output device 550 and a buffering component 562 for buffering the most recent audio output for a predefined duration in a buffer 563 .

[0112] The audio output providing system 560 may include a status component 564 for sending a status reply to the querying VCD to determine whether the audio output device 550 is currently emitting audio output.

[0113] The audio output providing system 560 may include an audio file component 565 for sending an audio file to the VCD when the audio output device currently sends an audio output. The audio file sent may be the latest buffered audio outputted from the buffer 563 of the audio output device 550.

[0114] Reference now Figure 6 , showing a mobile device (eg, Figure 1 Flowchart of an example method 600 for performing voice command filtering at a mobile device 150).

[0115] Method 600 begins at step 605, where a VCD (e.g., Figure 1 VCD 120 or Figure 4 This is shown in step 605. Communication with the VCD can be established in any suitable manner, including wired and wireless network connections.

[0116] VCD data is then received from the VCD. This is shown in step 610. The VCD data may include any data maintained in the VCD memory, including stored blocked directions, speech recognition data, trigger word data, audio characteristics of background noise (e.g., amplitude, pitch, etc. of the background noise), context information associated with the background noise (e.g., metadata associated with the background noise), and audio output status data (e.g., volume / power data of nearby output devices).

[0117] A command is then received at the mobile device. This is shown in step 615. The command may be received in response to the expression of the trigger word. A determination is then made as to whether the mobile device has direction analysis capabilities. This is shown in step 620. If the mobile device has direction analysis capabilities (e.g., the mobile device is configured to perform triangulation), a determination is made as to whether the command was received from a blocked direction. This is shown in step 625. This may be substantially similar to Figure 2B If the command is not received from the blocked direction, the command is executed at step 650. If the command is received from the blocked direction, a determination is made as to whether the command is in recognized speech. This is shown at step 645. Determining whether the command is in recognized speech can be substantially similar to Figure 2B The command is completed at step 254. If the command is a recognized voice, the command is executed at step 650. If the command is not a recognized voice, the command is ignored. This is shown at step 640.

[0118] If it is determined that the mobile device does not have a direction analysis capability, the sound is analyzed and compared with the details of the background noise source. This is shown in step 630. The VCD data received in step 610 may include the audio characteristics of the background noise that the VCD receives regularly. The audio characteristics of the background noise can be compared with the currently received audio data to determine whether there is a match. Determine whether background noise is possible based on the comparison completed in step 630. This is shown in step 635. For example, if the frequency, amplitude, pitch, etc. of the audio characteristics match the frequency, amplitude, pitch, etc. of the currently received audio, it can be determined that the audio is likely to be background noise.

[0119] In some embodiments, contextual information (e.g., time of day) can be considered when determining whether background noise is likely. For example, metadata associated with the background noise received at step 610 can be compared with metadata of audio currently received at the mobile device to determine whether background noise is likely. In some embodiments, contextual information can be considered together with audio data when determining whether background noise is likely.

[0120] If it is determined that background noise is possible (e.g., there is a substantial match, which may be based on a threshold), the command is ignored at step 640. If it is determined that background noise is not likely, a determination is made at step 645 whether the command was received in recognized speech. If the command is recognized speech, the command is executed at step 650. If the command is not recognized speech, the command is ignored at step 640.

[0121] The above operations may be completed in any order and are not limited to those described. In addition, some, all, or none of the above operations may be completed and still remain within the scope of the present disclosure.

[0122] Figure 7 is a flow chart illustrating an example method 700 for updating a VCD with a location of a mobile device so that the mobile device can be added to a blocked direction according to an embodiment of the present disclosure.

[0123] Method 700 begins at step 705 where communication is established with the VCD. An audio tone is then sent from the mobile device to notify the source of background noise. This is shown at operation 710. The tone may have any suitable amplitude or frequency. The VCD then performs a directional analysis so that the location of the mobile device may be stored as a blocked direction. This is shown at operation 715. This may be as described above with respect to Figure 1-6The direction analysis is performed. The mobile device then changes position. This is shown at operation 720. The audio tone is then re-issued to notify the new position as a potential background noise source. This is shown at operation 725. A determination is made as to whether a command is received from a mobile device at that position (e.g., the position at operation 720).

[0124] If it is received from this location, the command can be ignored. This is shown at operation 735. If the command is not received from this location, it is determined whether the mobile device leaves the vicinity. This is shown at operation 740. Determining whether the mobile device leaves the vicinity can be done based on the communication link cut off between the VCD and the mobile device. In some embodiments, determining whether the mobile device leaves the vicinity can be done based on the location data (e.g., global positioning system (GPS) data) of the mobile device. If it is determined that the device leaves the vicinity, the blocked direction is removed from the memory of the VCD. This is shown at operation 745. If it is determined that the device does not leave the vicinity of the VCD, method 700 returns to operation 730, where the monitoring command is continued at the VCD.

[0125] The above operations may be completed in any order and are not limited to those described. In addition, some, all, or none of the above operations may be completed and still remain within the scope of the present disclosure.

[0126] Reference now Figure 8 , shows an example computer system 801 (eg, Figure 1 VCD120, Figure 4 VCD 420, Figure 5 801 ) is a high-level block diagram of an example computer system that can be used to implement one or more of the methods, tools, and modules described herein and any related functionality (e.g., using one or more processor circuits of a computer or a computer processor). In some embodiments, the main components of the computer system 801 may include one or more CPUs 802, a memory subsystem 804, a terminal interface 812, a storage interface 814, an I / O (input / output) device interface 816, and a network interface 818, all of which may be directly or indirectly communicatively coupled for inter-component communication via a memory bus 803, an I / O bus 808, and an I / O bus interface unit 810.

[0127] Computer system 801 may include one or more general-purpose programmable central processing units (CPUs)

[0128] 802A, 802B, 802C, and 802D, collectively referred to herein as CPU 802. In some embodiments, computer system 801 may include multiple processors typical of relatively large systems; however, in other embodiments, computer system 801 may instead be a single CPU system. Each CPU 802 may execute instructions stored in memory subsystem 804 and may include one or more levels of onboard cache.

[0129] System memory 804 may include computer system readable media in the form of volatile memory, such as random access memory (RAM) 822 or cache memory 824. Computer system 801 may also include other removable / non-removable, volatile / non-volatile computer system storage media. By way of example only, storage system 826 may be provided for reading from and writing to non-removable, non-volatile magnetic media such as "hard drives". Although not shown, a disk drive for reading from and writing to a removable, non-volatile disk (e.g., "USB thumb drive" or "floppy disk") may be provided, or an optical drive for reading from or writing to a removable, non-volatile optical disk such as a CD-ROM, DVD-ROM or other optical media. In addition, memory 804 may include flash memory, such as a flash stick drive or a flash drive. The memory device may be connected to memory bus 803 via one or more data medium interfaces. Memory 804 may include at least one program product having a set (e.g., at least one) of program modules configured to perform the functions of different embodiments.

[0130] One or more programs / utilities 828, each having at least one set of program modules 830, may be stored in memory 804. Programs / utilities 828 may include a hypervisor (also referred to as a virtual machine monitor), one or more operating systems, one or more application programs, other program modules, and program data. Each of the operating system, one or more application programs, other program modules, and program data, or some combination thereof, may include an implementation of a networking environment. Programs 828 and / or program modules 830 generally perform the functions or methods of various embodiments.

[0131] In some embodiments, the program module 830 of the computer system 801 includes a voice command distinguishing module. The voice command distinguishing module can be configured to access one or more blocked directions of background voice noise from one or more audio output devices. The voice command distinguishing module can also be configured to receive voice input and determine whether the voice input is received from a blocked direction. The voice command distinguishing module can be configured to query the state of the audio device to determine whether it is emitting audio. In response to determining that the audio output device is emitting audio, the voice command distinguishing module can be configured to collect audio samples from the audio output device. Then, the speech recognition module can compare the audio sample and the voice command, and if the audio sample and the voice command are substantially matched, the voice command is ignored.

[0132] In an embodiment, the program module 830 of the computer system 801 includes a voice command filtering module. The voice command filtering module can be configured to receive VCD data from the VCD. The voice command module can be configured to determine whether a voice command is received from a blocked direction recorded in the VCD data. If a voice command is received from a blocked direction recorded in the VCD data, the voice command can be ignored.

[0133] Although the memory bus 803 is Figure 8 804 and I / O bus 808. In the embodiment of the present invention, the computer system 801 is shown as a single bus structure that provides a direct communication path between the CPU 802, the memory subsystem 804 and the I / O bus interface 810, but in some embodiments, the memory bus 803 may include multiple different buses or communication paths, which may be arranged in any of a variety of forms, such as hierarchical point-to-point links, star or mesh configurations, multi-layer buses, parallel and redundant paths, or any other suitable type of configuration. In addition, although the I / O bus interface 810 and the I / O bus 808 are shown as a single corresponding unit, in some embodiments, the computer system 801 may include multiple I / O bus interface units 810, multiple I / O buses 808, or both. In addition, although multiple I / O interface units are shown that separate the I / O bus 808 from the various communication paths to various I / O devices, in other embodiments, some or all of the I / O devices may be directly connected to one or more system I / O buses.

[0134] In some embodiments, computer system 801 may be a multi-user mainframe computer system, a single-user system, or a server computer or similar device with little or no direct user interface but receiving requests from other computer systems (clients). In addition, in some embodiments, computer system 801 may be implemented as a desktop computer, a portable computer, a laptop or notebook computer, a tablet computer, a pocket computer, a phone, a smart phone, a network switch or router, or any other suitable type of electronic device.

[0135] Notice, Figure 8 The main components of the exemplary computer system 801 are intended to be depicted. However, in some embodiments, the various components may have more Figure 8 The greater or lesser complexity represented in Figure 8 The components shown in or in addition to these components, and the number, type, and configuration of these components may vary.

[0136] It should be understood that although the present disclosure includes detailed descriptions about cloud computing, the implementation of the teachings set forth herein is not limited to cloud computing environments. Instead, the embodiments of the present disclosure can be implemented in conjunction with any other type of computing environment now known or later developed.

[0137] Cloud computing is a service delivery model for convenient, on-demand network access to a shared pool of configurable computing resources. Configurable computing resources are resources that can be quickly deployed and released with minimal management cost or interaction with the service provider, such as networks, network bandwidth, servers, processing, memory, storage, applications, virtual machines, and services. This cloud model can include at least five characteristics, at least three service models, and at least four deployment models.

[0138] Features include:

[0139] On-demand self-service: Cloud consumers can unilaterally and automatically deploy computing capabilities such as server time and network storage on demand without human interaction with the service provider.

[0140] Broad network access: Computing power can be accessed over the network through standard mechanisms that facilitate the use of the cloud through different types of thin-client or thick-client platforms (e.g., mobile phones, laptops, personal digital assistants (PDAs)).

[0141] Resource pool: The computing resources of the provider are grouped into resource pools and serve multiple consumers through a multi-tenant model, where different physical and virtual resources are dynamically allocated and reallocated on demand. Generally, consumers cannot control or even know the exact location of the provided resources, but can specify the location (such as country, state or data center) at a higher level of abstraction, so it is location-independent.

[0142] Rapid elasticity: The ability to quickly and elastically (sometimes automatically) deploy computing power to scale up quickly, and quickly release it to scale down quickly. To consumers, the computing power available for deployment often appears to be unlimited, and any amount of computing power can be accessed at any time.

[0143] Measurable services: Cloud systems automatically control and optimize resource usage by leveraging metering capabilities at a level of abstraction appropriate to the type of service (e.g., storage, processing, bandwidth, and active user accounts). Resource usage can be monitored, controlled, and reported, providing transparency to both service providers and consumers.

[0144] The service model is as follows:

[0145] Software as a Service (SaaS): The capability provided to consumers is to use the provider's applications running on a cloud infrastructure. Applications can be accessed from a variety of client devices through a thin client interface such as a web browser (e.g., web-based email). Aside from limited user-specific application configuration settings, consumers neither manage nor control the underlying cloud infrastructure including networks, servers, operating systems, storage, or even individual application capabilities.

[0146] Platform as a Service (PaaS): The capability provided to consumers is to deploy applications created or acquired by consumers on the cloud infrastructure, which are created using programming languages ​​and tools supported by the provider. Consumers neither manage nor control the underlying cloud infrastructure including networks, servers, operating systems or storage, but have control over the applications they deploy and may also have control over the configuration of the application hosting environment.

[0147] Infrastructure as a Service (IaaS): The capabilities provided to consumers are processing, storage, network, and other basic computing resources on which consumers can deploy and run arbitrary software including operating systems and applications. Consumers neither manage nor control the underlying cloud infrastructure, but have control over the operating system, storage, and applications they deploy, and may have limited control over selected network components (such as host firewalls).

[0148] The deployment model is as follows:

[0149] Private Cloud: Cloud infrastructure is run solely for an organization. The cloud infrastructure can be managed by the organization or a third party and can exist inside or outside the organization.

[0150] Community Cloud: A cloud infrastructure is shared by several organizations and supports a specific community with common interests (e.g., mission, security requirements, policies, and compliance considerations). A community cloud can be managed by multiple organizations within the community or by a third party and can exist inside or outside the community.

[0151] Public cloud: Cloud infrastructure is provided to the public or to large industry groups and is owned by an organization that sells cloud services.

[0152] Hybrid cloud: A cloud infrastructure consisting of two or more clouds (private, community or public) in a deployment model that remain distinct entities but are bound together by standardized or proprietary technologies (such as cloud burst sharing for load balancing between clouds) that enable data and application portability.

[0153] The cloud computing environment is service-oriented, with characteristics centered on statelessness, low coupling, modularity, and semantic interoperability. The core of cloud computing is the infrastructure that includes a network of interconnected nodes.

[0154] Reference now Fig. 9 , depicts an illustrative cloud computing environment 50. As shown, the cloud computing environment 50 includes one or more cloud computing nodes 10 with which a local computing device used by a cloud consumer can communicate, such as a personal digital assistant (PDA) (e.g., VCD 120 or VCD 420) or a cellular phone 54A (e.g., mobile device 150), a desktop computer 54B, a laptop computer 54C, and / or an automobile computer system 54N. The nodes 10 can communicate with each other. They can be physically or virtually grouped (not shown) in one or more networks, such as a private cloud, community cloud, public cloud, or hybrid cloud, or a combination thereof, as described above. This allows the cloud computing environment 50 to provide infrastructure, platform, and / or software as a service for which the cloud consumer does not need to maintain resources on a local computing device. It should be understood that Fig. 9 The types of computing devices 54A-N shown in are intended to be illustrative only, and computing node 10 and cloud computing environment 50 may communicate with any type of computerized device over any type of network and / or network addressable connection (eg, using a web browser).

[0155] Reference now Fig.10 , showing the cloud computing environment 50 ( Fig. 9 ) provides a set of functional abstraction layers. It should be understood in advance that Fig.10 The components, layers, and functions shown in are intended to be illustrative only, and the embodiments of the present disclosure are not limited thereto. As depicted, the following layers and corresponding functions are provided:

[0156] The hardware and software layer 60 includes hardware and software components. Examples of hardware components include: host 61; server 62 based on RISC (Reduced Instruction Set Computer) architecture; server 63; blade server 64; storage device 65; and network and network components 66. In some embodiments, software components include network application server software 67 and database software 68.

[0157] Virtualization layer 70 provides an abstraction layer from which the following examples of virtual entities can be provided: virtual servers 71; virtual storage 72; virtual networks 73, including virtual private networks; virtual applications and operating systems 74; and virtual clients 75.

[0158] In one example, the management layer 80 may provide the functionality described below. Resource provisioning 81 provides dynamic procurement of computing resources and other resources for performing tasks within a cloud computing environment. Metering and pricing 82 provides cost tracking when resources are utilized in a cloud computing environment, as well as billing or invoicing for consumption of these resources. In one example, these resources may include application software licenses. Security provides authentication for cloud consumers and tasks, as well as protection for data and other resources. User portal 83 provides access to the cloud computing environment for consumers and system administrators. Service level management 84 provides cloud computing resource allocation and management so that the required service levels are met. Service level agreement (SLA) planning and fulfillment 85 provides pre-scheduling and procurement of cloud computing resources, where future demand is anticipated based on the SLA.

[0159] The workload layer 90 provides examples of functions that can take advantage of the cloud computing environment. Examples of workloads and functions that can be provided from this layer include: mapping and navigation 91; software development and lifecycle management 92; virtual classroom education delivery 93; data analysis processing 94; transaction processing 95; and voice command processing 96.

[0160] As discussed in greater detail herein, it is contemplated that some or all of the operations of some embodiments of the methods described herein may be performed in an alternate order or may not be performed at all; furthermore, multiple operations may occur concurrently or as internal parts of a larger process.

[0161] The present disclosure may be a system, method, and / or computer program product. A computer program product may include a computer-readable storage medium (or multiple media) having computer-readable program instructions thereon for causing a processor to perform aspects of the present disclosure.

[0162] A computer-readable storage medium may be a tangible device capable of retaining and storing instructions used by an instruction execution device. A computer-readable storage medium may be, for example, but not limited to, an electronic storage device, a magnetic storage device, an optical storage device, an electromagnetic storage device, a semiconductor storage device, or any suitable combination of the foregoing. A non-exhaustive list of more specific examples of computer-readable storage media includes the following: a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), a static random access memory (SRAM), a portable compact disc read-only memory (CD-ROM), a digital versatile disk (DVD), a memory stick, a floppy disk, a mechanical encoding device such as a punch card or a raised structure in a groove on which instructions are recorded, and any suitable combination of the foregoing. As used herein, a computer-readable storage medium should not be interpreted as a temporary signal itself, such as a radio wave or other freely propagating electromagnetic wave, an electromagnetic wave propagated by a waveguide or other transmission medium (e.g., a light pulse by an optical fiber cable), or an electrical signal transmitted by a wire.

[0163] The computer-readable program instructions described herein can be downloaded from a computer-readable storage medium to a corresponding computing / processing device, or downloaded to an external computer or external storage device via a network, such as the Internet, a local area network, a wide area network, and / or a wireless network. The network can include copper transmission cables, optical transmission fibers, wireless transmissions, routers, firewalls, switches, gateway computers, and / or edge servers. The network adapter card or network interface in each computing / processing device receives the computer-readable program instructions from the network and forwards the computer-readable program instructions to be stored in a computer-readable storage medium within the corresponding computing / processing device.

[0164] The computer-readable program instructions for performing the operation of the present disclosure can be assembly instructions, instruction set architecture (ISA) instructions, machine-related instructions, microcode, firmware instructions, state setting data, or source code or object code written in any combination of one or more programming languages, including object-oriented programming languages ​​(such as Smalltalk, C++, etc.) and conventional procedural programming languages ​​(such as "C" programming language or similar programming languages). The computer-readable program instructions can be executed completely on the user's computer, partially on the user's computer, executed as an independent software package, partially on the user's computer and partially on a remote computer, or completely on a remote computer or server. In the latter case, the remote computer can be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or can be connected to an external computer (for example, using an Internet service provider through the Internet). In certain embodiments, the electronic circuit including, for example, a programmable logic circuit, a field programmable gate array (FPGA) or a programmable logic array (PLA) can execute the computer-readable program instructions to personalize the electronic circuit by utilizing the state information of the computer-readable program instructions, so as to perform aspects of the present disclosure.

[0165] Aspects of the present disclosure are described herein with reference to flowchart illustrations and / or block diagrams of methods, devices (systems) and computer program products according to embodiments of the present disclosure. It will be understood that each block of the flowchart and / or block diagram and the combination of blocks in the flowchart and / or block diagram can be implemented by computer-readable program instructions.

[0166] These computer-readable program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device create a device for implementing the functions / actions specified in one or more boxes of the flowchart and / or block diagram. These computer-readable program instructions can also be stored in a computer-readable storage medium, which can guide the computer, programmable data processing device, and / or other equipment to work in a specific manner, so that the computer-readable storage medium having the instructions stored therein includes an article of manufacture, which includes instructions for implementing various aspects of the functions / actions specified in one or more boxes of the flowchart and / or block diagram.

[0167] Computer-readable program instructions may also be loaded onto a computer, other programmable data processing apparatus, or other device, so that a series of operational steps are performed on the computer, other programmable apparatus, or other device to produce a computer-implemented process, so that the instructions executed on the computer, other programmable apparatus, or other device implement the functions / actions specified in one or more boxes of the flowchart and / or block diagram.

[0168] The flow chart and block diagram in the accompanying drawings show the possible architecture, function and operation of the system, method and computer program product according to various embodiments of the present disclosure.In this regard, each frame in the flow chart or block diagram can represent a module, segment or part of an instruction, which includes one or more executable instructions for realizing the specified logical function.In some alternative embodiments, the function noted in the frame may not occur in the order noted in the figure.For example, the two frames shown in succession can actually be performed substantially simultaneously, or these frames can sometimes be performed in reverse order, depending on the function involved.It will also be noted that the combination of each frame of the block diagram and / or flow chart illustration and the frame in the block diagram and / or flow chart illustration can be realized by a system based on dedicated hardware that performs a specified function or action or performs a combination of dedicated hardware and computer instructions.

[0169] The terms used herein are only used for the purpose of describing specific embodiments, and it is not intended to limit various embodiments. As used herein, the singular forms "one", "an" and "the" are intended to also include plural forms, unless the context clearly indicates otherwise. It will also be understood that the terms "including" and / or "comprising" when used in this specification specify the existence of stated features, integers, steps, operations, elements and / or components, but do not exclude the existence or addition of one or more other features, integers, steps, operations, elements, components and / or their groups. In the previous detailed description of the example embodiments of various embodiments, reference is made to the accompanying drawings (wherein the same reference numerals represent the same elements), which form a part of the present invention, and wherein specific example embodiments in which various embodiments can be practiced are illustrated by way of illustration. These embodiments are described in sufficient detail to enable those skilled in the art to practice these embodiments, but other embodiments may be used, and logical, mechanical, electrical and other changes may be made without departing from the scope of the various embodiments. In the previous description, many specific details are set forth to provide a thorough understanding of the various embodiments. However, various embodiments may be practiced without these specific details. In other examples, in order not to obscure the embodiments, known circuits, structures and techniques are not shown in detail.

[0170] Different instances of the word "embodiment" used in this specification do not necessarily refer to the same embodiment, but they may refer to the same embodiment. Any data and data structures shown or described herein are examples only, and in other embodiments, different data amounts, data types, number and type of fields, field names, number and type of rows, records, entries or data organizations may be used. In addition, any data may be combined with logic, so that a separate data structure may not be required. Therefore, the above detailed description should not be construed as restrictive.

[0171] The description of various embodiments of the present disclosure has been presented for illustrative purposes, but is not intended to be exhaustive or limited to the disclosed embodiments. Many modifications and variations will be apparent to those of ordinary skill in the art without departing from the scope of the described embodiments. The terms used herein are selected to best explain the principles of the embodiments, practical applications, or technical improvements existing in the market, or to enable other persons of ordinary skill in the art to understand the embodiments disclosed herein.

[0172] Although the present disclosure has been described in terms of specific embodiments, it is contemplated that changes and modifications thereof will become apparent to those skilled in the art. It is therefore intended that the appended claims be interpreted as covering all such changes and modifications that fall within the scope of the present disclosure.

Claims

1. A computer-implemented method for voice command filtering, the method comprising: include: establishing communication with a voice command device at a location; receiving data indicating a blocked direction from the voice command device; Receive voice commands; determining that the voice command was received from a blocked direction indicated in the data; querying the status of a plurality of audio output devices to determine which of the plurality of audio output devices are currently emitting audio output; obtaining an audio file from each of the audio output devices determined to be emitting audio output; comparing each of the obtained audio files to the received voice command; ignoring the received voice command in response to determining that the received voice command matches at least one of the obtained audio files to a degree exceeding a match certainty threshold; as well as In response to determining that the received voice command does not match each of the obtained plurality of audio files to a degree greater than a match certainty threshold, the voice command is executed.

2. The method according to claim 1, further comprising: include: receiving a second voice command; determining that the second voice command was received from the blocked direction indicated in the data; determining that the second voice command is a recognized voice; as well as The second voice command is executed.

3. The method of claim 1 or 2, wherein determining that the command is received from the blocked direction is determined by a time difference of arrival.

4. The method according to claim 1 or 2, further comprising: include: An audio tone is emitted to inform the voice command device of a second blocked direction, wherein the voice command device determines a location of the audio tone and stores it in the data indicating the blocked direction.

5. The method according to claim 1 or 2, further comprising: include: receiving a data set indicative of characteristics of background noise from the voice command device; receiving a second voice command; comparing the second voice command to the data set indicative of characteristics of background noise; Based on the comparing, determining that the second voice command matches a characteristic of the background noise; as well as In response to determining that the second voice command matches characteristics of the background noise, the second voice command is ignored.

6. The method according to claim 1 or 2, further comprising: include: receiving a data set indicating contextual data of background noise received at the voice command device; receiving a second voice command; comparing the context of the second voice command to context data of background noise received at the voice command device; determining, based on the comparing, that the context of the second voice command matches the context data of background noise received at the voice command device; as well as In response to determining that the context of the second voice command matches the context data of background noise received at the voice command device, ignoring the second voice command.

7. A system for voice command filtering, the system include: Memory for storing program instructions; as well as A processor configured to execute the program instructions to perform a method, the method comprising: establishing communication with a voice command device at a location; receiving data indicating a blocked direction from the voice command device; Receive voice commands; determining that the voice command was received from a blocked direction indicated in the data; querying the status of a plurality of audio output devices to determine which of the plurality of audio output devices are currently emitting audio output; obtaining an audio file from each of the audio output devices determined to be emitting audio output; comparing each of the obtained audio files to the received voice command; In response to determining that the received voice command matches at least one of the obtained audio files to a degree exceeding a match certainty threshold, ignoring the received voice command; and In response to determining that the received voice command does not match each of the obtained plurality of audio files to a degree greater than a match certainty threshold, the voice command is executed.

8. The system according to claim 7, in, The method performed by the processor further includes: receiving a second voice command; determining that the second voice command was received from the blocked direction indicated in the data; determining that the second voice command is a recognized voice; and The second voice command is executed.

9. A system according to claim 7 or 8, wherein determining that the command is received from the blocked direction is determined by a time difference of arrival.

10. The system according to claim 7 or 8, in, The method performed by the processor further includes: An audio tone is emitted to inform the voice command device of a second blocked direction, wherein the voice command device determines a location of the audio tone and stores it in the data indicating the blocked direction.

11. The system according to claim 7 or 8, in, The method performed by the processor further includes: receiving a data set indicative of characteristics of background noise from the voice command device; receiving a second voice command; comparing the second voice command to the data set indicative of characteristics of background noise; Based on the comparing, determining that the second voice command matches a characteristic of background noise; and In response to determining that the second voice command matches characteristics of the background noise, the second voice command is ignored.

12. The system according to claim 7 or 8, in, The method performed by the processor further includes: receiving a data set indicating contextual data of background noise received at the voice command device; receiving a second voice command; comparing the context of the second voice command to context data of background noise received at the voice command device; determining, based on the comparing, that the context of the second voice command matches the context data of background noise received at the voice command device; and In response to determining that the context of the second voice command matches the context data of background noise received at the voice command device, ignoring the second voice command.

13. A computer program product for voice command filtering, the computer program product comprising: A computer readable storage medium readable by a processing circuit and storing instructions for execution by the processing circuit for performing the method according to any one of claims 1 to 6.

Citation Information

Patent Citations

  • Electronic device microphone listening modes

    CN109479172A

  • Voice command filtering

    US20190250881A1