Activate voice recognition
By detecting the user's hands to activate the voice recognition system above the device, the problem of high power consumption and inconvenient user activation of the voice recognition system is solved, and a low-power consumption and safe voice activation method is achieved.
Patent Information
- Application Number
- CN202080052825.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2019-07-30
- Filing Date
- 2020-07-30
- Publication Date
- 2025-07-25
- Estimated Expiration
- 2040-07-30
AI Technical Summary
The continuous operation of existing voice recognition systems in mobile devices results in high power consumption, affecting battery life, and users need to say the activation keyword or press the button to activate voice recognition, which may be distracting.
By detecting the user's hand to generate an instruction above the device, the automatic voice recognition system is activated to process the audio signal, and the hand to leave the device indicates that the voice command is over.
It reduces the power consumption of the voice recognition system, reduces the false detection of incorrect activation and voice command ending, and provides a safe and convenient way to activate voice.
Smart Images

Figure CN114144831B_ABST
Abstract
Description
[0001] Claims priority under 35 U.S.C. § 119
[0002] This patent application claims the benefit of priority of non - provisional application No. 16 / 526,608, filed on July 30, 2019, entitled "ACTIVATING SPEECH RECOGNITION", which is assigned to the assignee of the present application, and the entire content thereof is hereby incorporated by reference in its entirety. Technical Field
[0003] Broadly speaking, the present disclosure relates to speech recognition, and more particularly, to activating a voice activation system. Background Art
[0004] Voice recognition is typically used to enable an electronic device to interpret spoken questions or commands from a user. Such spoken questions or commands can be recognized by analyzing an audio signal (e.g., microphone input) at an automatic speech recognition (ASR) engine, where the ASR engine generates a text output of the spoken question or command. "Always - on" ASR systems enable an electronic device to continuously scan the audio input to detect user commands or questions in the audio input. However, the continuous operation of the ASR system results in relatively high power consumption, which reduces the battery life when implemented in a mobile device.
[0005] In some devices, a spoken voice command is not recognized unless it is preceded by a spoken activation keyword. The recognition of the activation keyword enables such devices to activate the ASR engine to process the voice command. However, saying the activation keyword before each command takes additional time and requires the speaker to use the correct pronunciation and intonation. In other devices, a dedicated button is provided for the user to press to initiate speech recognition. However, in some situations (e.g., while driving a vehicle), locating and precisely pressing the button may cause the user's attention to be diverted from other tasks. Summary of the Invention
[0006] According to one embodiment of the present disclosure, a device for processing an audio signal representing an input sound includes a hand detector configured to generate a first indication in response to detecting at least a portion of a hand above at least a portion of the device. The device further includes an automatic speech recognition system configured to be activated in response to the first indication to process the audio signal.
[0007] According to another aspect of the present disclosure, a method of processing an audio signal representing an input sound includes: at a device, detecting at least a portion of a hand above at least a portion of the device. The method further includes: in response to detecting the portion of the hand above the portion of the device, activating an automatic speech recognition system to process the audio signal.
[0008] According to another aspect of the present disclosure, a non-transitory computer-readable medium including instructions that, when executed by one or more processors of a device, cause the one or more processors to perform operations for processing an audio signal representing an input sound. The operations include: detecting at least a portion of a hand above at least a portion of the device; and in response to detecting the portion of the hand above the portion of the device, activating an automatic speech recognition system to process the audio signal.
[0009] According to another aspect of the present disclosure, an apparatus for processing an audio signal representing an input sound includes: a unit for detecting at least a portion of a hand above at least a portion of a device; a unit for processing the audio signal. The processing unit is configured to be activated in response to detecting the portion of the hand above the portion of the device. BRIEF DESCRIPTION OF THE DRAWINGS
[0010] Figure 1 is a diagram of a particular illustrative embodiment of a system including a device operable to activate speech recognition.
[0011] Figure 2 is a diagram including a particular example of components that can be implemented in a Figure 1 device.
[0012] Figure 3 is Figure 1 a schematic diagram of another particular embodiment of a
[0013] Figure 4 device.
[0014] Figure 5 is a diagram of another particular illustrative embodiment of a device operable to activate speech recognition. Figure 1 is a diagram of a particular embodiment of a method of activating speech recognition that can be performed by a
[0015] Figure 6 device. Figure 1 is a diagram of another embodiment of a method of activating speech recognition that can be performed by a
[0016] Figure 7 device.
[0017] Figure 8 A diagram of a virtual reality or augmented reality head-mounted device operable to activate speech recognition.
[0018] Figure 9 A block diagram of a specific illustrative example of a device operable to activate speech recognition. DETAILED DESCRIPTION
[0019] Devices and methods for activating a speech recognition system are disclosed. Since always-on ASR systems that continuously scan audio input to detect user commands or questions in the audio input can result in relatively high power consumption, battery life is shortened when implementing an ASR engine in a mobile device. To reduce power consumption, some systems can use a reduced-capacity speech recognition processor that consumes less power than a full-power ASR engine to perform keyword detection on the audio input. When an activation keyword is detected, the full-power ASR engine can be activated to process the speech command after the activation keyword. However, requiring the user to say the activation keyword before each command is time-consuming and requires the speaker to use the correct pronunciation and intonation. Devices that require the user to press a dedicated button to initiate speech recognition may cause an unsafe diversion of the user's attention (e.g., while driving a vehicle).
[0020] As described herein, speech recognition is activated in response to detecting a hand above a portion of the device (e.g., the user's hand hovering over the screen of the device). The user can activate speech recognition for a speech command by placing the user's hand on the device without having to say an activation keyword or precisely locate and press a dedicated button. Removing the user's hand from the device can indicate that the user has finished speaking the speech command. Thus, speech recognition can be initiated conveniently and safely (e.g., while the user is driving a vehicle). Additionally, since placing the user's hand above the device can signal the device to initiate speech recognition and removing the user's hand from above the device indicates the end of the user's speech command, both incorrect activation of speech recognition and inaccurate detection of the end of a speech command can be reduced.
[0021] Unless explicitly limited by its context, the term "generate" is used to indicate any of its ordinary meanings, such as compute, produce, and / or provide. Unless explicitly limited by its context, the term "provide" is used to indicate any of its ordinary meanings, such as compute, generate, and / or produce. Unless explicitly limited by its context, the term "coupled" is used to indicate a direct or indirect electrical or physical connection. If the connection is indirect, there may be other blocks or components between the structures that are "coupled". For example, a speaker can be acoustically coupled to a nearby wall through an intermediate medium (e.g., air) that enables waves (e.g., sound) to propagate from the speaker to the wall and vice versa.
[0022] The term "configured" can be used with reference to a method, apparatus, device, system, or any combination thereof, as indicated by its particular context. When the term "comprising" is used in this specification and the claims, it does not exclude other elements or operations. The term "based on" (e.g., "A is based on B") is used to indicate any ordinary meaning, including case (i) "at least based on" (e.g., "A is at least based on B") and, if appropriate in a particular context, (ii) "equal to" (e.g., "A is equal to B"). In case (i) where A is based on B includes at least based on, this can include a configuration where A is coupled to B. Similarly, the term "responsive to" is used to indicate any ordinary meaning, including "at least responsive to." The term "at least one" is used to indicate any ordinary meaning, including "one or more." The term "at least two" is used to indicate any ordinary meaning, including "two or more."
[0023] Unless the particular context otherwise indicates, the terms "device" and "apparatus" are generally used interchangeably. Unless otherwise stated, any disclosure of the operation of a device with a particular feature is also expressly intended to disclose a method with a similar feature (and vice versa), and any disclosure of the operation of a device according to a particular configuration is also expressly intended to disclose a method according to a similar configuration (and vice versa). Unless the particular context otherwise indicates, the terms "method," "process," "procedure," and "technique" are used interchangeably and commonly. The terms "element" and "module" can be used to indicate a part of a larger configuration. The term "packet" can correspond to a data unit that includes a header portion and a payload portion. Incorporation of a portion of a document by reference shall also be understood to include the definition of a term or variable cited in that portion (if the definition appears elsewhere in the document), as well as any drawings cited in the incorporated portion.
[0024] As used herein, the term "communication device" refers to an electronic device that can be used for voice and / or data communication via a wireless communication network. Examples of communication devices include smart speakers, soundbars, cellular phones, personal digital assistants (PDAs), handheld devices, head-mounted devices, wireless modems, laptop computers, personal computers, and the like.
[0025] Figure 1FIG. 0 depicts a system 100 including a device 102 that is configured to activate an ASR system 140 to process an input sound 106 (e.g., a voice command) when at least a portion of a hand 190 is located above the device 102. The device 102 includes one or more microphones (represented as microphone 112), a screen 110, one or more sensors 120, a hand detector 130, and an ASR system 140. Perspective view 180 shows a hand 190 located above the device 102, and block diagram 182 shows the components of the device 102. In some embodiments, by way of illustrative and non-limiting example, the device 102 may include a portable communication device (e.g., a “smartphone”), a vehicle system (e.g., a voice interface for an automotive entertainment system, a navigation system, or an autonomous driving control system), a virtual reality or augmented reality head-mounted device, or a wireless speaker and voice command device with an integrated assistant application (e.g., a “smart speaker” device).
[0026] The microphone 112 is configured to generate an audio signal 114 in response to the input sound 106. In some embodiments, the microphone 112 is configured to be activated in response to an indication 132 to generate the audio signal 114, as further described in reference Figure 3 below.
[0027] One or more sensors 120 are coupled to the hand detector 130 and are configured to provide sensor data 122 to the hand detector 130. For example, the sensors 120 may include one or more cameras, such as a low-power ambient light sensor or a main camera, an infrared sensor, an ultrasonic sensor, one or more other sensors, or any combination thereof, as further described in reference Figure 2 below.
[0028] The hand detector 130 is configured to generate an indication 132 in response to detecting that at least a portion of a hand is above at least a portion of the device 102 (e.g., above the screen 110). As used herein, “at least a portion of a hand” may correspond to any part of the hand (e.g., one or more fingers, the thumb, the palm or the back of the hand, or any part thereof, or any combination thereof), or may correspond to the entire hand, by way of illustrative and non-limiting example. As used herein, “detecting a hand” is equivalent to “detecting at least a portion of a hand” and may include detecting two or more fingers, detecting at least one finger connected to a portion of the palm, detecting the thumb and at least one finger, detecting the thumb connected to at least a portion of the palm, or detecting the entire hand (e.g., four fingers, the thumb, and the palm), by way of illustrative and non-limiting example.
[0029] Although the hand 190 is described as being detected "above" the device 102, "above" the device 102 refers to being located at a specified relative position (or within a specified range of positions) relative to the position and orientation of one or more sensors 120. In an example where the device 102 is oriented such that the sensors 120 face upward (e.g., as Figure 1 shown), detecting the hand 190 above the device 102 indicates that the hand 190 is above the device 102. In an example where the device 102 is oriented such that the sensors 120 face downward, detecting the hand 190 above the device 102 means that the hand 190 is below the device 102.
[0030] The hand detector 130 is configured to process the sensor data 122 to determine whether the hand 190 is detected above the device 102. For example, as further described with reference to Figure 1 In some embodiments, the hand detector 130 processes image data to determine whether the shape of the hand has been captured by the camera, processes infrared data to determine whether the detected temperature of the hand 190 corresponds to the hand temperature, processes ultrasonic data to determine whether the distance between the hand 190 and the device 102 is within a specified range, or a combination thereof.
[0031] In some embodiments, the device 102 is configured to generate a notification for a user of the device 102 to indicate that voice recognition has been activated in response to detecting the hand 190 above the device 102, and may also be configured to generate a second notification to indicate that voice input for voice recognition will be deactivated in response to no longer detecting the hand 190 above the device 102. For example, the device 102 may be configured to generate an audio signal such as a ringtone or a voice message such as "ready", a visual signal such as a light that emits or flashes, a digital signal played by another device (e.g., played by an in-vehicle entertainment system that communicates with the device), or any combination thereof. Generating the notification enables the user to confirm that the device 102 is ready to receive voice commands and may further enable the user to detect and prevent false activations (e.g., caused by another object that may be misidentified as the hand 190), preventing missed activations due to the hand 190 being placed in an incorrect position. Since each activation of the ASR system 140 consumes power and uses processing resources, reducing false activations results in reduced power consumption and processing resource usage.
[0032] The ASR system 140 is configured to be activated in response to the indication 132 to process the audio signal 114. In an illustrative example, a specific bit of a control register represents the presence or absence of the indication 132, and a control circuit within or coupled to the ASR system 140 is configured to read the specific bit. A "1" value of the bit corresponds to the indication 132 and causes the ASR system 140 to be activated. In other embodiments, the indication 132 is implemented as a digital or analog signal on a bus or control line, an interrupt flag at an interrupt controller, or an optical or mechanical signal, as illustrative non-limiting examples.
[0033] When activated, the ASR system 140 is configured to process one or more portions (e.g., frames) of the audio signal 114 that includes the input sound 106. For example, the device 102 may buffer a series of frames of the audio signal 114 as sensor data 122 to be processed by the hand detector 130 such that when the indication 132 is generated, the ASR system 140 can process the buffered series of frames and generate an output representing the user's speech. The ASR system 140 may provide the recognized speech 142 as a text output of the speech content of the input sound 106 to another component of the device 102 (e.g., a "virtual assistant" application or other applications described in the reference Figure 3 to initiate an action based on the speech content.
[0034] When deactivated, the ASR system 140 does not process the audio signal 114 and consumes less power than when activated. For example, deactivation of the ASR system 140 may include: gating the input circuit of the ASR system 140 to prevent the audio signal 114 from being input into the ASR system 140, gating the clock signal to prevent the circuits within the ASR system 140 from switching, or both, thereby reducing dynamic power consumption. As another example, deactivation of the ASR system 140 may include: reducing the power supply to the ASR system 140 to reduce static power consumption without losing the state of circuit elements, removing power from at least a portion of the ASR system 140, or a combination thereof.
[0035] In some embodiments, the hand detector 130, the ASR system 140, or any combination thereof is implemented using dedicated circuitry or hardware. In some embodiments, the hand detector 130, the ASR system 140, or any combination thereof is implemented by the execution of firmware or software. For illustrative purposes, the device 102 may include: a memory configured to store instructions, and one or more processors configured to execute the instructions to implement the hand detector 130 and the ASR system 140, such as further described in the reference Figure 9 as further described.
[0036] During operation, the user can place the user's hand 190 above the device 102 before speaking a voice command. The hand detector 130 processes the sensor data 122 to determine that the hand 190 is above the device 102. In response to detecting that the hand 190 is above the device 102, the hand detector 130 generates an indication 132, which causes the activation of the ASR system 140. After the microphone 112 receives the voice command, the ASR system 140 processes the corresponding portion of the audio signal 114 to generate the recognized speech 142 indicating the voice command.
[0037] Activating the ASR system 140 when a hand is detected above the device 102 enables a user of the device 102 to activate speech recognition of a voice command by placing the user's hand 190 above the device without the user having to speak an activation keyword or having to precisely locate and press a dedicated button. Accordingly, speech recognition can be conveniently and safely initiated, such as when the user is driving a vehicle. Further, since placing the user's hand on the device signals the device to initiate speech recognition, incorrect activations of speech recognition can be reduced compared to systems that use keyword detection to activate speech recognition.
[0038] Figure 2 An example 200 depicting a further aspect of components that can be implemented in the Figure 1 device 102 is shown. As Figure 2 shown, the sensor 120 includes: one or more cameras 202 configured to provide image data 212 to the hand detector 130, an infrared (IR) sensor 208 configured to provide infrared sensor data 218 to the hand detector 130, and an ultrasonic sensor 210 configured to provide ultrasonic sensor data 220 to the hand detector. The image data 212, the infrared sensor data 218, and the ultrasonic sensor data 220 are included in the sensor data 122. The camera 202 includes a low-power ambient light sensor 204 configured to generate at least a portion of the image data 212, a main camera 206 configured to generate at least a portion of the image data 212, or both. Although the main camera 206 can capture image data with a higher resolution than the ambient light sensor 204, the ambient light sensor 204 can generate image data with sufficient resolution to perform hand detection and operates using less power than the main camera 206.
[0039] The hand detector 130 includes a hand pattern detector 230, a hand temperature detector 234, a hand distance detector 236, and an activation signal unit 240. The hand pattern detector 230 is configured to process the image data 212 to determine whether the image data 212 includes a hand pattern 232. In an exemplary embodiment, the hand pattern detector 230 uses a neural network trained to recognize the hand pattern 232 to process the image data 212. In another exemplary embodiment, the hand pattern detector 230 applies one or more filters to the image data 212 to identify the hand pattern 232. The hand pattern detector 230 is configured to send a first signal 231 to the activation signal unit 240 for indicating whether the hand pattern 232 is detected. Although a single hand pattern 232 is depicted, in other embodiments, multiple hand patterns representing different aspects of the hand may be included, such as a finger-closed pattern, a finger-spread pattern, a partial hand pattern, etc.
[0040] The hand temperature detector 234 is configured to process the infrared sensor data 218 from the infrared sensor 208 and send a second signal 235 to the activation signal unit 240, which indicates whether the infrared sensor data 218 indicates a temperature source having a temperature representing a human hand. In some embodiments, the hand temperature detector 234 is configured to determine whether at least a portion of the field of view of the infrared sensor 208 has a temperature source within the temperature range indicating a human hand. In some embodiments, the hand temperature detector 234 is configured to receive data indicating the hand position from the hand pattern detector 230 to determine whether the temperature source at the hand position matches the temperature range of a human hand.
[0041] The hand distance detector 236 is configured to determine the distance 250 between the hand 190 and at least a portion of the device 102. In one example, the hand distance detector 236 processes the ultrasonic sensor data 220 and generates a third signal 237 indicating whether the hand 190 is within a specified distance range 238. In some embodiments, the hand distance detector 236 receives data indicating the position of the hand 190 from the hand pattern detector 230, the hand temperature detector 234, or both, and uses the hand position data to determine the region in the field of view of the ultrasonic sensor 210 corresponding to the hand 190. In other embodiments, the hand distance detector 236 identifies the hand 190 by locating an object that is closest to the screen 110 and exceeds a specified portion (e.g., 25%) of the field of view of the ultrasonic sensor 210.
[0042] In certain embodiments, range 238 has a lower limit of 10 centimeters (cm) and an upper limit of 30 cm (i.e., range 238 includes distances greater than or equal to 10 cm and less than or equal to 30 cm). In other embodiments, range 238 is adjustable. For example, device 102 may be configured to perform an update operation where the user positions hand 190 in a preferred position relative to device 102 such that distance 250 can be detected and used to generate range 238 (e.g., by applying a lower offset to detected distance 250 to set the lower limit and applying an upper offset to detected distance 250 to set the upper limit).
[0043] Activation signal unit 240 is configured to generate indication 132 in response to first signal 231, second signal 235, and third signal 237: First signal 231 indicates detection of hand pattern 232 in image data 212, second signal 235 indicates detection of hand temperature within the human hand temperature range, and third signal 237 indicates detection of hand 190 within range 238 (e.g., hand 190 is at distance 250 between 10 centimeters and 30 centimeters from screen 110). For example, in an implementation where each of signals 231, 235, and 237 has a binary "1" value indicating detected and a binary "0" value indicating not detected, activation signal unit 240 may generate indication 132 as the logical AND of signals 231, 235, and 237 (e.g., in response to all three signals 231, 235, 237 having a value of 1, indication 132 has a value of 1). In another example, activation signal unit 240 is also configured to generate indication 132 having a value of 1 in response to any two of signals 231, 235, 237 having a value of 1.
[0044] In other embodiments, one or more of signals 231, 235, and 237 have a multi-bit value that indicates the likelihood of meeting the corresponding hand detection criteria. For example, first signal 231 may have a multi-bit value indicating the confidence in detecting the hand pattern, second signal 235 may have a multi-bit value indicating the confidence in detecting the hand temperature, and third signal 235 may have a multi-bit value indicating the confidence that the distance of hand 190 from device 102 is within range 238. Activation signal unit 240 may combine signals 231, 235, and 237 and compare the combined result to a threshold to generate indication 132. For example, activation signal unit 240 may apply a set of weights to determine the weighted sum of signals 231, 235, and 237. Activation signal unit 240 may output indication 132 having a value indicating hand detection in response to the weighted sum exceeding the threshold. The values of the weights and threshold may be hard-coded, or alternatively, the values of the weights and threshold may be adjusted dynamically or periodically based on user feedback regarding false positives and false negatives, as described further below.
[0045] In some embodiments, the hand detector 130 is further configured to generate a second indication 242 in response to detecting that the hand 190 is no longer above the device 102. For example, the hand detector may output the second indication 242 as having a value of 0 (indicating that no hand movement is detected) in response to detecting the hand 190, and may update the second indication 242 to have a value of 1 (e.g., to indicate a change from the "hand detected" state to the "hand not detected" state) in response to determining that the hand is no longer detected. The second indication 242 may correspond to an end-of-utterance signal for the ASR system 140, as further explained with reference to Figure 3 Further explanation.
[0046] Although Figure 2 depicts a plurality of sensors including an ambient light sensor 204, a main camera 206, an infrared sensor 208, and an ultrasonic sensor 210, in other embodiments, one or more of the ambient light sensor 204, the main camera 206, the infrared sensor 208, or the ultrasonic sensor 210 are omitted. For example, although the ambient light sensor 204 is capable of generating at least a portion of the image data 212 to detect the shape of the hand using lower power compared to using the main camera 206, in some embodiments, the ambient light sensor 204 is omitted and the main camera 206 is used to generate the image data. To reduce power, the main camera 206 may operate according to a duty cycle (e.g., at intervals of a quarter of a second) for hand detection. As another example, although the main camera 206 is capable of generating at least a portion of the image data 212 to detect the shape of the hand with higher resolution and thus higher accuracy (compared to using the ambient light sensor 204), in some embodiments, the main camera 206 is omitted and the ambient light sensor 204 is used to generate the image data 212.
[0047] As another example, although the infrared sensor 208 can generate infrared sensor data 218 to detect whether an object has a temperature matching that of a human hand, in other embodiments, the infrared sensor 208 is omitted and the device 102 performs hand detection without considering the temperature. As another example, although the ultrasonic sensor 210 can generate ultrasonic sensor data 220 to detect whether the distance of an object is within the range 238, in other embodiments, the ultrasonic sensor 210 is omitted and the device 102 performs hand detection without considering the distance from the device 102. Alternatively, one or more other mechanisms can be implemented for distance detection, such as by comparing the object positions in the image data of multiple cameras from the device 102 (e.g., parallax) or multiple cameras of different devices (e.g., the vehicle in which the device 102 is located), by estimating the distance 250 using the size of the hand detected in the image data 212 or the infrared sensor data 218, or by projecting structured light or other electromagnetic signals to estimate the object distance, as illustrative non-limiting examples.
[0048] Although increasing the number of sensors and the variety of sensor types generally improves the accuracy of hand detection, in some embodiments, two sensors or a single sensor provide sufficient accuracy for hand detection. As a non-limiting example, in some embodiments, the only sensor data used for hand detection is the image data 212 from the ambient light sensor 204. Although in some embodiments the sensors 120 are activated simultaneously, in other embodiments, one or more of the sensors 120 are controlled according to a "cascaded" operation, where power is saved by keeping one or more of the sensors 120 inactive until the hand detection criteria are met based on the sensor data from another sensor 120. For illustration purposes, the main camera 206, the infrared sensor 208, and the ultrasonic sensor 210 can be kept inactive until the hand pattern detector 230 detects a hand pattern 232 in the image data 212 generated by the ambient light sensor 204, in response to which one or more of the main camera 206, the infrared sensor 208, and the ultrasonic sensor 210 are activated to provide additional sensor data to improve the accuracy of hand detection.
[0049] Figure 3 Example 300 is depicted, which shows a further aspect of the components that can be implemented in the device 102. As Figure 3 shown, the activation circuit 302 is coupled to the hand detector 130 and the ASR system 140, and the ASR system includes a buffer 320 accessible by the ASR engine 330. The device 102 also includes a virtual assistant application 340 and a speaker 350 (e.g., the device 102 implemented as a wireless speaker and voice command device).
[0050] Activation circuit 302 is configured to activate the automatic speech recognition system 140 in response to receiving an indication 132. For example, activation circuit 302 is configured to generate an activation signal 310 in response to indication 132 transitioning to a state indicating hand detection (e.g., indication 132 transitions from a value of 0 indicating no hand detection to a value of 1 indicating hand detection). The activation signal 310 is provided to the ASR system 140 via signal 306 to activate the ASR system 140. Activating the ASR system 140 includes: starting buffering of the audio signal 114 at buffer 320 to generate buffered audio data 322. The activation signal 310 is also provided to the microphone 112 via signal 304 for activating the microphone 112, enabling the microphone to generate the audio signal 114.
[0051] Activation circuit 302 is further configured to generate an end-of-utterance signal 312. For example, activation circuit 302 is configured to generate the end-of-utterance signal 312 in response to a second indication 242 transitioning to a state indicating the end of hand detection (e.g., the second indication 242 transitions from a value of 0 (indicating no change in hand detection) to a value of 1 (indicating that the detected hand is no longer detected)). The end-of-utterance signal 312 is provided to the ASR system 140 via signal 308 to cause the ASR engine 330 to begin processing the buffered audio data 332.
[0052] Activation circuit 302 is configured to selectively activate one or more components of the ASR system 140. For example, activation circuit 302 may include or be coupled to a power management circuit, a clock circuit, a head switch or foot switch circuit, a buffer control circuit, or any combination thereof. Activation circuit 302 may be configured to initiate power-on of buffer 320, ASR engine 330, or both (e.g., by selectively applying or raising the voltage of the power supply for buffer 320, ASR engine 330, or both). As another example, activation circuit 302 may be configured to selectively gate or ungate the clock signal to buffer 320, ASR engine 330, or both (e.g., preventing circuit operation without removing power).
[0053] The recognized speech 142 output by the ASR system 140 is provided to the virtual assistant application 340. For example, the virtual assistant application 340 may be implemented by one or more processors executing instructions, as described in further detail, for example, with reference to Figure 9 The virtual assistant application 340 may be configured to perform one or more search queries, such as via a wireless connection to an Internet gateway, a search server, or other resources, search the local storage of the device 102, or a combination thereof.
[0054] For illustration, the audio signal 114 may represent the spoken question "What's the weather like today?" The virtual assistant application 340 may generate a query to access an Internet-based weather service to obtain a weather forecast for the geographic area where the device 102 is located. The virtual assistant application 340 is configured to generate an output (e.g., an output audio signal 342) that causes the speaker 350 to generate an auditory output, such as in a voice interface implementation. In other embodiments, the virtual assistant application 340 generates another output mode, such as a visual output signal that can be displayed by a screen or display integrated in or coupled to the device 102.
[0055] In some implementations, parameter values such as weights and thresholds used by the device 102 (e.g., in the hand detector 130) can be set by a manufacturer or vendor of the device 102. In some implementations, the device 102 is configured to adjust one or more such values based on detected false negatives, false activations, or a combination thereof associated with the ASR system 140 during the life of the device 102. For example, a history of false activations can be maintained by the device 102 such that features of the sensor data 122 that triggered false activations can be periodically used to automatically adjust one or more weights or thresholds, such as to emphasize the relative reliability of one sensor in hand detection relative to another sensor to reduce the likelihood of future false activations.
[0056] Despite Figures 1 - 3 Specific values are included in the description of , such as a value of "1" indicating a positive result (e.g., hand detection) and a value of "0" indicating a negative result, but it should be understood that these values are provided for illustration purposes only and are not limiting. For illustration purposes, in some embodiments, the indication 132 is indicated by a value of "0". As another example, in some embodiments, a value of "1" for the first signal 231 indicates a high likelihood that the hand pattern 232 is in the image data 212, while in other embodiments, a value of "0" for the first signal 231 indicates a low likelihood that the hand pattern 232 is in the image data 212. Similarly, in some embodiments, a value of "1" for the second signal 235, the third signal 237, or both indicates a high likelihood that the hand detection criteria are met, and in other embodiments, a value of "1" for the second signal 235, the third signal 237, or both indicates a high likelihood that the hand detection criteria are not met.
[0057] Figure 4 An implementation 400 of a device 402 is depicted that includes a hand detector 130 and an ASR system 140 that are integrated in discrete components such as a semiconductor chip or package, as shown in FIG. Figure 9is further described. Device 402 includes an audio signal input 410 (e.g., a first bus interface) to enable receipt of an audio signal 114 from a microphone external to device 402. Device 402 also includes a sensor data input 412 (e.g., a second bus interface) to enable receipt of sensor data 122 from one or more sensors external to device 402. Device 402 may also include one or more outputs to provide processing results (e.g., recognized speech 142 or output audio signal 342) to one or more external components (e.g., speaker 350). Device 402 is capable of implementing hand detection and speech recognition activation as components in a system that includes a microphone and other sensors (e.g., in a vehicle as depicted in Figure 7 as depicted), a virtual reality or augmented reality headset as depicted in Figure 8 , or a wireless communication device as depicted in Figure 9 .
[0058] Refer to Figure 5 , which depicts a particular implementation of method 500 for processing an audio signal representative of an input sound that may be performed by device 102 or device 402. The method begins at 502 and includes, at 504, determining whether a hand is above the screen of the device by processing sensor data 122, e.g., by hand detector 130. In response to detecting a hand above the screen, at 506, the microphone and buffer are activated. For example, the microphone 112 and buffer 320 in Figure 3 are activated via signals 304 and 306 by activation circuit 302.
[0059] In response to determining that the hand has been removed from the screen, at 508, method 500 includes, at 510, activating the ASR engine to process the buffered data. For example, the ASR engine 330 is activated by signal 308 generated by activation circuit 302 to process buffered audio data 322.
[0060] Activating ASR when a hand is detected above the screen enables a user to activate speech recognition for voice commands by positioning the user's hand, without having to speak an activation keyword or locate and press a dedicated button. Thus, speech recognition can be conveniently and securely activated, e.g., when the user is driving a vehicle. Additionally, since placing the user's hand on the screen initiates activation of components to receive voice commands for speech recognition and removing the user's hand from the screen initiates processing of the received voice commands, incorrect activation, deactivation, or both of speech recognition can be reduced compared to systems that use keyword detection to activate speech recognition.
[0061] Refer to Figure 6, as an illustrative and non - limiting example, depicts a specific embodiment of method 600 for processing an audio signal representing an input sound that can be performed by device 102 or device 402.
[0062] Method 600 begins at 602 and includes: at the device, detecting that at least a portion of a hand is above at least a portion of the device, at 604. For example, hand detector 130 detects hand 190 by processing sensor data 122 received from one or more sensors 120. In some embodiments, detecting a portion of the hand above a portion of the device includes: processing image data (e.g., image data 212) to determine whether the image data includes a hand pattern (e.g., hand pattern 232). In one example, image data is generated at a low - power ambient light sensor of the device (e.g., ambient light sensor 204). Detecting a portion of the hand above a portion of the device may further include: processing infrared sensor data (e.g., infrared sensor data 218) from an infrared sensor of the device. Detecting a portion of the hand above a portion of the device may also include: processing ultrasonic sensor data (e.g., ultrasonic sensor data 220) from an ultrasonic sensor of the device.
[0063] Method 600 includes: at 606, in response to detecting a portion of the hand above a portion of the device, activating an automatic speech recognition system to process the audio signal. For example, device 102 activates ASR system 140 in response to indication 132. In some embodiments, activating the automatic speech recognition system includes: starting a cache of the audio signal, e.g., device 102 (e.g., activation circuit 302) activates buffer 320 via signal 306. In some examples, in response to detecting a portion of the hand above a portion of the device (e.g., above the screen of the device), method 500 further includes: activating a microphone to generate an audio signal based on the input sound, e.g., device 102 (e.g., activation circuit 302) activates microphone 112 via signal 304.
[0064] In some embodiments, method 600 includes: at 608, detecting that the portion of the hand is no longer above the portion of the device, and in response to detecting that the portion of the hand is no longer above the portion of the device, at 610, providing an end - of - utterance signal to the automatic speech recognition system. In one example, hand detector 130 detects that the hand is no longer above the portion of the device, and activation circuit 302 provides end - of - utterance signal 312 to ASR engine 330 in response to second indication 242.
[0065] By activating the ASR system in response to detecting a hand above a portion of the device, method 600 enables a user to activate voice recognition of voice commands without having to speak an activation keyword or perform a locate-and-press of a dedicated button. Thus, voice recognition can be conveniently and securely initiated, such as when the user is driving a vehicle. Additionally, compared to systems that use keyword detection to activate voice recognition, false activations of the ASR system can be reduced.
[0066] Figure 5 Method 500, Figure 6 Method 600, or both, can be implemented by a field programmable gate array (FPGA) device, an application specific integrated circuit (ASIC), a processing unit such as a central processing unit (CPU), a DSP, a controller, another hardware device, a firmware device, or any combination thereof. By way of example, Figure 5 Method 500, Figure 6 Method 600, or both, can be executed by a processor executing instructions, such as described with reference to Figure 9 as described.
[0067] Figure 7 FIG. depicts an example of an embodiment 700 of a hand detector 130 and an ASR system 140 integrated into a vehicle dashboard device (e.g., automotive dashboard device 702). A visual interface device, such as a screen 110 (e.g., a touch screen display), is mounted within the automotive dashboard device 702 and is visible to the driver of the vehicle. A microphone 112 and one or more sensors 120 are also mounted within the automotive dashboard device 702, but in other embodiments, one or more of the microphone 112 and sensors 120 can also be located elsewhere in the vehicle, such as the microphone 112 in the steering wheel or near the driver's head. The hand detector 130 and ASR system 140, as shown by the dashed boundary, indicate that the hand detector 130 and ASR system 140 are not visible to the vehicle occupants. The hand detector 130 and ASR system 140 can be implemented within a device that also includes the microphone 112 and sensors 120 (e.g., within Figures 1 - 3 device 120), or the hand detector 130 and ASR system 140 can be separate and coupled to the microphone 112 and sensors 120 (e.g., within Figure 4 device 402).
[0068] In some embodiments, multiple microphones 112 and sensor sets 120 are integrated into a vehicle. For example, the microphone and sensor sets can be placed on each passenger seat (e.g., on an armrest control panel or a seatback display device) so that each passenger can use their hand to detect and input a voice command on the device. In some embodiments, the voice commands of each passenger can be routed to a common ASR system 140; in other embodiments, the vehicle includes multiple ASR systems 140 to be able to simultaneously process voice commands from multiple occupants of the vehicle.
[0069] Figure 8 An example of an embodiment 800 depicting a hand detector 130 and an ASR system 140 integrated into a head-mounted device 802 (e.g., a virtual reality or augmented reality head-mounted device) is shown. A screen 110 is positioned in front of the user's eyes so that when the head-mounted device 802 is worn, augmented reality or virtual reality images or scenes can be displayed to the user, and the sensor 120 is positioned to initiate ASR recognition when it detects the user's hand on it (e.g., in front of the screen 110). The microphone 112 is positioned to receive the user's voice when the head-mounted device 802 is worn. When the head-mounted device 802 is worn, the user can raise their hand in front of the screen 110 to indicate to the head-mounted device 802 that the user is about to speak a voice command to activate the ASR, and can lower their hand to indicate that the user has finished speaking the voice command.
[0070] Figure 9 A block diagram depicting a specific illustrative embodiment of a device 900 including a hand detector 130 and an ASR engine 330, such as in a wireless communication device implementation (e.g., a smartphone). In various embodiments, the device 900 can have more or fewer components than Figure 9 shown. In an illustrative implementation, the device 900 can correspond to the device 102. In an illustrative implementation, the device 900 can perform one or more operations described with reference to Figures 1 - 8 this specification.
[0071] In a particular embodiment, the device 900 includes a processor 906 (e.g., a central processing unit (CPU)). The device 900 can include one or more additional processors 910 (e.g., one or more DSPs). The processor 910 can include a voice and music codec (CODEC) 908 and a hand detector 130. The voice and music codec 908 can include a voice encoder (“vocoder”) encoder 936, a vocoder decoder 938, or both.
[0072] Device 900 may include a memory 986 and a CODEC 934. The memory 986 may include instructions 956 that may be executed by the one or more additional processors 910 (or processor 906) to implement the functions described with reference to the hand detector 130, the ASR engine 330, Figure 1 the ASR system 140, the activation circuit 302, or any combination thereof. Device 900 may include a wireless controller 940 coupled to an antenna 952 via a transceiver 950.
[0073] Device 900 may include a display 928 (e.g., screen 110) coupled to a display controller 926. A speaker 350 and a microphone 112 may be coupled to the CODEC 934. The CODEC 934 may include a digital-to-analog converter 902 and an analog-to-digital converter 904. In a particular embodiment, the CODEC 934 may receive an analog signal from the microphone 112, convert the analog signal into a digital signal using the analog-to-digital converter 904, and provide the digital signal to the voice and music codec 908. The voice and music codec 908 may process the digital signal, and the digital signal may be further processed by the ASR engine 330. In a particular embodiment, the voice and music codec 908 may provide the digital signal to the CODEC 934. The CODEC 934 may convert the digital signal into an analog signal using the digital-to-analog converter 902 and may provide the analog signal to the speaker 350.
[0074] In a particular embodiment, device 900 may be included in a system-in-package or system-on-chip device 922. In a particular embodiment, the memory 986, the processor 906, the processor 910, the display controller 926, the CODEC 934, and the wireless controller 940 are included in the system-in-package or system-on-chip device 922. In a particular embodiment, the input device 930 (e.g., one or more of the sensors 120) and the power supply 944 are coupled to the system-on-chip device 922. Additionally, in a particular embodiment, as Figure 9 shown, the display 928, the input device 930, the speaker 350, the microphone 112, the antenna 992, and the power supply 944 are external to the system-on-chip device 922. In a particular embodiment, each of the display 928, the input device 930, the speaker 350, the microphone 112, the antenna 992, and the power supply 944 may be coupled to a component (e.g., an interface or a controller) of the system-on-chip device 922.
[0075] Device 900 may include a smart speaker (e.g., processor 906 may execute instructions 956 to run a voice-controlled digital assistant application 340), a soundbar, a mobile communication device, a smartphone, a cellular phone, a laptop computer, a computer, a tablet device, a personal digital assistant, a display device, a television, a gaming console, a music player, a radio, a digital video player, a digital video disc (DVD) or Blu-ray disc player, a tuner, a camera, a navigation device, a virtual reality or augmented reality head-mounted device, an in-vehicle console device, or any combination thereof.
[0076] In combination with the described embodiments, an apparatus for processing an audio signal representative of an input sound includes: a unit for detecting at least a portion of a hand above at least a portion of the device. For example, the unit for detecting a portion of the hand may correspond to a hand detector 130, a hand pattern detector 230, a hand temperature detector 234, a hand distance detector 236, one or more other circuits or components configured to detect at least a portion of a hand above at least a portion of the device, or any combination thereof.
[0077] The apparatus further includes a unit for processing the audio signal. The processing unit is configured to be activated in response to detecting a portion of the hand above the portion of the device. For example, the unit for processing the audio signal may correspond to an ASR system 140, an ASR engine 330, a microphone 112, a CODEC 934, a voice and music codec 908, one or more other circuits or components configured to process the audio signal and be activated in response to detecting a portion of the hand above the portion of the device, or any combination thereof.
[0078] In some embodiments, the apparatus includes a unit for displaying information, and the detecting unit is configured to detect a portion of the hand above the unit for displaying information. For example, the unit for displaying information may include a screen 110, a display 928, a display controller 926, one or more other circuits or components configured to display information, or any combination thereof.
[0079] The apparatus may further include: a unit for generating an audio signal based on the input sound, the generating unit being configured to be activated in response to detecting a portion of the hand above the unit for displaying information. For example, the unit for generating the audio signal may correspond to a microphone 112, a microphone array, a CODEC 934, a voice and music codec 908, one or more other circuits or components configured to generate an audio signal based on the input sound and be activated in response to a first indication, or any combination thereof.
[0080] In some embodiments, the device includes a unit for generating image data, and the detection unit is configured to determine whether the image data includes a hand pattern, such as hand pattern detector 230. In some embodiments, the device includes at least one of the following: a unit for detecting the temperature associated with a part of the hand (e.g., hand temperature detector 234, infrared sensor 208, or a combination thereof), and a unit for detecting the distance between a part of the hand and the device (e.g., hand distance detector 236, ultrasonic sensor 210, camera array, structured light projector, one or more other mechanisms for detecting the distance between a part of the hand and the device, or any combination thereof).
[0081] In some embodiments, a non-transitory computer-readable medium (e.g., memory 986) includes instructions (e.g., instructions 956) that, when executed by one or more processors of the device (e.g., processor 906, processor 910, or any combination thereof), cause the one or more processors to perform operations for processing an audio signal representative of an input sound. These operations include detecting that at least a portion of a hand is above at least a portion of the device (e.g., at hand detector 130). For example, detecting that a part of the hand is above a part of the device may include: receiving sensor data 122, using one or more detectors (e.g., hand pattern detector 230, hand temperature detector 234, or hand distance detector 236) to process the sensor data 122 to determine whether one or more detection criteria are met, and generating an indication 132 (e.g., as described with reference to activation signal unit 240) at least in part in response to detecting that one or more criteria are met. For example, in some embodiments, processing the sensor data 122 to determine whether the detection criteria are met includes: applying a neural network classifier trained to identify hand pattern 232 (e.g., as described with reference to hand pattern detector 230) to process image data 212, or applying one or more filters to the image data 212 to detect hand pattern 232.
[0082] The operations further include: in response to detecting that a part of the hand is on the part of the device, activating an automatic speech recognition system to process the audio signal. For example, activating automatic speech recognition may include: detecting the indication 132 at the input of the ASR system 140, and in response to detecting the indication 132, performing at least one of powering on or clock activating at least one component of the ASR system (e.g., buffer 320, ASR engine 330).
[0083] Those of ordinary skill in the art should also understand that the various exemplary logic blocks, configurations, modules, circuits, and algorithmic steps described in connection with the disclosed embodiments can be implemented as electronic hardware, computer software executed by a processor, or a combination of both. The above general descriptions of the various exemplary components, blocks, configurations, modules, circuits, and steps are centered around their functions. Whether such a function is implemented as hardware or as processor-executable instructions depends on the specific application and the design constraints imposed on the overall system. Skilled technicians can implement the described functions in a flexible manner for each specific application. However, such implementation decisions should not be construed as departing from the scope of protection of the present disclosure.
[0084] The steps of a method or algorithm described in connection with the disclosed embodiments can be directly embodied as hardware, a software module executed by a processor, or a combination of both. The software module can be located in a random access memory (RAM), flash memory, read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), registers, a hard disk, a removable hard disk, a compact disc read-only memory (CD-ROM), or any other form of non-transitory storage medium known in the art. An exemplary storage medium can be connected to the processor such that the processor can read information from, and write information to, the storage medium. In an alternative scenario, the storage medium can also be a component of the processor. The processor and the storage medium can reside in an application-specific integrated circuit (ASIC). The ASIC can reside in a computing device or a user terminal. Alternatively, the processor and the storage medium can reside as discrete components in a computing device or a user terminal.
[0085] In order to enable those of ordinary skill in the art to implement or use the disclosed embodiments, the above descriptions are made in connection with the disclosed embodiments. For those of ordinary skill in the art, various modifications to these embodiments are obvious, and the principles defined herein can also be applied to other embodiments without departing from the scope of protection of the present disclosure. Therefore, the present disclosure is not limited to the embodiments shown herein, but is consistent with the broadest scope of the principles and novel features as defined in the appended claims.
Claims
1. An apparatus for processing an audio signal representing an input sound, the apparatus comprising: A hand detector configured to generate a first indication in response to detecting at least a part of a hand above at least a part of the apparatus; A plurality of sensors coupled to the hand detector and configured to provide sensor data to the hand detector; And An automatic speech recognition system configured to be activated in response to the first indication to process the audio signal; Wherein the hand detector is further configured to determine that at least a part of the hand is above at least a part of the apparatus in response to a weighted sum of detection results of the sensor data of the plurality of sensors exceeding a threshold, and wherein the threshold and the weights for weighted summing the detection results of the sensor data of the plurality of sensors are adjustable based on historical data.
2. The apparatus according to claim 1, further comprising: A screen, wherein the hand detector is configured to generate the first indication in response to detecting at least a part of the hand above the screen; and A microphone configured to be activated in response to the first indication to generate the audio signal based on the input sound.
3. The device according to claim 2, wherein The hand detector is configured to generate the first indication in response to detecting the part of the hand at a distance of 10 cm to 30 cm from the screen.
4. The device according to claim 1, wherein The plurality of sensors includes a camera configured to provide image data to the hand detector.
5. The device according to claim 4, wherein, The camera includes a low-power ambient light sensor configured to generate the image data.
6. The device according to claim 4, wherein The hand detector includes: a hand pattern detector configured to process the image data to determine whether the image data includes a hand pattern.
7. The device according to claim 6, wherein The plurality of sensors further includes an infrared sensor.
8. The apparatus according to claim 7, wherein, The hand detector further includes: a hand temperature detector configured to process infrared sensor data from the infrared sensor.
9. The device according to claim 1, further comprising: An activation circuit coupled to the hand detector and configured to activate the automatic speech recognition system in response to receiving the first indication.
10. The device according to claim 1, wherein, The automatic speech recognition system includes a buffer and an automatic speech recognition engine, and wherein activating the automatic speech recognition system includes starting caching of the audio signal at the buffer.
11. The apparatus according to claim 10, wherein The hand detector is further configured to: generate a second indication in response to detecting that the part of the hand is no longer located above the part of the apparatus, the second indication corresponding to an end-of-speech signal that causes the automatic speech recognition engine to start processing the audio data in the buffer.
12. The apparatus according to claim 1, wherein, The hand detector and the automatic speech recognition system are integrated in a vehicle.
13. The device according to claim 1, wherein, The hand detector and the automatic speech recognition system are integrated in a portable communication device.
14. The device according to claim 1, wherein, The hand detector and the automatic speech recognition system are integrated in a virtual reality or augmented reality head-mounted device.
15. A method for processing an audio signal representing an input sound, the method comprising: Obtaining sensor data from a plurality of sensors at an apparatus; Determine that at least a portion of the detected hand is above at least a portion of the device in response to a weighted sum of detection results of sensor data from the plurality of sensors exceeding a threshold, wherein the threshold and the weights for weighted summing the detection results of the sensor data from the plurality of sensors are adjustable based on historical data; and In response to determining that the detected portion of the hand is above the portion of the device, activate an automatic speech recognition system to process the audio signal.
16. The method according to claim 15, wherein, The portion of the device includes a screen of the device, and the method further includes: in response to detecting that the portion of the hand is above the screen, activating a microphone to generate the audio signal based on the input sound.
17. The method according to claim 15, further comprising: Detecting that the portion of the hand is no longer located above the portion of the device; and In response to detecting that the portion of the hand is no longer located above the portion of the device, providing an end-of-utterance signal to the automatic speech recognition system.
18. The method according to claim 15, wherein Activating the automatic speech recognition system includes: starting caching of the audio signal.
19. The method according to claim 15, wherein Determining that the detected portion of the hand is above the portion of the device includes: processing image data to determine whether the image data includes a hand pattern.
20. The method according to claim 19, wherein, The image data is generated at a low-power ambient light sensor of the device.
21. The method according to claim 19, wherein, Determining that the detected portion of the hand is above the portion of the device further includes: processing infrared sensor data from an infrared sensor of the device.
22. A non-transitory computer-readable medium including instructions that, when executed by one or more processors of a device, cause the one or more processors to perform operations for processing an audio signal representing an input sound, the operations including: Obtaining sensor data from a plurality of sensors; Determine that at least a portion of the detected hand is above at least a portion of the device in response to a weighted sum of detection results of sensor data from the plurality of sensors exceeding a threshold, wherein the threshold and the weights for weighted summing the detection results of the sensor data from the plurality of sensors are adjustable based on historical data; and In response to determining that the detected portion of the hand is above the portion of the device, activate an automatic speech recognition system to process the audio signal.
23. The non-transitory computer-readable medium according to claim 22, wherein, The portion of the device includes a screen of the device, and the operations further include: in response to determining that the detected portion of the hand is above the screen, activating a microphone to generate the audio signal based on the input sound.
24. The non-transitory computer-readable medium according to claim 22, the operations further including: Detecting that the portion of the hand is no longer located above the portion of the device; and In response to detecting that the portion of the hand is no longer located above the portion of the device, providing an end-of-utterance signal to the automatic speech recognition system.
25. The non-transitory computer-readable medium according to claim 22, wherein, Determining that the detected portion of the hand is above the portion of the device includes: processing sensor data to detect the shape of the hand.
26. A device for processing an audio signal representing an input sound, the device comprising: a unit for obtaining sensor data from a plurality of sensors; a unit for determining that at least a part of a detected hand is above at least a part of the device in response to a weighted sum of detection results of the sensor data of the plurality of sensors exceeding a threshold, wherein the threshold and the weights for weighted summing the detection results of the sensor data of the plurality of sensors are adjustable based on historical data; and a unit for processing the audio signal, the processing unit being configured to be activated in response to determining that the detected part of the hand is above the part of the device.
27. The device according to claim 26, further comprising: a unit for displaying information, wherein the detecting unit is configured to detect that the part of the hand is above the unit for displaying information; and a unit for generating the audio signal based on the input sound, the generating unit being configured to be activated in response to determining that the detected part of the hand is above the unit for displaying.
28. The apparatus according to claim 26, further comprising: A unit for generating image data, and wherein the detecting unit is configured to determine whether the image data includes a hand pattern.
29. The device according to claim 28, further comprising at least one of the following: a unit for detecting a temperature associated with the part of the hand, and a unit for detecting a distance between the part of the hand and the device.
Citation Information
Patent Citations
Control apparatus, vehicle, and portable terminal
CN103869967A
On -vehicle rear -view mirror of formula that triggers
CN207758675U
Mobile terminal and menu control method thereof
US20090253463A1
Electronic device controlled by a motion and controlling method thereof
US20120179472A1
Apparatus and method for speech recognition
US20130085757A1