Voice trigger for digital assistant

A low-power voice trigger system using duty cycles and multiple detectors efficiently initiates voice-based digital assistants with specific sounds, addressing the need for hands-free operation and reducing power consumption.

JP2025098048AActive Publication Date: 2025-07-01APPLE INC
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
JP2025033999
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2013-02-07
Filing Date
2025-03-04
Publication Date
2025-07-01
Estimated Expiration
2034-02-07

AI Technical Summary

Technical Problem

Existing voice-based digital assistants require tactile input to initiate, consuming power and inhibiting a hands-free experience, and existing voice triggers consume excessive power due to continuous speech processing.

Method used

A low-power voice trigger system using multiple sound detectors operating in duty cycles to detect specific sounds or phrases, such as 'Hey Siri', without requiring continuous processor activity, allowing hands-free initiation of voice-based services.

Benefits of technology

Enables efficient, hands-free operation of voice-based digital assistants by reducing power consumption through selective sound detection, maintaining a low-power state for most operations and activating only when specific sounds are detected.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025098048000001_ABST
    Figure 2025098048000001_ABST
Patent Text Reader

Abstract

To provide a method for operating a voice trigger, storage media, and an electronic apparatus.SOLUTION: A method performed by an electronic apparatus that includes one or more processors and a memory for storing instructions executed by the one or more processors includes receiving a sound input. The sound input may correspond to a spoken word or a phrase or a portion thereof. The method also includes the steps of: determining whether or not at least a portion of the sound input corresponds to a predetermined type of sound, such as a human voice; determining whether or not at least a portion of the sound input corresponds to a predetermined type; determining whether or not the sound input includes predetermined content, such as a predetermined trigger word or phrase; and initiating speech-based service, such as voice-based digital assistant, upon determining that the sound input includes the predetermined content.SELECTED DRAWING: Figure 5
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] <Cross - Reference to Related Applications> This application claims the benefit of U.S. Provisional Application No. 61 / 762,260, filed on February 7, 2013, entitled "VOICE TRIGGER FOR A DIGITAL ASSISTANT", and is hereby incorporated by reference in its entirety for all purposes.

[0002] <Technical Field> The disclosed embodiments generally relate to digital assistants, and more specifically, to methods and systems for voice triggers for digital assistants.

Background Art

[0003] In recent years, voice - based digital assistants, such as Apple's SIRI (registered trademark), for handling various tasks such as web search and navigation, have been introduced into the market. One advantage of such voice - based digital assistants is that the user can interact bidirectionally with the device in a hands - free state without operating or visually observing the device. Hands - free operation can be particularly useful when a person cannot or should not physically operate the device, such as during driving. However, to initiate a voice - based assistant, the user generally needs to press a button or select an icon on the touch screen. This tactile input inhibits the hands - free experience. Accordingly, it is advantageous to provide a method and system for enabling a voice - based digital assistant (or other speech - based services) using voice input or signals rather than tactile input.

[0004] Enabling a voice-based assistant using voice input requires monitoring the audio channel to detect voice input. This monitoring consumes power, which is a limited resource on battery-dependent handheld devices or portable devices where such voice-based digital assistants are often run. Therefore, it would be beneficial to provide an energy-efficient voice trigger that can be used to initiate voice-based services and / or speech-based services on a device. SUMMARY OF THE INVENTION

[0005] Accordingly, there is a need for a low-power voice trigger that can provide an "always listening" type voice trigger function without consuming excessive limited power resources. The embodiments described below provide a system and method for using a voice trigger in an electronic device to initiate a voice-based assistant. The two-way interaction with a voice-based digital assistant (or other speech-based services such as a speech-to-text rewriting service) often starts when the user presses an affordance (e.g., a button or icon) on the device to enable the digital assistant, and subsequently the device provides some indication to the user that the digital assistant is enabled and listening, such as light, sound (e.g., beep), or vocal output (e.g., "What can I do for you?"). As described herein, the voice trigger can also be implemented to become active in response to specific and predefined words, phrases, or sounds without requiring a physical two-way interaction by the user. For example, the user may be able to enable the SIRI digital assistant of an IPHONE (both provided by Apple Inc., the assignee of the present application) by calling out the phrase "Hey Siri". Accordingly, the device emits a beep, sound, or speech output (e.g., "What can I do for you?"), indicating to the user that the listening mode is active. Accordingly, the user can initiate a two-way interaction with the digital assistant without physically touching the device that provides the digital assistant function.

[0006] One technique for starting a speech-based service with a voice trigger is to have the speech-based service continuously listen for a predetermined trigger word, phrase, or sound (any of which may be referred to herein as a "trigger sound"). However, continuously operating a speech-based service (e.g., a voice-based digital assistant) requires significant speech processing and battery power. To reduce the power consumption associated with providing the voice trigger function, various techniques may be employed. In some embodiments, the main processor of the electronic device (i.e., the "application processor") is maintained in a low-power state or a no-power state while one or more low-power audio detectors that are (e.g., because they are independent of the application processor) are maintained in an active state. (When in the low-power state or no-power state, the application processor or any other processor, program, or module may be described as being in an inactive or standby mode.) For example, even when the application processor is inactive, a low-power audio detector is used to monitor the audio channel for the trigger sound. This audio detector is sometimes referred to herein as the trigger sound detector. In some embodiments, it is configured to detect specific sounds, phonemes, and / or words. The trigger sound detector (including hardware components and / or software components) is designed to recognize a specific word, sound, or phrase, but such a task requires significant computational and power resources and thus generally cannot provide or is not optimized for a full speech-to-text function. Thus, in some embodiments, the trigger sound detector recognizes whether the audio input contains (e.g., matches the sonic pattern of) a predefined pattern (e.g., the word "Hey Siri"), but cannot (or is not configured to) convert the audio input to text or recognize many other words. When the trigger sound is detected, the digital assistant is then caused to exit the standby mode so that the user can provide a voice command.

[0007] In some embodiments, the trigger sound detector is configured to detect various different trigger sounds, such as a set of words, phrases, sounds, and / or combinations thereof. The user can then use any of these sounds to initiate a speech-based service. In one example, the voice trigger is preconfigured to respond to phrases such as "Hey Siri", "Activate Siri", "Call digital assistant", or "Hello, HAL, can you hear me, HAL?". In some embodiments, the user needs to select one of the preconfigured trigger sounds as a single trigger sound. In some embodiments, the user selects a subset of the preconfigured trigger sounds so that the user can initiate a speech-based service with different trigger sounds. In some embodiments, all of the preconfigured trigger sounds are maintained as valid trigger sounds.

[0008] In some embodiments, another sound detector is used and the trigger sound detector can also be maintained in a low power mode or a no power mode for much of the time. For example, different types of sound detectors (e.g., those that use less power than the trigger sound detector) are used to monitor an audio channel and determine whether the sound input corresponds to a particular type of sound. Sounds are classified into different "types" based on certain distinguishable sound characteristics. For example, a sound belonging to the "human voice" type has certain spectral content, periodicity, fundamental frequency, etc. Other types of sounds (e.g., whistles, claps, etc.) have different characteristics. Different types of sounds are identified using audio processing techniques and / or signal processing techniques as described herein. This sound detector is sometimes referred to herein as the "sound type detector". For example, if a predetermined trigger phrase is "Hey Siri", the sound type detector determines whether the input roughly corresponds to a human voice. If the trigger sound is a non-vocal sound such as a whistle, the sound type detector determines whether the sound input roughly corresponds to a whistle. When an appropriate type of sound is detected, the sound type detector activates the trigger sound detector to further process and / or analyze the sound. Since the sound type detector requires less power than the trigger sound detector (e.g., uses a circuit with low power requirements and / or an audio processing algorithm that is more efficient than the trigger sound detector), the voice trigger function consumes less power than the trigger sound detector alone.

[0009] In some embodiments, yet another sound detector is used, and both the sound type detector and the trigger sound detector described above can be maintained in a low power mode or a no power mode for much of the time. For example, a sound detector that uses less power than the sound type detector is used to monitor the audio channel and determine whether the sound input meets a predetermined condition such as an amplitude threshold (e.g., volume). This sound detector may also be referred to herein as a noise detector. When the noise detector detects a sound that meets a predetermined threshold, the noise detector activates the sound type detector to further process and / or analyze the sound. Since the noise detector requires less power than the sound type detector or the trigger sound detector (e.g., by using a circuit with low required power and / or an efficient audio processing algorithm), the voice trigger function consumes less power than a combination of a sound type detector and a trigger sound detector without a noise detector.

[0010] In some embodiments, any one or more of the above sound detectors operate according to a duty cycle that cycles between "on" and "off". This further helps to reduce the power consumption of the voice trigger. For example, in some embodiments, the noise detector is "on" for 10 milliseconds (i.e., actively monitors the audio channel) and then "off" for 90 milliseconds. In this way, while still effectively providing a continuous noise detection function, the noise detector is "off" 90% of the time. In some embodiments, the on and off durations for each sound detector are selected such that all of the detectors are active while the trigger sound is still being input. For example, for the trigger phrase "Hey Siri", the sound detector may be configured to become active without delay and analyze a sufficient amount of the input, regardless of where in the duty cycle(s) the trigger phrase begins. For example, the trigger sound detector becomes active without delay and receives, processes, and analyzes sufficient sound "Hey Siri" to determine that the sound matches the trigger phrase. In some embodiments, the sound input is stored in memory as received and passed to an upstream detector so that most of the sound input can be analyzed. Accordingly, even if the trigger sound detector is not activated until after the trigger phrase has been spoken, the entire recorded trigger phrase can still be analyzed.

[0011] Some embodiments provide a method of operating a voice trigger. The method is executed on an electronic device including one or more processors and a memory storing instructions executed by the one or more processors. The method includes receiving an audio input. The method further includes determining whether at least a portion of the audio input corresponds to a predetermined type of sound. The method further includes determining whether the audio input includes predetermined content when it is determined that at least a portion of the audio input corresponds to the predetermined type. The method further includes starting a speech-based service when it is determined that the audio input includes the predetermined content. In some embodiments, the speech-based service is a voice-based digital assistant. In some embodiments, the speech-based service is a dictation service.

[0012] In some embodiments, determining whether the audio input corresponds to a predetermined type of sound is performed by a first sound detector, and determining whether the audio input includes predetermined content is performed by a second sound detector. In some embodiments, the first sound detector consumes less power during operation than the second sound detector. In some embodiments, the first sound detector performs a frequency domain analysis of the audio input. In some embodiments, determining whether the audio input corresponds to a predetermined type of sound is performed when it is determined that the audio input satisfies a predetermined condition (e.g., as determined by a third sound detector described below).

[0013] In some embodiments, the first sound detector periodically monitors an audio channel according to a duty cycle. In some embodiments, the duty cycle includes an on-time of about 20 milliseconds and an off-time of about 100 milliseconds.

[0014] In some embodiments, the predetermined type is a human voice and the predetermined content is one or more words. In some embodiments, determining whether at least a portion of the audio input corresponds to a predetermined type of sound includes determining whether at least a portion of the audio input includes frequency characteristics of a human voice.

[0015] In some embodiments, the second sound detector is activated in response to a determination by the first sound detector that the sound input corresponds to a predetermined type. In some embodiments, the second sound detector operates for at least a predetermined time after the determination by the first sound detector that the sound input corresponds to a predetermined type. In some embodiments, the predetermined time corresponds to the duration of predetermined content.

[0016] In some embodiments, the predetermined content is one or more predetermined phonemes. In some embodiments, the one or more predetermined phonemes constitute at least one word.

[0017] In some embodiments, the method includes determining whether the sound input meets a predetermined condition before determining whether the sound input corresponds to a predetermined type of sound. In some embodiments, the predetermined condition is an amplitude threshold. In some embodiments, determining whether the sound input meets the predetermined condition is performed by a third sound detector, and the third sound detector consumes less power during operation than the first sound detector. In some embodiments, the third sound detector periodically monitors the audio channel according to a duty cycle. In some embodiments, the duty cycle includes an on-time of about 20 milliseconds and an off-time of about 500 milliseconds. In some embodiments, the third sound detector performs a time-domain analysis of the sound input.

[0018] In some embodiments, the method includes storing at least a portion of the sound input in a memory and providing a portion of the sound input to a speech-based service when the speech-based service is started. In some embodiments, a portion of the sound input is stored in the memory using direct memory access.

[0019] In some embodiments, the method includes determining whether the audio input corresponds to the voice of a specific user. In some embodiments, the speech-based service is initiated when it is determined that the audio input includes predetermined content and that the audio input corresponds to the voice of a specific user. In some embodiments, the speech-based service is initiated in a limited access mode when it is determined that the audio input includes predetermined content and that the audio input does not correspond to the voice of a specific user. In some embodiments, the method includes outputting an audio prompt including the name of the specific user when it is determined that the audio input corresponds to the voice of the specific user.

[0020] In some embodiments, determining whether the audio input includes predetermined content includes comparing a representation of the audio input to a reference representation and determining that the audio input includes the predetermined content if the representation of the audio input matches the reference representation. In some embodiments, a match is determined if the representation of the audio input matches the reference representation with a predetermined confidence value. In some embodiments, the method includes receiving a plurality of audio inputs including the audio input and repeatedly adjusting the reference representation using each one of the plurality of audio inputs in response to determining that each audio input includes the predetermined content.

[0021] In some embodiments, the method includes determining whether the electronic device is in a predetermined orientation and enabling a predetermined mode of the voice trigger when it is determined that the electronic device is in the predetermined orientation. In some embodiments, the predetermined orientation corresponds to the device's display screen being substantially horizontal and facing downward, and the predetermined mode is the standby mode. In some embodiments, the predetermined orientation corresponds to the device's display screen being substantially horizontal and facing upward, and the predetermined mode is the listening mode.

[0022] Some embodiments provide a method of operating a voice trigger. The method is performed on an electronic device that includes one or more processors and a memory storing instructions executable by the one or more processors. The method includes operating the voice trigger in a first mode. The method further includes determining whether one or more of a microphone and a camera of the electronic device are blocked to determine whether the electronic device is in a substantially enclosed space. The method further includes switching the voice trigger to a second mode when it is determined that the electronic device is in a substantially enclosed space. In some embodiments, the second mode is a standby mode.

[0023] Some embodiments provide a method of operating a voice trigger. The method is performed on an electronic device that includes one or more processors and a memory storing instructions executable by the one or more processors. The method includes determining whether the electronic device is in a predetermined orientation and enabling a predetermined mode of the voice trigger when it is determined that the electronic device is in the predetermined orientation. In some embodiments, the predetermined orientation corresponds to the device's display screen being substantially horizontal and facing downward, and the predetermined mode is a standby mode. In some embodiments, the predetermined orientation corresponds to the device's display screen being substantially horizontal and facing upward, and the predetermined mode is a listening mode.

[0024] According to some embodiments, an electronic device includes a sound receiving unit configured to receive a sound input, and a processing unit connected to the sound receiving unit. The processing unit determines whether at least a part of the sound input corresponds to a predetermined type of sound. When it is determined that at least a part of the sound input corresponds to the predetermined type, the processing unit determines whether the sound input includes a predetermined content. When it is determined that the sound input includes the predetermined content, the processing unit is configured to start a speech-based service. In some embodiments, the processing unit is further configured to determine whether the sound input satisfies a predetermined condition before determining whether the sound input corresponds to a predetermined type of sound. In some embodiments, the processing unit is further configured to determine whether the sound input corresponds to the voice of a specific user.

[0025] According to some embodiments, an electronic device includes a voice trigger unit configured to operate a voice trigger in a first mode of a plurality of modes, and a processing unit connected to the voice trigger unit. In some embodiments, the processing unit determines whether one or more of the microphone and camera of the electronic device are blocked, thereby determining whether the electronic device is in a substantially enclosed space. When it is determined that the electronic device is in a substantially enclosed space, the processing unit is configured to switch the voice trigger to a second mode. In some embodiments, the processing unit determines whether the electronic device is in a predetermined orientation. When it is determined that the electronic device is in the predetermined orientation, the processing unit is configured to enable a predetermined mode of the voice trigger.

[0026] According to some embodiments, a computer-readable storage medium (e.g., a non-transitory computer-readable storage medium) is provided. The computer-readable storage medium stores one or more programs to be executed by one or more processors of an electronic device. The one or more programs include instructions for performing any of the methods described herein.

[0027] According to some embodiments, an electronic device (e.g., a portable electronic device) is provided that includes means for performing any of the methods described herein.

[0028] According to some embodiments, there is provided an electronic device (e.g., a portable electronic device) including a processing unit configured to perform any of the methods described herein.

[0029] According to some embodiments, there is provided an electronic device (e.g., a portable electronic device) including one or more processors and a memory storing one or more programs executed by the one or more processors, the one or more programs including instructions to perform any of the methods described herein.

[0030] According to some embodiments, there is provided an information processing apparatus for use within an electronic device, the information processing apparatus including means for performing any of the methods described herein.

Brief Description of the Drawings

[0031]

Figure 1

[0032]

Figure 2

[0033]

Figure 3A

[0034]

Figure 3B

[0035]

Figure 3C

[0036]

Figure 4

[0037]

Figure 5

Figure 6

Figure 7

[0038]

Figure 8

Figure 9

[0039] Like reference numerals refer to corresponding parts throughout the drawings.

DETAILED DESCRIPTION OF THE INVENTION

[0040] FIG. 1 is a block diagram of an operating environment 100 of a digital assistant according to some embodiments. The terms “digital assistant,” “virtual assistant,” “intelligent automated assistant,” “voice-based digital assistant,” or “automated digital assistant” refer to any information processing system that interprets natural language input in oral and / or text form to infer user intent (e.g., identify the type of task corresponding to the natural language input) and perform an action based on the inferred user intent (e.g., perform a task corresponding to the identified type of task). For example, to act based on the inferred user intent, the system can perform one or more of the following. Identify a task flow having steps and parameters designed to fulfill the inferred user intent (e.g., identify the type of task), input specific requirements from the inferred user intent into the task flow, execute the task flow by calling a program, method, service, API, or the like (e.g., send a request to a service provider), and generate an output response to the user in audible (e.g., conversation) and / or visual form.

[0041] Specifically, once started, the digital assistant system can accept user requests at least partially in the form of natural language commands, requests, statements, narratives, and / or queries. Generally, a user request is seeking either an answer providing information or the execution of a task by the digital assistant system. Generally, a satisfactory response to a user request is either the provision of the requested information answer, the execution of the requested task, or a combination of the two. For example, a user may ask the digital assistant system a question such as "Where am I now?" Based on the user's current location, the digital assistant may answer, "You are near the west gate inside Central Park." The user can also request the execution of a task, for example, by stating "I want you to invite my friends to my girlfriend's birthday party next week." In response, the digital assistant may confirm the request by generating an audio output such as "Yes, right away." and then send appropriate calendar invitations from the user's email address to each of the user's friends listed in the user's email address book or contact list. There are also many other ways to interact with the digital assistant to request information or the execution of various tasks. In addition to providing verbal responses and taking programmed actions, the digital assistant can also provide responses in other visual or audio formats (such as in the form of text, alerts, music, video, animation, etc.).

[0042] As shown in FIG. 1, in some embodiments, a digital assistant system is implemented according to a client-server model. The digital assistant system includes a client-side portion (e.g., 102a and 102b) (hereinafter, “digital assistant (DA) client 102”) that runs on user devices (e.g., 104a and 104b), and a server-side portion 106 (hereinafter “digital assistant (DA) server 106”) that runs on server system 108. The DA client 102 communicates with the DA server 106 through one or more networks 110. The DA client 102 provides client-side functions such as user interaction input and output processing, and communication with the DA server 106. The DA server 106 provides server-side functions for any number of DA clients 102 that are each resident on a respective user device 104 (also referred to as a client device or an electronic device).

[0043] In some embodiments, the DA server 106 includes a client-facing I / O interface 112, one or more processing modules 114, data and models 116, an I / O interface 118 to external services, a photo and tag database 130, and a photo tagging module 132. The client-facing I / O interface facilitates client-facing input and output processing for the digital assistant server 106. The one or more processing modules 114 utilize the data and models 116 to determine a user's intention based on natural language input and execute a task based on the estimated user intention. The photo and tag database 130 stores fingerprints of digital photos, and optionally the digital photos themselves, and tags associated with the digital photos. The photo tagging module 132 creates, stores, automatically tags, and links tags to locations within photos in relation to the photos and / or fingerprints.

[0044] In some embodiments, the DA server 106 communicates with external services 120 (e.g., navigation service(s) 122-1, messaging service(s) 122-2, information service(s) 122-3, calendar service 122-4, telephone service 122-5, photo service(s) 122-6, etc.) through network(s) 110 for task completion or information acquisition. The I / O interface 118 to the external services facilitates such communication.

[0045] Examples of the user device 104 include, but are not limited to, a handheld computer, a wireless personal digital assistant (PDA), a tablet computer, a laptop computer, a desktop computer, a cellular phone, a smartphone, an enhanced general packet radio service (EGPRS) mobile phone, a media player, a navigation device, a game console, a television, a remote control device, or any combination of two or more of these data processing devices, or any other suitable data processing device. Further details regarding the user device 104 are provided with respect to the exemplary user device 104 shown in FIG. 2.

[0046] Examples of the communication network(s) 110 include a local area network (LAN) and a wide area network (WAN) such as the Internet, for example. The communication network(s) 110 can be implemented using any well-known network protocol, including various wired or wireless protocols such as Ethernet (registered trademark), Universal Serial Bus (USB), FIREWIRE (registered trademark), Global System for Mobile Communications (GSM (registered trademark)) for mobile communication, Enhanced Data GSM Environment (EDGE), code division multiple access (CDMA), time division multiple access (TDMA), Bluetooth (registered trademark), Wi-Fi (registered trademark), voice over Internet Protocol (VoIP), Wi-MAX (registered trademark), or any other suitable communication protocol, etc.

[0047] The server system 108 can be implemented on at least one data processing device and / or a distributed network of computers. In some embodiments, the server system 108 also utilizes the services of various virtual devices and / or third-party service providers (e.g., third-party cloud service providers) to provide the basic computing resources and / or infrastructure resources of the server system 108.

[0048] The digital assistant system shown in FIG. 1 includes both a client-side portion (e.g., DA client 102) and a server-side portion (e.g., DA server 106). However, in some embodiments, the digital assistant system refers to only the server-side portion (e.g., DA server 106). In some embodiments, the functions of the digital assistant can be implemented as a stand-alone application installed on the user device. Additionally, the distribution of functions between the client portion and the server portion of the digital assistant can vary depending on the embodiment. For example, in some embodiments, DA client 102 is a thin client that provides only user-facing input and output processing functions and delegates all other functions of the digital assistant to DA server 106. For example, in some embodiments, DA client 102 is configured to execute or assist with one or more functions of DA server 106.

[0049] FIG. 2 is a block diagram of user device 104 according to some embodiments. User device 104 includes a memory interface 202, one or more processors 204, and a peripheral interface 206. The various components within user device 104 are coupled by one or more communication buses or signal lines. User device 104 includes various sensors, subsystems, and peripheral devices coupled to peripheral interface 206. The sensors, subsystems, and peripheral devices collect information and / or facilitate the various functions of user device 104.

[0050] For example, in some embodiments, a motion sensor 210 (e.g., an accelerometer), a light sensor 212, a GPS receiver 213, a temperature sensor, and a proximity sensor 214 are coupled to peripheral interface 206 to facilitate the functions of orientation, light, and proximity sensing. In some embodiments, other sensors 216, such as biometric sensors, barometers, etc., are connected to peripheral interface 206 to facilitate the relevant functions.

[0051] In some embodiments, user device 104 includes a camera subsystem 220 coupled to a peripheral interface 206. In some embodiments, an optical sensor 222 of camera subsystem 220 facilitates camera functions such as taking photos and recording video clips. In some embodiments, user device 104 includes one or more wired and / or wireless communication subsystems 224 that provide communication functions. Communication subsystem 224 typically includes various communication ports, radio frequency receivers and transmitters, and / or optical (e.g., infrared) receivers and transmitters. In some embodiments, user device 104 includes an audio subsystem 226 coupled to one or more speakers 228 and one or more microphones 230 to facilitate audio-enabled functions such as speech recognition, voice response, digital recording, and telephone functions. In some embodiments, audio subsystem 226 is coupled to a voice trigger system 400. In some embodiments, voice trigger system 400 and / or audio subsystem 226 includes a low-power audio circuit and / or program (i.e., including hardware and / or software) for receiving and / or analyzing audio inputs, such as, for example, one or more analog-to-digital converters, digital signal processors (DSPs), voice detectors, memory buffers, codecs, etc. In some embodiments, the low-power audio circuit provides a voice (or sound) trigger function for one or more aspects of user device 104, such as a voice-based digital assistant or other speech-based service (alone or in addition to other components of user device 104). In some embodiments, the low-power audio circuit provides a voice trigger function even when other components of user device 104, such as processor(s) 204, I / O subsystem 240, memory 250, etc., are powered down and / or in standby mode. This voice trigger system 400 is described in further detail with respect to FIG. 4.

[0052] In some embodiments, I / O subsystem 240 is also coupled to peripheral interface 206. In some embodiments, user device 104 includes touch screen 246, and I / O subsystem 240 includes touch screen controller 242 coupled to touch screen 246. When user device 104 includes touch screen 246 and touch screen controller 242, touch screen 246 and touch screen controller 242 are typically configured to detect contact and movement or the interruption thereof using any of a plurality of touch sensing technologies, such as, for example, capacitive, resistive, infrared, surface acoustic wave technology, proximity sensor arrays, and the like. In some embodiments, user device 104 includes a display that does not include a touch sensitive surface. In some embodiments, user device 104 includes a separate touch sensitive surface. In some embodiments, user device 104 includes other input controller(s) 244. When user device 104 includes other input controller(s) 244, other input controller(s) 244 are typically coupled to other input / control devices 248, such as one or more buttons, rocker switches, thumb wheels, infrared ports, USB ports, and / or pointer devices such as a stylus.

[0053] The memory interface 202 is coupled to the memory 250. In some embodiments, the memory 250 includes a persistent computer-readable medium such as a high-speed random access memory and / or a non-volatile memory (e.g., one or more magnetic disk storage devices, one or more flash memory devices, one or more optical storage devices, and / or other non-volatile solid-state storage devices). In some embodiments, the memory 250 stores an operating system 252, a communication module 254, a graphical user interface module 256, a sensor processing module 258, a telephone module 260, and an application 262, as well as subsets or supersets thereof. The operating system 252 includes instructions for processing basic system services and instructions for performing hardware-dependent tasks. The communication module 254 facilitates communication with one or more additional devices, one or more computers, and / or one or more servers. The graphical user interface module 256 facilitates graphical user interface processing. The sensor processing module 258 facilitates sensor-related processing and functions (e.g., processing of audio input received using one or more microphones 228). The telephone module 260 facilitates telephone-related processes and functions. The application module 262 facilitates various functions of user applications such as electronic messaging, web browsing, media processing, navigation, imaging, and / or other processes and functions. In some embodiments, the user device 104 stores one or more software applications 270-1 and 270-2 associated with at least one of external service providers within the memory 250.

[0054] As described above, in some embodiments, the memory 250 also stores client - side digital assistant instructions (e.g., within the digital assistant client module 264) and various user data 266 (e.g., user - specific vocabulary data, preference data, and / or other data such as the user's electronic address book or contact list, to - do list, shopping list, etc.) to provide client - side functions of the digital assistant.

[0055] In some embodiments, the digital assistant client module 264 can receive voice input, text input, touch input, and / or gesture input through various user interfaces of the user device 104 (e.g., the I / O subsystem 244). The digital assistant client module 264 can also provide output in audio, visual, and / or tactile forms. For example, the output can be provided as voice, sound, alert, text message, menu, graphic, video, animation, vibration, and / or a combination of two or more of the above. During operation, the digital assistant client module 264 communicates with a digital assistant server (e.g., the digital assistant server 106, FIG. 1) using the communication subsystem 224.

[0056] In some embodiments, the digital assistant client module 264 collects additional information from the surrounding environment of the user device 104 using various sensors, subsystems, and peripheral devices to establish the context associated with the user input. In some embodiments, the digital assistant client module 264 provides context information or a subset thereof along with the user input to a digital assistant server (e.g., the digital assistant server 106, FIG. 1) to assist in inferring the user's intent.

[0057] In some embodiments, the context information obtainable with user input includes sensor information, such as illumination of the surrounding environment, ambient noise, ambient temperature, images or videos, etc. In some embodiments, the context information also includes the physical state of the device, such as the orientation of the device, the location of the device, device temperature, power level, speed, acceleration, motion pattern, cellular signal strength, etc. In some embodiments, information related to the software state of the user device 106, such as the processes running on the user device 104, installed programs, past and current network activities, background services, error logs, resource usage, etc., is also provided as context information related to user input to the digital assistant server (e.g., digital assistant server 106, FIG. 1).

[0058] In some embodiments, the DA client module 264 selectively provides information stored on the user device 104 (e.g., at least a portion of the user data 266) in response to a request from the digital assistant server. In some embodiments, the digital assistant client module 264 also solicits additional input from the user via a natural language dialog or other user interface in response to a request by the digital assistant server 106 (FIG. 1). The digital assistant client module 264 passes the additional input to the digital assistant server 106 to assist the digital assistant server 106 in estimating the user intention represented by the user request and / or achieving the user intention.

[0059] In some embodiments, the memory 250 may include additional instructions or fewer instructions. Further, various functions of the user device 104 may be implemented in the form of hardware and / or firmware, including in the form of one or more signal processing and / or application-specific integrated circuits, and thus, the user device 104 need not include all of the modules and applications shown in FIG. 2.

[0060] Figure 3A is a block diagram of an exemplary digital assistant system 300 (also referred to as a digital assistant) according to some embodiments. In some embodiments, the digital assistant system 300 is implemented on a stand-alone computer system. In some embodiments, the digital assistant system 300 is distributed across multiple computers. In some embodiments, some of the modules and functions of the digital assistant are split into a server portion and a client portion. The client portion resides on a user device (e.g., user device 104) and communicates with the server portion (e.g., server system 108) through one or more networks, as shown, for example, in FIG. 1. In some embodiments, the digital assistant system 300 is an embodiment of the server system 108 (and / or digital assistant server 106) shown in FIG. 1. In some embodiments, the digital assistant system 300 is implemented within a user device (e.g., user device 104, FIG. 1), thereby eliminating the need for a client-server system. The digital assistant system 300 is merely an example of a digital assistant system, and it should be noted that the digital assistant system 300 may have more or fewer components than shown, may combine two or more components, or may have a different configuration or arrangement of components. The various components shown in FIG. 3A may be implemented in the form of hardware, software, firmware, or a combination thereof, including one or more signal processing and / or application specific integrated circuits.

[0061] The digital assistant system 300 includes a memory 302, one or more processors 304, an input / output (I / O) interface 306, and a network communication interface 308. These components communicate with each other through one or more communication buses or signal lines 310.

[0062] In some embodiments, the memory 302 includes a persistent computer-readable medium such as a high-speed random access memory and / or a non-volatile computer-readable storage medium (e.g., one or more magnetic disk storage devices, one or more flash memory devices, one or more optical storage devices, and / or other non-volatile solid memory devices).

[0063] The I / O interface 306 couples the input / output devices 316 of the digital assistant system 300, such as a display, a keyboard, a touch screen, and a microphone, to the user interface module 322. The I / O interface 306 works in cooperation with the user interface module 322 to receive user inputs (e.g., voice input, keyboard input, touch input, etc.) and process them as appropriate. In some embodiments, when the digital assistant is implemented on a stand-alone user device, the digital assistant system 300 includes the components described with respect to the user device 104 in FIG. 2 and either an I / O and communication interface (e.g., one or more microphones 230). In some embodiments, the digital assistant system 300 represents the server portion of the digital assistant implementation and interacts with the user through a client-side portion resident on a user device (e.g., the user device 104 shown in FIG. 2).

[0064] In some embodiments, the network communication interface 308 includes a wired communication port(s) 312 and / or a wireless transceiver circuit 314. The wired communication port(s) receive and transmit communication signals via one or more wired interfaces such as Ethernet, Universal Serial Bus (USB), FIREWIRE®, etc. The wireless circuit 314 typically receives and transmits RF signals and / or optical signals to / from a communication network and other communication devices. The wireless communication can use any of a plurality of communication standards, protocols, and technologies such as GSM®, EDGE, CDMA, TDMA, Bluetooth®, Wi-Fi®, VoIP, Wi-MAX®, or any other suitable communication protocol. The network communication interface 308 enables communication between the digital assistant system 300 and a network such as the Internet, an intranet, etc., and / or a wireless network such as a cellular telephone network, a wireless local area network (LAN), etc., and / or a metropolitan area network (MAN), and other devices.

[0065] In some embodiments, the persistent computer-readable storage medium of the memory 302 stores programs, modules, instructions, and data structures that include all or a subset of an operating system 318, a communication module 320, a user interface module 322, one or more applications 324, and a digital assistant module 326. The one or more processors 304 execute these programs, modules, instructions, and perform read / write operations to / from the data structures.

[0066] The operating system 318 (e.g., an embedded operating system such as Darwin (registered trademark), RTXC (registered trademark), LINUX (registered trademark), UNIX (registered trademark), OS X (registered trademark), iOS (registered trademark), Windows (registered trademark), or VxWorks) includes various software components and / or drivers for controlling and managing general system tasks (e.g., memory management, storage device control, power management, etc.), and facilitates communication between various hardware, firmware, and software components.

[0067] The communication module 320 facilitates communication between the digital assistant system 300 and other devices through the network communication interface 308. For example, the communication module 320 can communicate with the communication module 254 of the device 104 shown in FIG. 2. The communication module 320 also includes various software components for processing data received by the wireless circuit 314 and / or the wired communication port 312.

[0068] In some embodiments, the user interface module 322 receives commands and / or inputs from the user via the I / O interface 306 (e.g., from a keyboard, touch screen, and / or microphone), and provides user interface objects on the display.

[0069] The application 324 includes programs and / or modules configured to be executed by one or more processors 304. For example, when the digital assistant system is implemented on a stand-alone user device, the application 324 may include user applications such as games, calendar applications, navigation applications, or email applications. When the digital assistant system 300 is implemented on a server farm, the application 324 may include, for example, resource management applications, diagnostic applications, or scheduling applications.

[0070] Memory 302 also stores a digital assistant module (i.e., the server part of the digital assistant) 326. In some embodiments, digital assistant module 326 includes the following sub-modules, or subsets or supersets thereof. That is, an input / output processing module 328, a speech-to-text (STT) processing module 330, a natural language processing module 332, a dialog flow processing module 334, a task flow processing module 336, a service processing module 338, and a photo module 132. Each of these processing modules has access to one or more of the following data and models of digital assistant 326, or subsets or supersets thereof. Ontology 360, vocabulary index 344, user data 348, classification module 349, disambiguation module 350, task flow model 354, service model 356, photo tagging module 358, search module 360, and local tag / photo storage 362.

[0071] In some embodiments, using the processing modules (e.g., input / output processing module 328, STT processing module 330, natural language processing module 332, dialog flow processing module 334, task flow processing module 336, and / or service processing module 338), data, and models implemented within digital assistant module 326, digital assistant system 300 performs at least some of the following. Identifying the user's intention expressed in natural language input received from the user, actively eliciting and obtaining the information necessary to fully infer the user's intention (e.g., by disambiguating words, names, meanings, etc.), determining a task flow to satisfy the inferred intention, and executing the task flow to satisfy the inferred intention. In some embodiments, the digital assistant also takes appropriate actions if a satisfactory response was not provided to the user or could not be provided for various reasons.

[0072] In some embodiments, as described below, the digital assistant system 300 processes natural language input to identify a user's intent and tag digital photos, tagging the digital photos with appropriate information. In some embodiments, the digital assistant system 300 also performs other tasks related to photos, such as searching for digital photos using natural language input, automatic tagging of photos, and the like. As shown in FIG. 3B, in some embodiments, the I / O processing module 328 acts bi-directionally with the user through the I / O device 316 of FIG. 3A or acts bi-directionally with a user device (e.g., the user device 104 of FIG. 1) through the network communication interface 308 of FIG. 3A, obtains user input (e.g., speech input), and provides a response to the user input. The I / O processing module 328 optionally obtains context information associated with the user input from the user device upon receiving the user input or immediately thereafter. The context information includes user-specific data, vocabulary, and / or preferences related to the user input. In some embodiments, the context information also includes the software and hardware state of the device (e.g., the user device 104 in FIG. 1) at the time the user request is received and / or information about the user's surrounding environment at the time the user request is received. In some embodiments, the I / O processing module 328 also sends additional questions to the user about the user request and receives answers from the user. In some embodiments, when the user request is received by the I / O processing module 328 and the user request includes voice input, the I / O processing module 328 transfers the voice input to the speech-to-text (STT) processing module 330 for voice-to-text conversion.

[0073] In some embodiments, the speech-to-text processing module 330 receives voice input (e.g., a user's utterance captured in a voice recording) through the I / O processing module 328. In some embodiments, the speech-to-text processing module 330 uses various acoustic and language models to recognize the voice input as a sequence of phonemes and ultimately as a sequence of words or tokens written in one or more languages. The speech-to-text processing module 330 is implemented using any suitable speech recognition techniques, acoustic models, and language models, such as hidden Markov models, dynamic time warping (DTW)-based speech recognition, and other statistical and / or analytical techniques. In some embodiments, the speech-to-text processing can be performed at least in part by a third-party service or on the user's device. When the speech-to-text processing module 330 obtains the result of the speech-to-text processing (e.g., a sequence of words or tokens), it passes the result to the natural language processing module 332 for intent estimation. The natural language processing module 332 (the "natural language processor") of the digital assistant 326 obtains the sequence of words or tokens (the "token sequence") generated by the speech-to-text processing module 330 and attempts to associate the token sequence with one or more "executable intents" recognized by the digital assistant. As used herein, an "executable intent" represents a task having an associated task flow that can be executed by the digital assistant 326 and / or the digital assistant system 300 (FIG. 3A) and is implemented within the task flow model 354. The associated task flow is a series of programmed actions and steps that the digital assistant system 300 takes to execute the task. The scope of the capabilities of the digital assistant system depends on the number and types of task flows implemented and stored within the task flow model 354, or, in other words, on the number and types of "executable intents" recognized by the digital assistant system 300. However, the effectiveness of the digital assistant system 300 also depends on the ability of the digital assistant system to infer the correct "executable intent(s)" from the user's request expressed in natural language.

[0074] In some embodiments, in addition to the sequence of words or tokens obtained from the speech-to-text processing module 330, the natural language processor 332 also receives context information associated with the user request (e.g., from the I / O processing module 328). The natural language processor 332 optionally uses the context information to clarify, complement, and / or further clarify the information contained within the token sequence received from the speech-to-text processing module 330. The context information includes, for example, user preferences, the state of the user device's hardware and / or software, sensor information collected before, during, or immediately after the user request, previous interactions (e.g., dialogs) between the digital assistant and the user, and the like.

[0075] In some embodiments, natural language processing is based on Ontology 360. Ontology 360 is a hierarchical structure that includes a plurality of nodes, and each node represents either a "plurality of executable intent groups" or one or more of other "plurality of attributes", such as an "executable intent" or an "attribute". As described above, an "executable intent" represents a task that the digital assistant system 300 has the ability to execute (e.g., a task that is "executable" or can be the target of execution). An "attribute" represents a parameter that is associated with an executable intent or a sub-aspect of another attribute. The link between the executable intent node and the attribute node in Ontology 360 defines how the parameter represented by the attribute node is related to the task represented by the executable intent node. In some embodiments, Ontology 360 is composed of executable intent nodes and attribute nodes. Within Ontology 360, each executable intent node is directly linked to one or more attribute nodes, or through one or more intermediate attribute nodes. Similarly, each attribute node is directly linked to one or more executable intent nodes, or through one or more intermediate attribute nodes. For example, the Ontology 360 shown in FIG. 3C includes an "Restaurant Reservation" node, which is an executable intent node. The attribute nodes "Restaurant", "Date / Time" (for reservation), and "Number of Persons" are directly connected to the "Restaurant Reservation" node (i.e., the executable intent node), respectively. Further, the attribute nodes "Dish", "Price Range", "Phone Number", and "Location" are sub-nodes of the attribute node "Restaurant" and are connected to the "Restaurant Reservation" node through the intermediate attribute node "Restaurant", respectively. For another example, the Ontology 360 shown in FIG. 3C also includes a "Reminder Setting" node, which is another executable intent node. The attribute nodes "Date / Time" (for reminder setting) and "Theme" (for reminder) are connected to the "Reminder Setting" node, respectively. Since the attribute node "Date / Time" is related to both the task of making a restaurant reservation and the task of setting a reminder, the attribute node "Date / Time" is connected to both the "Restaurant Reservation" node and the "Reminder Setting" node in Ontology 360.

[0076] An actionable intent node, together with its connected concept nodes, can be described as a "domain". In this description, each domain is associated with a respective actionable intent and refers to a group of nodes (and the relationships between them) associated with a particular actionable intent. For example, the ontology 360 shown in FIG. 3C includes an example of a restaurant reservation domain 362 and an example of a reminder domain 364 within the ontology 360. The restaurant reservation domain includes the actionable intent node "restaurant reservation", the attribute nodes "restaurant", "date / time", and "number of persons involved", and the sub-attribute nodes "dish", "price range", "phone number", and "location". The reminder domain 364 includes the actionable intent node "reminder setting", and the attribute nodes "theme" and "date / time". In some embodiments, the ontology 360 is composed of a number of domains. Each domain can share one or more other domains and one or more attribute nodes. For example, the attribute node "date / time" can be associated with many other domains (e.g., a scheduling domain, a travel reservation domain, a movie ticket domain, etc.) in addition to the restaurant reservation domain 362 and the reminder domain 364. FIG. 3C shows two exemplary domains within the ontology 360, but the ontology 360 may include other domains (i.e., actionable intents) such as "initiate a call", "find a route", "schedule a meeting", "send a message", and "provide an answer to a question", "tag a photo", etc. For example, the domain of "send a message" is associated with the actionable intent node of "send a message" and can further include attribute nodes such as "recipient(s)", "message type", and "message body". The attribute node "recipient" can be further defined by sub-attribute nodes such as "recipient name" and "message address", for example.

[0077] In some embodiments, Ontology 360 includes all domains (and thus actionable intents) that a digital assistant can understand and act upon. In some embodiments, Ontology 360 may be modified, such as by adding or removing domains or nodes, or changing the relationships between nodes within Ontology 360.

[0078] In some embodiments, nodes associated with multiple related actionable intents may be clustered under a "superordinate domain" within Ontology 360. For example, a "travel" superordinate domain may include a cluster of attribute nodes and actionable intent nodes related to travel. Actionable intent nodes related to travel may include, for example, "flight reservation", "hotel reservation", "car rental", "find directions", "find attractions", etc. Actionable intent nodes under the same superordinate domain (e.g., the "travel" superordinate domain) may share many attribute nodes. For example, actionable intent nodes for "flight reservation", "hotel reservation", "car rental", "find directions", "find attractions" may potentially share one or more of the attribute nodes "departure location", "destination", "departure date / time", "arrival date / time", and "number of persons involved".

[0079] In some embodiments, each node within the ontology 360 is associated with a set of words and / or phrases related to the attributes or actionable intents represented by that node. Each respective set of words and / or phrases associated with a node is the so-called "vocabulary" associated with that node. Each respective set of words and / or phrases associated with a node can be stored within the vocabulary index 344 (FIG. 3B) in relation to the attributes or actionable intents represented by that node. For example, returning to FIG. 3B, the vocabulary associated with the node for the attribute of "restaurant" may include words such as "food", "drink", "dish", "hunger", "eat", "pizza", "fast food", "meal", etc. As another example, the vocabulary associated with the node for the actionable intent of "initiate a phone call" may include words and phrases such as "call", "phone", "dial", "ring", "call this number", "make a call to", etc. The vocabulary index 344 optionally includes words and phrases in different languages. In some embodiments, the natural language processor 332 shown in FIG. 3B receives a token sequence (e.g., a text string) from the speech-to-text processing module 330 and determines which nodes are implied by the words within the token sequence. In some embodiments, if it is determined that a word or phrase within the token sequence is associated with one or more nodes within the ontology 360 (via the vocabulary index 344), then that word or phrase will "trigger" or "activate" those nodes. If multiple nodes are "triggered", based on the amount and / or relative importance of the activated nodes, the natural language processor 332 will select one of the actionable intents as the task (or type of task) that the user intended to have the digital assistant perform. In some embodiments, the domain having the most "triggered" nodes is selected.In some embodiments, the domain with the highest confidence value (e.g., based on the relative importance of its various triggered nodes) is selected. In some embodiments, the domain is selected based on a combination of the number and importance of the triggered nodes. In some embodiments, when selecting a node, additional factors such as whether the digital assistant system 300 has accurately interpreted similar requests from the user previously are also considered.

[0080] In some embodiments, the digital assistant system 300 also stores the names of specific entities in the vocabulary index 344. Therefore, when one of these names is detected in the user request, the natural language processor 332 will be able to recognize that the name refers to a specific instance of an attribute or sub-attribute within the ontology. In some embodiments, the names of specific entities are the names of companies, restaurants, people, movies, and the like. In some embodiments, the digital assistant system 300 can search for and identify specific entity names from other data sources such as the user's address book, contact list, movie database, musician database, and / or restaurant database. In some embodiments, when the natural language processor 332 identifies that a certain word in the token sequence is the name of a specific entity (such as a name in the user's address book or contact list), that word is given additional importance when selecting the actionable intent within the ontology for the user request. For example, if the word "Mr. Santo" is recognized from the user request and the surname "Santo" is found as one of the contacts in the user's contact list in the vocabulary index 344, then at that time, the user request is likely to correspond to the "send a message" or "initiate a call" domain. As another example, if the word "ABC Cafe" is found in the user request and the term "ABC Cafe" is found as the name of a specific restaurant in the user's city in the vocabulary index 344, then at that time, the user request is likely to correspond to the "restaurant reservation" domain.

[0081] User data 348 includes user-specific information such as the user's unique vocabulary, user preferences, user address, the user's default language and second language, the user's contact list, and other short-term or long-term information about each user. The natural language processor 332 can use the user-specific information to complement the information contained in the user input and further clarify the user's intention. For example, for the user request "I want to invite my friends to my birthday party", instead of asking the user to explicitly provide such information in the user's request to determine who the "friends" are and when and where the "birthday party" will be held, the natural language processor 332 can access the user data 348.

[0082] In some embodiments, the natural language processor 332 includes a classification module 349. In some embodiments, the classification module 349 determines whether each of one or more terms within a text string (e.g., corresponding to voice input associated with a digital photograph) is any of an entity, an action, or a location, as described in more detail below. In some embodiments, the classification module 349 classifies each term of the one or more terms as being one of an entity, an action, or a location. When the natural language processor 332 identifies an actionable intent (or domain) based on a user request, the natural language processor 332 generates a structured query to represent the identified actionable intent. In some embodiments, the structured query includes parameters for one or more nodes within a domain related to the actionable intent, and at least some of the parameters are augmented with specific information and requirements specified within the user request. For example, a user may say "Please make a dinner reservation for me at the sushi restaurant at 7:00." In this case, the natural language processor 332 may be able to accurately identify the actionable intent as "restaurant reservation" based on the user input. According to the ontology, a structured query for the "restaurant reservation" domain may include parameters such as {dish}, {time}, {date}, {number of people}, and the like. Based on the information included within the user's utterance, the natural language processor 332 may generate a partial structured query for the restaurant reservation domain. Here, the partial structured query includes the parameters {dish = "sushi"} and {time = "7:00 PM"}. However, in this example, the user's utterance does not include sufficient information to complete the structured query associated with the domain. Therefore, other required parameters such as {number of people} and {date} are not specified within the structured query based on the currently available information. In some embodiments, the natural language processor 332 adds received context information to some of the parameters of the structured query.For example, when a user requests a sushi restaurant "near me", the natural language processor 332 may add GPS coordinates from the user device 104 to the {location} parameter in the structured query.

[0083] In some embodiments, the natural language processor 332 passes a structured query (including any completed parameters) to a task flow processing module 336 (the "task flow processor"). The task flow processor 336 is configured to perform one or more of receiving a structured query from the natural language processor 332, completing the structured query, and performing actions required to "complete" the user's ultimate request. In some embodiments, various procedures necessary to complete these tasks are provided within the task flow model 354. In some embodiments, the task flow model 354 includes procedures for obtaining additional information from the user and a task flow for performing actions associated with actionable intents. As described above, in order to complete the structured query, the task flow processor 336 may need to initiate an additional dialog with the user to obtain additional information and / or resolve the ambiguity of potentially ambiguous utterances. When such an interaction is required, the task flow processor 336 invokes a dialog processing module 334 (dialog processor) to engage in the interaction with the user. In some embodiments, the dialog processing module 334 determines how (and / or when) to request additional information from the user, receives user responses, and processes them. In some embodiments, questions are provided to the user through the I / O processing module 328 and responses are received from the user. For example, the dialog processing module 334 presents dialog output to the user via audio and / or visual output and receives input from the user via oral or physical (e.g., touch gesture) responses. Continuing with the above example, when the task flow processor 336 invokes the dialog processor 334 to determine the "number of people" and "date" information for a structured query associated with the domain "restaurant reservation", the dialog processor 334 generates questions such as "How many people?" and "Which day would you like?" to pass to the user.When receiving an answer from the user, the dialog processing module 334 passes the information to the task flow processor 336 to add the missing information to the structured query or to complete the missing information from the structured query.

[0084] In some cases, the task flow processor 336 may receive a structured query with one or more ambiguous attributes. For example, a structured query for the "send message" domain may indicate that the intended recipient is "Bob", and the user may have multiple contacts with the name "Bob". The task flow processor 336 will request that the dialog processor 334 remove the ambiguity of this attribute of the structured query. As a result, the dialog processor 334 may ask the user "Which Bob?", and display (or read out) a list of contacts with the name "Bob" that the user can choose from.

[0085] In some embodiments, the dialog processor 334 includes an ambiguity removal module 350. In some embodiments, the ambiguity removal module 350 removes the ambiguity of one or more ambiguous terms (e.g., one or more ambiguous terms of a text string corresponding to voice input associated with a digital photo). In some embodiments, for a first term of one or more terms that has multiple candidate meanings, the ambiguity removal module 350 prompts the user for additional information about the first term, receives the additional information corresponding to the prompt from the user, and identifies the entity, action, and location associated with the first term according to the additional information.

[0086] In some embodiments, the ambiguity removal module 350 removes the ambiguity of pronouns. In such embodiments, the ambiguity removal module 350 identifies one of the one or more terms as a pronoun and determines the noun that the pronoun refers to. In some embodiments, the ambiguity removal module 350 uses a contact list associated with the user of the electronic device to determine the noun that the pronoun refers to. Alternatively, or additionally, the ambiguity removal module 350 determines the noun that the pronoun refers to as the name of an entity, action, or location identified in a previous voice input associated with a previously tagged digital photograph. Alternatively, or additionally, the ambiguity removal module 350 determines the noun that the pronoun refers to as the name of a person identified based on a previous voice input associated with a previously tagged digital photograph. In some embodiments, the ambiguity removal module 350 accesses information obtained from one or more sensors (e.g., proximity sensor 214, light sensor 212, GPS receiver 213, temperature sensor 215, motion sensor 210) of a handheld electronic device (e.g., user device 104) to determine one or more meanings of the terms. In some embodiments, the ambiguity removal module 350 identifies two terms respectively associated with any of an entity, an action, or a location. For example, the first of the two terms refers to a person, and the second of the two terms refers to a location. In some embodiments, the ambiguity removal module 350 identifies three terms respectively associated with any of an entity, an action, or a location.

[0087] When the task flow processor 336 completes a structured query for a feasible intention, the task flow processor 336 proceeds to execute the final task associated with the feasible intention. Accordingly, the task flow processor 336 executes steps and instructions within the task flow model according to the specific parameters included within the structured query. For example, a task flow model for the feasible intention of "restaurant reservation" may include steps and instructions for contacting the restaurant and actually making a reservation for a specific number of guests at a specific time. For example, using a structured query such as {restaurant reservation, restaurant = ABC Cafe, date = 3 / 12 / 2012, time = 7:00 PM, number of guests = 5}, the task flow processor 336 may (1) log in to a restaurant reservation system configured to accept reservations to the server of ABC Cafe or multiple restaurants such as ABC Cafe, (2) enter date, time, and number of guests information into a form on the website, (3) submit the form, and (4) make a calendar entry for the reservation in the user's calendar. In another example, although described in more detail below, the task flow processor 336 may execute steps and instructions related to tagging or searching digital photos in response to voice input, for example, in cooperation with the photo module 132. In some embodiments, the task flow processor 336 uses the assistance of the service processing module 338 ("service processor") to complete the task requested by the user input or to provide an answer to the information requested by the user input. For example, the service processor 338 can make a phone call, set a calendar item, call a map search, call or interact bidirectionally with other user applications installed on the user device, and call or interact bidirectionally with third-party services (e.g., restaurant reservation portals, social network websites or services, banking portals, etc.) instead of the task flow processor 336.In some embodiments, the protocols and application programming interfaces (APIs) required by each service can be specified by each service model in service model 356. The service processor 338 accesses an appropriate service model for the service and generates service requests according to the protocols and APIs required by the service associated with the service model.

[0088] For example, if a restaurant enables an online reservation service, the restaurant can present a service model that specifies the parameters required for making a reservation and the API for communicating the values of the required parameters to the online reservation service. When requested by the task flow processor 336, the service processor 338 uses the web address stored in the service model 356 to establish a network connection with the online reservation service and can send the required reservation parameters (e.g., time, date, number of persons involved) to the online reservation interface in a format compliant with the API of the online reservation service.

[0089] In some embodiments, the natural language processor 332, the dialog processor 334, and the task flow processor 336 are used jointly and iteratively to infer and clarify the user's intention, obtain information to further clarify and narrow down the user's intention, and finally generate a response that achieves the user's intention (e.g., provide output to the user or complete a task).

[0090] In some embodiments, after all the tasks necessary to fulfill the user's request have been executed, the digital assistant 326 formulates a confirmation response and sends the response back to the user through the I / O processing module 328. If the user request is for an answer to information, the confirmation response presents the requested information to the user. In some embodiments, the digital assistant also asks the user to indicate whether the user is satisfied with the response created by the digital assistant 326.

[0091] Now, attention is drawn to FIG. 4, which is a block diagram showing the components of the voice trigger system 400 according to some embodiments. (The voice trigger system 400 is not limited to voice, and the embodiments described herein apply equally to non-voice.) The voice trigger system 400 is configured within the electronic device 104 with various components, modules, and / or software programs. In some embodiments, the voice trigger system 400 includes a noise detector 402, a voice type detector 404, a trigger voice detector 406, a speech-based service 408, and an audio subsystem 226, each of which is connected to an audio bus 401. In some embodiments, more or fewer of these modules are used. The voice detectors 402, 404, and 406 may be referred to as modules and may include hardware (e.g., circuits, memory, processors, etc.), software (e.g., programs, software on a chip, firmware, etc.), and / or any combination thereof for performing the functions described herein. In some embodiments, as shown by the dashed lines in FIG. 4, the voice detectors are communicatively, programmatically, physically, and / or operably (e.g., via a communication bus) connected to each other. (For simplicity of explanation, FIG. 4 shows each voice detector connected only to adjacent voice detectors. It will be understood that each voice detector may similarly be connected to any of the other voice detectors.)

[0092] In some embodiments, the audio subsystem 226 includes a codec 410, an audio digital signal processor (DSP) 412, and a memory buffer 414. In some embodiments, this audio subsystem 226 is connected to one or more microphones 230 (FIG. 2) and one or more speakers 228 (FIG. 2). The audio subsystem 226 provides audio input to the sound detectors 402, 404, 406 and the speech-based service 408 (similarly, other components or modules such as a telephone and / or the baseband subsystem of a telephone) for processing and / or analysis. In some embodiments, the audio subsystem 226 is connected to an external audio system 416 that includes at least one microphone 418 and at least one speaker 420.

[0093] In some embodiments, the speech-based service 408 is a voice-based digital assistant and corresponds to one or more components or functions of the digital assistant system described above in connection with FIGS. 1-3C. In some embodiments, the speech-based service is a speech-to-text service, a dictation service, etc. In some embodiments, the noise detector 402 monitors the audio channel and determines whether the audio input from the audio subsystem 226 meets a predetermined condition such as an amplitude threshold. The audio channel corresponds to a stream of voice information received by one or more sound collection devices such as one or more microphones 230 (FIG. 2). The audio channel refers to the voice information regardless of its processing state, or refers to the specific hardware that is processing and / or transmitting the voice information. For example, the audio channel may refer to the analog electrical impulses from the microphone 230 (and / or the circuit through which they are propagated), as well as the digitally encoded audio stream as a result of the processing of the analog electrical impulses (by, for example, the audio subsystem 226 and / or any other audio processing system of the electronic device 104).

[0094] In some embodiments, the predetermined condition is whether the audio input exceeds a specific volume at a predetermined time. In some embodiments, the noise detector uses time-domain analysis of the audio input, which requires relatively few computational resources and battery resources compared to other types of analysis (such as performed by the audio type detector 404, trigger word detector 406, and / or speech-based service 408). In some embodiments, other types of signal processing and / or audio analysis, including, for example, frequency-domain analysis, are used. When the noise detector 402 determines that the audio input meets the predetermined condition, it activates an upstream audio detector such as the audio type detector 404 (e.g., by providing a control signal to start one or more processing routines and / or by providing power to the upstream audio detector). In some embodiments, the upstream audio detector is activated according to other satisfied conditions. For example, in some embodiments, the upstream audio detector is activated in response to determining that the device is not stored in an enclosed space (e.g., based on a light detector that detects a threshold level of light).

[0095] The sound type detector 404 monitors the audio channel and determines whether the audio input corresponds to a specific type of sound, such as a sound unique to a human voice, a whistle, applause, etc. The types of sounds configured to be recognized by the sound type detector 404 correspond to the specific trigger sound(s) configured to be recognized by the voice trigger. In embodiments where the trigger sound is speech or a phrase, the sound type detector 404 includes a "voice activity detector" (VAD). In some embodiments, the sound type detector 404 uses frequency domain analysis of the audio input. For example, the sound type detector 404 generates a spectrogram of the received audio input (e.g., using a Fourier transform) to analyze the spectral components of the audio input and determine whether the audio input is likely to correspond to a specific sound type or classification (e.g., a human voice). Thus, in embodiments where the trigger sound is speech or a phrase, if the audio channel picks up background noise (e.g., traffic noise) instead of a human voice, the VAD does not activate the trigger sound detector 406. In some embodiments, the sound type detector 404 is maintained active as long as a predetermined condition of any downstream sound detector (e.g., the noise detector 402) is met. For example, in some embodiments, the sound type detector 404 is maintained active as long as the audio input includes sound above a predetermined amplitude threshold (as determined by the noise detector 402), and becomes inactive when the sound drops below the predetermined threshold. In some embodiments, once activated, the sound type detector 404 is maintained active until conditions such as the expiration of a timer (e.g., 1, 2, 5, or 10 seconds, or any other appropriate duration), the completion of a specific on / off cycle count of the sound type detector 404, or the occurrence of an event (e.g., the amplitude of the sound drops below a second threshold as determined by the noise detector 402 and / or the sound type detector 404) are met.

[0096] As described above, when the sound type detector 404 determines that the audio input corresponds to a predetermined type of sound, it activates an upstream sound detector such as the trigger sound detector 406 (e.g., by providing a control signal to start one or more processing routines and / or by providing power to the upstream sound detector).

[0097] The trigger sound detector 406 is configured to determine whether the audio input includes at least a portion of a specific predetermined content (e.g., at least a portion of a trigger word, phrase, or sound). In some embodiments, the trigger sound detector 406 compares a representation of the audio input (the "input representation") to one or more reference representations of the trigger word. If the input representation matches at least one of the one or more reference representations with an acceptable confidence value, the trigger sound detector 406 starts the speech-based service 408 (e.g., by providing a control signal to start one or more processing routines and / or by providing power to an upstream sound detector). In some embodiments, the input representation and the one or more reference representations are spectrograms (or mathematical representations thereof), which represent how the spectral density of a signal changes over time. In some embodiments, the representation is another type of audio signature or voiceprint. In some embodiments, starting the speech-based service 408 includes exiting one or more circuits, programs, and / or processors from standby mode and invoking a sound-based service. The sound-based service then prepares to provide more comprehensive speech recognition, speech-to-text processing, and / or natural language processing. In some embodiments, the voice trigger system 400 includes a voice authentication function so that it can determine whether the audio input corresponds to the voice of a specific person, such as the owner / user of the device. For example, in some embodiments, the sound type detector 404 uses voiceprinting technology to determine that the audio input was spoken by an authorized user. Voice authentication and voiceprinting are described in more detail in U.S. Patent Application No. 13 / 053,144, which is assigned to the assignee of the present application and is hereby incorporated by reference in its entirety. In some embodiments, the voice authentication is included in any of the sound detectors described herein (e.g., the noise detector 402, the sound type detector 404, the trigger sound detector 406, and / or the speech-based service 408).In some embodiments, voice authentication is implemented as a module separate from the above-described sound detector (e.g., as voice authentication module 428, FIG. 4), and may be operably arranged after noise detector 402, after sound type detector 404, after trigger sound detector 406, or at any other suitable location.

[0098] In some embodiments, trigger sound detector 406 is maintained active as long as the conditions of any downstream sound detector(s) (e.g., noise detector 402 and / or sound type detector 404) are met. For example, in some embodiments, trigger sound detector 406 is maintained active as long as the sound input includes sound above a predetermined threshold (as detected by noise detector 402). In some embodiments, it is maintained active as long as the sound input includes a particular type of sound (as detected by sound type detector 404). In some embodiments, it is maintained active as long as both of the aforementioned conditions are met.

[0099] In some embodiments, once activated, the trigger sound detector 406 is maintained active until conditions are met such as the expiration of a timer (e.g., 1, 2, 5, or 10 seconds, or any other appropriate duration), the completion of a specific number of on / off cycles of the trigger sound detector 406, or the occurrence of an event (e.g., the sound amplitude falling below a second threshold). In some embodiments, when one sound detector activates another detector, both sound detectors are maintained active. However, the sound detectors may be enabled or disabled multiple times, and it is not necessary to enable all of the downstream (e.g., lower power and / or miniaturized) sound detectors (or for each condition to be met) in order to enable the upstream sound detector. For example, in some embodiments, the noise detector 402 and the sound type detector 404 determine that their respective conditions are met, and after the trigger sound detector 406 is activated, one or both of the noise detector 402 and the sound type detector 404 become disabled and / or enter a standby mode during the operation of the trigger sound detector 406. In other embodiments, both (or one or the other) of the noise detector 402 and the sound type detector 404 are maintained active during the operation of the trigger sound detector 406. In various embodiments, different combinations of sound detectors become active at different times, and whether one becomes enabled or disabled may depend on the state of other sound detectors or may be independent of the state of other sound detectors.

[0100] FIG. 4 illustrates three individual sound detectors, each configured to detect a different aspect of the sound input, and in various embodiments of the voice trigger, more or fewer sound detectors may be used. For example, in some embodiments, only the trigger sound detector 406 is used. In some embodiments, the trigger sound detector 406 is used in combination with either the noise detector 402 or the sound type detector 404. In some embodiments, all of the detectors 402 - 406 are used. In some embodiments, additional sound detectors are also included.

[0101] Moreover, different combinations of sound detectors may be used at different times. For example, a particular combination of sound detectors and how they act bidirectionally may depend on one or more conditions, such as context or the operating state of the device. As one specific example, when the device is connected to a power source (and thus does not rely solely on battery power), the trigger sound detector 406 is active while the noise detector 402 and the sound type detector 404 are maintained inactive. In another example, when the device is in a pocket or backpack, all sound detectors are inactive. A power-saving voice trigger function can be provided by cascading sound detectors as described above, where detectors that require a lot of power are called only when needed by detectors that require less power. As described above, further power savings are achieved by operating one or more of the sound detectors according to a duty cycle. For example, in some embodiments, the noise detector 402 operates according to a duty cycle so as to perform efficient continuous noise detection even when the noise detector is at least temporarily inactive. In some embodiments, the noise detector 402 is on for 10 milliseconds and off for 90 milliseconds. In some embodiments, the noise detector 402 is on for 20 milliseconds and off for 500 milliseconds. Other on and off durations are also possible.

[0102] In some embodiments, if the noise detector 402 detects noise during its "on" interval, the noise detector 402 is maintained on and further processes and / or analyzes the sound input. For example, the noise detector 402 may be configured to activate an upstream sound detector when it detects a sound above a predetermined amplitude for a predetermined time (e.g., 100 milliseconds). Thus, if the noise detector 402 detects a sound above a predetermined amplitude during its 10-millisecond "on" interval, it does not immediately enter the "off" interval. Instead, the noise detector 402 is maintained active, continues to process the sound input, and determines whether it exceeds a threshold for a predetermined total duration (e.g., 100 milliseconds).

[0103] In some embodiments, the sound type detector 404 operates according to a duty cycle. In some embodiments, the sound type detector 404 is on for 20 milliseconds and off for 100 milliseconds. Other on and off durations are also possible. In some embodiments, the sound type detector 404 can determine whether the sound input corresponds to a predetermined type of sound during the "on" interval of its duty cycle. Thus, if the sound type detector 404 determines during its "on" interval that the sound is of a particular type, the sound type detector 404 activates the trigger sound detector 406 (or any other upstream sound detector). Alternatively, in some embodiments, when the sound type detector 404 detects a sound that can correspond to a predetermined type during the "on" interval, the detector does not immediately enter the "off" interval. Instead, the sound type detector 404 remains active and continues to process the sound input to determine whether it corresponds to a predetermined type of sound. In some embodiments, when the sound detector determines that a predetermined type of sound has been detected, it activates the trigger sound detector 406, further processes the sound input, and determines whether a trigger sound has been detected. Similar to the noise detector 402 and the sound type detector 404, in some embodiments, the trigger sound detector 406 operates according to a duty cycle. In some embodiments, the trigger sound detector 406 is on for 50 milliseconds and off for 50 milliseconds. Other on and off durations are also possible. When the trigger sound detector 406 detects during its "on" interval that there is a sound that can correspond to the trigger sound, the detector does not immediately enter the "off" interval. Instead, the trigger sound detector 406 remains active and continues to process the sound input to determine whether it includes the trigger sound. In some embodiments, when such a sound is detected, the trigger sound detector 406 remains active and processes the sound for a predetermined duration, such as 1, 2, 5, or 10 seconds, or any other appropriate duration. In some embodiments, the duration is selected based on the length of the particular trigger word or sound that is configured to be detected. For example, if the trigger phrase is "Hey Siri", the trigger word detector operates for about 2 seconds to determine whether the sound input includes the phrase.

[0104] In some embodiments, some of the sound detectors are operated according to a duty cycle, and others operate continuously when enabled. For example, in some embodiments, only the first sound detector operates according to a duty cycle (e.g., the noise detector 402 in FIG. 4), and the upstream sound detectors operate continuously once activated. In some other embodiments, the noise detector 402 and the sound type detector 404 operate according to a duty cycle, while the trigger sound detector 406 operates continuously. Whether a particular sound detector operates continuously or according to a duty cycle depends on one or more conditions, such as the context or the operating state of the device. In some embodiments, when the device is connected to a power source and does not rely solely on battery power, all of the sound detectors operate continuously once activated. In other embodiments, when the device is in a pocket or backpack (as determined by, e.g., sensors and / or microphone signals), the noise detector 402 (or any of the sound detectors) operates according to a duty cycle, but operates continuously when it is determined that the device may not be stored. In some embodiments, whether a particular sound detector operates continuously or according to a duty cycle depends on the battery charge level of the device. For example, the noise detector 402 operates continuously when the battery charge is more than 50%, and operates according to a duty cycle when the battery charge is less than 50%. In some embodiments, the voice trigger includes noise, echo, and / or sound cancellation functions (collectively referred to as noise cancellation). In some embodiments, the noise cancellation is performed by the audio subsystem 226 (e.g., by the audio DSP 412). The noise cancellation reduces or removes unwanted noise or sound from the audio input before it is processed by the sound detector. In some cases, the unwanted noise is background noise from the user's environment, such as clicks from a fan or keyboard operation. In some embodiments, the unwanted noise is any sound above or below a predetermined amplitude or frequency.For example, in some embodiments, sounds above the vocal range of an average person (e.g., 3,000 Hz) are filtered out or removed from the signal. In some embodiments, multiple microphones (e.g., microphone 230) are used to help determine which components of the received sound should be reduced and / or removed. For example, in some embodiments, audio subsystem 226 uses beamforming techniques to identify each portion of the sound or audio input that originates from a single point in space (e.g., the user's mouth). Audio subsystem 226 then focuses on this sound by removing from the audio input sounds that are equally received by all microphones (e.g., background sounds that do not originate from any particular direction).

[0105] In some embodiments, DSP 412 is configured to cancel or remove from the audio input sounds that are being output by the device on which the digital assistant is operating. For example, if audio subsystem 226 is outputting music, radio, a podcast, voice output, or any other audio content (e.g., via speaker 228), DSP 412 removes any of the output sounds that are picked up by the microphone and included in the audio input. Accordingly, the audio input does not include (or at least includes less of) this output voice. In response, the audio input provided to the sound detector is cleaner and a more accurate trigger. The noise cancellation aspects are described in further detail in U.S. Patent No. 7,272,224, which is assigned to the assignee of the present application and is hereby incorporated by reference in its entirety.

[0106] In some embodiments, different sound detectors require the sound input to be filtered and / or pre - processed in different ways. For example, in some embodiments, the noise detector 402 is configured to analyze the time - domain audio signal between 60 and 20,000 Hz, and the sound type detector is configured to perform a frequency - domain analysis of the audio between 60 and 3,000 Hz. Thus, in some embodiments, the audio DSP 412 (and / or other audio DSPs of device 104) pre - processes the received audio according to the needs of each sound detector. In some embodiments, on the other hand, the sound detectors are configured to filter and / or pre - process the audio from the audio subsystem 226 according to their specific needs. In such cases, the audio DSP 412 may still perform noise cancellation before providing the sound input to the sound detectors. In some embodiments, the context of the electronic device is used to help determine whether and how the voice trigger is operating. For example, when the device is in a pocket, wallet, or backpack, it is unlikely that the user will invoke a speech - based service such as a voice - based digital assistant. Also, during a loud rock concert, it is unlikely that the user will invoke a speech - based service. Some users are unlikely to invoke speech - based services at a particular time (e.g., late at night). On the other hand, there are also contexts where the user is likely to use the voice trigger to invoke speech - based services. For example, some users are likely to use the voice trigger while driving, when alone, at work, etc. Various techniques are used to determine the context of the device. In various embodiments, the device uses information from one or more of the following components or information sources to determine the context of the device. That is, a GPS receiver, an optical sensor, a microphone, a proximity sensor, an orientation sensor, an inertial sensor, a camera, a communication circuit and / or antenna, a charging circuit and / or a power supply circuit, a switch position, a temperature sensor, a compass, an accelerometer, a calendar, user preferences, etc.The context of the device can subsequently be used to adjust whether and how to operate the voice trigger. For example, in certain contexts, the voice trigger is disabled (or operates in a different mode) as long as the context is maintained. For example, in some embodiments, the voice trigger is disabled when the phone is in a predetermined orientation (e.g., placed face down on a surface), during a predetermined period (e.g., between 10:00 PM and 8:00 AM), when the phone is in the "silent" or "do not disturb" mode (e.g., based on a switch position, mode setting, or user preference), when the device is in a substantially enclosed space (e.g., a pocket, bag, wallet, drawer, or glove box), when the device is near another device having a voice trigger and / or speech-based service (e.g., based on a proximity sensor, voice communication / wireless communication / infrared communication), etc. In some embodiments, instead of being disabled, the voice trigger system 400 operates in a low power mode (e.g., by operating the noise detector 402 according to a duty cycle with a 10 millisecond "on" interval and a 5 second "off" interval). In some embodiments, the audio channel is monitored less frequently when the voice trigger system 400 is operating in the low power mode. In some embodiments, the voice trigger uses a different sound detector or combination of sound detectors when in the low power mode than when in the normal mode. (The voice trigger may allow for many different modes or operating states, each of which may use a different amount of power, and different embodiments use them according to their specific designs.).

[0107] On the other hand, if the device is in some other context, the voice trigger remains active (or operates in a different mode) as long as the context is maintained. For example, in some embodiments, the voice trigger is active when the device is connected to power, the phone is in a predetermined orientation (e.g., placed face-up on a surface), during a predetermined period (e.g., between 8:00 AM and 10:00 PM), the device is in motion and / or in a vehicle (e.g., based on GPS signal, Bluetooth® connection, or connection to a vehicle, etc.). The manner of detecting the verification when the device is in a vehicle is further described in more detail in U.S. Provisional Patent Application No. 61 / 657,744, which is hereby incorporated by reference in its entirety and belongs to the assignee of the present application. Various specific examples of methods for determining a particular context are provided below. In various embodiments, different technologies and / or information sources are used to detect these and other contexts.

[0108] As described above, whether the voice trigger system 400 is active (e.g., during listening) can depend on the physical orientation of the device. In some embodiments, the voice trigger is active when the device is placed "face up" on a surface (e.g., with the display and / or touch screen surface visible), and / or inactive when "face down". This provides an easy way for the user to enable and / or disable the voice trigger without the need to operate a settings menu, switch, or button. In some embodiments, the device detects whether it is placed face up or face down on a surface using a light sensor (e.g., based on the difference in incident light on the front and back faces of device 104), proximity sensor, magnetic sensor, accelerometer, gyroscope, tilt sensor, camera, etc. In some embodiments, other operating modes, settings, parameters, or preferences are affected by the orientation and / or position of the device. In some embodiments, the specific trigger sound, word, or phrase that the voice trigger is listening for depends on the orientation and / or position of the device. For example, in some embodiments, the voice trigger listens for a first trigger word, phrase, or sound when the device is in one orientation (e.g., face up on a surface), and a different trigger word, phrase, or sound when the device is in a different orientation (e.g., face down). In some embodiments, the trigger phrase for face down is longer and / or more complex than that for face up. Thus, the user can place the device face down when other people are around or in a noisy environment, reducing the likelihood of unauthorized activation for shorter or simpler trigger words while still enabling the voice trigger to operate. As one specific example, the face up trigger phrase could be "Hey Siri", while the face down trigger phrase could be "Hey Siri, this is Andrew, please wake up". The longer trigger phrase also provides a longer audio sample for processing and / or analysis by the sound detector and / or voice authenticator, thus increasing the accuracy of the voice trigger and reducing unauthorized activation.

[0109] In some embodiments, the device 104 detects whether the device is inside a vehicle (e.g., an automobile). The voice trigger is particularly useful for calling speech-based services when the user is inside the vehicle to help reduce the physical two-way actions required to operate the device and / or speech-based services. In fact, one of the advantages of a voice-based digital assistant is that it can be used to perform tasks when it is not possible or dangerous to touch and operate the device. Thus, the voice trigger may be used when the device is inside the vehicle so that the user does not need to touch the device to call the digital assistant. In some embodiments, the device determines that it is inside the vehicle by detecting that it is connected to and / or paired with the vehicle through BLUETOOTH® communication (or other wireless communication) or through something such as a docking connector or cable. In some embodiments, the device determines that it is inside the vehicle by determining the location and / or speed of the device (e.g., using a GPS receiver, an accelerometer, and / or a gyroscope). For example, if it is determined that the device is moving at over 20 miles per hour and moving along a road, it is determined that the device is likely inside the vehicle, and the voice trigger is subsequently maintained as enabled and / or maintained in a high-power state or a high-sensitivity state.

[0110] In some embodiments, the device detects whether the device is stored (e.g., in a pocket, wallet, bag, drawer, etc.) by determining whether it is in a substantially enclosed space. In some embodiments, the device uses an optical sensor (e.g., a dedicated ambient light sensor and / or a camera) to determine that it is stored. For example, in some embodiments, the device is likely stored if the optical sensor detects weak light or no light. In some embodiments, the time and / or the location of the device are also considered. For example, if a high light level is expected (e.g., during the day) and the optical sensor detects a low light level, the device is stored and the voice trigger system 400 may not be needed. Accordingly, this voice trigger system 400 enters a low power state or a standby state. In some embodiments, the difference in light detected by sensors located on opposite sides of the device can be used to determine its position and thus whether it is stored. Specifically, when the device is not stored in a pocket or bag and is placed on a table or surface, the user may attempt to enable the voice trigger. When the device is placed face down (or face up) on a surface such as a table or desk, one side of the device is blocked and the other surface is exposed to ambient light, while only weak light or no light hits that side. Thus, if the front and back light sensors of the device detect significantly different light levels, the device is determined not to be stored. On the other hand, if the light sensors on opposite sides detect the same or similar light levels, the device is determined to be stored in a substantially enclosed space. Also, if both light sensors detect a low light level during the day (or if the device expects the phone to be in a bright environment), the device is determined to be stored with a high confidence value.

[0111] In some embodiments, other techniques are used to determine (instead of or in addition to the optical sensor) whether the device is being stored. For example, in some embodiments, the device emits one or more sounds (e.g., tones, clicks, pings, etc.) from a speaker or transducer (e.g., speaker 228), monitors one or more microphones or transducers (e.g., microphone 230), and detects the echo of the omitted sound(s). (In some embodiments, the device emits inaudible signals, such as sounds outside the human audible range.) From the echo, the device determines the characteristics of the surrounding environment. For example, a relatively large environment (e.g., indoors or in a vehicle) reflects sound differently than a relatively narrow, enclosed environment (e.g., a pocket, wallet, bag, drawer, etc.).

[0112] In some embodiments, the voice trigger system 400 operates differently when it is close to other devices (such as other devices having voice trigger and / or speech-based services) than when it is away from other devices. This can be useful, for example, in stopping or reducing the sensitivity of the voice trigger system 400 so that when multiple devices are close to each other, if a person utters a trigger word, other surrounding devices are not similarly triggered. In some embodiments, the device uses RFID, proximity communication, infrared signals / acoustic signals, etc. to determine proximity to other devices. As described above, voice trigger is particularly useful when the device is operating in hands-free mode, such as when the user is driving. In such cases, the user often uses an external audio system such as a wired headset or wireless headset, a wristwatch with a speaker and / or microphone, a vehicle built-in microphone and speaker, etc., and does not need to hold the device close to the face to make a call or dictate text input. For example, a wireless headset and a vehicle audio system may be connected to the electronic device using BLUETOOTH (registered trademark) communication, or any other suitable wireless communication. However, this can be inefficient for a voice trigger that monitors the voice received via a wireless audio accessory, due to the power required to maintain an open audio channel with the wireless accessory. In particular, a wireless headset can hold sufficient power in its battery to provide several hours of continuous talk time, and thus, rather than being used to simply monitor ambient voice and wait for a potential trigger sound, it is suitable for storing the battery for when the headset is actually needed for communication. Moreover, a wired external headset accessory may require more power than an on-board microphone alone, and keeping the headset microphone active consumes the device's battery charging power. This is particularly true considering that the ambient voice received by a wireless headset or wired headset usually consists mostly of silence or irrelevant sounds.Thus, in some embodiments, the voice trigger system 400 monitors the voice from the microphone 230 on the device even if the device is connected to an external microphone (wired or wireless). Subsequently, when the voice trigger detects the trigger word, the device initiates an active audio link with the external microphone and subsequently receives audio inputs (such as commands to the voice-based digital assistant) via the external microphone rather than the microphone 230 on the device. If certain conditions are met, an active communication link can be maintained between the device and an external audio system 416 (which may be communicatively coupled to the device 104 via wired or wireless means), and the voice trigger system 400 can listen for the trigger sound via the external audio system 416 instead of (or in addition to) the microphone 230 on the device. For example, in some embodiments, the movement characteristics of the electronic device and / or the external audio system 416 (determined, for example, by accelerometers, gyroscopes, etc. on each device) are used to determine whether the voice trigger system 400 should monitor background sound using the microphone 230 on the device or the external microphone 418. Specifically, the difference in movement between the device and the external audio system 416 provides information about whether the external audio system 416 is actually in use. For example, if both the device and the wireless headset are moving substantially equally (or not moving), it may be determined that the headset is not in use or not being worn. This can occur, for example, because both devices are close to each other and in an idle state (e.g., placed on a table or in a pocket, bag, wallet, drawer, etc.). Accordingly, under these conditions, the voice trigger system 400 monitors the microphone on the device because it is unlikely that the headset is actually being used. If there is a difference in movement between the wireless headset and the device, it is determined that the user is wearing the headset.These conditions may occur, for example, while the headset is worn on the user's head (where at least some movement may occur even if the wearer is relatively stationary) because the device is placed, e.g., on a surface or inside a bag. Under these conditions, since the headset is considered to be worn, the voice trigger system 400 maintains an active communication link and monitors the headset microphone 418 instead of (or in addition to) the microphone 230 on the device. This technique focuses on the differences in movement between the device and the headset, so common movement of both devices cancels out. This can be useful, for example, when the user is using the headset in a moving vehicle where the device (e.g., a mobile phone) is in a cup holder, on an empty seat, or in the user's pocket and the headset is worn on the user's head. When common movement of both devices cancels out (e.g., the movement of the vehicle), relative movement of the headset as compared to the device can be determined (if any), and it can be determined whether the headset is likely in use (or whether the headset is not worn). The above description refers to wireless headsets, but similar techniques apply equally to wired headsets.

[0113] Since human voices vary widely, it may be necessary or useful to tune the voice trigger to improve its accuracy in recognizing the voice of a specific user. Also, a person's voice can change over time due to, for example, natural voice changes due to illness, aging, or hormonal changes, etc. Thus, in some embodiments, the voice trigger system 400 can adapt to the voice and / or speech recognition profile of a particular user or user group. As described above, the sound detector (e.g., the sound type detector 404 and / or the trigger sound detector 406) may be configured to compare a representation of the sound input (e.g., a sound or utterance provided by the user) to one or more reference representations. For example, if the input representation matches the reference representation at a predetermined confidence level, the sound detector determines that the sound input corresponds to a predetermined type of sound (e.g., the sound type detector 404) or that the sound input contains a predetermined content (e.g., the trigger sound detector 406). To tune the voice trigger system 400, in some embodiments, the device adjusts the reference representation to which the input representation is compared. In some embodiments, the reference representation is adjusted (or created) as part of a voice registration procedure or "training" procedure, where the user outputs the trigger sound several times so that the device can adjust (or create) the reference representation. The device then creates a reference representation using the person's actual voice.

[0114] In some embodiments, the device uses the trigger sound received under normal usage conditions to adjust the reference representation. (For example, when a voice input that meets all of the triggering criteria is found) After a normal voice triggering event, for example, the device uses the information from the voice input to adjust and / or tune the reference representation. In some embodiments, only the voice input that is determined to meet all or part of the triggering criteria at a particular confidence value level is used to adjust the reference representation. Thus, if the voice trigger has low reliability in terms of the voice input corresponding to or including the trigger sound, that voice input may be ignored for the purpose of adjusting the reference representation. On the other hand, in some embodiments, the voice input that meets the voice trigger system 400 at a low confidence value is used to adjust the reference representation.

[0115] In some embodiments, as more audio inputs are received, device 104 repeatedly adjusts the reference representation (using these or other techniques) to adapt to slight changes in the user's voice over time. For example, in some embodiments, device 104 (and / or related devices or services) adjusts the reference representation after each normal triggering event. In some embodiments, device 104 analyzes the audio input associated with each normal triggering event, determines whether the reference representation should be adjusted based on that input (e.g., if certain conditions are met), and adjusts the reference representation only if it is appropriate to do so. In some embodiments, device 104 maintains a moving average of the reference representation over a long period of time. In some embodiments, voice trigger system 400 detects sounds that do not meet one or more of the triggering criteria (as determined by, for example, one or more of the sound detectors), although this may actually be attempted by a legitimate user. For example, voice trigger system 400 may be configured to respond to a trigger phrase such as "Hey Siri", but if the user's voice changes (e.g., due to illness, aging, accent / tone changes, etc.), voice trigger system 400 may not recognize the user's attempt to activate the device. (This may also occur if voice trigger system 400 is set to default conditions and / or the user has not initialized or performed a training procedure to customize voice trigger system 400 for that particular user's voice, etc., such that voice trigger system 400 is not properly tuned for that user's voice.) If voice trigger system 400 does not respond to the user's first attempt to activate the voice trigger, the user will probably repeat the trigger phrase. The device detects that these repeated audio inputs are similar to each other and / or similar to the trigger phrase (even if they are not similar enough to cause voice trigger system 400 to enable the speech-based service).When such conditions are met, the device determines that the audio input corresponds to a legitimate attempt to enable the voice trigger system 400. In response, in some embodiments, the voice trigger system 400 uses those received audio inputs to adjust the voice trigger system 400 in one or more ways so that similar utterances by the user are approved as legitimate triggers in the future. In some embodiments, these audio inputs are used to adapt the voice trigger system 400 only when certain conditions or combinations of conditions are met. For example, in some embodiments, the audio input is used to adapt the voice trigger system 400 when a predetermined number of audio inputs are received continuously (e.g., 2, 3, 4, 5, or any other suitable number), when the audio input is sufficiently similar to a reference representation, when the audio inputs are sufficiently similar to each other, when the audio inputs are close to each other (e.g., received within a predetermined period of time and / or at a predetermined interval and / or in the vicinity thereof), and / or in any combination of these or other conditions. In some cases, the voice trigger system 400 may detect one or more audio inputs that do not meet one or more of the triggering criteria, where manual initiation of a speech-based service (e.g., by pressing a button or icon) continues. In some embodiments, the voice trigger system 400 determines that the audio input actually corresponds to a failed voice triggering attempt when a speech-based service is started shortly after the audio input is received. In response, the voice trigger system 400 uses those received audio inputs to adjust the voice trigger system 400 in one or more ways so that utterances by the user are approved as legitimate triggers in the future, as described above.

[0116] While the adaptation techniques described above refer to adjusting the reference representation, other aspects of the trigger sound detection technique may be adjusted in the same or a similar manner in addition to, or instead of, adjusting the reference representation. For example, in some embodiments, the device adjusts how the sound input is filtered, such as by concentrating on and / or reducing certain frequencies or frequency ranges of the sound input, and / or what filters are applied to the sound input. In some embodiments, the device adjusts the algorithm used to compare the input representation and the reference representation. For example, in some embodiments, one or more terms of the mathematical function used to determine the difference between the input representation and the reference representation are changed, added, or removed, or replaced with a different mathematical function. In some embodiments, adaptation techniques such as those described above require more resources than the voice trigger system 400 can provide or is configured to provide. In particular, the sound detector may not have the amount or type of processor, data, or memory, or access thereto, necessary to perform iterative adaptation of the reference representation and / or the sound detection algorithm (or any other suitable aspect of the voice trigger system 400). Thus, in some embodiments, one or more of the adaptation techniques described above are performed by a more powerful processor, such as an application processor (e.g., processor(s) 204), or by a different device (e.g., server system 108). However, the voice trigger system 400 is designed to operate even when the application processor is in standby mode. Thus, the sound input used for adaptation of the voice trigger system 400 is received when the application processor is not available and cannot process the sound input. Accordingly, in some embodiments, the sound input is stored by the device so that it can be further processed and / or analyzed after being received. In some embodiments, the sound input is stored in the memory buffer 414 of the audio subsystem 226.In some embodiments, the audio input is stored in system memory (e.g., memory 250, FIG. 2) using direct memory access (DMA) techniques (e.g., including using a DMA engine to copy or move data without the need to wake the application processor). The stored audio input is subsequently provided to or accessed by the application processor (or server system 108, or another suitable device) such that, after startup, the application processor can execute one or more of the adaptation techniques described above. In some embodiments.

[0117] Figures 5 to 7 are flow diagrams representing a method for operating a voice trigger according to a particular embodiment. This method is optionally managed by instructions stored in a computer memory or a persistent computer-readable storage medium (e.g., the memory 250 of the client device 104, the memory 302 associated with the digital assistant system 300), and is executed by one or more processors of one or more computer systems of a digital assistant system, including but not limited to the server system 108 and / or the user device 104a. The computer-readable storage medium may include a magnetic or optical disk storage device, a solid-state storage device such as a flash memory, or other non-volatile memory device(s). The computer-readable instructions stored on the computer-readable storage medium may include one or more of source code, assembly language code, object code, or other instruction formats interpreted and executed by one or more processors. In various embodiments, some operations of each method shown in each figure may be combined, and / or the order of some operations may be changed from the order. Also, in some embodiments, the operations shown in and / or described in connection with individual figures and / or individual methods may be combined to form other methods, and the operations shown in and / or described in connection with the same figure and / or the same method may be divided into different methods. Moreover, in some embodiments, one or more operations in the method are executed by modules of a digital assistant system 300 and / or an electronic device (e.g., the user device 104), including, for example, a natural language processing module 332, a dialog flow processing module 334, an audio subsystem 226, a noise detector 402, a sound type detector 404, a trigger sound detector 406, a speech-based service 408, and / or any of their sub-modules. Figure 5 shows a method 500 for operating a voice trigger system according to some embodiments (e.g., the voice trigger system 400 of Figure 4, Figure 4). In some embodiments, the method 500 is executed on an electronic device including one or more processors and a memory storing instructions executed by one or more processors (e.g., the electronic device 104).This electronic device receives a voice input (502). This voice input may correspond to speech (e.g., words, phrases, or sentences), human pronunciation (e.g., whistling, tongue clicking, finger snapping, clapping, etc.), or any other sound (e.g., electronically generated chirping sounds, mechanical noisemakers, etc.). In some embodiments, the electronic device receives the voice input via an audio subsystem 226 (e.g., codec 410, audio DSP 412, and buffer 414, as well as microphones 230 and 418 described in connection with FIG. 4).

[0118] In some embodiments, the electronic device determines (504) whether the voice input meets a predetermined condition. In some embodiments, the electronic device applies time-domain analysis to the voice input to determine whether the voice input meets a predetermined condition. For example, the electronic device analyzes the voice input over a period of time to determine whether the sound amplitude reaches a predetermined level. In some embodiments, the threshold is met when the amplitude of the voice input (e.g., volume) meets and / or exceeds a predetermined threshold. In some embodiments, it is met when the voice input meets and / or exceeds a predetermined threshold for a predetermined time. As will be described in more detail below, in some embodiments, determining (504) whether the voice input meets a predetermined condition is performed by a third sound detector (e.g., noise detector 402). (The third sound detector is used in this case to distinguish this sound detector from other sound detectors (e.g., the first and second sound detectors described below), and does not necessarily indicate any operating position or order of the sound detectors.)

[0119] The electronic device determines whether the sound input corresponds to a predetermined type of sound (506). As described above, sounds are classified into various "types" based on specific distinguishable sound characteristics. Determining whether the sound input corresponds to a predetermined type includes determining whether the sound input includes or indicates the characteristics of a specific type. In some embodiments, the predetermined type of sound is a human voice. In such embodiments, determining whether the sound input corresponds to a human voice includes determining whether the sound input includes the frequency characteristics of a human voice (508). As will be described in more detail below, in some embodiments, determining whether the sound input corresponds to a predetermined type of sound (506) is performed by a first sound detector (e.g., sound type detector 404). When it is determined that the sound input corresponds to a predetermined type of sound, the electronic device determines whether the sound input includes predetermined content (510). In some embodiments, the predetermined content corresponds to one or more predetermined phonemes (512). In some embodiments, the one or more predetermined phonemes constitute at least one word. In some embodiments, the predetermined content is a sound (e.g., a whistle, a click, or a clap). In some embodiments, as will be described below, determining whether the sound input includes predetermined content (510) is performed by a second sound detector (e.g., trigger sound detector 406).

[0120] When it is determined that the voice input includes predetermined content, the electronic device starts a speech-based service (514). In some embodiments, the speech-based service is a voice-based digital assistant as detailed above. In some embodiments, the speech-based service is a dictation service, where the speech input is converted to text and included in and / or displayed in a text input field (such as in an email, text message, word processing, or note-taking application, etc.). In an embodiment where the speech-based service is a voice-based digital assistant, when the voice-based digital assistant is started, a prompt (such as a sound or speech prompt) is issued to the user, indicating that the user can provide voice input and / or commands to the digital assistant. In some embodiments, starting the voice-based digital assistant includes enabling an application processor (such as processor(s) 204, FIG. 2), starting one or more programs or modules (such as digital assistant client module 264, FIG. 2), and / or creating a connection to a remote server or device (such as digital assistant server 106, FIG. 1).

[0121] In some embodiments, the electronic device determines (516) whether the voice input corresponds to the voice of a particular user. For example, one or more voice authentication techniques are applied to the voice input to determine whether it corresponds to the voice of an authorized user of the device. The voice authentication techniques are described in detail above. In some embodiments, the voice authentication is performed by one of the voice detectors (e.g., trigger voice detector 406). In some embodiments, the voice authentication is performed by a dedicated voice authentication module (including any suitable hardware and / or software). In some embodiments, the voice-based service is initiated in response to determining that the voice input contains predetermined content and that the voice input corresponds to the voice of a particular user. Thus, for example, a voice-based service (e.g., a voice-based digital assistant) is initiated only when a trigger word or phrase is spoken by an authorized user. This can be particularly useful in reducing the likelihood that the service can be invoked by an unauthorized user and in preventing the utterance of a trigger voice by one user from activating the voice trigger of another user when multiple electronic devices are in proximity.

[0122] In some embodiments where the speech-based service is a voice-based digital assistant, in response to determining that the audio input contains predetermined content but does not correspond to the voice of a particular user, the voice-based digital assistant is started in a limited access mode. In some embodiments, the limited access mode enables the digital assistant to access only a subset of the data, services, and / or functions that the digital assistant would otherwise be able to provide. In some embodiments, the limited access mode corresponds to a write-only mode (e.g., such that an unauthorized user of the digital assistant cannot access data from calendars, to-do lists, contacts, photos, emails, text messages, etc.). In some embodiments, the limited access mode corresponds to a sandboxed instance of the speech-based service, such that the speech-based service does not read from or write to the user's data, such as user data 266 on device 104 (FIG. 2), or any other device (e.g., user data 348 of FIG. 3A, which may be stored on a remote server such as server system 108 of FIG. 1).

[0123] In some embodiments, in response to determining that the audio input includes predetermined content and corresponds to the voice of a specific user, the voice-based digital assistant outputs a prompt that includes the name of the specific user. For example, when a specific user is identified via voice authentication, the voice-based digital assistant may output a prompt such as "Mr. Peter, what can I do for you?" instead of a more general prompt such as a tone, beep, or non-proprietary voice prompt. As described above, in some embodiments, the first sound detector determines whether the audio input corresponds to a predetermined type of sound (at step 506), and the second sound detector determines whether the sound detector includes predetermined content (at step 510). In some embodiments, the first sound detector consumes less power during operation than the second sound detector, for example, because the first sound detector uses a technique with less processor load than the second sound detector. In some embodiments, the first sound detector is the sound type detector 404, and the second sound detector is the trigger sound detector 406, both of which are described above in connection with FIG. 4. In some embodiments, during these operations, the first sound detector and / or the second sound detector periodically monitors the audio channel according to a duty cycle, as described above in connection with FIG. 4.

[0124] In some embodiments, the first sound detector and / or the sound detector performs frequency domain analysis of the audio input. For example, these sound detectors perform a Laplace transform, Z transform, or Fourier transform to generate a frequency spectrum or determine the spectral density of the audio input or a part thereof. In some embodiments, the first sound detector is a voice activity detector configured to determine whether the audio input includes a frequency that is a characteristic of human voice (or other features, aspects, or traits of the audio input that are characteristics of human voice).

[0125] In some embodiments, the second sound detector is off or inactive until the first sound detector detects a sound input of a predetermined type. Accordingly, in some embodiments, method 500 includes activating the second sound detector in response to determining that the sound input corresponds to a predetermined type. (In other embodiments, the second sound detector is activated according to other conditions, or operates continuously regardless of the determination from the first sound detector.) In some embodiments, activating the second sound detector includes enabling hardware and / or software (including, for example, circuits, processors, programs, memories, etc.). In some embodiments, the second sound detector operates for at least a predetermined time after activation (e.g., is enabled and monitors an audio channel). For example, when the first sound detector determines that the sound input corresponds to a predetermined type (e.g., includes a human voice), the second sound detector is activated to determine whether the sound input also includes a predetermined content (e.g., a trigger word). In some embodiments, the predetermined time corresponds to the duration of the predetermined content. Thus, if the predetermined content is the phrase "Hey Siri", the predetermined time is long enough to determine whether the phrase has been spoken (e.g., 1 or 2 seconds, or any other appropriate duration). If the predetermined content is longer, such as the phrase "Hey Siri, start and help me", the predetermined time is longer (e.g., 5 seconds, or another appropriate duration). In some embodiments, the second sound detector operates as long as the first sound detector detects a sound corresponding to a predetermined type. In such embodiments, for example, as long as the first sound detector detects a human voice in the sound input, the second sound detector processes the sound input and determines whether it includes a predetermined content.

[0126] As described above, in some embodiments, a third sound detector (e.g., noise detector 402) determines whether a sound input meets a predetermined condition (at step 504). In some embodiments, the third sound detector consumes less power during operation than the first sound detector. In some embodiments, the third sound detector periodically monitors an audio channel according to a duty cycle, as described above with respect to FIG. 4. Also, in some embodiments, the third sound detector performs a time domain analysis of the sound input. In some embodiments, the time domain analysis has a lower processor load than the frequency domain analysis applied by the second sound detector, so the third sound detector consumes less power than the first sound detector.

[0127] Similar to the above description regarding activating a second sound detector (e.g., trigger sound detector 406) in response to a determination by a first sound detector (e.g., sound type detector 404), in some embodiments, the first sound detector is activated in response to a determination by a third sound detector (e.g., noise detector 402). For example, in some embodiments, the sound type detector 404 is activated in response to a determination by the noise detector 402 that a sound input meets a predetermined condition (e.g., exceeds a specific volume for a sufficient duration). In some embodiments, activating the first sound detector includes enabling hardware and / or software (e.g., including circuits, processors, programs, memories, etc.). In other embodiments, the first sound detector is activated or continuously operated in response to other conditions. In some embodiments, the device stores (518) at least a portion of the sound input in a memory. In some embodiments, the memory is the buffer 414 of the audio subsystem 226 (FIG. 4). The stored sound input enables non-real-time processing of the sound input by the device. For example, in some embodiments, one or more of the sound detectors read and / or receive the stored sound input and process this stored sound input. This can be particularly useful when an upstream sound detector (e.g., trigger sound detector 406) is not activated until partway through the receipt of the sound input by the audio subsystem 226. In some embodiments, the stored portion of the sound input is provided to a speech-based service when the speech-based service is started (520). Thus, even if the speech-based service is not fully operational until a portion of the sound input is received, the speech-based service can copy, process, or otherwise operate on the stored portion of the sound input. In some embodiments, the stored portion of the sound input is provided to an adaptation module of the electronic device.

[0128] In various embodiments, steps (516)-(520) are performed at different positions within method 500. For example, in some embodiments, one or more of steps (516)-(520) are performed between steps (502) and (504), between steps (510) and (514), or at any other suitable position.

[0129] FIG. 6 shows a method 600 for operating a voice trigger system according to some embodiments (e.g., voice trigger system 400 of FIG. 4, FIG. 4). In some embodiments, method 600 is executed on an electronic device including one or more processors and a memory storing instructions executable by the one or more processors (e.g., electronic device 104). The electronic device determines (602) whether it is in a predetermined orientation. In some embodiments, the electronic device uses an optical sensor (including a camera), a microphone, a proximity sensor, a magnetic sensor, an accelerometer, a gyroscope, a tilt sensor, etc. to detect its orientation. For example, the electronic device compares the amount or luminance of light incident on the sensor of the front camera with the amount or luminance of light incident on the sensor of the rear camera to determine whether it is placed face down or face up on a surface. If the amount and / or luminance detected by the front camera is sufficiently greater than that detected by the rear camera, the electronic device determines that it is face up. On the other hand, if the amount and / or luminance detected by the rear camera is sufficiently greater than that of the front camera, the device determines that it is face down. When the electronic device determines that it is in the predetermined orientation, the electronic device enables a predetermined mode of the voice trigger (604). In some embodiments, the predetermined orientation corresponds to the device's display screen being substantially horizontal and face down, and the predetermined mode is the standby mode. (606). For example, in some embodiments, when a smartphone or tablet is placed on a table or desk such that the screen is face down, the voice trigger enters the standby mode (e.g., power off) to prevent unintentional activation of the voice trigger.

[0130] On the one hand, in some embodiments, the predetermined orientation corresponds to the device's display screen being substantially horizontal and facing upward, and the predetermined mode is the listening mode (608). Thus, for example, when a smartphone or tablet is placed on a table or desk such that the screen faces upward, the voice trigger enters the listening mode and can respond to the user when a trigger is detected.

[0131] FIG. 7 shows a method 700 of operating a voice trigger according to some embodiments (e.g., voice trigger system 400, FIG. 4). In some embodiments, method 700 is executed on an electronic device that includes one or more processors and a memory storing instructions executed by the one or more processors (e.g., electronic device 104). The electronic device operates the voice trigger (e.g., voice trigger system 400) in a first mode (702). In some embodiments, the first mode is a normal listening mode.

[0132] The electronic device determines whether it is in a substantially enclosed space by detecting that one or more of the electronic device's microphone and camera are blocked (704). In some embodiments, substantially enclosed spaces include pockets, wallets, bags, drawers, glove boxes, briefcases, and the like.

[0133] As described above, in some embodiments, the device emits one or more sounds (e.g., tones, clicks, pings, etc.) from a speaker or transducer, monitors one or more microphones or transducers, and detects echoes of the omitted sound(s) to detect that the microphone is blocked. For example, a relatively large environment (e.g., indoors or inside a vehicle) reflects sound differently than a relatively narrow, substantially enclosed environment (e.g., a wallet or a pocket). Thus, when the device detects that the microphone (or the speaker that emitted the sound) is blocked based on the echo (or lack of echo), the device determines that it is in a substantially enclosed space. In some embodiments, the device detects that the microphone is blocked by detecting that the microphone picks up sounds characteristic of an enclosed space. For example, when the device is in a pocket, the microphone can detect a characteristic soft noise due to contact with or proximity to the fibers of the pocket. In some embodiments, the device detects that the camera is blocked based on the light reception level by a sensor or by determining whether a focused image can be obtained. For example, if the camera sensor detects a low level of light at a time when a high level of light is expected (e.g., during the day), the device determines that the camera is blocked and that the device is in a substantially enclosed space. As another example, the camera may attempt to acquire a focused image on its sensor. Typically, this becomes difficult when the camera is in a very dark place (e.g., a pocket or a backpack) or when it is too close to the subject it is trying to focus on (e.g., inside a wallet or a backpack). Thus, if the camera cannot acquire a focused image, the device determines that it is in a substantially enclosed space.

[0134] When it is determined that the electronic device is in a substantially enclosed space, the electronic device switches the voice trigger to a second mode (706). In some embodiments, the second mode is a standby mode (708). In some embodiments, when in the standby mode, the voice trigger system 400 continues to monitor ambient voices but does not respond to received sounds, regardless of whether the voice trigger system 400 would otherwise be activated. In some embodiments, in the standby mode, the voice trigger system 400 is disabled and does not process voices to detect trigger sounds. In some embodiments, the second mode includes operating one or more sound detectors of the voice trigger system 400 according to a duty cycle different from that of the first mode. In some embodiments, the second mode includes operating a different combination of sound detectors from that of the first mode.

[0135] In some embodiments, the second mode corresponds to a more sensitive monitoring mode, enabling the voice trigger system 400 to detect and respond to trigger sounds even when in a substantially enclosed space. In some embodiments, when the voice trigger switches to the second mode, the device periodically determines whether the electronic device is still in a substantially enclosed space by detecting whether one or more of the electronic device's microphones and cameras are blocked (e.g., using any of the techniques described above with respect to step (704)). If the device is still in a substantially enclosed space, the voice trigger system 400 remains in the second mode. In some embodiments, when the device is moved out of the substantially enclosed space, the electronic device returns the voice trigger to the first mode.

[0136] According to some implementations, FIG. 8 shows a functional block diagram of an electronic device 800 configured in accordance with the principles of the present invention as described above. The functional blocks of this device can be implemented by hardware, software, or a combination of hardware and software to execute the principles of the present invention. Those skilled in the art will understand that the functional blocks described in FIG. 8 can be combined or divided into sub-blocks to implement the principles of the present invention as described above. Therefore, the description herein can support any possible combination or division, or the definition of further functional blocks described herein.

[0137] As shown in FIG. 8, the electronic device 800 includes a sound receiving unit 802 configured to receive a sound input. The electronic device 800 also includes a processing unit 806 connected to the speech receiving unit 802. In some embodiments, the processing unit 806 includes a noise detection unit 808, a sound type detection unit 810, a trigger sound detection unit 812, a service start unit 814, and a voice authentication unit 816. In some embodiments, the noise detection unit 808 corresponds to the noise detector 402 described above and is configured to perform any of the above-described operations related to the noise detector 402. In some embodiments, the sound type detection unit 810 corresponds to the sound type detector 404 described above and is configured to perform any of the above-described operations related to the sound type detector 404. In some embodiments, the trigger sound detection unit 812 corresponds to the trigger sound detector 406 described above and is configured to perform any of the above-described operations related to the trigger sound detector 406. In some embodiments, the voice authentication unit 816 corresponds to the voice authentication module 428 described above and is configured to perform any of the above-described operations related to the voice authentication module 428. The processing unit 806 is configured to determine whether at least a part of the sound input corresponds to a predetermined type of sound (e.g., by the sound type detection unit 810), and when it is determined that at least a part of the sound input corresponds to a predetermined type, determine whether the sound input contains a predetermined content (e.g., by the trigger sound detection unit 812), and when it is determined that the sound input contains a predetermined content, start a speech-based service (e.g., by the service start unit 814).

[0138] In some embodiments, the processing unit 806 is also configured to determine whether the audio input meets a predetermined condition (e.g., by the noise detection unit 808) before determining whether the audio input corresponds to a predetermined type of sound. In some embodiments, the processing unit 806 is also configured to determine whether the audio input corresponds to the voice of a specific user (e.g., by the voice authentication unit 816).

[0139] According to some implementations, FIG. 9 shows a functional block diagram of an electronic device 900 configured in accordance with the principles of the present invention as described above. The functional blocks of this device can be implemented by hardware, software, or a combination of hardware and software for executing the principles of the present invention. It will be understood by those skilled in the art that the functional blocks described in FIG. 9 can be combined or divided into sub-blocks to implement the principles of the present invention as described above. Thus, the description herein supports any possible combination or division, or the definition of further functional blocks described herein.

[0140] As shown in FIG. 9, the electronic device 900 includes a voice trigger unit 902. The voice trigger unit 902 can be operated in various different modes. In the first mode, the voice trigger unit receives a voice input and determines whether a specific criterion is met (e.g., listening mode). In the second mode, the voice trigger unit 902 does not receive and / or process a voice input (e.g., standby mode). The electronic device 900 also includes a processing unit 906 connected to the voice trigger unit 902. In some embodiments, the processing unit 906 may include one or more sensors (e.g., including a microphone, a camera, an accelerometer, a gyroscope, etc.) and a mode switching unit 910 and / or may be in contact with an environment detection unit 908. In some embodiments, the processing unit 906 determines whether the electronic device is in a substantially enclosed space by detecting that one or more of the microphone and the camera of the electronic device are blocked (e.g., by the environment detection unit 908). When it is determined that the electronic device is in a substantially enclosed space, the voice trigger is configured to be switched to the second mode (e.g., by the mode switching unit 910).

[0141] In some embodiments, the processing unit determines whether the electronic device is in a predetermined orientation (e.g., by the environment detection unit 908). When it is determined that the electronic device is in a predetermined orientation, the processing unit is configured to enable a predetermined mode of the voice trigger (e.g., by the mode switching unit 910).

[0142] According to some implementations, FIG. 10 shows a functional block diagram of an electronic device 1000 configured in accordance with the principles of the present invention as described above. The functional blocks of this device can be implemented by hardware, software, or a combination of hardware and software to execute the principles of the present invention. Those skilled in the art will understand that the functional blocks described in FIG. 10 can be combined or divided into sub-blocks to implement the principles of the present invention as described above. Therefore, the description herein supports any possible combination or division, or the definition of further functional blocks described herein.

[0143] As shown in FIG. 10, the electronic device 1000 includes a voice trigger unit 1002. The voice trigger unit 1002 can be operated in various different modes. In the first mode, the voice trigger unit receives a voice input and determines whether it meets certain criteria (e.g., listening mode). In the second mode, the voice trigger unit 1002 does not receive and / or process a voice input (e.g., standby mode). The electronic device 1000 also includes a processing unit 1006 connected to the voice trigger unit 1002. In some embodiments, the processing unit 1006 may include a microphone and / or a camera, and a mode switching unit 1010, and / or may include an environment detection unit 1008 that may also serve as a contact.

[0144] The processing unit 1006 is configured to determine whether the electronic device is in a substantially enclosed space (e.g., by the environment detection unit 1008) by detecting that one or more of the microphone and camera of the electronic device are blocked. When it is determined that the electronic device is in a substantially enclosed space, the voice trigger is switched to the second mode (e.g., by the mode switching unit 1010). The above description has been made with reference to specific embodiments for illustrative purposes. However, the above exemplary description is not intended to be exhaustive or to limit the disclosed embodiments to the exact form. Many modifications and variations are possible in view of the above teachings. The embodiments have been chosen and described in order to best explain the principles of the disclosed ideas and their practical application, thereby enabling those skilled in the art to make various changes suitable for the particular uses contemplated and to best utilize them.

[0145] It should be understood that, of course, terms such as "first" and "second" can be used herein to describe various elements, but these elements should not be limited by these terms. These terms are only used to distinguish one element from another. For example, as long as the name is consistently changed for all occurrences of the "first sound detector" and the name is consistently changed for all occurrences of the "second sound detector", the first sound detector can be called the second sound detector and, similarly, the second sound detector can be called the first sound detector without changing the meaning of the description. The first sound detector and the second sound detector are both sound detectors, but they are not the same sound detector.

[0146] The terms used in this specification are for the purpose of describing particular embodiments and are not intended to limit the scope of the claims. As used in the description of the embodiments described and the appended claims, the singular forms "a", "an", and "the" are intended to include the plural forms as well, unless the context clearly indicates otherwise. When used in this specification, the term "and / or" is also to be understood to refer to and include any and all possible combinations of one or more of the associated listed items. When the terms "comprises" and / or "comprising" are used in this specification, they specify the presence of the stated features, integers, steps, operations, elements, and / or components, but do not preclude the presence or addition of one or more other features, integers, steps, operations, elements, components, and / or groups thereof. When used in this specification, the term "if" can be interpreted to mean "when" or "upon" or "in response to a determination that" or "in accordance with a determination that" or "in response to a detection that" the previously stated condition is true, depending on the context. Similarly, the phrase "[when it is determined that the previously stated condition is true]" or "[if the previously stated condition is true]" or "[when the previously stated condition is true]" can be interpreted to mean "upon a determination that" or "in the determination that" or "in response to a determination that" or "in accordance with a determination that" or "upon a detection that" or "in response to a detection that" the previously stated condition is true.

Claims

1. 1. A method for operating a voice trigger, executed on an electronic device including one or more processors and a memory storing instructions for execution by the one or more processors, comprising: receiving an audio input; determining whether at least a portion of the sound input corresponds to a predetermined type of sound; upon determining that at least a portion of the sound input corresponds to the predetermined type, determining whether the sound input includes predetermined content; initiating a speech-based service upon determining that the sound input includes the predetermined content; The method according to claim 1, further comprising:

2. 2. The method of claim 1, wherein the step of determining whether the sound input corresponds to a predetermined type of sound is performed by a first sound detector and the step of determining whether the sound input includes predetermined content is performed by a second sound detector, the first sound detector consuming less power in operation than the second sound detector.

3. 3. The method of claim 2, wherein the second sound detector is activated in response to the first sound detector determining that the sound input corresponds to the predetermined type.

4. 3. The method of claim 2, wherein the second sound detector is operated for at least a predetermined time period after the first sound detector determines that the sound input corresponds to the predetermined type.

5. 2. The method of claim 1, wherein the predetermined type is a human voice and the predetermined content is one or more words.

6. 2. The method of claim 1, wherein the predetermined content is one or more predetermined phonemes.

7. 7. The method of claim 6, wherein the one or more predetermined phonemes comprise at least one word.

8. 10. The method of claim 1, further comprising the step of determining whether the sound input satisfies a predetermined condition prior to determining whether the sound input corresponds to a predetermined type of sound.

9. The method of claim 8 , wherein the predetermined condition is an amplitude threshold.

10. 9. The method of claim 8, wherein the step of determining whether the sound input satisfies a predetermined condition is performed by a third sound detector, the third sound detector consuming less power in operation than the first sound detector.

11. storing at least a portion of the sound input in a memory; providing said portion of said sound input to said speech-based service when said speech-based service is initiated; The method of claim 1 further comprising:

12. 10. The method of claim 1, further comprising the step of determining whether the sound input corresponds to the voice of a particular user.

13. 13. The method of claim 12, wherein upon determining that the sound input includes the predetermined content and that the sound input corresponds to the voice of the particular user, the speech-based service is initiated.

14. 14. The method of claim 13, wherein upon determining that the sound input includes the predetermined content and that the sound input does not correspond to the voice of the particular user, the speech-based service is initiated in a limited access mode.

15. 14. The method of claim 13, further comprising the step of outputting a voice prompt including a name of the particular user upon determining that the sound input corresponds to the voice of the particular user.

16. determining whether the electronic device is in a predetermined orientation; upon determining that the electronic device is in the predetermined orientation, enabling a predetermined mode of the voice trigger; The method of claim 1 further comprising:

17. 1. A method for operating a voice trigger, executed on an electronic device including one or more processors and a memory storing instructions for execution by the one or more processors, comprising: operating a voice trigger in a first mode; determining whether the electronic device is within a substantially enclosed space by detecting that one or more of a microphone and a camera of the electronic device are obstructed; switching the voice trigger to a second mode upon determining that the electronic device is within a substantially enclosed space; The method according to claim 1, further comprising:

18. 20. The method of claim 17, wherein the second mode is a standby mode.

19. 18. The method according to claim 17, characterized in that the first mode is a listening mode.

20. 1. A method for operating a voice trigger, executed on an electronic device including one or more processors and a memory storing instructions for execution by the one or more processors, comprising: determining whether the electronic device is in a predetermined orientation; enabling a predetermined mode of a voice trigger upon determining that the electronic device is in the predetermined orientation; The method according to claim 1, further comprising:

21. 21. The method of claim 20, wherein the predetermined orientation corresponds to a display screen of the device being substantially horizontal and facing downwards, and the predetermined mode is a standby mode.

22. 21. The method of claim 20, wherein the predetermined orientation corresponds to a display screen of the device being substantially horizontal and facing upwards, and the predetermined mode is a listening mode.

23. 1. A computer-readable storage medium storing one or more programs for execution by one or more processors of an electronic device, the one or more programs comprising: instructions for receiving audio input; instructions for determining whether at least a portion of the sound input corresponds to a predetermined type of sound; instructions for determining if the sound input includes predetermined content upon determining that at least a portion of the sound input corresponds to the predetermined type; instructions for initiating a speech-based service upon determining that the sound input includes the predetermined content; A computer-readable storage medium comprising:

24. a sound receiving unit configured to receive a sound input; a processing unit coupled to the sound receiving unit; The electronic device comprises: determining whether at least a portion of the sound input corresponds to a predetermined type of sound; determining whether the sound input includes predetermined content upon determining that at least a portion of the sound input corresponds to the predetermined type; Initiating a speech-based service upon determining that the sound input includes the predetermined content. An electronic device characterized by being configured as follows.

25. 25. The electronic device of claim 24, wherein the processing unit is further configured to determine whether the sound input satisfies a predetermined condition prior to determining whether the sound input corresponds to a predetermined type of sound.

Citation Information

Patent Citations

  • Voice input device, voice recognition system and voice recognition method

    JP2010217754A

  • mobile communication terminal

    JP4319573B2

  • Mobile personal audio device

    US20090318198A1