Implementation for voice assistant on device

A device-agnostic audio assistant library addresses the challenge of providing a consistent voice assistant experience across multiple devices by enabling local and cloud-based processing, supporting device-specific features, and decoupling voice assistant innovation from device innovation.

JP2025081345APending Publication Date: 2025-05-27GOOGLE LLC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2025014013
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2016-05-13
Filing Date
2025-01-30
Publication Date
2025-05-27

AI Technical Summary

Technical Problem

Existing voice assistants lack the ability to provide a consistent user experience across multiple devices while supporting device-specific features and allowing for innovation in voice assistant functionality to be decoupled from device innovation.

Method used

A device-agnostic audio assistant library that enables voice assistant functionality across various devices, allowing for local processing of audio data, connectivity to the cloud, and asynchronous updates, while maintaining a consistent user experience and supporting device-specific features.

Benefits of technology

The solution provides a consistent voice assistant experience across diverse devices, decouples innovation in voice assistant functionality from device innovation, and allows for seamless integration of device-specific features, enhancing user interaction and device control.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025081345000001_ABST
    Figure 2025081345000001_ABST
Patent Text Reader

Abstract

To provide implementation for a voice assistant on a device.SOLUTION: A method at an electronic device with an audio input system includes: receiving a verbal input at the device; processing the verbal input; transmitting a request to a remote system, the request including information determined based on the verbal input; receiving a response to the request, in which the response is generated by the remote system in accordance with the information based on the verbal input; and performing an operation in accordance with the response. One or more of the receiving, processing, transmitting, receiving, and performing are performed by one or more voice processing modules of a voice assistant library being executed on the electronic device.SELECTED DRAWING: Figure 5
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Technical Field The present application generally relates to computer technology including, but not limited to, voice assistants for devices and related libraries.

Background Art

[0002] Background Along with the development of the Internet and cloud computing, the popularity of voice-based assistants that interact with users through audio / voice input / output is increasing. These assistants provide an interface for consuming digital media and can provide various types of information such as, for example, news, sports scores, weather, and stocks.

[0003] Users will likely have multiple devices with desirable voice-based assistant functionality. It is desirable to have a voice-based assistant that can be implemented and used across various devices, provide a consistent experience across these various devices, and support device-specific features.

Summary of the Invention

Means for Solving the Problems

[0004] Summary The implementations described herein are directed to incorporating or including a voice assistant in an embedded system and / or device in a way that enables control of local devices for a variety of operating system platforms.

[0005] According to some implementations, the device-side library, which is thin and has low resource usage, has features including local processing of audio data, listening for wake words or hot words, and sending user requests. Further features include connectivity to the cloud brain, an extensible voice operation control system, a portability layer that enables integration into many diverse operating environments, and the ability to be updated asynchronously with the rest of the client software.

[0006] The described implementations have the advantage of providing a similar user experience for interacting with a voice assistant across many different devices.

[0007] The described implementations also have the additional advantage of decoupling innovation in the voice assistant functionality from the innovation available on the device itself. For example, if an improved recognition pipeline is created, while the recognition results are output to the device, the device manufacturer can continue to benefit from previous voice commands without having to do anything to receive the recognition results.

[0008] According to some implementations, a method in an electronic device having an audio input system, one or more processors, and a memory storing one or more programs executed by the one or more processors includes receiving an oral input at the device, processing the oral input, sending a request including information determined based on the oral input to a remote system, receiving a response to the request generated by the remote system in response to the information based on the oral input, and performing an operation in response to the response. One or more of the steps of receiving, processing, transmitting, receiving, and executing are performed by one or more audio processing modules of an audio assistant library running on an electronic device, and the audio processing module provides a plurality of audio processing operations accessible to one or more application programs and / or operating software running or executable on the electronic device.

[0009] In some implementations, a device-agnostic audio assistant library for an electronic device with an audio input system includes one or more audio processing modules configured to operate on a common operating system implemented on multiple different electronic device types, and the audio processing module provides a plurality of audio processing operations accessible to application programs and operating software running on the electronic device, thereby enabling the portability of an audio-responsive application configured to interact with one or more of the audio processing operations.

[0010] In some implementations, the electronic device includes an audio input system, one or more processors, and a memory storing one or more programs executed by the one or more processors. The one or more programs include instructions for receiving verbal input at the device, instructions for processing the verbal input, instructions for transmitting a request including information determined based on the verbal input to a remote system, instructions for receiving a response to the request generated by the remote system in response to the information based on the verbal input, and instructions for performing an operation in response to the response, and one or more of receiving, processing, transmitting, receiving, and executing are performed by one or more audio processing modules of an audio assistant library running on the electronic device, and the audio processing module provides a plurality of audio processing operations accessible to one or more application programs and / or operating software running or executable on the electronic device.

[0011] In some implementations, the non-transitory computer-readable storage medium stores one or more programs. The one or more programs include instructions that, when executed by an electronic device having an audio input system and one or more processors, cause the electronic device to receive a verbal input at the device, process the verbal input, send a request including information determined based on the verbal input to a remote system, receive a response to the request generated by the remote system in response to the information based on the verbal input, perform an operation according to the response, and one or more of receiving, processing, sending, receiving, and performing are executed by one or more audio processing modules of an audio assistant library running on the electronic device, and the audio processing modules provide a plurality of audio processing operations accessible to one or more application programs and / or operating software running or executable on the electronic device. BRIEF DESCRIPTION OF THE DRAWINGS

[0012]

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5

[0013] Throughout the drawings, the same reference numerals refer to corresponding parts. Description of Implementations Here, various implementation forms will be described in detail. Examples of these implementation forms are shown in the accompanying drawings. In the following detailed description, many specific details will be described in order to provide a thorough understanding of the present invention and the described implementation forms. However, the present invention can be implemented even without these specific details. In other cases, well-known methods, procedures, components, and circuits will not be described in detail so as not to obscure the aspects of the implementation forms unnecessarily.

[0014] In some implementation forms, the purpose of the voice assistant is to provide the user with a personalized voice interface that is available across various devices and enables a wide variety of use cases, providing a consistent experience throughout the user's day. The voice assistant and / or related functions may be integrated into first-party products and devices, as well as third-party products and devices.

[0015] Use case examples include media. Voice commands can be used to initiate the playback and control of music, radio, podcasts, news, and other audio media by voice. For example, the user can issue voice commands such as (e.g., "Play jazz music", "Play FM107.5", "Skip to the next song", "Play 'continuously'") to play or control various types of audio media. Furthermore, such commands can be used to play audio media from various sources such as over-the-air radio broadcast station online streaming, music subscription services, local storage, remote storage, etc. Additionally, the voice assistant may support additional content by leveraging integrations that can be used with casting devices.

[0016] Another use case example includes remote playback. The user may issue a voice command to a casting device that includes a voice assistant function, and in response to the voice command, media is played (e.g., cast) on a device included in the group of one or more specified devices, on a device specified in the command, or on one or more devices in the area specified in the command. Also, the user can specify a general category or specific content in the command, and appropriate media is played according to the category or content specified in the command.

[0017] Yet another use case example includes non-media such as productivity-enhancing functions (e.g., timer, alarm clock, calendar), home automation, question and answer leveraging search engine technology (e.g., search query), entertainment (e.g., assistant personality, jokes, games, Easter eggs), and daily tasks (e.g., means of transportation, travel, food, finance, gifts, etc.).

[0018] In some implementations, the voice assistant is provided as an optional function of the casting device, and the voice assistant function may be updated as part of the casting device.

[0019] In some implementations, the detection of hotwords or keywords included in voice commands and the user's verbal input is performed by an application processor (e.g., in the client device or casting device to which the user speaks the voice command or verbal input). In some implementations, the detection of hotwords is performed by an external digital signal processor (in contrast to the client device or casting device to which the user speaks the voice command or verbal input, for example, by the server system processing the voice command). is performed by the server system processing the voice command).

[0020] In some implementations, a device having a voice assistant function includes one or more of remote field support, "push-to-assist" or "push-to-talk" (e.g., a button to initiate the voice assistant function), and AC power.

[0021] In some implementations, the voice assistant includes an audio input device (e.g., a microphone, an ongoing playback media loopback), the state of the microphone (e.g., on / off), ducking (e.g., lowering the volume of all outputs when the assistant is triggered (triggered) by a hotword or push-to-talk) ), and an application programming interface (API) for one or more of new assistant events and status messages (e.g., the assistant has been activated (e.g., heard a hotword, the assistant button has been pressed), listening to voice, waiting on the server, responding, response has ended, an alarm / timer is ringing).

[0022] In some implementations, a device having a voice assistant function may communicate with another device (e.g., a settings application on a smartphone) for configuration purposes to enable or facilitate the voice assistant function on the device (e.g., set up the voice assistant function on the device, provide a tutorial to the user). The configuration or setup may include specifying the location of the device, associating with a user account, opting in to voice control by the user, linking to media services (e.g., video streaming services, music streaming services) and prioritizing media services, home automation settings, and the like.

[0023] In some implementations, a device having a voice assistant may include one or more user interface elements or displays to the user. One or more of the user interface elements are physical elements (e.g., a pattern of light displayed using one or more LEDs, a sound pattern output by a speaker), “push-to-assist” or “push-to-talk” triggers that are not influenced by hotwords, “mute microphone” triggers and visual status displays, a visual display of “hotword waiting status”, a visual display of “detecting hotword”, a visual display of “the assistant is actively listening” that can be visually recognized from a short distance (e.g., 15 feet), a visual display of “the assistant is working / thinking”, a visual display of “there is a voice message / notification”, a control method and status indicator of “volume level”, and one or more of a “pause / resume” control method. In some implementations, these physical user interface elements are provided by a client device or a casting device. In some implementations, the voice assistant supports a common set of user interface elements or displays across different devices so that the experience is consistent across devices with different experiences.

[0024] In some implementations, the voice assistant supports device-specific commands and / or hotwords, as well as a defined standard set of commands and / or hotwords.

[0025] FIG. 1 is a diagram showing a network environment 100 according to some implementations. The network environment 100 includes a casting device 106 and / or a voice assistant client device 104. The casting device 106 (e.g., CHROMECAST by GOOGLE INC.) includes an audio input device 108 (e.g., , is directly or communicably connected to a microphone and an audio output device 110 (e.g., one or more speakers). In some implementations, the audio input device 108 and the audio output device 110 are components of a device (e.g., a speaker system, a television, a sound bar) communicably connected to the casting device 106. In some implementations, the audio input device 108 is a component of the casting device 106, the audio output device 110 is a component of a device communicably connected to the casting device 106, or the audio output device 110 is a component of the casting device 106 and the audio input device 108 is a component of a device communicably connected to the casting device 106. In some implementations, the audio input device 108 and the audio output device 110 are components of the casting device 106.

[0026] In some implementations, the casting device 106 is communicably connected to the client 102. The client 102 may include an application or module (e.g., a casting device settings app) that facilitates the configuration of the casting device 106, including a voice assistant function.

[0027] In some implementations, the casting device 106 is connected to the display 144.

[0028] In some implementations, the casting device 106 includes one or more visual indicators 142 (e.g., LED lights).

[0029] In some implementations, the casting device 106 includes a receiving module 146. In some implementations, the receiving module 146 operates the casting device 106 and operates on, for example, hardware functions and communication with content sources. In some implementations, in the casting device 106, there are different receiving modules 146 for different content sources. In some implementations, the receiving module 146 includes sub-modules for different content sources, respectively.

[0030] The voice assistant client device 104 (e.g., a smartphone, laptop or desktop computer, tablet computer, voice command device, mobile device, or in-vehicle system having GOOGLE ASSISTANT by GOOGLE INC., GOOGLE HOME by GOOGLE INC.) includes an audio input device 132 (e.g., a microphone) and an audio output device 134 (e.g., one or more speakers, headphones). In some implementations, the voice assistant client device 104 (e.g., a voice command device, mobile device, or in-vehicle system having GOOGLE ASSISTANT by GOOGLE INC., GOOGLE HOME by GOOGLE INC.) is communicatively connected to a client 140 (e.g., a smartphone, tablet device). The client 140 may include an application or module (e.g., a voice command device setting application) that facilitates the setting of the voice assistant client device 104, including a voice assistant function.

[0031] In some implementations, the voice assistant client device 104 includes one or more visual indicators 152 (e.g., LED lights). An example of a voice assistant client device having a visual indicator (e.g., an LED light) is shown in FIG. 4A of U.S. Provisional Application No. 62 / 336,566, filed May 13, 2016, entitled "LED Design Language for Visual Affordance of Voice User Interfaces" (incorporated herein by reference). which is incorporated herein by reference and shown in FIG. 4A of U.S. Provisional Application No. 62 / 336,566, filed May 13, 2016, entitled "LED Design Language for Visual Affordance of Voice User Interfaces".

[0032] The casting device 106 and the voice assistant client device 104 each include an instance of a voice assistant module or library 136. The voice assistant module / library 136 is a module / library that implements a voice assistant function across various devices (e.g., the casting device 106, the voice assistant client device 104). The voice assistant function is consistent across various devices while continuing to permit device-specific features (e.g., support for controlling device-specific features by the voice assistant). In some implementations, the voice assistant module / library 136 is the same or similar across devices, and instances of the same library can be included in various devices.

[0033] In some implementations, depending on the type of device, the voice assistant module / library 136 is included in an application installed on the device or in the device's operating system, or is embedded in the device (e.g., embedded in firmware).

[0034] In some implementations, the voice assistant module / library 136-1 in the casting device 106 communicates with the receiving module 146 to perform voice assistant operations.

[0035] In some implementations, the voice assistant module / library 136-1 in the casting device 106 can control or affect the visual indicator 142.

[0036] In some implementations, the voice assistant module / library 136-2 in the voice assistant client device 104 can control or affect the visual indicator 152.

[0037] The casting device 106 and the voice assistant client device 104 are communicatively connected to the server system 114 through one or more communication networks 112 (e.g., local area network, wide area network, Internet). The voice assistant module / library 136 detects (e.g., receives) the verbal input picked up (e.g., captured) by the audio input device 108 / 132, processes the verbal input (e.g., to detect a hotword), and transmits the processed verbal input or the encoded processed verbal input to the server 114. The server 114 receives the processed verbal input or its encoded version, processes the received verbal input, and determines an appropriate response to the verbal input. The appropriate response may be content, information, or an instruction, command, or metadata for the casting device 106 or the voice assistant client device 104 to cause a function or operation to be executed on the casting device 106 or the voice assistant client device 104. The server 114 sends the response to the casting device 106 or the voice assistant client device 104 where the content or information is output (e.g., output from the audio output device 110 / 134) and / or the function is executed. As part of the processing, the server 114 may communicate with one or more content / information sources 138 and obtain or refer to content or information for the response. In some implementations, the content / information source 138 may be a search engine, a database, information associated with the user's account (For example, calendars, task lists, email), websites, and media streaming services, etc. may be mentioned. In some implementation forms, the voice assistant client device 104 and the casting device 106 may communicate or interact with each other. Examples of such communication or interaction, and examples of the operation of the voice assistant client device 104 (for example, GOOGLE HOME by GOOGLE INC.) were filed on May 13, 2016, and titled "LED Design Language for Visual Affordance of Voice User Interfaces (Design Language for Visual Affordance of Voice User Interfaces)", U.S. Provisional Application No. 62 / 336,566, filed on May 13, 2016, and titled "Voice-Controlled Closed Caption Display (Closed Caption Display Controlled by Voice)", U.S. Provisional Application No. 62 / 336,569, and filed on May 13, 2016, and titled "Media Transfer among Media Output Devices (Media Transfer between Media Output Devices)", U.S. Provisional Application No. 62 / 336,565. All of these applications are incorporated herein by reference. In some implementation forms, the voice assistant module / library 136 receives the verbal input captured by the audio input device 108 / 132 and transmits the verbal input (without or with little processing) or the encoded version thereof to the server 114. The server 114 processes the verbal input, detects the hot word, determines an appropriate response, and sends this response to the casting device 106 or the voice assistant client device 104.

[0038]

[0039] If the server 114 determines that the verbal input includes a command for the casting device 106 or the voice assistant client device 104 to execute a function, the server 114 transmits a response including an instruction or metadata instructing the casting device 106 or the voice assistant client device 104 to execute the function. The function may be specific to the device, and a function for supporting such a function in the voice assistant may be included in the casting device 106 or the client 104 as a custom module or function added to or linked to the voice assistant module / library 136.

[0040] In some implementations, the server 114 includes or is connected to a voice processing backend 148 that performs the processing operations of the verbal input and determines the response to the verbal input.

[0041] In some implementations, the server 114 includes a downloadable voice assistant library 150. The downloadable voice assistant library 150 (e.g., the same as or an updated version of the voice assistant library 136) may include new features or functions, or updates, and can be downloaded to add a voice assistant library to the device or update the voice assistant library 136.

[0042] FIG. 2 is a block diagram showing an example of a voice assistant client device 104 or a casting device 106 of the network environment 100 according to some implementation forms. Examples of the voice assistant client device 104 include a mobile phone, a tablet computer, a laptop computer, a desktop computer, a wireless speaker (e.g., GOOGLE HOME by GOOGLE INC.), a voice command device (e.g., GOOGLE HOME by GOOGLE INC.), a television, a sound bar, a casting device (e.g., CHROMECAST by GOOGLE INC.), a media streaming device, a home appliance, a household electronic device, an in-vehicle Examples include, but are not limited to, systems and wearable personal devices. A voice assistant client device 104 (e.g., GOOGLE HOME by GOOGLE INC., a mobile device with the GOOGLE ASSISTANT function) or a casting device 106 (e.g., CHROMECAST by GOOGLE INC.) typically includes one or more processing units (CPUs) 202, one or more network interfaces 204, a memory 206, and one or more communication buses 208 (sometimes called a chipset) for connecting these components to each other. The voice assistant client device 104 or the casting device 106 includes one or more input devices 210 to facilitate user input. The one or more input devices 210 include an audio input device 108 or 132 (e.g., a voice command input section or a microphone), and optionally other input devices such as a keyboard, a mouse, a touch screen display, a touch input pad, a gesture capture camera, or other input buttons or controls. In some implementations, the voice assistant client device 102 uses a microphone and voice recognition, or a camera and gesture recognition, to assist or replace a keyboard. Also, the voice assistant client device 104 or the casting device 106 includes one or more output devices 212. The one or more output devices 212 include an audio output device 110 or 134 (e.g., one or more speakers, headphones, etc.), and optionally one or more display devices (e.g., display 144) and / or one or more visual indicators 142 or 152 (e.g., LEDs) that enable the presentation of a user interface and display content and information. Optionally, the voice assistant client device 104 or the casting device 106 includes a position detection unit 214, such as a GPS (Global Positioning Satellite) or other geolocation receiver, to identify the location of the voice assistant client device 104 or the casting device 106.In addition, the voice assistant client device 104 or the casting device 106 may optionally include a proximity detection device 215, such as an IR sensor, for determining the proximity of the voice assistant client device 104 or the casting device 106 to other objects (e.g., the user / wearer in the case of a wearable personal device). Optionally, the voice assistant client device 104 or the casting device 106 includes one or more sensors 213 (e.g., an accelerometer, a gyroscope, etc.).

[0043] The memory 206 includes a high-speed random access memory such as DRAM, SRAM, DDR RAM, or other random access solid state storage devices, and, if necessary, non-volatile memory such as one or more magnetic disk storage devices, one or more optical disk storage devices, one or more flash memory devices, or one or more other non-volatile solid state storage devices. The memory 206 includes, if necessary, one or more storage devices located remotely from one or more processing devices 202. The memory 206, or the non-volatile memory within the memory 206, includes a non-transitory computer-readable storage medium. In some implementations, the memory 206, or the non-transitory computer-readable storage medium of the memory 206, stores the following programs, modules, and data structures, or subsets or supersets thereof.

[0044] ● An operating system 216 that includes procedures for handling various basic system services and for performing hardware-dependent tasks.

[0045] ● One or more network interfaces 204 (wired or wireless), and the voice assistant client device 104 or the casting device 106 via one or more networks 112 such as the Internet, other wide area networks, local area networks, metropolitan area networks, etc. to other A network communication module 218 for connecting to a device (e.g., server system 114, client 102, 140, other voice assistant client devices 104 or casting device 106).

[0046] ●A user interface module 220 for enabling the presentation of information on a voice assistant client device 104 or casting device 106 via one or more output devices 212 (e.g., a display, speakers, etc.).

[0047] ●An input processing module 222 for processing one or more user inputs or the dialogues captured or received by one or more input devices 210 and interpreting the said inputs and dialogues.

[0048] ●A voice assistant module 136 for processing oral inputs, providing the said oral inputs to the server 114, receiving responses from the server 114, and outputting the said responses.

[0049] ●Client data 226 for storing data associated with at least the voice assistant module 136, including the following.

[0050] ○Voice assistant settings 228 for storing the settings and configurations of the voice assistant module 136 and information associated with the voice assistant function.

[0051] ○Content source / information source 230 and content category / information category 232 for storing defined and / or user-specified sources and categories of content or information.

[0052] ○Usage history 234 for storing information (e.g., logs) related to the operations and uses of the voice assistant module 136, such as received commands and requests, responses to commands and requests, and actions taken in response to commands and requests.

[0053] ○ User account and authorization 236 for storing authorization and authentication information of one or more users for accessing each user's account in the content source / information source 230 and the account information of these authorized accounts.

[0054] ○ Receiving module 146 for operating the casting function of the casting device 106, including communicating with the content source.

[0055] In some implementations, the voice assistant client device 104 or the casting device 106 includes one or more libraries and one or more application programming interfaces (APIs) for the voice assistant and related functions. These libraries may be included in the voice assistant module 136 or the receiving module 146, or may be linked to each other by the voice assistant module 136 or the receiving module 146. The libraries include modules associated with the voice assistant function or other functions that facilitate the voice assistant function. The API provides an interface to the hardware and other software (e.g., operating system, other applications) that facilitate the voice assistant function. For example, the voice assistant client library 240, the debugging library 242, the platform API 244, and the POSIX API 246 may be stored in the memory 206. These libraries and APIs will be described in more detail below with reference to FIG. 4.

[0056] In some implementations, the voice assistant client device 104 or the cast The ting device 106 includes a voice application 250 that utilizes the modules and functions of the voice assistant client library 240, and optionally includes a debugging library 242, a platform API 244, and a POSIX API 246. In some implementations, the voice application 250 is a first-party or third-party application that becomes voice-enabled by using the voice assistant client library 240, and so on.

[0057] Each of the above elements may be stored in one or more of the aforementioned storage devices and corresponds to an instruction set for executing the above functions. The above modules or programs (i.e., instruction sets) do not need to be implemented as separate software programs, procedures, modules, or data structures, so various subsets of these modules may be combined or rearranged in various implementations. In some implementations, the memory 206 stores a subset of the above modules and data structures if necessary. Further, the memory 206 stores additional modules and data structures not described above if necessary.

[0058] FIG. 3 is a block diagram showing an example of the server system 114 of the network environment 100 according to some implementation forms. The server 114 typically includes one or more processing devices (CPUs) 302, one or more network interfaces 304, a memory 306, and one or more communication buses 308 for connecting these components (which may also be referred to as a chipset) to each other. The server 114 optionally includes one or more input devices 310 that facilitate user input, such as a keyboard, a mouse, a voice command input unit or a microphone, a touch screen display, a touch input pad, a gesture capture camera, or other input buttons or control units. Additionally, the server 114 may use a microphone and voice recognition, or a camera and gesture recognition to assist or replace the keyboard. In some implementation forms, the server 114 optionally includes one or more cameras, scanners, or optical sensor units for photographing, for example, a series of graphic codes printed on an electronic device. Further, the server 114 optionally includes one or more output devices 312 that include one or more speakers and / or one or more display devices that enable the presentation of a user interface and display content.

[0059] The memory 306 includes high-speed random access memories such as DRAM, SRAM, DDR RAM, or other random access solid-state storage devices, and, if necessary, non-volatile memories such as one or more magnetic disk storage devices, one or more optical disk storage devices, one or more flash memory devices, or one or more other non-volatile solid-state storage devices. The memory 306 includes, if necessary, one or more storage devices located remotely from one or more processing devices 302. The memory 306, or the non-volatile memory within the memory 306, includes a non-transitory computer-readable storage medium. In some implementation forms, the memory 306, or the non-transitory computer-readable storage medium of the memory 306, stores the following programs, modules, and data structures, or subsets or supersets thereof.

[0060] ● An operating system 316 including procedures for processing various basic system services and for performing hardware-dependent tasks.

[0061] ● One or more processing devices 304 (wired or wireless), and a network communication module 318 for connecting the server system 114 to other devices (e.g., voice assistant client device 104, casting device 106, client 102, client 140) via one or more networks 112 such as the Internet, other wide area networks, local area networks, metropolitan area networks. ● A proximity / position identification module 320 for identifying the proximity and / or position of the voice assistant client device 104 or casting device 106 based on the location information of the client device 104 or casting device 106.

[0062] ● A voice assistant backend 116 for processing voice inputs (e.g., voice inputs received from the voice assistant client device 104 and casting device 106) including at least one or more of the following.

[0063] ○ A voice input processing module 324 for processing the voice input and identifying the commands and requests included in the voice input.

[0064] ○ A content / information collection module 326 for collecting content and information responses to the commands and requests.

[0065] ○ A response generation module 328 for generating a voice output in response to the commands and requests and adding the voice output together with the content and information that is the response.

[0066] ○ A response generation module 328 for generating a voice output in response to the commands and requests and adding the voice output together with the content and information that is the response.

[0067] ● A server system data 330 that stores data related to the operation of at least a voice assistant platform, including the following.

[0068] ○ User data 332 for storing information associated with a user of the voice assistant platform, including the following.

[0069] - User voice assistant settings 334 for storing voice assistant setting information corresponding to the voice assistant settings 228 and information corresponding to the content source / information source 230 and the content category / information category 232.

[0070] - User history 336 for storing the user's history (e.g., log) about the voice assistant, including the history of commands and requests and the corresponding responses.

[0071] - User account and authorization 338 for storing the user's authorization and authentication information for accessing each of the user's accounts in the content source / information source 230 and the account information of these authorized accounts corresponding to the user account and authorization 236.

[0072] Each of the above elements may be stored in one or more of the aforementioned storage devices and corresponds to a set of instructions for executing the above functions. The above modules or programs (i.e., sets of instructions) do not need to be implemented as separate software programs, procedures, modules, or data structures, so various subsets of these modules may be combined or rearranged in various implementation forms. In some implementation forms, the memory 206 stores a subset of the above modules and data structures if necessary. Further, the memory 206 stores additional modules and data structures not described above if necessary.

[0073] In some implementations, the voice assistant module 136 (FIG. 2) includes one or more libraries. Each library includes a module or sub-module that executes a respective function. For example, the voice assistant client library includes a module that executes functions of the voice assistant and. Also, the voice assistant module 136 may include one or more application programming interfaces (APIs) for cooperating with specific hardware (e.g., hardware on a client device or a casting device), specific operating software, or a remote system.

[0074] In some implementations, the library includes a module that supports audio signal processing operations, including, for example, bandpass processing, filtering processing, cancellation processing, and hotword detection. In some implementations, the library includes a module for connecting to a backend (e.g., server-based) voice processing system. In some imple mentations, the library includes a module for debugging (e.g., debugging of speech recognition, debugging of hardware problems, automated testing).

[0075] FIG. 4 is a diagram showing libraries and APIs that can be stored in the voice assistant client device 104 or the casting device 106 and can be executed by the voice assistant module 136 or another application. The libraries and APIs may include a voice assistant client library 240, a debugging library 242, a platform API 244, and a POSIX API 246. An application (e.g., the voice assistant module 136, other applications that may want to support collaboration with the voice assistant) in the voice assistant client device 104 or the casting device 106 may include or be linked to these libraries and APIs and execute these libraries and APIs to provide or support voice assistant functionality in the application. In some implementations, the voice assistant client library 240 and the debugging library 242 are separate libraries. Separating the voice assistant client library 240 and the debugging library 242 separately facilitates different release and update procedures taking into account the different security impacts of these libraries.

[0076] In some implementations, these libraries are flexible. The libraries may be used across multiple device types and may incorporate the same voice assistant functionality.

[0077] In some implementations, the libraries are compatible with different operating systems or platforms that utilize these standard shared objects (e.g., various Linux distributions and flavors of embedded Linux) because the libraries rely on standard shared objects (e.g., standard Linux® shared objects).

[0078] In some implementations, the POSIX API 246 provides a standard API for compatibility with various operating systems. Thus, the voice assistant client library 240 may be included in devices of different operating systems that conform to POSIX, and the POSIX API 246 provides a compatibility interface between the voice assistant client library 240 and different operating systems.

[0079] In some implementations, the library includes modules to support and facilitate base use cases that are available across different types of devices (e.g., timers, alarms, volume controls) that implement a voice assistant.

[0080] In some implementations, the voice assistant client library 240 includes a controller interface 402 that includes functions or modules for starting, configuring, and interacting with the voice assistant. In some implementations, the controller interface 402 includes a "Start()" function or module 404 for starting the voice assistant on the device, a "RegisterAction()" function or module 406 for registering actions with the voice assistant (e.g., so that actions can be made executable via the voice assistant), a "Reconfigure()" function 408 for reconfiguring the voice assistant with updated settings, and a "RegisterEventObserver()" function 410 for registering a set of functions for basic events with the assistant.

[0081] In some implementations, the voice assistant client library 240 includes a plurality of functions or modules associated with specific voice assistant functions. For example, the hotword detection module 412 processes voice input to detect hotwords. The voice processing module 414 processes the voice included in the voice input and converts the voice to text or converts text to voice (e.g., identification of words and expressions, conversion from voice to text data, conversion from text data to voice). The action processing module 416 performs actions and operations in response to verbal input. The local timer / alarm / volume adjustment module 418 facilitates the alarm clock, timer, and volume adjustment functions in the device, as well as their control by voice input (e.g., manages the timer, clock, and alarm clock in the device). The logging / evaluation metric module 420 records voice input and responses (e.g., takes logs) and determines and records relevant evaluation metrics (e.g., response time, idle time, etc.). The audio input processing module 422 processes the audio of the voice input. The MP3 decoding module 424 decodes audio encoded in MP3. The audio input module 426 captures audio from an audio input device (e.g., a microphone). The audio output module 428 outputs audio from an audio output device (e.g., a speaker). An event queuing / status tracking module 430 for queuing events associated with the voice assistant in the device and tracking the state of the voice assistant in the device.

[0082] In some implementations, the debugging library 242 provides modules and functions for debugging. For example, the HTTP server module 432 facilitates debugging of connectivity issues, and the debug server / audio streaming module 434 debugs audio issues.

[0083] In some implementations, the platform API 244 provides an interface between the voice assistant client library 240 and the hardware capabilities of the device. For example, the platform API includes a button input interface 436 for capturing button inputs to the device, a loopback audio interface 438 for capturing loopback audio, a logging / metric interface 440 for logging and judging metrics, an audio input interface 442 for capturing audio inputs, an audio output interface 444 for outputting audio, and an authentication interface 446 for authenticating the user using other services that can interact with the voice assistant. The advantage of the voice assistant client library composition shown in FIG. 4 is that it can provide the same or similar voice processing functions across various voice assistant device types having a consistent API and a set of functions for the voice assistant. This consistency supports the portability of voice assistant applications and the consistency of voice assistant operations, facilitating familiarity with consistent user interactions as well as voice assistant applications and functions operating on different device types. In some implementations, the voice All or part of the voice assistant client library 240 may be provided at the server 114 to support a server-based voice assistant application (e.g., an operating server application for voice inputs sent to the server 114 for processing).

[0084] Code examples of classes and functions corresponding to the controller 402 ("Controller") and related classes are shown below. These classes and functions can be adopted by applications executable on various devices via a common API.

[0085] The following class "ActionModule" facilitates an application to register the module of the application in order to process commands provided by a voice assistant server.

[0086]

Number

[0087] The following class "BuildInfo" may be used to describe the application in which the voice assistant client library 240 is running or the voice assistant client device 104 itself (for example, using an identifier or version number of the application, platform, and / or device).

[0088]

Number

[0089] The following class "EventDelegate" defines functions associated with basic events such as the start of speech recognition, the start and completion of the output of the response of the voice assistant.

[0090]

Number

[0091] The following class "DefaultEventDelegate" defines an overridden function that does nothing for a specific event.

[0092]

Number

[0093] The following class "Settings" defines settings (e.g., locale, geographical location, file system directory) that can be provided to the controller 402.

[0094]

Number

[0095] The following class "Controller" corresponds to the controller 402, and the functions Start(), Reconfigure(), RegisterAction(), and RegisterEventObserver() correspond to the functions Start() 404, Reconfigure() 408, RegisterAction() 406, and RegisterEventObserver() 410, respectively.

[0096]

Number

[0097] In some implementations, the voice assistant client device 104 or the casting device 106 implements the platform (e.g., a set of interfaces for communicating with other devices using the same platform, and an operating system configured to support the set of interfaces). The following code example shows functions associated with an interface for the voice assistant client library 402 to interact with the platform.

[0098] The following class "Authentication" defines an authentication token for authenticating a user of a voice assistant having a specific account.

[0099]

Number

[0100] The following class "OutputStreamType" defines the type of audio output stream.

[0101]

Number

[0102] The following class "SampleFormat" defines the format of the supported audio samples (e.g., PCM format).

[0103]

Number

[0104] The following "BufferFormat" defines the format of the data stored in the audio buffer of the device.

[0105]

Number

[0106] The following class "AudioBuffer" defines a buffer for audio data.

[0107]

Number

[0108] The following class "AudioOutput" defines an interface for audio output.

[0109]

Number

[0110] The following class "AudioInput" defines an interface for capturing audio input.

[0111] [Number]

[0112] The following class "Resources" defines access to system resources.

[0113] [Number]

[0114] The following class "PlatformApi" specifies a platform API for the voice assistant client library 240 (for example, platform API 244).

[0115] [Number]

[0116] In some implementations, volume adjustment may be processed outside the voice assistant client library 240. For example, the system volume may be managed by a device that is not controlled by the voice assistant client library 240. As another example, the voice assistant client library 240 may continue to support volume adjustment, but requests for volume adjustment with respect to the voice assistant client library 240 are directed to the device.

[0117] In some implementations, the alarm and timer functions included in the voice assistant client library 240 may be disabled by the user, or may be disabled when implementing the library on the device.

[0118] Also, in some implementations, the voice assistant client library 240 supports an interface to the LEDs on the device and facilitates the display of LED animations on the device's LEDs.

[0119] In some implementations, the voice assistant client library 240 may be included in or linked to a casting reception module (e.g., reception module 146) in the casting device 106. The link between the voice assistant client library 240 and the reception module 146 may include, for example, support for further actions (e.g., local media playback) and support for controlling the LEDs on the casting device 106.

[0120] FIG. 5 is a flowchart of a method 500 for processing verbal input on a device, according to some implementations. The method 500 includes an audio input system (e.g., audio input device 108 / 132), one or more processors (e.g., processing device(s) 202), and a memory (e.g., memory 206) storing one or more programs to be executed by the one or more processors, in an electronic device (e.g., voice assistant cl It is executed in the client device 104 and the casting device 106. In some implementations, the electronic device includes an audio input system (e.g., the audio input device 108 / 132), one or more processors (e.g., the processing device(s) 202), and a memory (e.g., the memory 206) storing one or more programs executed by the one or more processors, and the one or more programs include instructions for executing the method 500. In some implementations, a non-transitory computer-readable storage medium includes the one or more programs, and the one or more programs include instructions that, when executed by an electronic device having an audio input system (e.g., the audio input device 108 / 132) and one or more processors (e.g., the processing device(s) 202), cause the electronic device to execute the method 500. The program or instructions for executing the method 500 may be included in the modules, libraries, etc. described above with reference to FIGS. 2-4.

[0121] The device receives a verbal input at the device (502). The client device 104 / casting device 106 captures the verbal input (e.g., voice input) uttered by the user.

[0122] The device processes the verbal input (504). The client device 104 / casting device 106 processes the verbal input. The processing may include hotword detection, conversion to text data, and identification of words and expressions corresponding to commands, requests, and / or parameters provided by the user. In some implementations, this processing may be minimal or there may be no processing at all. For example, this processing may include encoding the verbal input audio for transmission to the server 114, or may include providing the raw audio captured of the verbal input for transmission to the server 114.

[0123] The device transmits a request (506) that includes information determined based on the verbal input to a remote system. The client device 104 / casting device 106 processes the verbal input and determines the request from the verbal input by identifying the request and one or more related parameters from the verbal input. The client device 104 / casting device 106 transmits the determined request to a remote system (e.g., server 114). The remote system determines and generates a response to the request. In some implementations, the client device 104 / casting device 106 transmits the verbal input to the server 114 (e.g., as encoded audio, as raw audio data), and the server 114 processes the verbal input and determines the request and related parameters.

[0124] The device receives a response to the request (508). The response may be generated by the remote system in response to information based on the verbal input. The remote system (e.g., server 114) determines and generates a response to the request and transmits this response to the client device 104 / casting device 106.

[0125] The device executes an operation in response to the response (510). The client device 104 / casting device 106 executes one or more operations in response to the received response. For example, if the response is a command to cause the device to output specific information as audio, the client device 104 / casting device 106 extracts this information, converts this information into voice audio output, and outputs the voice audio from a speaker. As another example, if the response is a command to cause the device to play media content, the client device 104 / casting device 106 extracts the media content and plays the media content.

[0126] One or more of the above-mentioned receiving, processing, transmitting, receiving, and executing are performed by one or more voice processing modules of a voice assistant library running on an electronic device, and the voice processing modules provide a plurality of voice processing operations accessible to one or more application programs and / or operating software running on or executable on the electronic device (512). The client device 104 / casting device 106 may have a voice assistant client library 240 including functions and modules for performing one or more of the above-mentioned receiving step, processing step, transmitting step, receiving step, and executing step. The modules of the voice assistant client library 240 provide a plurality of voice processing operations and assistant operations accessible to applications, operating systems, and platform software in the client device 104 / casting device 106 that include or link to the library 240 (e.g., execute the library 240 and related APIs).

[0127] In some implementations, at least some of the voice processing operations associated with the voice processing module may be performed on a remote system interconnected with the electronic device via a wide area network. For example, processing oral input to determine a request may be performed by a server 114 connected to the client device 104 / casting device 106 through the network(s) 112.

[0128] In some implementations, the voice assistant library is executable on a common operating system that can operate on multiple different device types, thereby enabling the portability of voice-enabled applications configured to interact with one or more of the voice processing operations. The voice assistant client library 240 (as well as related libraries and APIs, such as the debugging library 242, the platform API 244, the POSIX API 246) utilizes standard elements (such as objects) of a defined operating system (such as Linux), and thus is operable on a variety of devices that execute a defined distribution or flavor of the operating system (such as different Linux or Linux-based distributions or flavors). In this way, the voice assistant function is available on a variety of devices, and the voice assistant experience is consistent across these various devices.

[0129] In some implementations, requests and responses may be processed at the device. For example, for basic functions that may be local to the device, such as a timer, an alarm clock, a clock, and volume adjustment, the client device 104 / casting device 106 processes the verbal input, determines that a request corresponds to one of these basic functions, determines the response at the device, and may perform one or more operations in response to the response. The device may continue to report requests and responses to the server 114 for logging purposes.

[0130] In some implementations, a device-agnostic voice assistant library for an electronic device with an audio input system includes one or more voice processing modules configured to execute on a common operating system implemented on multiple different electronic device types. The voice processing modules provide a plurality of voice processing operations that are accessible to application programs and operating software executing on the electronic device, thereby enabling the portability of voice-enabled applications configured to interact with one or more of the voice processing operations. The voice assistant client library 240 is a library that can be executed on various devices that share the same defined operating system base (e.g., the library and the device's operating system are Linux-based). Since it is a library that can be executed on various devices that share the same defined operating system base (e.g., the library and the device's operating system are Linux-based), this library is device-agnostic. Library 240 provides a plurality of modules for voice assistant functions accessible to applications across various devices.

[0131] In some implementations, at least some of the voice processing operations associated with the voice processing module are executed on a backend server interconnected with the electronic device via a wide area network. For example, library 240 communicates with server 114 and includes a module that sends oral input to server 114 for processing and determines requests.

[0132] In some implementations, the voice processing operations include device-specific operations configured to control a device (e.g., directly or communicably) connected to the electronic device. Library 240 may include functions or modules for controlling other devices (e.g., wireless speakers, smart TVs, etc.) connected to client device 104 / casting device 106.

[0133] In some implementations, the audio processing operation includes an information / media request operation configured to provide requested information and / or media content to a user of the electronic device or on a device (e.g., directly or communicatively) connected to the electronic device. The library 240 may include functions or modules for retrieving information or media and providing the information or media on the client device 104 / casting device 106 or on a connected device (e.g., reading an email aloud, reading a newspaper article aloud, playing streaming music).

[0134] For purposes of describing various elements, terms such as "first", "second", etc. may be used herein, but it will be understood that the elements should not be limited by these terms. These terms are only used to distinguish one element from another. For example, the first contact can be referred to as the second contact and vice versa without changing the meaning of the description, provided that the names of all the first contacts are changed consistently and the names of all the second contacts are changed consistently. The first contact and the second contact are both contacts, but they are not the same contact.

[0135] The terms used herein are for the purpose of describing particular implementations only and are not intended to limit the scope of the claims. The singular forms "a", "an", and "the" used in the description of the implementations and the appended claims are intended to include the plural forms as well, unless the context clearly indicates otherwise. The term "and / or" as used herein refers to and encompasses any one or more of the associated listed items and all possible combinations thereof. The terms "comprises" and / or "comprising" as used herein are, in this specification It will be understood that when used in the above description, it specifically recites the presence of the stated features, integers, steps, operations, elements, and / or components, but does not preclude the presence or addition of one or more other features, integers, steps, operations, elements, components, and / or groups thereof.

[0136] As used herein, the term "if" means "when," "upon," "in response to determining," "in accordance with a determination," or "in response to detecting" the stated condition precedent is true, depending on the context. Similarly, the phrase "if the stated antecedent condition is determined to be true" may be interpreted as meaning "if the stated antecedent condition is determined to be true" in the "it is determined [that a stated condition precedent is true]," "if [a stated condition precedent is true]," and "when [a stated condition precedent is true]" may be interpreted to mean "upon determining," "in response to determining," "in accordance with a determination," "upon detecting," or "in response to detecting" that a stated condition precedent is true, depending on the context.

[0137] Various implementation forms are referred to in detail, and examples thereof are shown in the accompanying drawings. In the following detailed description, numerous specific details are set forth for a thorough understanding of the present invention and the described implementation forms. However, the present invention can be practiced without these specific details. In other instances, well-known methods, procedures, components, and circuits have not been described in detail so as not to obscure aspects of the implementation forms unnecessarily.

[0138] The above description has been presented for purposes of illustration and is not intended to be exhaustive or to limit the invention to the precise form disclosed. Many modifications and variations are possible in light of the above teaching. The implementation forms have been chosen and described in order to best explain the principles of the invention and its practical application to enable one of ordinary skill in the art to utilize the invention and various implementation forms with various modifications suitable for the particular application contemplated.

Claims

1. 1. An electronic device comprising an audio input system, one or more processors, and a memory storing one or more programs executed by the one or more processors, receiving verbal input at the device; processing the verbal input; sending a request to a remote system including information determined based on said verbal input; receiving a response to the request generated by the remote system in response to information based on the verbal input; performing an action responsive to the response; The method, wherein one or more of the receiving, processing, transmitting, receiving, and executing steps are performed by one or more voice processing modules of a voice assistant library running on the electronic device, the voice processing modules providing a plurality of voice processing operations accessible to one or more application programs and / or operating software running or executable on the electronic device.

2. The method of claim 1 , wherein at least some voice processing operations associated with the voice processing module are performed on the remote system that is interconnected with the electronic device via a wide area network.

3. 10. The method of any one of the preceding claims, wherein the voice assistant library is executable on a common operating system operable on a number of different device types, thereby enabling portability of voice-enabled applications configured to interact with one or more of the voice processing operations.

4. 1. A device-agnostic voice assistant library for an electronic device with an audio input system, comprising:

1. A voice assistant library comprising: one or more voice processing modules configured to operate on a common operating system implemented on a plurality of different electronic device types, the voice processing modules providing a plurality of voice processing operations accessible to application programs and operating software executing on the electronic device, thereby enabling portability of voice-enabled applications configured to interact with one or more of the voice processing operations.

5. The voice assistant library of any one of the preceding claims, wherein at least some voice processing operations associated with the voice processing module are executed on a back-end server interconnected with the electronic device via a wide area network.

6. 2. The voice assistant library of claim 1, wherein the voice processing operations include device specific operations configured to control devices connected to the electronic device.

7. 2. The voice assistant library of any one of the preceding claims, wherein the voice processing operations include information / media request operations configured to provide requested information and / or media content to a user of the electronic device or on a device connected to the electronic device.

8. an audio input system; one or more processors; a memory storing one or more programs to be executed by the one or more processors, the one or more programs comprising: instructions for receiving verbal input at the device; instructions for processing the verbal input; instructions for transmitting a request to a remote system including information determined based on the verbal input; instructions for receiving a response to the request, the response being generated by the remote system in response to information based on the verbal input; instructions for performing an action responsive to the response; An electronic device, wherein one or more of the receiving, processing, transmitting, receiving, and executing are performed by one or more voice processing modules of the voice assistant library running on the electronic device, the voice processing modules providing a plurality of voice processing operations accessible to one or more application programs and / or operating software running or executable on the electronic device.

9. The device of claim 8 , wherein at least some voice processing operations associated with the voice processing module are performed on the remote system, which is interconnected with the electronic device via a wide area network.

10. 10. The device of any one of the preceding claims, wherein the voice assistant library is capable of executing on a common operating system operable on a number of different device types, thereby enabling portability of voice-enabled applications configured to interact with one or more of the voice processing operations.

11. A non-transitory computer-readable storage medium having stored thereon one or more programs, the one or more programs including instructions that, when executed by an electronic device having an audio input system and one or more processors, cause the electronic device to: receiving a verbal input at the device; Processing said verbal input; transmitting a request to a remote system including information determined based on said verbal input; receiving a response to the request generated by the remote system in response to information based on the verbal input; performing an action responsive to said response; A non-transitory computer-readable storage medium, wherein one or more of the receiving, processing, transmitting, receiving, and executing are performed by one or more voice processing modules of the voice assistant library running on the electronic device, the voice processing modules providing a plurality of voice processing operations accessible to one or more application programs and / or operating software running or executable on the electronic device.

12. 12. The computer-readable storage medium of claim 11, wherein at least some voice processing operations associated with the voice processing module are performed on the remote system that is interconnected with the electronic device via a wide area network.

13. The voice assistant library is capable of running on a common operating system that can run on multiple different device types, thereby allowing the voice processing operations to be 13. A computer readable storage medium according to any one of the preceding claims, enabling portability of voice-enabled applications configured to interact with one or more of the following:

14. an audio input system; one or more processors; and a memory storing one or more programs executed by said one or more processors, said one or more programs including instructions for executing the method according to any one of claims 1 to 3.

15. A non-transitory computer readable storage medium having stored thereon one or more programs, the one or more programs comprising instructions that, when executed by an electronic device having an audio input system and one or more processors, cause the electronic device to perform the method of any one of claims 1 to 3.