Implementation of a voice assistant on a device

A device-side library for voice assistants provides a consistent experience across devices by enabling local processing and cloud connectivity, allowing device-specific enhancements to be independent, thus enhancing user interaction.

JP7870243B2Active Publication Date: 2026-06-04GOOGLE LLC

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
GOOGLE LLC
Filing Date
2022-12-15
Publication Date
2026-06-04

AI Technical Summary

Technical Problem

Existing voice-based assistants lack the ability to provide a consistent experience across multiple devices and support device-specific functions effectively.

Method used

A device-side library for voice assistants that enables local processing of audio data, supports connectivity to the cloud, and provides a portable voice control system, allowing integration into various operating environments and asynchronous updates.

Benefits of technology

Enables a consistent user experience across different devices while allowing innovations in voice assistant functionality to be separate from device-specific advancements, facilitating seamless integration and updates.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007870243000018
    Figure 0007870243000018
  • Figure 0007870243000019
    Figure 0007870243000019
  • Figure 0007870243000020
    Figure 0007870243000020
Patent Text Reader

Abstract

Provided are a voice assistant library, a voice processing method, a device, and a storage medium that can provide a consistent experience across various devices and support functions specialized for a particular device. [Solution] A voice processing module that provides multiple voice processing operations accessible to one or more application programs and / or operating software includes steps of receiving verbal input at a device, processing the verbal input, sending a request to a remote system that includes information determined based on the verbal input, receiving a response to the request generated by the remote system, and performing an action according to the response.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Technical Field This application generally relates to computer technology including, but not limited to, voice assistants for devices and related libraries.

Background Art

[0002] Background Along with the development of the Internet and cloud computing, the popularity of voice-based assistants that interact with users through audio / voice input / output has been increasing. These assistants provide an interface for consuming digital media and can provide various types of information such as, for example, news, sports scores, weather, and stocks.

[0003] Users will likely have multiple devices with desired voice-based assistant functionality. It is desirable to have a voice-based assistant that can be implemented and used across various devices, provide a consistent experience across these various devices, and support functions specific to a particular device.

Summary of the Invention

Means for Solving the Problems

[0004] Summary The implementations described herein are directed to incorporating or including a voice assistant in an embedded system and / or device in a way that enables controlling local devices for a variety of operating system platforms.

[0005] According to several implementations, a thin, resource-efficient device-side library features local processing of audio data, listening for wake words or hot words, and sending user requests. Further features include connectivity to the cloud brain, an extensible voice control system, a portability layer enabling integration into many diverse operating environments, and the ability to update asynchronously with the rest of the client software.

[0006] The implementation described has the advantage of providing a similar user experience for interacting with the voice assistant across many different devices.

[0007] The described implementation has another advantage: it allows innovations in voice assistant functionality to be separated from innovations available from the device itself. For example, if an improved recognition pipeline is created, the recognition results are output to the device, while the device manufacturer can continue to benefit from previous voice commands without having to do anything to receive the recognition results.

[0008] According to some implementations, a method in an electronic device having an audio input system, one or more processors, and a memory storing one or more programs executed by the one or more processors includes the steps of: receiving oral input in the device; processing the oral input; sending a request to a remote system containing information determined based on the oral input; receiving a response to the request generated by the remote system in accordance with the information based on the oral input; and performing an action in accordance with the response. One or more of the steps of receiving, processing, transmitting, receiving, and executing are performed by one or more voice processing modules of a voice assistant library running on an electronic device, and the voice processing modules provide multiple voice processing operations that are accessible to one or more application programs and / or operating software running or executable on the electronic device.

[0009] In some implementations, a device-independent voice assistant library for electronic devices with an audio input system includes one or more voice processing modules configured to run on a common operating system implemented on multiple different electronic device types, the voice processing modules providing multiple voice processing operations accessible to application programs and operating software running on the electronic device, thereby enabling the portability of voice-enabled applications configured to interact with one or more of these voice processing operations.

[0010] In some implementations, the electronic device comprises an audio input system, one or more processors, and memory storing one or more programs executed by the one or more processors. The one or more programs include instructions for receiving oral input on the device, instructions for processing the oral input, instructions for sending a request containing information determined based on the oral input to a remote system, instructions for receiving a response to the request generated by the remote system in response to the information based on the oral input, and instructions for performing actions in response to the response. One or more of receiving, processing, transmitting, receiving, and executing are performed by one or more speech processing modules of a speech assistant library running on the electronic device, and the speech processing modules provide multiple speech processing operations accessible to one or more application programs and / or operating software running or executable on the electronic device.

[0011] In some implementations, a non-temporary computer-readable storage medium stores one or more programs. The one or more programs include instructions, which, when executed by an electronic device having an audio input system and one or more processors, cause the electronic device to receive oral input on the device, process the oral input, send a request to a remote system containing information determined based on the oral input, receive a response to the request generated by the remote system in accordance with the information based on the oral input, and perform an action in accordance with the response. One or more of receiving, processing, sending, receiving, and performing are performed by one or more speech processing modules of a speech assistant library running on the electronic device, and the speech processing modules provide a number of speech processing operations accessible to one or more application programs and / or operating software running or executable on the electronic device. [Brief explanation of the drawing]

[0012] [Figure 1] This is a block diagram showing examples of network environments related to several implementation forms. [Figure 2] This figure shows examples of voice assistant client devices in several implementation forms. [Figure 3] This figure shows examples of server systems in several implementation forms. [Figure 4] This is a block diagram showing the functional views of a voice assistant library across several implementation configurations. [Figure 5] This is a flowchart illustrating methods for processing verbal input on a device, relating to several implementation configurations. [Modes for carrying out the invention]

[0013] Throughout the drawing, the same reference number refers to the corresponding part. Description of implementation Here, various implementation configurations are described in detail. Examples of these configurations are shown in the attached drawings. In the following detailed description, many specific details are provided to give a full understanding of the present invention and the described configurations. However, the present invention can be carried out without these specific details. In other cases, well-known methods, procedures, components, and circuits are not described in detail so as not to unnecessarily obscure the configuration of the implementation.

[0014] In some implementations, the goal of a voice assistant is to provide users with a personalized voice interface that is usable across various devices and enables a wide range of use cases, providing a consistent experience throughout the user's day. The voice assistant and / or related functionality may be integrated into first-party products and devices, as well as third-party products and devices.

[0015] Use cases include media. Voice commands may be used to initiate playback and control of music, radio, podcasts, news, and other audio media. For example, a user can play or control various types of audio media by issuing voice commands such as (e.g., "Play jazz music," "Play FM107.5," "Skip to the next song," "Play 'Continuous'"). Furthermore, such commands may be used to play audio media from various sources, such as online streaming of terrestrial radio stations, music subscription services, local storage, and remote storage. In addition, the voice assistant may support further content by utilizing integrations available on the casting device.

[0016] Another use case example involves remote playback. A user may issue voice commands to a casting device that includes voice assistant functionality, and in response to the voice commands, media will be played (for example, cast) on the device specified in the command, on a device included in a specified group of one or more devices, or on one or more devices in an area specified in the command. The user can also specify a general category or specific content in the command, and appropriate media will be played according to the category or content specified in the command.

[0017] Further use case examples include non-media features such as productivity-enhancing functions (e.g., timers, alarm clocks, calendars), home automation, search engines (e.g., search queries), questions and answers that leverage technology, entertainment (e.g., assistant personalities, jokes, games, Easter eggs), and daily tasks (e.g., transportation, travel, food, finance, gifts, etc.).

[0018] In some implementations, the voice assistant is offered as an optional feature of the casting device, and the voice assistant functionality may be updated as part of the casting device.

[0019] In some implementations, the detection of hotwords or keywords included in voice commands and user input is performed by the application processor (for example, at the client device or casting device to which the user speaks the voice command or input). In some implementations, hotword detection is performed by an external digital signal processor (when the user speaks the voice command or This is done, for example, by a server system processing voice commands, as opposed to the client device or casting device to which the voice input is spoken.

[0020] In some implementations, a device having a voice assistant function includes one or more of remote assistance, "push-to-assist" or "push-to-talk" (e.g., a button to initiate the voice assistant function), and an AC power supply.

[0021] In some implementations, the voice assistant includes an audio input device (e.g., a microphone, a media loopback of ongoing playback), the state of the microphone (e.g., on / off), ducking (e.g., lowering the volume of all outputs when the assistant is triggered (triggered) by a hotword or push-to-talk) ), as well as an application programming interface (API) for one or more of new assistant events and status messages (e.g., the assistant is activated (e.g., heard a hotword, the assistant button was pressed), listening to voice, waiting on the server, responding, response completed, alarm / timer ringing).

[0022] In some implementations, a device having a voice assistant function may communicate with another device (e.g., a settings application on a smartphone) for settings purposes to enable or facilitate the voice assistant function on the device. Settings or setup may include specifying the location of the device, associating with a user account, opting in the user to voice control, linking to media services (e.g., video streaming service, music streaming service) and prioritizing media services, home automation settings, etc.

[0023] In some implementations, a device with a voice assistant may include one or more user interface elements or displays for the user. One or more of these user interface elements may be physical elements (e.g., a light pattern displayed using one or more LEDs, a sound pattern emitted by a speaker) and may include one or more of the following: a hotword-independent "push-to-assist" or "push-to-talk" trigger, a "mute microphone" trigger and visual status indicator, a visual "waiting for hotword status," a visual "hotword detected," a visual "assistant is actively listening" visible from a short distance (e.g., 15 feet), a visual "assistant is working / thinking," a visual "voice message / notification available," a "volume level" control method and status indicator, and a "pause / resume" control method. In some implementations, these physical user interface elements are provided by a client device or a casting device. In some implementations, the voice assistant supports a common set of user interface elements or displays across different devices to ensure a consistent experience across different devices.

[0024] In some implementations, the voice assistant supports device-specific commands and / or hotwords, as well as a predefined standard set of commands and / or hotwords.

[0025] Figure 1 shows a network environment 100 relating to several implementation forms. The network environment 100 includes a casting device 106 and / or a voice assistant client device 104. The casting device 106 (for example, Chromecast by Google Inc.) is connected to an audio input device 108 (for example). The audio input device 108 and audio output device 110 (e.g., one or more speakers) are connected directly or communicatively to the microphone and audio output device 110. In some implementations, the audio input device 108 and audio output device 110 are components of a device communicatively connected to the casting device 106 (e.g., a speaker system, television, soundbar). In some implementations, the audio input device 108 is a component of the casting device 106 and the audio output device 110 is a component of the device communicatively connected to the casting device 106, or the audio output device 110 is a component of the casting device 106 and the audio input device 108 is a component of the device communicatively connected to the casting device 106. In some implementations, the audio input device 108 and audio output device 110 are components of the casting device 106.

[0026] In some implementations, the casting device 106 is communicatively connected to the client 102. The client 102 may include an application or module (for example, a casting device configuration app) that facilitates the configuration of the casting device 106, including voice assistant functionality.

[0027] In some implementations, the casting device 106 is connected to the display 144.

[0028] In some implementations, the casting device 106 includes one or more visual indicators 142 (e.g., LED lights).

[0029] In some implementations, the casting device 106 includes a receiving module 146. In some implementations, the receiving module 146 operates the casting device 106, including communication with hardware functions and content sources, for example. In some implementations, the casting device 106 has different receiving modules 146 for different content sources. In some implementations, the receiving module 146 includes submodules for each different content source.

[0030] The voice assistant client device 104 (for example, a smartphone, laptop or desktop computer, tablet computer, voice command device, mobile device, or in-vehicle system with Google Assistant or Google Home by Google Inc.) comprises an audio input device 132 (for example, a microphone) and an audio output device 134 (for example, one or more speakers or headphones). In some implementations, the voice assistant client device 104 (for example, a voice command device, mobile device, or in-vehicle system with Google Assistant or Google Home by Google Inc.) is communicatively connected to a client 140 (for example, a smartphone or tablet device). The client 140 may include an application or module (for example, a voice command device configuration app) that facilitates the configuration of the voice assistant client device 104, including voice assistant functionality.

[0031] In some implementations, the voice assistant client device 104 includes one or more visual indicators 152 (e.g., LED lights). An example of a voice assistant client device having visual indicators (e.g., LED lights) was filed on May 13, 2016, in the "LED Design Language for Visual Affordance of Voice User Interfaces" patent application. This is shown in Figure 4A, which illustrates U.S. Provisional Application No. 62 / 336,566, titled "Design Language" (as incorporated herein by reference).

[0032] The casting device 106 and the voice assistant client device 104 each contain instances of the voice assistant module or library 136. The voice assistant module / library 136 is a module / library that implements voice assistant functionality across different devices (e.g., the casting device 106, the voice assistant client device 104). The voice assistant functionality is consistent across different devices while still allowing device-specific features (e.g., support for controlling device-specific features by the voice assistant). In some implementations, the voice assistant module / library 136 is the same or similar across devices, and instances of the same library may be included in different devices.

[0033] In some implementations, depending on the device type, the voice assistant module / library 136 may be included in an application installed on the device, in the device's operating system, or embedded in the device (for example, embedded in the firmware).

[0034] In some implementations, the voice assistant module / library 136-1 in the casting device 106 communicates with the receiving module 146 to perform voice assistant operations.

[0035] In some implementations, the voice assistant module / library 136-1 in the casting device 106 can control or influence the visual indicator 142.

[0036] In some implementations, the voice assistant module / library 136-2 in the voice assistant client device 104 can control or influence the visual indicator 152.

[0037] The casting device 106 and the voice assistant client device 104 are communicably connected to the server system 114 via one or more communication networks 112 (e.g., a local area network, a wide area network, or the Internet). The voice assistant module / library 136 detects (e.g., receives) spoken input picked up (e.g., captured) by the audio input devices 108 / 132, processes the spoken input (e.g., to detect hotwords), and sends the processed spoken input or an encoded version of the processed spoken input to the server 114. The server 114 receives the processed spoken input or its encoded version, processes the received spoken input, and determines an appropriate response to the spoken input. The appropriate response may be content, information, or instructions, commands, or metadata to cause the casting device 106 or the voice assistant client device 104 to perform a function or action. Server 114 sends a response to the casting device 106 or voice assistant client device 104 from which content or information is output (for example, from audio output devices 110 / 134) and / or which function is executed. As part of the process, Server 114 may communicate with one or more content / information sources 138 to retrieve or reference content or information for the response. In some implementations, the content / information source 138 may be a search engine, a database, or information associated with a user's account. Examples include calendars, task lists, email, websites, and media streaming services. In some implementations, the voice assistant client device 104 and the casting device 106 may communicate or interact with each other. Examples of such communication or interaction, and examples of the operation of the voice assistant client device 104 (for example, Google Home by Google Inc.) are described in U.S. Provisional Application No. 62 / 336,566, filed on May 13, 2016, entitled "LED Design Language for Visual Affordance of Voice User Interfaces," and also filed on May 13, 2016, entitled "Voice-Controlled Closed Caption Display." U.S. Provisional Application No. 62 / 336,569, titled "Display of Closed Captions", and U.S. Provisional Application No. 62 / 336,569, filed on 13 May 2016, titled "Media Transfer among Media Output Devices". Disclosed in Patent No. 565. All of these applications are incorporated herein by reference.

[0038] In some implementations, the voice assistant module / library 136 receives spoken input captured by the audio input devices 108 / 132 and sends the spoken input (with little or no processing) or an encoded version thereof to the server 114. The server 114 processes the spoken input to detect hotwords, determines an appropriate response, and sends this response to the casting device 106 or the voice assistant client device 104.

[0039] If the server 114 determines that the oral input contains a command for the casting device 106 or the voice assistant client device 104 to execute a function, the server 114 sends a response containing an instruction or metadata telling the casting device 106 or the voice assistant client device 104 to execute the function. The function may be device-specific, and functionality to support such a function in the voice assistant may be included in the casting device 106 or client 104 as a custom module or function added to or linked to the voice assistant module / library 136.

[0040] In some implementations, the server 114 includes, or is connected to, a speech processing backend 148 that performs processing operations for oral input and determines the response to said oral input.

[0041] In some implementations, the server 114 includes a downloadable voice assistant library 150. The downloadable voice assistant library 150 (for example, the same as or an updated version of voice assistant library 136) may include new features or functions, or updates, and can be downloaded to add a voice assistant library to a device or to update voice assistant library 136.

[0042] Figure 2 is a block diagram showing examples of voice assistant client devices 104 or casting devices 106 in a network environment 100, relating to several implementation forms. Examples of voice assistant client devices 104 include mobile phones, tablet computers, laptop computers, desktop computers, wireless speakers (e.g., Google Home by Google Inc.), voice command devices (e.g., Google Home by Google Inc.), televisions, soundbars, casting devices (e.g., Chromecast by Google Inc.), media streaming devices, home appliances, home electronics, and in-vehicle devices. Examples include, but are not limited to, systems and wearable personal devices. A voice assistant client device 104 (e.g., Google Home by Google Inc., a mobile device with Google Assistant functionality) or a casting device 106 (e.g., Chromecast by Google Inc.) typically comprises one or more processing units (CPUs) 202, one or more network interfaces 204, memory 206, and one or more communication buses 208 (sometimes called chipsets) for connecting these components to each other. The voice assistant client device 104 or casting device 106 also comprises one or more input devices 210 to facilitate user input. One or more input devices 210 include audio input devices 108 or 132 (e.g., a voice command input unit or microphone), and, if necessary, other input devices such as a keyboard, mouse, touchscreen display, touch input pad, gesture capture camera, or other input buttons or controls). In some implementations, the voice assistant client device 102 uses a microphone and voice recognition, or a camera and gesture recognition, to assist or replace a keyboard. The voice assistant client device 104 or casting device 106 also includes one or more output devices 212. One or more output devices 212 include audio output devices 110 or 134 (e.g., one or more speakers, headphones, etc.) and, if necessary, one or more display devices (e.g., displays 144) and / or one or more visual indicators 142 or 152 (e.g., LEDs) that enable the presentation of a user interface and display content and information. If necessary, the voice assistant client device 104 or casting device 106 includes a location detection unit 214, such as a GPS (Global Positioning Satellite) or other geolocation receiver, for determining the location of the voice assistant client device 104 or casting device 106.Additionally, the voice assistant client device 104 or casting device 106 may optionally include a proximity detection device 215, such as an IR sensor, for determining the proximity of the voice assistant client device 104 or casting device 106 to other objects (for example, the user / wearer in the case of a wearable personal device). Optionally, the voice assistant client device 104 or casting device 106 may include sensors 213 (for example, accelerometers, gyroscopes, etc.).

[0043] Memory 206 includes high-speed random-access memory such as DRAM, SRAM, DDR RAM, or other random-access solid-state memory, and, if necessary, one or more magnetic disk storage devices, one or more optical disk storage devices, one or more flash memory devices, or one or more other non-volatile solid-state memory devices. Memory 206 also includes one or more storage devices located away from one or more processing units 202, if necessary. Memory 206, or the non-volatile memory within Memory 206, includes a non-temporary computer-readable storage medium. In some implementations, Memory 206, or the non-temporary computer-readable storage medium within Memory 206, stores the following programs, modules, and data structures, or subsets or supersets thereof:

[0044] ● An operating system 216 that includes procedures for handling various basic system services and for performing hardware-dependent tasks.

[0045] ●One or more network interfaces 204 (wired or wireless), and one or more networks 112 such as the Internet, other wide area networks, local area networks, metropolitan area networks, etc., to connect the voice assistant client device 104 or casting device 106 to other A network communication module 218 for connecting to a device (for example, a server system 114, clients 102, 140, or other voice assistant client devices 104 or casting device 106).

[0046] ● A user interface module 220 for enabling the presentation of information on a voice assistant client device 104 or casting device 106 via one or more output devices 212 (e.g., a display, speaker, etc.).

[0047] ● An input processing module 222 for processing and interpreting interactions captured or received by one or more user inputs or one or more input devices 210.

[0048] ● A voice assistant module 136 for processing verbal input, providing said verbal input to server 114, receiving a response from server 114, and outputting said response.

[0049] ● Client data 226 for storing data associated with at least the voice assistant module 136, including the following:

[0050] ○Voice assistant settings 228 for storing information related to the settings and configuration of the voice assistant module 136 and the voice assistant function.

[0051] Content source / information source 230 and content category / information category 232 for storing predefined and / or user-specified sources and categories of content or information.

[0052] ○ Usage history 234 for storing information (e.g., logs) associated with the operation and use of the voice assistant module 136, such as received commands and requests, responses to commands and requests, and actions performed in response to commands and requests.

[0053] User accounts and authorizations 236 for storing authorization and authentication information for one or more users to access each user's account in the content source / information source 230 and the account information of those authorized accounts.

[0054] ○ A receiving module 146 for operating the casting function of the casting device 106, including communicating with a content source.

[0055] In some implementations, the voice assistant client device 104 or casting device 106 includes one or more libraries and one or more application programming interfaces (APIs) for the voice assistant and related functions. These libraries may be included in the voice assistant module 136 or receiving module 146, or may be linked to each other by the voice assistant module 136 or receiving module 146. The libraries include modules associated with voice assistant functions or other functions that facilitate voice assistant functions. The APIs provide interfaces to hardware and other software (e.g., operating systems, other applications) that facilitate voice assistant functions. For example, the voice assistant client library 240, the debugging library 242, the platform API 244, and the POSIX API 246 may be stored in memory 206. These libraries and APIs are described in more detail below with reference to Figure 4.

[0056] In some implementations, the voice assistant client device 104 or CAST The tapping device 106 includes a voice application 250 that utilizes modules and functions of the voice assistant client library 240, and optionally includes a debugging library 242, platform API 244, and POSIX API 246. In some implementations, the voice application 250 is a first-party or third-party application that becomes voice-enabled by using the voice assistant client library 240.

[0057] Each of the above elements may be stored in one or more of the aforementioned memory devices and corresponds to an instruction set for executing the above function. Since the above modules or programs (i.e., instruction sets) do not need to be implemented as separate software programs, procedures, modules, or data structures, various subsets of these modules may be combined or rearranged in various implementations. In some implementations, memory 206 stores a subset of the above modules and data structures, if necessary. Furthermore, memory 206 stores additional modules and data structures not described above, if necessary.

[0058] Figure 3 is a block diagram showing examples of a server system 114 in a network environment 100, relating to several implementation configurations. The server 114 typically comprises one or more processing units (CPUs) 302, one or more network interfaces 304, memory 306, and one or more communication buses 308 for connecting these components (sometimes called chipsets) to each other. The server 114 optionally includes one or more input devices 310 to facilitate user input, such as a keyboard, mouse, voice command input unit or microphone, touchscreen display, touch input pad, gesture capture camera, or other input buttons or control units. Furthermore, the server 114 may use a microphone and voice recognition, or a camera and gesture recognition, to assist or replace the keyboard. In some implementation configurations, the server 114 optionally includes one or more cameras, scanners, or optical sensors for capturing, for example, graphic series codes printed on electronic devices. The server 114 also optionally includes one or more output devices 312, including one or more speakers and / or one or more display devices, which enable the presentation of a user interface and display content.

[0059] Memory 306 includes high-speed random-access memory such as DRAM, SRAM, DDR RAM, or other random-access solid-state memory, and, if necessary, one or more magnetic disk memory devices, one or more optical disk memory devices, one or more flash memory devices, or one or more other non-volatile solid-state memory devices. Memory 306 also includes one or more storage devices located away from one or more processing units 302, if necessary. Memory 306, or the non-volatile memory within Memory 306, includes a non-temporary computer-readable storage medium. In some implementations, Memory 306, or the non-temporary computer-readable storage medium within Memory 306, stores the following programs, modules, and data structures, or subsets or supersets thereof:

[0060] ● An operating system 316 that includes procedures for handling various basic system services and for performing hardware-dependent tasks.

[0061] ●One or more processing units 304 (wired or wireless), and one or more networks 112 such as the Internet, other wide area networks, local area networks, metropolitan area networks, etc., connect the server system 114 to other devices (for example, voice assistant client devices 104, casting A network communication module 318 for connecting to the client devices 106, 102, and 140.

[0062] ● Proximity / location module 320 for determining the proximity and / or location of the voice assistant client device 104 or casting device 106 based on the location information of the client device 104 or casting device 106.

[0063] ● A voice assistant backend 116 for processing voice assistant oral input (for example, oral input received from a voice assistant client device 104 and a casting device 106), including at least one of the following:

[0064] ○ A verbal input processing module 324 for processing verbal input and identifying the commands and requests contained in the verbal input.

[0065] ○ A content / information gathering module 326 for collecting content and informational responses to commands and requests.

[0066] ○ A response generation module 328 for generating audio output in response to commands and requests, and for adding said audio output along with the response content and information.

[0067] ● Server system data 330 that stores data associated with the operation of the voice assistant platform, including the following:

[0068] ○ User data 332 for storing information associated with a user of the voice assistant platform, including the following:

[0069] - User voice assistant settings 334 for storing voice assistant setting information corresponding to voice assistant setting 228, and information corresponding to content source / information source 230 and content category / information category 232.

[0070] - User history 336 for storing the user's history (e.g., logs) about the voice assistant, including a history of commands and requests, as well as corresponding responses.

[0071] - User accounts and authorizations 338 for storing user authorization and authentication information for accessing each user's account in content source / information source 230, and account information for those authorized accounts corresponding to user accounts and authorizations 236.

[0072] Each of the above elements may be stored in one or more of the aforementioned memory devices and corresponds to an instruction set for executing the above function. Since the above modules or programs (i.e., instruction sets) do not need to be implemented as separate software programs, procedures, modules, or data structures, various subsets of these modules may be combined or rearranged in various implementations. In some implementations, memory 206 stores a subset of the above modules and data structures, if necessary. Furthermore, memory 206 stores additional modules and data structures not described above, if necessary.

[0073] In some implementations, the voice assistant module 136 (Figure 2) includes one or more libraries. A library contains modules or submodules that perform their respective functions. For example, the voice assistant client library includes the voice assistant It includes a module that executes the function. The voice assistant module 136 may also include one or more application programming interfaces (APIs) for interacting with specific hardware (e.g., hardware on a client device or casting device), specific operating software, or a remote system.

[0074] In some implementations, the library includes modules that support audio signal processing operations, such as bandpassing, filtering, erasure, and hotword detection. In some implementations, the library includes modules for connecting to a backend (e.g., server-based) audio processing system. In its installed form, the library includes modules for debugging (e.g., debugging speech recognition, debugging hardware problems, automated testing).

[0075] Figure 4 shows libraries and APIs that may be stored in the voice assistant client device 104 or the casting device 106 and that can be executed by the voice assistant module 136 or another application. The libraries and APIs may include the voice assistant client library 240, the debugging library 242, the platform API 244, and the POSIX API 246. An application in the voice assistant client device 104 or the casting device 106 (for example, the voice assistant module 136, or any other application that would like to support collaboration with the voice assistant) may include, link to, and execute these libraries and APIs in order to provide or support voice assistant functionality in the application. In some implementations, the voice assistant client library 240 and the debugging library 242 are separate libraries. Keeping the voice assistant client library 240 and the debugging library 242 separate facilitates different release and update procedures that take into account the different security implications of these libraries.

[0076] In some implementations, these libraries offer flexibility. They can be used across multiple device types and may incorporate the same voice assistant functionality.

[0077] In some implementations, the library relies on standard shared objects (e.g., standard Linux® shared objects), making it compatible with different operating systems or platforms that utilize these standard shared objects (e.g., various Linux distributions and flavors of embedded Linux).

[0078] In some implementations, POSIX API 246 provides a standard API for compatibility with various operating systems. Therefore, the voice assistant client library 240 may be included in devices with different POSIX-compliant operating systems, and POSIX API 246 provides a compatibility interface between the voice assistant client library 240 and different operating systems.

[0079] In some implementations, the library includes modules to support and facilitate base use cases available across different types of devices that implement voice assistants (e.g., timers, alarms, volume controls).

[0080] In some implementations, the voice assistant client library 240 is a voice assistant The controller interface 402 includes functions or modules for starting, configuring, and interacting with the voice assistant. In some implementations, the controller interface 402 includes a "Start()" function or module 404 for starting the voice assistant on the device, a "RegisterAction()" function or module 406 for registering actions with the voice assistant (for example, so that actions can be performed via the voice assistant), a "Reconfigure()" function 408 for reconfiguring the voice assistant with updated settings, and a "RegisterEventObserver()" function 410 for registering a set of functions for basic events with the assistant.

[0081] In some implementations, the voice assistant client library 240 includes multiple functions or modules associated with specific voice assistant functions. For example, the hotword detection module 412 processes voice input to detect hotwords. The voice processing module 414 processes the voice contained in the voice input, converting voice to text or text to voice (e.g., word and expression identification, voice-to-text data conversion, text-to-speech conversion). The action processing module 416 performs actions and behaviors in response to verbal input. The local timer / alarm / volume control module 418 facilitates alarm clock, timer, and volume control functions on the device, as well as their control via voice input (e.g., managing the timer, clock, and alarm clock on the device). The logging / evaluation metrics module 420 records voice input and responses (e.g., logging) and determines and records relevant evaluation metrics (e.g., response time, idle time, etc.). The audio input processing module 422 processes the audio of the voice input. The MP3 decoding module 424 decodes audio encoded in MP3. The audio input module 426 captures audio from an audio input device (e.g., a microphone). The audio output module 428 outputs audio from an audio output device (e.g., a speaker). An event queuing / state tracking module 430 queues events associated with the voice assistant on the device and tracks the state of the voice assistant on the device.

[0082] In some implementations, the debugging library 242 provides modules and functions for debugging. For example, the HTTP server module 432 facilitates debugging connectivity issues, and the debug server / audio streaming module 434 debugs audio issues.

[0083] In some implementations, the platform API 244 provides an interface between the voice assistant client library 240 and the device's hardware functions. For example, the platform API includes a button input interface 436 for capturing button inputs to the device, a loopback audio interface 438 for capturing loopback audio, a logging / metrics interface 440 for logging and determining evaluation metrics, an audio input interface 442 for capturing audio input, an audio output interface 444 for outputting audio, and an authentication interface 446 for authenticating the user using other services that can interact with the voice assistant. The advantage of the voice assistant client library organization shown in Figure 4 is that it can provide the same or similar voice processing functionality on various voice assistant device types, each with a consistent API and set of voice assistant functions. This consistency supports the portability of voice assistant applications and the consistency of voice assistant operation, facilitating consistent user interaction and familiarity with voice assistant applications and functions operating on different device types. All or part of the assistant client library 240 may be provided on server 114 to support server-based voice assistant applications (for example, a server application that operates on voice input sent to server 114 for processing).

[0084] The following are code examples of the classes and functions corresponding to Controller 402 ("Controller"), as well as related classes. These classes and functions can be adopted by applications that can run on various devices via a common API.

[0085] The following class, "ActionModule," facilitates the registration of an application's module to process commands provided by a voice assistant server.

[0086]

number

[0087] The following class "BuildInfo" may be used to describe the application running the voice assistant client library 240 or the voice assistant client device 104 itself (for example, using an identifier or version number for the application, platform, and / or device).

[0088]

number

[0089] The class "EventDelegate" below defines functions associated with basic events, such as the start of speech recognition, and the start and completion of output from the voice assistant.

[0090]

number

[0091] The class "DefaultEventDelegate" below defines an override function that does nothing for a specific event.

[0092]

number

[0093] The following class "Settings" defines the settings that may be provided to controller 402 (for example, locale, geographical location, file system directory).

[0094]

number

[0095] The class "Controller" below corresponds to controller 402, and the Start(), Reconfigure(), RegisterAction(), and RegisterEventObserver() functions correspond to the Start()404, Reconfigure()408, RegisterAction()406, and RegisterEventObserver()410 functions, respectively.

[0096]

number

[0097] In some implementations, the voice assistant client device 104 or casting device 106 implements a platform (for example, a set of interfaces for communicating with other devices using the same platform, and an operating system configured to support that set of interfaces). The following code example shows functions associated with the interfaces for the voice assistant client library 402 to interact with the platform.

[0098] The following class, "Authentication," defines an authentication token for authenticating a voice assistant user who has a specific account.

[0099]

number

[0100] The class "OutputStreamType" below defines the type of audio output stream.

[0101]

number

[0102] The following class, "SampleFormat," defines the supported audio sample formats (for example, PCM format).

[0103]

number

[0104] The "BufferFormat" described below defines the format of the data stored in the device's audio buffer.

[0105]

number

[0106] The following class, "AudioBuffer," defines an audio data buffer.

[0107]

number

[0108] The following class, "AudioOutput," defines an interface for audio output.

[0109]

number

[0110] The following class, "AudioInput," defines an interface for capturing audio input.

[0111]

number

[0112] The following class, "Resources," defines access to system resources.

[0113]

number

[0114] The class "PlatformApi" below specifies the platform API for the voice assistant client library 240 (for example, platform API 244).

[0115]

number

[0116] In some implementations, volume control may be handled outside of the voice assistant client library 240. For example, system volume may be managed by a device not controlled by the voice assistant client library 240. In another example, the voice assistant client library 240 may still support volume control, but requests for volume control from the voice assistant client library 240 are directed to the device.

[0117] In some implementations, the alarm and timer functions included in the voice assistant client library 240 may be disabled by the user or disabled when the library is implemented on the device.

[0118] In addition, in some implementations, the voice assistant client library 240 supports interfacing with LEDs on the device, facilitating the display of LED animations on the device's LEDs.

[0119] In some implementations, the voice assistant client library 240 may be included in or linked to a casting receiver module (e.g., receiver module 146) in the casting device 106. The link between the voice assistant client library 240 and the receiver module 146 may include, for example, support for further actions (e.g., local media playback) and support for controlling LEDs on the casting device 106.

[0120] Figure 5 is a flowchart of method 500 for processing oral input on a device, relating to several implementation forms. Method 500 is an electronic device (e.g., voice assistant) having an audio input system (e.g., audio input device 108 / 132), one or more processors (e.g., processing unit(s) 202), and memory (e.g., memory 206) that stores one or more programs executed by the one or more processors. The method is executed in the input device 104, the casting device 106). In some implementations, the electronic device comprises an audio input system (e.g., an audio input device 108 / 132), one or more processors (e.g., a processing unit 202), and a memory (e.g., a memory 206) storing one or more programs executed by the one or more processors, the one or more programs including instructions for executing method 500. In some implementations, a non-temporary computer-readable storage medium includes one or more programs, the one or more programs including instructions, which, when executed by an electronic device having an audio input system (e.g., an audio input device 108 / 132) and one or more processors (e.g., a processing unit 202), cause the electronic device to execute method 500. The programs or instructions for executing method 500 may be included in the modules, libraries, etc. described above with reference to Figures 2 to 4.

[0121] The device receives verbal input on the device (502). The client device 104 / casting device 106 captures the verbal input (e.g., voice input) made by the user.

[0122] The device processes the oral input (504). The client device 104 / casting device 106 processes the oral input. Processing may include hotword detection, conversion to text data, and identification of words and expressions corresponding to user-provided commands, requests, and / or parameters. In some implementations, this processing may be minimal or nonexistent. For example, this processing may include encoding the oral input audio for transmission to the server 114, or preparing captured raw audio of the oral input for transmission to the server 114.

[0123] The device sends a request containing information determined based on the verbal input to a remote system (506). The client device 104 / casting device 106 processes the verbal input and determines the request from the verbal input by identifying the request and one or more associated parameters from the verbal input. The client device 104 / casting device 106 sends the determined request to a remote system (e.g., server 114). The remote system determines and generates a response to the request. In some implementations, the client device 104 / casting device 106 sends the verbal input to the server 114 (e.g., as encoded audio, as raw audio data), and the server 114 processes the verbal input and determines the request and associated parameters.

[0124] The device receives a response to the request (508). The response may be generated by a remote system based on information from the verbal input. The remote system (e.g., server 114) determines and generates a response to the request and sends this response to the client device 104 / casting device 106.

[0125] The device performs an action in response (510). The client device 104 / casting device 106 performs one or more actions in response to the received response. For example, if the response is a command to cause the device to output specific information by audio, the client device 104 / casting device 106 retrieves this information, converts it to an audio output, and outputs the audio through the speaker. As another example, if the response is a command to cause the device to play media content, the client device 104 / casting device 106 retrieves the media content and plays the media content.

[0126] One or more of the aforementioned receiving, processing, transmitting, receiving, and executing actions are performed by one or more voice processing modules of a voice assistant library running on an electronic device, and the voice processing modules provide a plurality of voice processing actions that are accessible to one or more application programs and / or operating software running or executable on the electronic device (512). The client device 104 / casting device 106 may have a voice assistant client library 240 which includes functions and modules for performing one or more of the aforementioned receiving, processing, transmitting, receiving, and executing steps. The modules of the voice assistant client library 240 provide a plurality of voice processing actions and assistant actions that are accessible to applications, operating systems, and platform software in the client device 104 / casting device 106 which includes or links to the library 240 (for example, which executes the library 240 and associated APIs).

[0127] In some implementations, at least some speech processing operations associated with the speech processing module may be performed on a remote system connected to each other via a wide area network. For example, processing oral input to determine a request may be performed by a server 114 connected to a client device 104 / casting device 106 via a network(s) 112.

[0128] In some implementations, the voice assistant library can run on a common operating system that can operate on multiple different device types, thereby enabling the portability of voice-enabled applications configured to interact with one or more voice processing operations. The voice assistant client library 240 (and related libraries and APIs, e.g., debugging library 242, platform API 244, POSIX API 246) utilizes standard elements (e.g., objects) of a predefined operating system (e.g., Linux), so that it can run on various devices running different distributions or flavors of that predefined operating system (e.g., different Linux or Linux-based distributions or flavors). Thus, voice assistant functionality is available to various devices, and the voice assistant experience is consistent across those various devices.

[0129] In some implementations, requests and responses may be processed at the device level. For example, for basic functions that may be local to the device, such as timers, alarm clocks, clocks, and volume controls, the client device 104 / casting device 106 may process the verbal input, determine that the request corresponds to one of these basic functions, determine the response at the device level, and perform one or more actions in response to the response. The device may then report the request and response to the server 114 for logging purposes.

[0130] In some implementations, a device-independent voice assistant library for electronic devices with an audio input system includes one or more voice processing modules configured to run on a common operating system implemented on multiple different electronic device types, the voice processing modules providing multiple voice processing operations accessible to application programs and operating software running on the electronic device, thereby enabling the portability of voice-enabled applications configured to interact with one or more of these voice processing operations. The voice assistant client library 240 uses the same predefined operating system base as the library (for example, the operating systems of the library and the device are Linux-based). Because it is a library that can be run on various devices shared as such, this library is device-independent. Library 240 provides multiple modules for voice assistant functionality that allows applications to be accessed across various devices.

[0131] In some implementations, at least some speech processing operations associated with the speech processing module are performed on backend servers connected to each other via a wide area network with electronic devices. For example, library 240 includes a module that communicates with server 114, sends it to server 114 for processing spoken input, and determines the request.

[0132] In some implementations, the audio processing operations include device-specific operations configured to control devices connected to the electronic device (e.g., directly or communicatively). Library 240 may also include functions or modules for controlling other devices connected to the client device 104 / casting device 106 (e.g., wireless speakers, smart TVs, etc.).

[0133] In some implementations, the audio processing operation includes an information / media request operation configured to provide the requested information and / or media content to the user of the electronic device, or to provide it on a device connected to the electronic device (e.g., directly or communicatively). Library 240 may include functions or modules for retrieving the information or media and providing the information or media on the client device 104 / casting device 106 or on a connected device (e.g., reading an email aloud, reading a newspaper article aloud, playing streaming music).

[0134] Terms such as "first," "second," etc., may be used herein to describe various elements, but it will be understood that elements should not be limited by these terms. These terms are used merely to distinguish one element from another. For example, the first contact may be referred to as the second contact without changing the meaning of the description, only if the names of the first contact and the second contact are all changed without contradiction. The first contact and the second contact are both contacts, but they are not the same contact.

[0135] The terms used herein are for the sole purpose of describing specific implementations and are not intended to limit the scope of the claims. The singular forms “a,” “an,” and “the” used in descriptions of implementations and the appended claims are intended to include the plural form unless the context clearly indicates otherwise. The terms “and / or” used herein are understood to refer to and encompass any one or more of the items described relating to the invention, and all possible combinations thereof. The terms “comprises” and / or “comprising” are defined herein. When used in this context, it will be understood that the existence of the described features, integers, steps, actions, elements, and / or components is specifically mentioned, but not excluded from the existence or addition of one or more other features, integers, steps, actions, elements, components, and / or groups thereof.

[0136] As used herein, the term "if" means, depending on the context, "when," "upon," "in response to determining," "in accordance with a determination," or "in response to detecting" that the preceding condition is true. It can be interpreted as having a taste. Similarly, the expression "if the stated prior conditions are determined to be true (if "It is determined that a stated condition precedent is true," "if a stated condition precedent is true," and "when a stated condition precedent is true" can be interpreted, depending on the context, as meaning "upon determining," "in response to determining," "in accordance with a determination," "upon detecting," or "in response to detecting."

[0137] Various implementation configurations are referenced in detail, with examples shown in the accompanying drawings. In the following detailed description, many specific details are provided for a thorough understanding of the present invention and the described implementation configurations. However, the present invention can be carried out even without these specific details. In other cases, well-known methods, procedures, components, and circuits are not described in detail to avoid unnecessarily obscuring the nature of the implementation configurations.

[0138] The above description includes specific implementation examples for the sake of clarity. However, the above illustrative description is not intended to be exhaustive or to limit the invention to any strict form of disclosure. In view of the above teachings, many modifications and variations are possible. To enable those skilled in the art to make the most of the invention and various implementations using various modifications suitable for specific conceivable applications, the implementations have been selected and described to best illustrate the principles of the invention and their practical applications.

Claims

1. A method for an electronic device comprising an audio input system, one or more processors, and a memory storing one or more programs executed by the one or more processors, The process includes the step of downloading a voice assistant library containing multiple voice processing modules from a server, wherein the multiple voice processing modules are: Multiple modules configured to run on a common operating system implemented on multiple different types of electronic devices, and to provide multiple common speech processing operations accessible to application programs and operating systems running or executable on said different types of electronic devices, The method further includes one or more custom modules for providing device-specific voice processing operations configured to control a device connected to the electronic device, and the method further includes, The steps include implementing the voice assistant library to run on the electronic device by including it in an application installed on the electronic device, including it in the operating system of the electronic device, or embedding it in the firmware of the electronic device, The steps include receiving oral input in the aforementioned electronic device, The steps include processing the aforementioned oral input, The steps include sending a request to a remote system that includes information determined based on the aforementioned verbal input, The steps include receiving a response to the request generated by the remote system in accordance with the information based on the verbal input, A method comprising the step of performing an action in response to the response using one or more of the voice processing modules of the voice assistant library running on the electronic device.

2. The method according to claim 1, wherein at least some of the voice processing operations associated with the voice processing module are performed on the remote system which is connected to the electronic devices via a wide area network.

3. The method according to claim 1 or 2, wherein the voice assistant library is executable on a common operating system that can run on multiple different types of electronic devices, thereby enabling the portability of voice-enabled applications configured to utilize one or more of the voice processing operations.

4. A voice assistant library for multiple electronic devices, each equipped with an audio input system, It comprises one or more audio processing modules, and the one or more audio processing modules are It can run on a common operating system that can operate on multiple different types of electronic devices, The voice assistant library is configured to be implemented on each of the plurality of electronic devices such that it is included in an application installed on the electronic device, included in the operating system of the electronic device, or embedded in the firmware of the electronic device. It is configured to provide one or more common speech processing operations accessible to application programs and operating systems running or executable on the aforementioned multiple different types of electronic devices, The voice assistant library further comprises one or more custom modules for providing device-specific voice processing operations configured to control devices connected to the electronic device, The voice assistant library further comprises one or more application programming interfaces (APIs) configured to provide an interface between one or more voice processing operations and the hardware and / or software of the electronic device, A voice assistant library that enables the portability of the one or more voice processing modules and APIs between the multiple different electronic device types of voice-enabled applications configured to utilize the one or more voice processing operations.

5. The voice assistant library according to claim 4, wherein at least some voice processing operations associated with the voice processing module are performed on backend servers connected to each other with the electronic devices via a wide area network.

6. The voice assistant library according to claim 4 or 5, wherein the voice processing operation includes an information / media request operation configured to provide requested information and / or media content to the user of the electronic device or on a device connected to the electronic device.

7. An electronic device, wherein the electronic device is Audio input system and, One or more processors, The system comprises a memory that stores one or more programs executed by the one or more processors, and the one or more programs are The instructions include a voice assistant library containing multiple voice processing modules, which are configured to run on a common operating system implemented on multiple different electronic device types, and include multiple modules for providing multiple common voice processing operations accessible to application programs and operating systems running or executable on the different electronic device types, and one or more custom modules for providing device-specific voice processing operations configured to control devices connected to the electronic devices, the one or more programs further include Instructions for implementing the voice assistant library to run on the electronic device by being included in an application installed on the electronic device, included in the operating system of the electronic device, or embedded in the firmware of the electronic device, A command for receiving oral input in the aforementioned electronic device, Commands for processing the aforementioned oral input, A command for sending a request to a remote system, which includes information determined based on the aforementioned verbal input, A command generated by the remote system in response to the information based on the verbal input, for receiving a response to the request, An electronic device comprising one or more of the voice processing modules of the voice assistant library running on the electronic device, which include instructions for performing an action in response to the response.

8. The device according to claim 7, wherein at least some of the voice processing operations associated with the voice processing module are performed on the remote system which is connected to the electronic device via a wide area network.

9. The device according to claim 7 or 8, wherein the voice assistant library is executable on a common operating system that can run on multiple different types of electronic devices, thereby enabling the portability of voice-enabled applications configured to utilize one or more of the voice processing operations.

10. Audio input system and, One or more processors, An electronic device comprising: a memory storing one or more programs executed by the one or more processors, wherein the one or more programs include instructions for performing the method according to any one of claims 1 to 3.

11. A program that causes a computer to perform the method described in any one of claims 1 to 3.