Deliver compatible supplemental content via digital assistants

Through the digital assistant controlling the delivery of supplementary content, and using the unused display device space to provide supplementary content, it solves the problem that the digital assistant GUI only utilizes part of the screen space, achieving more efficient screen use and battery management.

CN115812193BActive Publication Date: 2025-05-06GOOGLE LLC
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202080102670.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2020-10-21
Publication Date
2025-05-06
Estimated Expiration
2040-10-21

AI Technical Summary

Technical Problem

In the prior art, the graphical user interface (GUI) of the digital assistant only utilizes part of the display device, resulting in wasting the remaining part, which in turn wastes the actual screen usage area.

Method used

Through the digital assistant, the transfer of supplementary content is controlled by using an unused part of the display device to provide supplementary content, satisfying the content parameters that reduce battery consumption of mobile computing devices and reducing empty display space.

Benefits of technology

Effectively utilize the remaining space of the display device, reduce battery consumption and computing resources, and improve the actual use efficiency of the screen.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115812193B_ABST
    Figure CN115812193B_ABST
Patent Text Reader

Abstract

Controlling the delivery of supplemental content via a digital assistant is provided. A system receives a data packet including voice input detected by a microphone of the computing device from a computing device. The system processes the data packet to generate an action data structure. The system generates content selection criteria for input into a content selector component to select supplemental content. The system sends a request to the content selector component to select supplemental content. The system receives a supplemental content item selected by the content selector component based on the content selection criteria generated by the digital assistant component. The system provides the action data structure in response to the voice input in a first graphical user interface slot via a display device. The system provides the supplemental content item in a second graphical user interface slot via the display device.
Need to check novelty before this filing date? Find Prior Art

Description

Background Art

[0001] Different computing devices may have different capabilities, such as different input or output interfaces or computing performance capabilities. Summary of the invention

[0002] The present disclosure generally relates to controlling the delivery of supplementary content via a digital assistant. The digital assistant may be invoked in response to a trigger word detected in a voice input. The digital assistant may process the voice input to generate an output provided via a graphical user interface ("GUI") outputted via a display device. However, the GUI of the digital assistant may utilize only a portion of the display device, while the remainder of the display device remains empty or unused by the digital assistant. Leaving the remainder of the display device empty results in wasted screen real estate. The present technical solution may control the delivery of supplementary content to be provided via a portion of the display device that is not used by the digital assistant GUI invoked in response to the voice input. The present technical solution may select supplementary content that meets content parameters configured to reduce battery consumption of a mobile computing device, thereby allowing the digital assistant to reduce wasted or empty display space without requiring excessive computation or energy utilization. The present technical solution may pre-process or filter supplementary content items based on content parameters to validate the supplementary content, thereby providing it via a GUI slot established by the digital assistant.

[0003] At least one aspect relates to a system for controlling the delivery of supplementary content via a digital assistant. The system may include a data processing system including a memory and one or more processors. The system may include a digital assistant component executed by the data processing system. The digital assistant component may receive a data packet from a computing device via a network. The data packet may include a voice input detected by a microphone of the computing device. The digital assistant component may process the data packet to generate an action data structure in response to the voice input. The digital assistant component may generate content selection criteria for input into a content selector component to select supplementary content provided by a third-party content provider. The content selection criteria may include one or more keywords and a digital assistant content type. The digital assistant component may send a request to the content selector component to select supplementary content based on the content selection criteria. The digital assistant component may receive, in response to the request, a supplementary content item selected by the content selector component based on the content selection criteria generated by the digital assistant component. The digital assistant component may provide an action data structure in response to the voice input in a first graphical user interface slot via a display device coupled to the computing device. The digital assistant component may provide a supplementary content item selected by the content selector component in a second graphical user interface slot via a display device coupled to the computing device.

[0004] At least one aspect relates to a method for controlling the delivery of supplementary content via a digital assistant. The method may be performed by a data processing system including a memory and one or more processors. The method may be performed by a digital assistant component executed by the data processing system. The method may include the data processing system receiving a data packet from a computing device via a network. The data packet may include a voice input detected by a microphone of the computing device. The method may include the data processing system processing the data packet to generate an action data structure in response to the voice input. The method may include the data processing system generating content selection criteria for input into a content selector component to select supplementary content provided by a third-party content provider. The content selection criteria may include one or more keywords and a digital assistant content type. The method may include the data processing system sending a request to the content selector component to select supplementary content based on the content selection criteria. The method may include the data processing system receiving, in response to the request, a supplementary content item selected by the content selector component based on the content selection criteria generated by the digital assistant component. The method may include the data processing system providing an action data structure in response to the voice input in a first graphical user interface slot via a display device coupled to the computing device. The method may include the data processing system providing the supplemental content item selected by the content selector component to be provided in a second graphical user interface slot via a display device coupled to the computing device.

[0005] These and other aspects and implementations are discussed in detail below. The foregoing information and the following detailed description include illustrative examples of the various aspects and implementations, and provide an overview or framework for understanding the nature and character of the claimed aspects and implementations. The accompanying drawings provide illustration and further understanding of the various aspects and implementations, and are incorporated into and constitute a part of this specification. BRIEF DESCRIPTION OF THE DRAWINGS

[0006] The drawings are not drawn to scale. The same reference numbers and names in different drawings represent the same elements. For clarity, not every component is labeled in every drawing. In the drawings:

[0007] Figure 1 is an illustration of an example system for controlling delivery of supplemental content via a digital assistant, according to one implementation;

[0008] Figure 2 is an illustration of an example user interface having slots according to one implementation;

[0009] Figure 3 is an illustration of an example method of controlling the delivery of supplemental content via a digital assistant, according to one implementation;

[0010] Figure 4is an illustration of an example method for validating supplemental content based on content type according to one implementation; and

[0011] Figure 5 is to illustrate that the systems and methods described and illustrated herein (including, for example, Figure 1 Depicted system, Figure 2 Depicted user interface and Figure 3 and Figure 4 A block diagram of the general architecture of a computer system that depicts the elements of the method. DETAILED DESCRIPTION

[0012] The following is a more detailed description of various concepts related to methods, apparatuses, and systems for controlling the delivery of supplemental content via a digital assistant and their implementation.The various concepts introduced above and discussed in more detail below can be implemented in any of a number of ways.

[0013] The present technical solution is generally directed to controlling the delivery of supplementary content via a digital assistant. A computing device with a display device can provide one or more graphical user interfaces ("GUIs") via the display device. The computing device can provide a main screen or main GUI. When the computing device receives a voice input, the computing device can process the voice input to detect a trigger keyword or hot word configured to call a digital assistant. When a trigger word or hot word is detected, the computing device can call a digital assistant to process the voice input and perform an action in response to the voice input. The digital assistant can process the voice input to generate an output. The digital assistant can provide the output via the digital assistant GUI. However, the digital assistant GUI may not consume the entire main screen or light track area (footprint) of the display device. That is, the digital assistant GUI can only utilize a portion of the display device, while the remaining portion of the display device or the main screen can remain empty or unused by the digital assistant. Leaving the remaining portion of the main screen or display device may result in a waste of the actual screen area. Even if this portion of the display is not used by the GUI of the digital assistant, the computing device may still consume battery resources or other resources to provide a portion of the main screen in this remaining space. Thus, the present technical solution can control the delivery of supplemental content to be provided via a portion of a display device that is not used by a digital assistant GUI invoked in response to voice input. The present technical solution can select supplemental content that meets content parameters established to reduce battery consumption of a mobile computing device, thereby allowing a digital assistant to reduce wasted or empty GUI space without requiring excessive computation or energy utilization.

[0014] For example, a digital assistant may generate a GUI to display output in response to voice input. The digital assistant GUI may be designed or constructed as a compact GUI that does not utilize the entire portion of the display device that is available for providing GUI output. In such a configuration, the digital assistant may not fully utilize the available digital screen real estate. The present technical solution may utilize the empty screen space generated by the compact digital assistant GUI to provide supplemental content that is related to or enhances the output of the digital assistant.

[0015] In an illustrative example, the voice input may include a query "jobs near me". A preprocessor or other component of the computing device may detect the input query and call a digital assistant to perform an action based on the input query. The digital assistant may engage in a conversation with the user, perform a search, or perform another action in response to the input query. However, such a conversation or action may not occupy the entire screen accessible to the computing device for output. The digital assistant may determine that there is wasted or unused screen space. In response to determining that there is wasted or unused screen space, the digital assistant may generate content selection criteria. The content selection criteria may be based on the input query, the action to be performed based on the input query, or an attribute associated with the computing device or the available screen space. When generating the content selection criteria, the digital assistant may send a content request based on the content selection criteria. The digital assistant may send a request for supplementary content to a content selector component configured to perform a real-time content selection process. In response to the request, the digital assistant may receive the selected supplementary content item. The digital assistant may provide supplementary content items to be provided or presented in a graphical user interface slot located or positioned in a portion of the screen that is not used by the digital assistant for a user conversation with the computing device.

[0016] The digital assistant can automatically create a GUI slot for the supplementary content item. The digital assistant can configure the GUI slot for the supplementary content item. The digital assistant can establish the size of the slot. The digital assistant can limit the type of supplementary content items that can be provided via the slot. For example, the digital assistant can limit the type of supplementary content items to assistant type content items. Assistant type content can refer to content items compatible with the slot generated by the assistant. The digital assistant can control the format of the supplementary content items provided via the slot. The format can include, for example, application installation content, image content, video content, or audio content. The digital assistant can establish actions for the slot, such as fixing the supplementary content item to the home screen, muting the slot, adjusting the parameters of the slot, or otherwise interacting with the slot. The digital assistant can configure the slot based on a quality signal or content parameter. Content parameters can include brightness level, file size, number of frames, resolution, or audio output level. In some cases, the digital assistant can generate content selection criteria based on one or more attributes or configurations of the input query and the slot using a machine learning model or engine.

[0017] Therefore, by allowing the digital assistant to automatically establish and configure GUI slots for supplementary content items that meet content parameters, the present technical solution can reduce wasted screen space without using too many computing resources. Content parameters can be configured to reduce computing and energy consumption by, for example, limiting the brightness level or file size of supplementary content items. Although the selection of supplementary content items can occur in parallel through separate systems or cloud-based computing infrastructures, the server digital assistant can control the delivery of supplementary content items based on the delivery of action data structures in response to voice input. For example, the server digital assistant can combine or package the supplementary content items with the action data structure to reduce data transmission on the network. The server digital assistant can add a buffer to the delivery so that the action data structure in response to voice input is presented before the supplementary content item, thereby avoiding the introduction of delays or delays when presenting the action data structure. In addition, by automatically generating requests for supplementary content, the digital assistant can actively enhance or improve the action data structure without the need for separate downstream requests.

[0018] Figure 1 An example system 100 for controlling the delivery of supplementary content via a digital assistant is shown. The system 100 may include a content selection infrastructure. The system 100 may include a data processing system 102. The data processing system 102 may communicate with one or more of a computing device 140, a service provider device 154, or a supplementary digital content provider device 152 via a network 105. The network 105 may include a computer network such as the Internet, a local area network, a wide area network, a metropolitan area network, or other regional network, an intranet, a satellite network, and other communication networks such as a voice or data mobile phone network. The network 105 may be used to access information resources such as web pages, websites, domain names, or uniform resource locators, which may be provided, rendered, presented, or displayed on at least one local computing device 140 (such as a laptop, desktop computer, tablet, digital assistant device, smart phone, mobile telecommunication device, portable computer, or speaker). For example, via the network 105, a user of the local computing device 140 may access information or data provided by the supplementary digital content provider device 152. The computing device 140 may or may not include a display; for example, the computing device may include a limited type of user interface, such as a microphone and a speaker. In some cases, the primary user interface of computing device 140 may be a microphone and speaker, or a voice interface. In some cases, computing device 140 includes a display device 150 coupled to computing device 140 , and the primary user interface of computing device 140 may utilize display device 150 .

[0019] A local computing device 140 may refer to a computing device 140 used by a user or owned by a user. A local computing device 140 may refer to a computing device or client device located in a public place such as a hotel, office, restaurant, retail store, mall, park, or a private place such as a residence. The term "local" may refer to a computing device located where a user can interact with the computing device using voice input or other input. The local computing device may be remote from a remote server such as data processing system 102. Thus, a local computing device 140 may be located in a hotel room, mall, den, or other building or residence where a user can interact with the local computing device 140 using voice input, while data processing system 102 may be remotely located, for example, in a data center. A local computing device 140 may be referred to as a digital assistant device.

[0020] Network 105 may include or constitute a display network, such as a subset of information resources available on the Internet associated with a content placement or search engine results system, or a subset of information resources that are eligible to include a third-party digital component as part of a digital component placement campaign. Data processing system 102 may use network 105 to access information resources, such as web pages, websites, domain names, or uniform resource locators that may be provided, output, submitted, or displayed by local client computing device 140. For example, via network 105, a user of local client computing device 140 may access information or data provided by supplemental digital content provider device 152 or service provider computing device 154.

[0021] The network 105 may be any type or form of network, and may include any of the following: a point-to-point network, a broadcast network, a wide area network, a local area network, a telecommunications network, a data communications network, a computer network, an ATM (asynchronous transfer mode) network, a SONET (synchronous optical network) network, an SDH (synchronous digital hierarchy) network, a wireless network, and a wired network. The network 105 may include a wireless link, such as an infrared channel or a satellite frequency band. The topology of the network 105 may include a bus, star, or ring network topology. The network may include a mobile phone network using any protocol for communicating between mobile devices, including the Advanced Mobile Phone Protocol ("AMPS"), Time Division Multiple Access ("TDMA"), Code Division Multiple Access ("CDMA"), Global System for Mobile Communications ("GSM"), General Packet Radio Service ("GPRS"), or Universal Mobile Telecommunications System ("UMTS"). Different types of data may be transmitted via different protocols, or the same type of data may be transmitted via different protocols.

[0022] System 100 may include at least one data processing system 102. Data processing system 102 may include at least one logical device, such as a computing device having a processor that communicates with, for example, computing device 140, supplementary digital content provider device 152 (or third-party content provider device, content provider device) or service provider device 154 (or third-party service provider device) via network 105. Data processing system 102 may include at least one computing resource, server, processor or memory. For example, data processing system 102 may include multiple computing resources or servers located in at least one data center. Data processing system 102 may include multiple logically grouped servers and facilitate distributed computing technology. Logical groups of servers may be referred to as data centers, server clusters or machine clusters. Servers may also be geographically dispersed. A data center or machine cluster may be managed as a single entity, or a machine cluster may include multiple machine clusters. The servers in each machine cluster may be heterogeneous, that is, one or more of the servers or machines may operate according to one or more types of operating system platforms.

[0023] The servers in the machine cluster can be stored in a high-density rack system along with the associated storage systems and located in an enterprise data center. For example, merging the servers in this manner can improve system manageability, data security, physical security of the system, and system performance by locating the servers and high-performance storage systems on a local high-performance network. Centralization of all or some of the data processing system 102 components, including servers and storage systems, and coupling them with advanced system management tools, allows for more efficient use of server resources, which saves power and processing requirements, and reduces bandwidth usage.

[0024] System 100 may include, access, or otherwise interact with at least one third-party device, such as service provider device 154 or supplemental digital content provider device 152. Service provider device 154 may include at least one logical device, such as a computing device having a processor that communicates with, for example, computing device 140, data processing system 102, or supplemental digital content provider device 152 via network 105. Service provider device 154 may include at least one computing resource, server, processor, or memory. For example, service provider device 154 may include a plurality of computing resources or servers located in at least one data center.

[0025] The supplementary digital content provider device 152 can provide an audio-based digital component for display as an audio output digital component by the local computing device 140. The digital component can be referred to as a sponsored digital component because it is provided by a third-party sponsor. The digital component can include an order for goods or services, such as a voice-based message stating: "Do you want me to call a taxi for you?" For example, the supplementary digital content provider device 152 can include a memory storing a series of audio digital components, which can be provided in response to voice-based queries. The supplementary digital content provider device 152 can also provide an audio-based digital component (or other digital component) to the data processing system 102, in which these audio-based digital components can be stored in the data repository 124. The data processing system 102 can select the audio digital component and provide (or instruct the supplementary digital content provider device 152 to provide) these audio digital components to the client computing device 140. The audio-based digital component can be audio only or can be combined with text, image or video data.

[0026] The service provider device 154 may include, interface with, or otherwise communicate with the data processing system 102. The service provider device 154 may include, interface with, or otherwise communicate with the local computing device 140. The service provider device 154 may include, interface with, or otherwise communicate with the computing device 140, which may be a mobile computing device. The service provider device 154 may include, interface with, or otherwise communicate with the supplemental digital content provider device 152. For example, the service provider device 154 may provide the digital component to the local computing device 140 for execution by the local computing device 140. The service provider device 154 may provide the digital component to the data processing system 102 for storage by the data processing system 102. The service provider device 154 may provide the rules or parameters associated with the digital component to the data processing system 102 for storage in the content data 126 data structure.

[0027] The local computing device 140 may include, interface with or otherwise communicate with at least one sensor 144, transducer 146, audio driver 148, or local digital assistant 142. The local computing device 140 may include a display device 150, such as a light indicator, a light emitting diode ("LED"), an organic light emitting diode ("OLED"), or other visual indicator configured to provide a visual or optical output. The sensor 144 may include, for example, an ambient light sensor, a proximity sensor, a temperature sensor, an accelerometer, a gyroscope, a motion detector, a GPS sensor, a position sensor, a microphone, or a touch sensor. The transducer 146 may include a speaker or a microphone. The audio driver 148 may provide a software interface to the hardware transducer 146. The audio driver may execute an audio file or other instructions provided by the data processing system 102 to control the transducer 146 to generate a corresponding sound wave or sound waves.

[0028] The local digital assistant 142 may include one or more processors (e.g., processor 510), logic arrays, or memory. The local digital assistant 142 may detect keywords and perform actions based on the keywords. The local digital assistant 142 may filter out one or more terms or modify the terms before transmitting the terms to the data processing system 102 (e.g., server digital assistant component 104) for further processing. The local digital assistant 142 may convert analog audio signals detected by the microphone into digital audio signals and transmit one or more data packets carrying the digital audio signals to the data processing system 102 via the network 105. In some cases, the local digital assistant 142 may transmit data packets carrying some or all of the input audio signal in response to detecting an instruction to perform such a transmission. The instruction may include, for example, a trigger keyword or other keyword or approval to transmit the data packet including the input audio signal to the data processing system 102.

[0029] The local digital assistant 142 can pre-filter or pre-process the input audio signal to remove certain frequencies of the audio. Pre-filtering can include filters such as low-pass filters, high-pass filters, or band-pass filters. The filter can be applied in the frequency domain. The filter can be applied using digital signal processing techniques. The filter can be configured to maintain frequencies corresponding to human voice or human speech while eliminating frequencies that fall outside the typical frequencies of human speech. For example, a band-pass filter can be configured to remove frequencies below a first threshold (e.g., 70Hz, 75Hz, 80Hz, 85Hz, 90Hz, 95Hz, 100Hz, or 105Hz) and above a second threshold (e.g., 200Hz, 205Hz, 210Hz, 225Hz, 235Hz, 245Hz, or 255Hz). Applying a band-pass filter can reduce computing resource utilization in downstream processing. In some cases, the local digital assistant 142 on the computing device 140 can apply a band-pass filter before transmitting the input audio signal to the data processing system 102, thereby reducing network bandwidth utilization. However, based on the computing resources available to computing device 140 and the available network bandwidth, it may be more efficient to provide the input audio signal to data processing system 102 to allow data processing system 102 to perform filtering.

[0030] The local digital assistant 142 may apply additional pre-processing or pre-filtering techniques, such as noise reduction techniques, to reduce ambient noise levels that may interfere with the natural language processor. The noise reduction techniques may improve the accuracy and speed of the natural language processor, thereby improving the performance of the data processing system 102 and managing the presentation of a graphical user interface provided via the display device 150. The local digital assistant 142 may filter the input audio signal to create a filtered input audio signal, convert the filtered input audio signal into a data packet, and transmit the data packet to a data processing system including one or more processors and memory.

[0031] The local digital assistant 142 may determine to call or launch an application on the computing device 140. The local digital assistant 142 may receive instructions or commands from the server digital assistant component 104 to call or launch an application on the computing device 140. The local digital assistant 142 may receive a deep-link or other information to facilitate launching of the application on the computing device 140 or execution of the application on the computing device 140.

[0032] The local client computing device 140 may be associated with an end user who enters a voice query as audio input into the local client computing device 140 (via the sensor 144) and receives audio output in the form of computer-generated speech, which may be provided to the local client computing device 140 from the data processing system 102 (or the supplemental digital content provider device 152 or the service provider computing device 154) and output from the transducer 146 (e.g., a speaker). The computer-generated speech may include a recording of speech from a real person or generated by a computer.

[0033] The data repository 124 may include one or more local or distributed databases, and may include a database management system. The data repository 124 may include a computer data storage device or memory, and may store one or more of content data 126, attributes 128, content parameters 130, templates 156, or accounts 158. The content data 126 may include or refer to information related to supplementary content items or sponsored content items. The content data may include content items, digital component objects, content selection criteria, content types, identifiers of content providers, keywords, or other information associated with the content items, which the content selector component 120 may use to perform a real-time content selection process.

[0034] Attributes 128 may include or refer to qualities, features, or characteristics associated with a GUI slot. Example attributes may include the dimensions of the GUI slot, the location of the GUI slot. Attributes 128 of a GUI slot may define aspects of a sponsored content item provided via the GUI slot, such as a brightness level, a file size of the sponsored content item, a duration of the sponsored content item (e.g., video or audio duration).

[0035] Content parameters 130 can point to a standard quality signal for a certain type of sponsored content item or supplementary content item. The type of supplementary content item can include digital assistant, search, context, video or streamlining. Different types of sponsored content items can include different content parameters, which are configured to improve the presentation of content items in a certain type of slot or in a certain type of computing device. Content parameters can include, for example, brightness level, file size, processor utilization, duration or other quality signals. For example, the content parameters of a digital assistant content item can include a first brightness threshold value that is lower than a second brightness threshold value set for a contextual content item.

[0036] Template 156 may include fields in a structured data set that may be populated by direct action API 110 to facilitate an operation requested via input audio. Template 156 data structure may include different types of templates for different actions. Account 158 ​​may include information associated with an electronic account. An electronic account may be associated with computing device 140 or a user thereof. Account 158 ​​may include, for example, an identifier, historical network activity, preferences, profile information, or other information that may be helpful in generating an action data structure or selecting a supplemental content item.

[0037] The data processing system 102 may include a content placement system having at least one computing resource or server. The data processing system 102 may include, interface with, or otherwise communicate with at least one interface 106. The data processing system 102 may include, interface with, or otherwise communicate with at least one natural language processor component 108. The data processing system 102 may include, interface with, or otherwise communicate with at least one direct action application programming interface (“API”) 110. The data processing system 102 may include, interface with, or otherwise communicate with at least one query generator component 112. The data processing system 102 may include, interface with, or otherwise communicate with at least one socket injector component 114. Data processing system 102 may include, interface with, or otherwise communicate with at least one transmit controller component 116. Interface 106, natural language processor component 108, direct action API 110, query generator component 112, slot injector component 114, or transmit controller component 116 may form a server digital assistant component 104. Data processing system 102 may include, interface with, or otherwise communicate with at least one server digital assistant component 104. Server digital assistant component 104 may communicate or interface with one or more voice-based interfaces or various digital assistant devices or surfaces to provide data or receive data or perform other functions.

[0038] The data processing system 102 may include, interface with, or otherwise communicate with at least one verification component 118. The data processing system 102 may include, interface with, or otherwise communicate with at least one content selector component 120. The data processing system 102 may include, interface with, or otherwise communicate with at least one data repository 124.

[0039] The server digital assistant component 104, interface 106, NLP component 108, direct action API 110, query generator component 112, slot injector component 114, transmission controller component 116, verification component 118, or content selector component 120 may each include at least one processing unit or other logic device, such as a programmable logic array engine, or a module configured to communicate with a data repository 124 or database. The server digital assistant component 104, interface 106, NLP component 108, direct action API 110, query generator component 112, slot injector component 114, transmission controller component 116, verification component 118, content selector component 120, and data repository 124 may be separate components, a single component, or part of the data processing system 102. The system 100 and its components, such as the data processing system 102, may include hardware elements, such as one or more processors, logic devices, or circuits.

[0040] The data processing system 102 can obtain anonymous computer network activity information associated with multiple local computing devices 140 (or computing devices or digital assistant devices). The user of the local computing device 140 or the mobile computing device can affirmatively authorize the data processing system 102 to obtain the network activity information corresponding to the local computing device 140 or the mobile computing device. For example, the data processing system 102 can prompt the user of the computing device 140 to agree to obtain one or more types of network activity information. The local computing device 140 may include a mobile computing device such as a smart phone, a tablet, a smart watch, or a wearable device. The identity of the user of the local computing device 140 can remain anonymous, and the computing device 140 can be associated with a unique identifier (e.g., a unique identifier of the user or computing device provided by the user of the data processing system or computing device). The data processing system can associate each observation with a corresponding unique identifier.

[0041] The data processing system 102 may include an interface 106 (or interface component) designed, configured, constructed, or operated to receive and send information using, for example, data packets. The interface 106 may receive and send information using one or more protocols, such as a network protocol. The interface 106 may include a hardware interface, a software interface, a wired interface, or a wireless interface. The interface 106 may facilitate the conversion or formatting of data from one format to another. For example, the interface 106 may include an application programming interface that includes definitions for communicating between various components, such as software components. The interface 106 may communicate with one or more of the local computing device 140, the supplemental digital content provider device 152, or the service provider device 154 via the network 105.

[0042] The data processing system 102 may interface with an application, script, or program (such as an app) installed on the local client computing device 140 to transmit the input audio signal to the interface 106 of the data processing system 102 and drive the components of the local client computing device to submit the output audio signal. The data processing system 102 may receive a data packet or other signal including or identifying an audio input signal.

[0043] The data processing system 102 or the server digital assistant component 104 may include a natural language processor ("NLP") component 108. For example, the data processing system 102 may execute or run the NLP component 108 to receive or obtain an audio signal and parse the audio signal. For example, the NLP component 108 may provide interaction between humans and computers. The NLP component 108 may be configured with techniques for understanding natural language and allowing the data processing system 102 to derive meaning from human or natural language input. The NLP component 108 may include or be configured with techniques based on machine learning (such as statistical machine learning). The NLP component 108 may utilize a decision tree, a statistical model, or a probabilistic model to parse the input audio signal. The NLP component 108 can perform, for example, such as named entity recognition (e.g., given a stream of text, determine which items in the text map to appropriate names, such as people or places, and what the type of each such name is, such as person, location, or organization), natural language generation (e.g., converting information from a computer database or semantic intent into understandable human language), natural language understanding (e.g., converting text into a more formal representation, such as a first-order logic structure that a computer module can manipulate), machine translation (e.g., automatically translating text from one human language to another), morpheme segmentation (e.g., breaking words into individual morphemes and identifying categories of morphemes, which is challenging based on considerations of the morphological or structural complexity of the words of a language), question answering (e.g., determining answers to human language questions, which can be specific or open-ended), and semantic processing (e.g., processing that can occur after words are identified and their meanings are encoded in order to associate the identified words with other words of similar meaning).

[0044] The NLP component 108 can convert the audio input signal into recognized text by comparing the input signal with a representative audio waveform set stored (e.g., in a data repository 124) and selecting the closest match. The audio waveform set can be stored in the data repository 124 or other databases accessible to the data processing system 102. The representative waveform is generated across a large group of users and can then be expanded with speech samples from the users. After the audio signal is converted into recognized text, the NLP component 108 uses the actions that the data processing system 102 can provide, such as through cross-user training or by manual specification, to match the text with associated words. Various aspects or functions of the NLP component 108 can be performed by the data processing system 102 or a local computing device 140. For example, a local NLP component can be executed on a local computing device 140 to perform various aspects of converting the input audio signal into text and transmitting the text to the data processing system 102 via a data packet for further natural language processing.

[0045] The audio input signal may be detected by a sensor 144 or transducer 146 (e.g., a microphone) of the local client computing device 140. Via the transducer 146, the audio driver 148, or other components, the local client computing device 140 may provide the audio input signal to the data processing system 102 (e.g., via the network 105), where it may be received (e.g., through the interface 106) and provided to the NLP component 108 or stored in the data repository 124.

[0046] The local computing device 140 may include an audio driver 148, a transducer 146, a sensor 144, and a local digital assistant 142. The sensor 144 may receive or detect an input audio signal (e.g., a voice input). The local digital assistant 142 may be coupled to the audio driver, the transducer, and the sensor. The local digital assistant 142 may filter the input audio signal (e.g., by removing certain frequencies or suppressing noise) to create a filtered input audio signal. The local digital assistant 142 may convert the filtered input audio signal into a data packet (e.g., using a software or hardware digital-to-analog converter). In some cases, the local digital assistant 142 may convert the unfiltered input audio signal into a data packet and transmit the data packet to the data processing system 102. The local digital assistant 142 may transmit the data packet to the data processing system 102, which includes one or more processors and memory that execute a natural language processor component, an interface, a speaker recognition component, and a direct action application programming interface.

[0047] The data processing system 102 may receive a data packet from the local digital assistant 142 via an interface, the data packet including a filtered (or unfiltered) input audio signal detected by a sensor. The data processing system 102 may process the data packet to perform an action or otherwise respond to the voice input. In some cases, the data processing system 102 may identify an acoustic signature from the input audio signal. The data processing system 102 may identify an electronic account corresponding to the acoustic signature based on a lookup in a data repository (e.g., querying a database). In response to the identification of the electronic account, the data processing system 102 may establish a session and an account to be used in the session. The account may include a profile with one or more policies. The data processing system 102 may parse the input audio signal to identify a request and a trigger keyword corresponding to the request.

[0048] The data processing system 102 may provide the status to a local digital assistant 142 of the local computing device 140. The local computing device 140 may receive an indication of the status. The audio driver may receive an indication of the status of the profile and generate an output signal based on the indication. The audio driver may convert the indication into an output signal, such as a sound signal or an acoustic output signal. The audio driver may drive a transducer 146 (e.g., a speaker) to generate sound based on the output signal generated by the audio driver.

[0049] In some cases, local computing device 140 may include a light source. The light source may include one or more LEDs, lights, displays, or other components or devices configured to provide an optical or visual output. Local digital assistant 142 may cause the light source to provide a visual indication corresponding to the state. For example, the visual indication may be a status indicator light that turns on, a change in light color, a light pattern having one or more colors, or a visual display of text or an image.

[0050] The NLP component 108 may obtain an input audio signal. In response to the local digital assistant 142 detecting a trigger keyword, the NLP component 108 of the data processing system 102 may receive a data packet having a voice input or an input audio signal. The trigger keyword may be a wake-up signal or a hot word that instructs the local computing device 140 to convert the subsequent audio input into text and transmit the text to the data processing system 102 for further processing.

[0051] After receiving the input audio signal, the NLP component 108 can identify at least one request or at least one keyword corresponding to the request. The request can indicate the intention or subject matter of the input audio signal. The keyword can indicate the type of action that may be taken. For example, the NLP component 108 can parse the input audio signal to identify at least one request to leave home for dinner and watch a movie at night. The trigger keyword can include at least one word, phrase, root or partial word or derivative indicating the action to be taken. For example, the trigger keyword "go" or "to go to" from the input audio signal can indicate that transportation is needed. In this example, the input audio signal (or the identified request) does not directly express the intention of transportation, but the trigger keyword indicates that the transport is an auxiliary action of at least one other action indicated by the request. In another example, the voice input can include a search query, such as "find jobs near me".

[0052] The NLP component 108 can parse the input audio signal to identify, determine, retrieve or otherwise obtain a request and one or more keywords associated with the request. For example, the NLP component 108 can apply semantic processing techniques to the input audio signal to identify keywords or requests. The NLP component 108 can apply semantic processing techniques to the input audio signal to identify keywords or phrases including one or more keywords (such as a first keyword and a second keyword). For example, the input audio signal may include the sentence "I want to buy an audio book." The NLP component 108 can apply semantic processing techniques or other natural language processing techniques to the data packet including the sentence to identify the keywords or phrases "want to buy" and "audio book." The NLP component 108 can further identify multiple keywords, such as buy and audio book. For example, the NLP component 108 can determine that the phrase includes a first keyword and a second keyword.

[0053] The NLP component 108 may filter the input audio signal to identify trigger keywords. For example, a data packet carrying the input audio signal may include "It would be great if I could find someone who could help me get to the airport," in which case the NLP component 108 may filter out one or more of the following terms: "if," "I," "can," "find," "can," "help," "of," "people," "just," or "great." By filtering out these terms, the NLP component 108 may more accurately and reliably identify trigger keywords, such as "go to the airport," and determine that this is a request for a taxi or ride sharing service.

[0054] In some cases, the NLP component may determine that the data packet carrying the input audio signal includes one or more requests. For example, the input audio signal may include the statement "I would like to purchase an audio book and a monthly subscription to movies." The NLP component 108 may determine that this is a request for an audio book and a streaming multimedia service. The NLP component 108 may determine whether this is a single request or multiple requests. The NLP component 108 may determine that these are two requests: a first request to a service provider that provides audio books, and a second request to a service provider that provides movie streaming. In some cases, the NLP component 108 may combine multiple determined requests into a single request and transmit the single request to the service provider device 154. In some cases, the NLP component 108 may transmit separate requests to another service provider device, or transmit two requests separately to the same service provider device 154.

[0055] The data processing system 102 may include a direct action API 110 that is designed and constructed to generate an action data structure responsive to a request based on one or more keywords in the voice input. The processor of the data processing system 102 may call the direct action API 110 to execute a script that generates a data structure to be provided to the service provider device 154 or other service provider to obtain a digital component, subscription service, or product (such as a car or audiobook from a car share service). The direct action API 110 may obtain data from the data repository 124, as well as data received from the local client computing device 140 with the consent of the end user, to determine location, time, user account, logistical, or other information to allow the service provider device 154 to perform an operation, such as booking a car from a car share service. Using the direct action API 110, the data processing system 102 may also communicate with the service provider device 154 to complete the conversion by making a car share ride reservation in this example.

[0056] The direct action API 110 can perform a specified action to satisfy the end user's intent, as determined by the data processing system 102. The action can include performing a search using a search engine, launching an application, ordering a product or service, providing requested information, or controlling a network-connected device (e.g., an Internet of Things device). Depending on the action specified in its input and the parameters or rules in the data repository 124, the direct action API 110 can execute code or a dialog script that identifies the parameters required to satisfy the user's request. Such code can, for example, look up additional information in the data repository 124 (such as the name of a home automation service or a third-party service), or it can provide an audio output for presentation at the local client computing device 140 to inquire the end user, such as the intended destination of the requested taxi. The direct action API 110 can determine the parameters and can package the information into an action data structure, which can then be sent to another component such as the content selector component 120 or the service provider computing device 154 to be implemented.

[0057] The direct action API 110 may receive instructions or commands from the NLP component 108 or other components of the data processing system 102 to generate or construct an action data structure. The direct action API 110 may determine the type of action in order to select a template from the templates 156 stored in the data repository 124. The type of action may include, for example, services, products, reservations, tickets, multimedia content, audiobooks, managing subscriptions, adjusting subscriptions, transferring digital currency, purchasing or music. The type of action may also include the type of service or product. For example, the type of service may include a car sharing service, a meal delivery service, a laundry service, a maid service, a repair service, a housekeeping service, an equipment automation service, or a media streaming service. The type of product may include, for example, clothes, shoes, toys, electronic products, computers, books, or jewelry. The type of reservation may include, for example, a dinner reservation or a salon appointment. The type of ticket may include, for example, a movie ticket, a stadium ticket, or an airline ticket. In some cases, the type of service, product, reservation, or ticket may be classified based on price, location, transportation type, availability, or other attributes.

[0058] The NLP component 108 can parse the input audio signal to identify a request and a trigger keyword corresponding to the request, and provide the request and the trigger keyword to the direct action API 110 so that the direct action API generates a first action data structure in response to the request based on the trigger keyword. After identifying the type of the request, the direct action API 110 can access the corresponding template in the template 156. The template can include fields in the structured data set that can be filled by the direct action API 110 to advance the operation requested by the input audio detected by the local computing device 140 of the service provider device 154 (such as dispatching a taxi to pick up the end user at the pickup location and transport the end user to the destination location). The direct action API 110 can perform a search in the template 156 to select a template that matches the trigger keyword and one or more characteristics of the request. For example, if the request corresponds to a request for a car or ride to a destination, the data processing system 102 can select a car sharing service template. The car sharing service template can include one or more of the following fields: device identifier, pickup location, destination location, number of passengers, or service type. The direct action API 110 can fill the field with a value. To populate the fields with values, the direct action API 110 may ping, poll, or otherwise obtain information from one or more sensors 144 of the computing device 140 or a user interface of the device 140. For example, the direct action API 110 may use a location sensor such as a GPS sensor to detect the source location. The direct action API 110 may obtain additional information by submitting a survey, prompt, or query to an end user of the computing device 140. The direct action API may submit the survey, prompt, or query via the interface 106 of the data processing system 102 and a user interface (e.g., an audio interface, a voice-based user interface, a display, or a touch screen) of the computing device 140. Thus, the direct action API 110 may select a template for an action data structure based on a trigger keyword or request, populate one or more fields in the template with information detected by one or more sensors 144 or obtained via a user interface, and generate, create, or otherwise construct an action data structure to facilitate the service provider device 154 to perform an operation.

[0059] To construct or generate an action data structure, the data processing system 102 may identify one or more fields in the selected template to fill with values. These fields may be filled with numerical values, strings, Unicode values, Boolean logic, binary values, hexadecimal values, identifiers, location coordinates, geographic regions, timestamps, or other values. The fields or the data structure itself may be encrypted or masked to maintain data security.

[0060] After determining the fields in the template, the data processing system 102 can identify values ​​for the fields to populate the fields of the template to create an action data structure. The data processing system 102 can obtain, retrieve, determine or otherwise identify values ​​for the fields by performing a lookup or other query operation on the data repository 124.

[0061] In some cases, data processing system 102 may determine that information or values ​​for a field do not exist in data repository 124. Data processing system 102 may determine that information or values ​​stored in data repository 124 are out-of-date, stale, or inappropriate for the purpose of constructing an action data structure responsive to the trigger keywords and requests identified by NLP component 108 (e.g., the location of local client computing device 140 may be an old location instead of the current location; an account may have expired; the target restaurant may have moved to a new location; physical activity information; or transportation methods).

[0062] If data processing system 102 determines that it does not currently have access to the value or information for a field in the memory of data processing system 102 for the template, data processing system 102 may obtain the value or information. Data processing system 102 may obtain or acquire the information by querying or polling one or more available sensors of local client computing device 140, prompting the end user of local client computing device 140 for the information, or accessing an online network-based resource using HTTP protocol. For example, data processing system 102 may determine that it does not have the current location of local client computing device 140, which may be a required field for the template. Data processing system 102 may query local client computing device 140 for location information. Data processing system 102 may request local client computing device 140 to provide location information using one or more location sensors 144 (such as global positioning system sensors, WIFI triangulation, cellular tower triangulation, Bluetooth beacons, IP addresses, or other location sensing technologies).

[0063] In some cases, data processing system 102 may use a second profile to generate an action data structure. Data processing system 102 may then determine whether the action data structure generated using the second profile complies with the first profile. For example, the first profile may include a policy that blocks a type of action data structure (such as purchasing a product from an electronic online retailer via a local computing device 140). The input audio detected by local computing device 140 may have included a request to purchase a product from an electronic online retailer. Data processing system 102 may have used a second profile to identify account information associated with an electronic online retailer, and then generated an action data structure to purchase the product. The action data structure may include an account identifier corresponding to an electronic account, which is associated with an acoustic signature identified by data processing system 102.

[0064] The data processing system 102 can identify an account associated with a user providing voice input. The data processing system can receive an audio input signal detected by a local computing device 140, identify an acoustic signature, and identify an electronic account corresponding to the acoustic signature. The data processing system 102 can identify an electronic account corresponding to the acoustic signature based on a search in an account 158 ​​in a data repository 124. The data processing system 102 can access an acoustic signature stored in a data repository 124 (such as in an account 158). The data processing system 102 can be configured with one or more speaker recognition techniques, such as pattern recognition. The data processing system 102 can be configured with a text-independent speaker recognition process. In a text-independent speaker recognition process, the text used to establish an electronic account can be different from the text used to identify a speaker later. Therefore, the data processing system 102 can perform speaker recognition or speech recognition to identify an electronic account corresponding to the signature of the input audio signal.

[0065] For example, the data processing system 102 can identify acoustic features that differ between input speech sources in the input audio signal. Acoustic features can reflect physical or learned patterns that can correspond to unique input speech sources. Acoustic features can include, for example, pitch or speaking style. Techniques for identifying, processing, and storing signatures can include frequency estimation (e.g., instantaneous fundamental frequency or discrete energy separation algorithms), hidden Markov models (e.g., stochastic models for modeling randomly changing systems, where future states depend on current states, and where the system is modeled as having unobserved states), Gaussian mixture models (e.g., parameter probability density functions represented as weighted sums of Gaussian component densities), pattern matching algorithms, neural networks, matrix representations, vector quantization (e.g., quantization techniques in signal processing that allow probability density functions to be modeled by the distribution of prototype vectors), or decision trees. Other techniques can include anti-speaker techniques, such as cohort models and world models. The data processing system 102 can be configured with machine learning models to facilitate pattern recognition or adapt to speaker characteristics.

[0066] The data processing system 102 may include a query generator component 112 that is designed, constructed, and operated to generate queries or content selection criteria based on speech input. The query generator component 112 may be part of or interact with the server digital assistant component 104. For example, the server digital assistant component 104 may include a query generator component 112. The query generator component 112 may generate content selection criteria for input to the content selector component 120. The content selector component 120 may use the content selection criteria to select supplementary content provided by a third-party content provider (e.g., a supplementary digital content provider device 152). The query generator component 112 may generate content selection criteria having one or more keywords and a digital assistant content type.

[0067] The query generator component 112 may receive from the NLP component 108 an indication of one or more keywords or requests identified by the NLP component 108 in an input audio signal received from the computing device 140. The query generator component 112 may process the one or more keywords and requests to determine whether to generate a request for supplemental content from the supplemental digital content provider device 152. Supplemental content may refer to or include sponsored content. For example, a supplemental digital content provider may bid on a content item to be selected in a real-time content selection process to be provided to a user via the computing device 140.

[0068] The query generator component 112 can generate content selection criteria based on one or more keywords identified in the speech input. The content selection criteria can include keywords in the speech input. The query generator component 112 can further expand or broaden the keywords to select additional keywords.

[0069] The query generator component 112 can identify additional content selection criteria based on the profile information stored in the account 158. For example, the query generator component 112 can receive an indication of an account associated with the user or voice input and retrieve account information, such as historical information associated with network activity (e.g., search history, clicks, or selections performed). The query generator component 112 can generate keywords or other content selection criteria based on the historical information associated with the network activity. For example, if the search history includes a search query for "jobs near me", the query generator component 112 can generate keywords based on the term "jobs" and keywords based on the location of the computing device 140.

[0070] The query generator component 112 can access information associated with the sensors 144 of the computing device 140 to generate the content selection criteria. The query generator component 112 can ping or poll one or more sensors 144 of the computing device 140 to obtain information such as location information, motion information, or any other information that can help generate the content selection criteria.

[0071] The query generator component 112 can interrogate or poll the computing device 140 to obtain information associated with the computing device 140, such as device configuration information. The device configuration information can include the type of device, available user interfaces of the device, remaining battery information, network connection information, display device 150 information (e.g., the size of the display or the resolution of the display), or the manufacturer and model of the computing device 140. The query generator component 112 can use the device configuration to generate content selection criteria for input into the content selection component 120. By generating content selection criteria based on the device configuration information, the data processing system 102 can facilitate the selection of supplemental or sponsored content items based on the device configuration information associated with the computing device 140. Therefore, the data processing system 102 can select supplemental content items that are compatible with the device configuration information or otherwise optimized for the device configuration (e.g., reducing energy or computing resource utilization of the computing device 140).

[0072] The query generator component 112 can generate content selection criteria that indicate the type of sponsored or supplemental content selected to be provided via the computing device 140. The type of sponsored content can include, for example: digital assistant, search, streaming video, streaming audio, or contextual content. Digital assistant content can refer to content that is configured to be provided via a graphical user interface slot generated by a local digital assistant 142 on the computing device 140. The digital assistant content can meet certain quality signals or content parameters that optimize the digital assistant content to be provided in a graphical user interface slot generated by the local digital assistant 142. These content parameters can include, for example, image size, brightness level, file size, network bandwidth utilization, video duration, or audio duration. Examples of digital assistant content can include assistant applications, chatbots, image content, video content, or audio content.

[0073] Search-type content may refer to sponsored content that is configured to be provided with search results on a web browser. Search content may include text-only content or text and image content. Search content may include hyperlinks to web pages.

[0074] Streaming video type content may refer to video content that is configured to be played before a video starts, when a video ends, or during a video interruption. Streaming video type content may include video and audio. Streaming audio type content items may refer to content items that are played before, after, or during a streaming audio interruption. Contextual content may refer to content items that are displayed together with a web page published by a content publisher, such as a news article, blog, or other online document.

[0075] The query generator component 112 can generate content selection criteria based on attributes associated with a graphical user interface ("GUI") slot in which the supplemental or sponsored content is to be provided. In some cases, the attributes of the content slot can be predetermined or preconfigured. The attributes of the content slot can be stored in attributes 128 in the data repository 124. The attributes can include, for example, the size of the slot, the resolution of the display device 150, or other information. The attributes can be set based on the type of computing device 140. For example, a laptop computing device can have one or more attributes that are different from a smartphone computing device; or a smartphone computing device can have one or more attributes that are different from a desktop computing device.

[0076] The data processing system 102 may include a slot injector component 114 that is designed, constructed, and operable to generate, create, or otherwise provide a graphical user interface slot on a display device 150 of a computing device 140. The local digital assistant 142 may provide sponsored or supplemental content items in the GUI slot generated by the slot injector component 114. The server digital assistant component 104 may include the slot injector component 114, or otherwise interface or communicate with the slot injector component 114.

[0077] The slot injector component 114 can establish one or more properties for the GUI slot on the computing device 140. The slot injector component 114 can use various techniques to establish the properties. For example, the slot injector component 114 can obtain the properties that are preconfigured or assigned to the computing device 140 from the properties 128. The properties can be preconfigured or assigned based on the type of the computing device 140. For example, the properties 128 can include a mapping of the slot size to the size of the display device 150. Therefore, the slot injector component 114 can identify the size of the display device 150 and select the corresponding slot size. The slot injector component 114 can determine the size of the display device 150 based on the device configuration information received from the computing device 140. The slot injector component 114 can determine the size established for the slot based on the manufacturer or model of the computing device 140. For example, the properties 128 can include a mapping of the device manufacturer and model to the slot size.

[0078] When providing a content item via a GUI slot on computing device 140, slot injector component 114 may determine a property for the GUI slot that reduces energy consumption. Properties that may reduce energy consumption may include, for example, a brightness level of the content item, a file size of the content item, or a duration of the content item (e.g., the length of a video or audio content item). In some cases where the content item is an application (such as an assistant application or a chatbot application), the property may include a processor utilization level of the application.

[0079] Socket injector component 114 may determine energy consumption related attributes based on characteristics of computing device 140. For example, if computing device 140 is powered by a battery and is not currently being charged, socket injector component 114 may select attributes for the socket that reduce battery utilization and energy consumption. However, if computing device 140 is connected to a power source such as an electrical outlet, socket injector component 114 may select attributes that allow selection of content items that can consume more energy.

[0080] Data processing system 102 may receive an indication that the computing device is receiving power from a battery and is not connected to a charger. In response to the indication, data processing system 102 may generate a graphical user interface slot (e.g., Figure 2 , the graphical user interface slot having a first attribute that reduces energy consumption relative to a second attribute that consumes more energy. For example, the second attribute may have a higher brightness level. The second attribute may be used to generate a GUI slot for an action data structure (e.g., Figure 2 Attributes of the digital assistant content type may reduce the energy used to present the supplemental content via the computing device relative to attributes of different types of supplemental content (e.g., search content, video streaming content, or contextual content) that cause the computing device to use more energy. These attributes may relate to duration, file size, brightness level, or other characteristics of the digital content that affect computing resources, network, or energy consumption.

[0081] Thus, data processing system 102 (e.g., via slot injector component 114) can generate a second graphical user interface slot. Slot injector component 114 can generate a GUI slot in a home screen of computing device 140 (such as a smartphone home screen). A home screen can refer to a primary screen, a start screen, or a default GUI output by computing device 140 via display device 150. The home screen can display links, settings, and notifications to applications. Data processing system 102 can generate a graphical user interface slot with properties configured for a smartphone display.

[0082] The query generator component 112 can generate content selection criteria based on the properties of the slot established by the slot injector component 114. The query generator component 112 can include the properties of the slot in the content selection criteria. The content selection criteria can include the properties of the slot. The query generator component 112 can select the type of content based on the properties of the slot. For example, the properties established by the slot injector component 114 can indicate the type of content that is compatible with the slot. For example, if the slot injector component 114 is called by the server digital assistant component 104 to generate a GUI slot to provide supplemental content together with an action data structure generated in response to voice input, the content type can be a digital assistant type.

[0083] The server digital assistant component 104 can send a request for content to the content selector component 120. The server digital assistant component 104 can transmit a request for supplementary or sponsored content to a third-party content provider. The server digital assistant component 104 can transmit the content selection criteria generated by the query generator component 112 to the content selector component 120. The content selector component 120 can perform a content selection process to select a supplementary content item or a sponsored content item. The content item can be a sponsored or supplementary digital component object. The content item can be provided by a third-party content provider (such as a supplementary digital content provider device 152). The supplementary content item can include an advertisement for a product or service. The content selector component 120 can select a content item using the content selection criteria in response to receiving a request for content from the server digital assistant component 104.

[0084] The server digital assistant component 104 can receive supplementary or sponsored content items from the content selector component 120. The server digital assistant component 104 can receive content items in response to the request. The server digital assistant component 104 can provide supplementary content items to the supplementary digital content provider device 152 to provide via the GUI slot generated by the slot injector component 114. However, in some cases, before the direct action API 110 generates an action data structure in response to the voice input audio signal received from the computing device 140, the server digital assistant component 104 can receive content items from the content selector component 120. Compared with the content selector component 120 using the content selection criteria to select the content item in the real-time content selection process, the direct action API 110 may spend more time to generate the action data structure. For example, the amount of calculation made by the direct action API 110 to generate the action data structure can be greater than the amount of calculation made by the content selector component 120 to select the content item. In another example, relative to the computing infrastructure used by the direct action API 110, the computing infrastructure used by the content selector component 120 can be more powerful, robust or scalable. In yet another example, the content selector component 120 can be more efficient in selecting content items compared to the direct action API 110 selecting action data structures.

[0085] Therefore, due to various technical challenges, the time that direct action API 110 generates action data structure may be longer than the time that content selector component 120 selects content project.In addition, due to the time delay in generating action data structure, data processing system 102 can transfer content project to computing device 140 for presentation before action data structure.However, transmitting content project before action data structure can introduce further delay when presenting action data structure, because computing device 140 may submit and provide content project by computing resources, and this can delay the presentation of action data structure.Relative to first providing action data structure or providing action data structure with content project simultaneously, presenting sponsorship or supplementary content project before action data structure can provide poor user experience or cause poor user interface.For example, in order to provide improved user experience, data processing system 102 can provide both action data structure and supplementary content project simultaneously (for example, in 0.01 second, 0.05 second, 0.1 second, 0.2 second, 0.3 second, 0.4 second or 0.5 second).Another technical challenge comprises sending multiple transmissions to computing device 140, or performing multiple remote procedure calls. Sending the action data structure in one transmission and the content item in a second transmission may result in redundancy or inefficiency in the network, and the computing device 140 or local digital assistant 142 must receive and process multiple transmissions.

[0086] Therefore, the present technical solution includes a transmission controller component 116 that is designed, constructed and operated to receive action data structures and content items and control the delivery or transmission to reduce network utilization, energy consumption or processor utilization while improving the user interface and user experience of the computing device 140. The server digital assistant component 104 may include, access, interface with or otherwise communicate with the transmission controller component 116. The transmission controller component 116 may be part of the server digital assistant component 104. To address one or more technical challenges caused by transmitting action data structures and content items, the transmission controller component 116 may control the transmission of the action data structure or content items.

[0087] For example, the transmission controller component 116 may receive both the action data structure and the supplemental content item. The transmission controller component 116 may receive the action data structure from the direct action API 110. The transmission controller component 116 may receive the supplemental content item from the content selector component 120. The transmission controller component 116 may determine not to transmit the content item before the action data structure. The transmission controller component 116 may determine not to transmit the content item in a transmission separate from the action data structure. Instead, for example, the transmission controller component 116 may determine to combine the action data structure with the content item to generate a combined data packet. The transmission controller component 116 may transmit the combined data packet as a set of data packets to the computing device 140 via the network 105 in one transmission. By generating a combined data packet to be transmitted to the computing device 140, the transmission controller component 116 may improve the user interface or user experience by facilitating the computing device 140 to simultaneously present the content item and the action data structure. By generating a combined data packet, the transmission controller component 116 may reduce the number of separate transmissions to the computing device 140 via the network 105, thereby improving network transmission efficiency. Additionally, a single transmission may also improve network security compared to multiple transmissions. Thus, the transmission controller component 116 may be configured with a strategy to wait to transmit the action data structure or content item until both are received, and then generate a combined data packet for transmission to the computing device 140.

[0088] The combined data packet may include a data packet having a header, and the payload may include both the action data structure and the content item. The combined data packet may include instructions for computing device 140 on how to provide the action data structure and the content item. For example, the combined data packet may include instructions for local digital assistant 142 that cause local digital assistant 142 to provide the action data structure in a first GUI slot generated by local digital assistant 142 for the action data structure and to provide a supplemental or sponsored content item in a second GUI slot generated by local digital assistant 142 for the content item.

[0089] In some cases, the transmission controller component 116 can transmit the action data structure and the content item in the separated data transmission, but add delay or buffering in the transmission to reduce the delay between the transmission of the action data structure and the transmission of the content item. For example, if the transmission controller component 116 receives the content item from the content selector component 120 before receiving the action data structure from the direct action API 110, the transmission controller component 116 can add delay or buffering to the content item transmission. The transmission controller component 116 can be pre-configured with buffering or delay, such as 0.01 seconds, 0.02 seconds, 0.05 seconds, 0.1 seconds or other time intervals that are convenient to reduce the delay or time difference between the transmission of the content item and the transmission of the action data structure. The transmission controller component 116 can automatically determine buffering or delay based on the history execution of the content selector component 120 and the direct action API 110. For example, relative to the direct action API 110 generating the action data structure in response to voice input, the transmission controller component 116 can determine that the content selector component 120 is 0.1 seconds faster on average when selecting the sponsored content item via the real-time content selection process. The transmission controller component 116 may add a 0.1 second buffer or delay to the transmission of the sponsored content item to reduce the time difference between the transmission of the action data structure and the transmission of the content item.

[0090] In some cases, the transmission controller component 116 may wait to transmit the content item until the action data structure has been generated. The transmission controller component 116 may transmit the action data structure and the content item in separate transmissions, but reorder the transmissions so that the action data structure is transmitted before the content item.

[0091] The transmission controller component 116 can transmit the content item without waiting to receive the action data structure. The local digital assistant 142 can be configured to wait to submit or provide the content item until the local digital assistant 142 receives the action data structure. The transmission controller component 116 can transmit with the content item an instruction to instruct the local digital assistant 142 not to submit the content item upon receiving the content item, but to wait until the action data structure is generated and transmitted, so that the local digital assistant 142 can submit or provide the action data structure before or simultaneously with the content item.

[0092] Thus, for example, the server digital assistant component 104 may send a request to select supplemental content to the content selector component 120 in a manner that overlaps with generating an action data structure in response to a voice input. The server digital assistant component 104 may send the request while the server digital assistant component 104 is still generating the action data structure. For example, the query generator component 112 may generate content selection criteria and a request for content before the direct action API 110 has completed generating the action data structure. The query generator component 112 may continue to send requests for content and the generated content selection criteria in a manner that overlaps with the direct action API 110 generating the action data structure. The query generator component 112 may continue to send requests for content without waiting for the direct action API 110 to generate the action data structure. The server digital assistant component 104 may receive supplemental content items from the content selector component 120 before the direct action API 110 generates the action data structure. The transmission controller component 116 may determine to delay the delivery or transmission of the supplemental content items to the computing device 140 until the generation of the action data structure is completed. Transmit controller component 116 may, in response to the generation of the action data structure, provide the action data structure and the supplemental content item to the computing device for presentation.

[0093] If the query generator component 112 sends a request to select supplemental content to the content selector component 120 in a manner that overlaps with the generation of the action data structure, the transmission controller component 116 can receive the selected supplemental content item from the content selector component 120 before the direct action API 110 generates the action data structure. The transmission controller component 116 can instruct the computing device 140 to provide the supplemental content item in the second graphical user interface slot in response to presenting the action data structure in the first graphical user interface slot. For example, the transmission controller component 116 can transmit the supplemental content item independently when receiving the supplemental content item from the content selector component 120, but includes instructions for the local digital assistant 142 to present the supplemental content item only in the second GUI slot in response to presenting the action data structure in the first GUI slot, after presenting the action data structure in the first GUI slot, or simultaneously.

[0094] The data processing system 102 may include a content selector component 120 that is designed, constructed, or operated to select supplementary content items (or sponsored content items or digital component objects). In order to select sponsored content items or digital components, the content selector component 120 may use the generated content selection criteria to select matching sponsored content items based on broad matches, exact matches, or phrase matches. For example, the content selector component 120 may analyze, parse, or otherwise process the subject of a candidate sponsored content item to determine whether the subject of the candidate sponsored content item corresponds to the subject of a keyword or phrase of the content selection criteria generated by the query generator component 112. The content selector component 120 may use image processing techniques, character recognition techniques, natural language processing techniques, or database searches to identify, analyze, or recognize the voice, audio, terminology, character, text, symbol, or image of a candidate digital component. The candidate sponsored content item may include metadata indicating the subject of the candidate digital component, in which case the content selector component 120 may process the metadata to determine whether the subject of the candidate digital component corresponds to the input audio signal. The content actions provided by the supplemental digital content provider device 152 may include content selection criteria that the data processing system 102 may match with the criteria indicated in the second profile layer or the first profile layer.

[0095] The supplementary digital content provider device 152 can provide additional indicators when establishing a content campaign that includes a digital component. The supplementary digital content provider device 152 can provide information at the content campaign or content group level that the content selector component 120 can identify by performing a search using information about candidate digital components. For example, the candidate digital component can include a unique identifier that can be mapped to a content group, content campaign, or content provider. The content selector component 120 can determine information about the supplementary digital content provider device 152 based on information in the content data 126 stored in the data repository 124.

[0096] In response to the request, the content selector component 120 can select a digital component object from a data repository 124 or a database associated with a supplementary digital content provider device 152. The supplementary digital content can be provided by a supplementary digital content provider device different from a service provider device 154. The supplementary digital content can correspond to a service type (e.g., a taxi service versus a meal delivery service) different from the service type of the action data structure. The computing device 140 can interact with the supplementary digital content. The computing device 140 can receive an audio in response to the digital component. The computing device 140 can receive an indication of selecting a hyperlink or other button associated with the digital component object, making or allowing the computing device 140 to identify the supplementary digital content provider device 152 or the service provider device 154, requesting a service from the supplementary digital content provider device 152 or the service provider device 154, indicating that the supplementary digital content provider device 152 or the service provider device 154 performs a service, sending information to the supplementary digital content provider device 152 or the service provider device 154, or asking the supplementary digital content provider device 152 or the service provider device 154.

[0097] Supplementary digital content provider device 152 can establish electronic content actions. Electronic content actions can be stored in data repository 124 as content data 126. Electronic content actions can refer to one or more content groups corresponding to a common theme. Content actions can include a hierarchical data structure, which includes content groups, digital component data objects, and content selection criteria provided by a content provider. The content selection criteria provided by supplementary digital content provider device 152 can be compared with the content selection criteria generated by query generator component 112 to identify matching supplementary content items for transmission to computing device 140. The content selection criteria provided by supplementary digital content provider device 152 can include the type of content, such as digital assistant content type, search content type, streaming video content type, streamlined audio content type, or contextual content type. In order to create a content action, supplementary digital content provider device 152 can specify the value of the action level parameter of the content action. Action level parameters may include, for example, the name of the action, the preferred content network for placing the digital component object, the value of resources to be used for the content action, the start and end dates of the content action, the duration of the content action, the schedule for digital component object placement, the language, the geographic location, the type of computing device on which the digital component object is provided. In some cases, an impression may refer to when a digital component object is fetched from its source (e.g., data processing system 102 or supplemental digital content provider device 152) and is countable. In some cases, robot activity may be filtered and excluded as an impression due to the possibility of click fraud. Thus, in some cases, an impression may refer to a measurement of a web server's response to a page request from a browser, filtered from robot activity and error codes, and recorded as close as possible to the point of opportunity for submitting a digital component object for display on computing device 140. In some cases, an impression may refer to a visible or audible impression; for example, a digital component object is at least partially (e.g., 20%, 30%, 30%, 40%, 50%, 60%, 70%, or more) visible on display device 150 of client computing device 140, or audible via a speaker (e.g., transducer 146) of computing device 140. A click or selection may refer to a user interaction with a digital component object, such as voice, mouse click, touch interaction, gesture, shake, audio interaction, or keyboard click in response to an audible impression. A conversion may refer to a user taking a desired action on a digital component object; for example, purchasing a product or service, completing a survey, visiting a physical store corresponding to a digital component, or completing an electronic transaction.

[0098] The supplemental digital content provider device 152 may further establish one or more content groups for a content action. A content group includes one or more digital component objects and corresponding content selection criteria, such as keywords, words, terms, phrases, geographic locations, types of computing devices, time of day, interests, topics, or verticals. Content groups under the same content action may share the same action level parameters, but may have customized specifications for specific content group level parameters, such as keywords, negative keywords (e.g., blocking the placement of digital components in the presence of negative keywords on the primary content), bids on keywords, or parameters associated with bids or content actions.

[0099] In order to create a new content group, the content provider can provide the value of the content group level parameter for the content group. The content group level parameters include, for example, the content group name or content group theme, and the bid or result (e.g., click, impression or conversion) for different content placement opportunities (e.g., automatic placement or managed placement). The content group name or content group theme can be one or more terms that the supplementary digital content provider device 152 can use to grasp the topic or theme selected for display of the digital component object of its content group. For example, a car dealer can create different content groups for each brand of vehicle it represents, and can further create different content groups for each model of vehicle it represents. Examples of content group themes that car dealers can use include, for example, "Manufacturer A Sports Car", "Manufacturer B Sports Car", "Manufacturer C Sedan", "Manufacturer C Truck", "Manufacturer C Hybrid Car" or "Manufacturer D Hybrid Car". For example, an example content action theme can be "Hybrid Car", and includes content groups for both "Manufacturer C Hybrid Car" and "Manufacturer D Hybrid Car".

[0100] The supplementary digital content provider device 152 can provide one or more keywords and digital component objects to each content group. Keywords can include terms related to the product or service associated with the digital component object or identified by the digital component object. Keywords can include one or more terms or phrases. For example, a car dealer can include "sports car", "V-6 engine", "four-wheel drive", "fuel efficiency" as keywords for content groups or content actions. In some cases, content providers can specify negative keywords to avoid, prevent, block or disable content placement about certain terms or keywords. Content providers can specify matching types for selecting digital component objects, such as exact matching, phrase matching or broad matching.

[0101] The supplemental digital content provider device 152 may provide one or more keywords for the data processing system 102 to use to select digital component objects provided by the supplemental digital content provider device 152. The supplemental digital content provider device 152 may identify one or more keywords to bid on and further provide bid amounts for various keywords. The supplemental digital content provider device 152 may provide additional content selection criteria used by the data processing system 102 to select digital component objects. Multiple supplemental digital content provider devices 152 may bid on the same or different keywords, and the data processing system 102 may run a content selection process or an advertising auction in response to receiving an indication of the keywords of the electronic message.

[0102] The supplementary digital content provider device 152 can provide one or more digital component objects for the data processing system 102 to select. The data processing system 102 can (e.g., via the content selector component 120) select the digital component object when the content placement opportunity of matching resource allocation, content scheduling, maximum bid, keywords and other selection criteria specified for the content group becomes available.Different types of digital component objects (such as voice digital components, audio digital components, text digital components, image digital components, video digital components, multimedia digital components, digital component links or assistant application components) can be included in the content group.Digital component objects (or digital components, supplementary content items or sponsored content items) can include, for example, content items, online documents, audio, images, videos, multimedia content, sponsored content or assistant applications.After selecting the digital component, the data processing system 102 can (e.g., via the transmission controller component 116) transmit the digital component object to submit on the display device 150 of the computing device 140 or the computing device 140.Submission can include displaying the digital component on the display device, executing applications such as chatbots or dialogue robots, or playing the digital component via the speakers of the computing device 140. Data processing system 102 may provide instructions to computing device 140 to render the digital component objects. Data processing system 102 may instruct computing device 140 or an audio driver 148 of computing device 140 to generate audio signals or sound waves.

[0103] The data processing system 102 may pre-process or otherwise analyze the received content items when receiving the supplemental content items from the supplemental digital content provider device 152 to verify the content items for delivery and presentation. The data processing system 102 may analyze, evaluate, verify or otherwise process the content items to identify errors, vulnerabilities, malicious code or features or quality issues. For example, in order to reduce wasteful energy consumption, network bandwidth utilization, computing resource utilization and latency, the data processing system 102 may include a verification component 118 that is designed, constructed and operated to verify the supplemental content items before authorizing the supplemental content items for selection by the content selector component 120. Due to the additional technical challenges of presenting the supplemental content items together with the action data structure generated by the server digital assistant component 104, the data processing system 102 may verify the content items based on the content item type.

[0104] For example, the supplementary digital content provider device 152 can provide supplementary digital content items. The supplementary digital content items can be marked, marked or otherwise indicated as type. The type can be an assistant type content item. The data processing system 102 can access the content parameter data structure 130 stored in the data repository 124, including content parameters or quality signals established for assistant type content. The content parameters can include, for example, the brightness level of the content item, the file size of the content item, the duration of the content item, the processor utilization of the content item, or other quality signals. The content parameters can include thresholds for brightness, file size, duration, or processor utilization. The verification component 118 can simulate the content item submitted for reception to measure the brightness level, processor utilization, file size, duration, or other content parameters. The verification component 118 can compare the simulated measurement with the threshold value stored in the content parameter data structure 130 to determine whether the content item meets the threshold value. Meeting the threshold value can refer to or include that the measured quality signal is less than or equal to the threshold value.

[0105] For example, verification component 118 may determine the perceived brightness of an image corresponding to a supplemental content item based on simulating an image or otherwise processing the image. Verification component 118 may compare the perceived brightness to a brightness threshold established for an assistant-type content item. If the determined perceived brightness of the image is greater than the brightness threshold, verification component 118 may reject the supplemental content item. In some cases, verification component 118 may not reject the supplemental content item, but may remove a flag or indication that the supplemental content item qualifies as an assistant-type content item. Data processing system 102 may include an assistant-type flag or remove an assistant-type flag based on whether the content item meets the threshold.

[0106] Thus, the verification component 118 can help optimize the performance of the system 100 by weighting the content items to increase or decrease the likelihood that the content items are selected by the content selector component 120. If the verification component 118 verifies the content item based on the content parameters (e.g., the content item meets the content parameter threshold), the verification component 118 can indicate that the content item is valid. A valid content item can refer to a content item that is marked as an assistant type, and the verification component indicates that the content item meets the content parameters of the assistant type content item.

[0107] The content selector component 120 can select content items based on content selection criteria. For example, if the content selection criteria generated by the query generator component 112 includes assistant-type content, then the verification assistant content type can be weighted more heavily than content items that are not assistant-type. Therefore, when the content selection criteria include assistant-type content, the assistant-type content items can have a higher likelihood of being selected by the content selector component 120. The digital assistant content type can define attributes of the content item (e.g., content parameters such as a brightness level).

[0108] For example, the content selector component 120 may perform a real-time content selection process in response to a request from the query generator component 112 or receiving content selection criteria from the query generator component 112. Real-time content selection may refer to or include performing content selection in response to a request. Real-time may refer to or include selecting content within 0.2 seconds, 0.3 seconds, 0.4 seconds, 0.5 seconds, 0.6 seconds, or 1 second of receiving the request. Real-time may refer to selecting content in response to receiving an input audio signal from the computing device 140.

[0109] The content selector component 120 can identify a plurality of candidate supplemental content items corresponding to the digital assistant content type. The content selector component 120 can use the content selection criteria to identify candidate supplemental content items that match the digital assistant content type and one or more keywords generated by the query generator component 112 for the content selection criteria. The content selector component 120 can determine a score or ranking for each of the plurality of candidate supplemental content items to select the highest ranked supplemental content items to provide to the computing device 140. The content selector component 120 can select supplemental content items that have been verified as being of the digital assistant type.

[0110] In some cases, the content selector component 120 may weight different content types differently. The content selector component 120 may apply a higher weight to a content type that matches the content type indicated in the content selection criteria generated by the query generator component 112. The query generator component 112 may generate a content selection criteria that includes a desired content type or an optimal content type. In some cases, the query generator component 112 may require that the selected content item matches the desired content type, while in other cases, the query generator component 112 may indicate an increase in the possibility that a content item with a matching content type is selected. For example, the content selector component 120 may identify a first plurality of candidate supplementary content items corresponding to a digital assistant content type based on one or more keywords. The content selector component 120 may identify a second plurality of candidate supplementary content items having a second content type different from the digital assistant content type based on one or more keywords. The content selector component 120 may increase the weight of the first plurality of candidate supplementary content items to increase the possibility of selecting one of the first plurality of candidate supplementary content items relative to selecting one of the second plurality of candidate supplementary content items. The content selector component 120 can then determine the total score of each of the first plurality of content items and the second plurality of content items, wherein the score of the first plurality of content items increases based on the weight. The content selector component 120 can select the content item of the highest ranking or the highest score, which can be one of the first plurality of content items or one of the second plurality of content items. Therefore, the content selector component 120 can select a content item with a content type different from the digital assistant type in the second plurality of content items, even if the weight of the digital assistant content is greater.

[0111] The content selector component 120 can receive one or more supplementary content items from a third-party content provider (e.g., a supplementary digital content provider device 152). For one or more of the supplementary content items, the content selector component 120 can receive an indication from the third-party content provider that the one or more supplementary content items correspond to a digital assistant content type. The content selector component 120 can call a verification component 118 to identify content parameters for the digital assistant content type. The verification component 118 can perform a verification process on the one or more supplementary content items in response to the indication of the digital assistant content type to identify one or more valid supplementary content items that meet the content parameters established for the digital assistant content type. The content selector component 120 can select a supplementary content item from the one or more valid supplementary content items.

[0112] Figure 2140. The user interface 200 may be provided by a computing device 140. The user interface 200 may be provided via one or more systems or components of the system 100 (including, for example, a data processing system, a server digital assistant, or a local digital assistant). The user interface 200 may be output by a display device communicatively coupled to the computing device 140. The user interface 200 may include a home screen 222 (e.g., a primary screen, a start screen, or a default screen). The home screen 222 may include one or more applications, such as App_A 214, App_B 216, App_C 218, or App_D 220. Applications 214-220 may refer to phone applications, web browsers, contacts, calculators, text messaging applications, or other applications.

[0113] The user may provide voice input or input audio signals via the home screen 222. The user may invoke the digital assistant and provide voice input. When the local digital assistant is invoked, the computing device 140 may provide an indication that the local digital assistant is active or invoked. The indication may include, for example, a microphone icon 212. In some cases, selecting the microphone icon 212 may invoke the digital assistant.

[0114] The user can provide voice input, such as a query. For example, the query can be "jobs near me". The local digital assistant can display a text box 210 with the voice input query "jobs near me". The local digital assistant can send a data packet, an input audio signal, or a voice input query to the server digital assistant for further processing. Upon receiving the input audio signal or the voice input query, the server digital assistant can generate an action data structure in response to the voice input query. The action data structure can include search results, jobs near the computing device. The server digital assistant can provide the action data structure to the computing device using instructions to submit or provide the action data structure in the second GUI slot 204. The data processing system can provide the local digital assistant with instructions to establish a second GUI slot 204 based on one or more attributes associated with the computing device 140 or the action data structure.

[0115] The data processing system, upon receiving the voice input query, may generate content selection criteria and request supplemental content. The data processing system may generate content selection criteria based on the voice input query. The data processing system may select sponsored or supplemental content. The data processing system may transmit sponsored content for provision via the user interface 200. For example, the data processing system may provide instructions to the local digital assistant to generate a first GUI slot 202 in which the sponsored content item 208 is provided. The first GUI slot 202 may be constructed or injected on the main screen 222. The first GUI slot 202 may be separate or independent from the second GUI slot 204 in which the action data structure is provided.

[0116] The first GUI slot 202 may be established using one or more actions 206. The actions 206 may include, for example, pinning, moving, resizing, hiding, or minimizing. Pinning may refer to pinning the sponsored content item 208 or the first GUI slot 202 to the home screen so that the sponsored content item 208 stays on the home screen 222 of the user interface 200. Moving may refer to moving the first GUI slot 202 to another location or position on the home screen 222. For example, from the top of the home screen 222 to the middle or bottom of the home screen 222 or any other location on the home screen 222. Resizing may refer to changing the size of the first GUI slot 202. Resizing the first GUI slot 202 may cause the local digital assistant to resize the sponsored content item 208 presented via the first GUI slot 202. Hiding may refer to removing the first GUI slot 202 or making the first GUI slot 202 no longer visible via the home screen 222.

[0117] The local digital assistant can remember the action 206 selected by the user and update the properties or configuration of the first GUI slot 202 based on the selected action. For example, moving the first GUI slot 202 or resizing the first GUI slot 202 can cause the first GUI slot 202 to be moved for subsequent sponsored content items 208 provided in the first GUI slot 202. However, locking can refer to locking a specific sponsored content item 208. For example, the sponsored content item 208 selected in response to the "jobs near me" voice input query can include an advertisement for a clothing retailer that sells business suits. When the user sees the sponsored content item 208, he or she can determine to lock the sponsored content item 208 on the home screen 222.

[0118] Figure 3is an illustration of an example method for controlling delivery of supplemental content via a digital assistant, according to one implementation. Method 300 may be performed by, for example, one or more of a computing device, a data processing system, a local digital assistant, or a server digital assistant. At act 302, method 300 may include receiving voice input. The data processing system may receive a data packet including voice input or an input audio signal. The voice input may include a voice input query provided by a user or other speaker and detected by a microphone of a computing device, such as a smartphone or tablet computing device.

[0119] At action 304, the data processing system may process the speech input to generate an action data structure. The data processing system (e.g., direct action API) may use natural language processing techniques to process or parse the speech input and generate the action data structure. The action data structure may be a response to the speech input.

[0120] In action 306, the data processing system may generate content selection criteria. The data processing system may generate the content selection criteria based on the voice input. The data processing system may generate the content selection criteria based on keywords in the voice input. The data processing system may generate the content selection criteria based on a profile or account information associated with a computing device that detects the voice input. The data processing system may generate the content selection criteria including one or more keywords. The content selection criteria may include or indicate a content type, such as a digital assistant content type.

[0121] The data processing system may generate a request for content and send the request for content together with the content selection criteria to a content selector component. The data processing system may include a content selector component. The data processing system may select a supplementary content item based on the content selection criteria. The data processing system may select the content item using a real-time content selection process. The data processing system may select the content item using an online auction. The sponsorship or supplementary content item selected by the data processing system may be different from the action data structure generated by the content item. The action data structure may be generated in response to voice input. The supplementary content item may be selected using an online auction-based system in which a third-party content provider may bid on the supplementary content item in order to win the auction.

[0122] In action 310, the data processing system may receive the selected supplemental content item. The data processing system may receive the supplemental content item from the content selector component. In some cases, the data processing system may receive the supplemental content item before the data processing system generates the action data structure. For example, the hardware or computing infrastructure that selects the supplemental content item may select the supplemental content item before the direct action API generates the action data structure.

[0123] In decision block 312, the data processing system may determine whether to control transmission. Control transmission may refer to or include adding buffering, combining action data structures and supplementary content items, or otherwise controlling how action data structures and content items are provided relative to each other. The data processing system may determine control transmission based on a policy or rule. For example, if the data processing system receives the supplementary content items before generating the action data structure, the data processing system may determine control transmission. If the data processing system has historically received the amount of time of the content items before the action data structure is greater than a threshold, the data processing system may determine control transmission. The data processing system may determine control transmission to reduce energy consumption, network bandwidth utilization, or computing resource utilization. For example, if the computing device has limited battery resources or is on a network with limited bandwidth, the data processing system may determine control transmission to reduce network bandwidth utilization or battery energy utilization.

[0124] If the data processing system determines to control the transmission, the data processing system can continue to action 314 to execute the transmission control protocol. The data processing system can determine to buffer the content item or otherwise delay the transmission of the content item to transmit the action data structure before or simultaneously with the content item. For example, if the data processing system takes 0.4 seconds to generate the action data structure, but receives the supplemental content item within 0.1 seconds, the data processing system can add 0.3 seconds of buffering or delay to the content item transmission so that the action data structure and the supplemental content item are transmitted to the computing device simultaneously for presentation.

[0125] The data processing system can control the transmission by providing instructions to the computing device to present the content item after presenting the action data structure. The data processing system can provide instructions to the computing device to present the content item while presenting the action data structure. Thus, the data processing system can control the presentation or submission of the action data structure and the supplemental content item relative to each other.

[0126] The data processing system may control the transmission by combining the action data structure and the supplemental content item. The data processing system may generate a data packet including the combination of the action data structure and the supplemental content item. The combined data packet may include instructions for causing the computing device to submit the action data structure in a different GUI slot than the supplemental content item.

[0127] However, if the data processing system determines not to control the transmission, the data processing system can continue to actions 316 and 318 to transmit the supplemental content item and the action data structure to the computing device for presentation. The data processing system can transmit the supplemental content item in action 318 and transmit the action data structure in action 316 according to the transmission control policy (if selected).

[0128] Figure 4 100 , which is a diagram of an example method for verifying supplementary content based on content type according to an implementation. Method 400 may be performed by one or more systems or components depicted in Figure 100 (including, for example, a data processing system, a content selector component, or a verification component). In action 402, method 400 may include receiving a supplementary content item from a content provider. The content provider may be a third-party content provider that provides sponsored content. The data processing system may receive supplementary content or sponsored content as part of a content action. The data processing system may provide a user interface to allow a content provider to transmit a content item. The data processing system may provide an action establishment user interface or a graphical user interface in which a content provider may upload or otherwise transmit an electronic file containing a supplementary or sponsored content item.

[0129] In action 404, the data processing system can identify the content type. The content provider can indicate the content type of each content item in the content items uploaded to the data processing system. The content type can indicate where the content item can be provided. For example, the content type can include digital assistant, search, context, streaming video or streaming audio. For different digital media, different content types can be optimized or preferred. For example, search content can include plain text content. Context content can include text and images. Streaming audio content can only include audio. Streaming video content can include both audio and video content. For example, digital assistant content can include text, images, audio, video or assistant applications.

[0130] Different types of content may include different content parameters. The content parameters may refer to the duration of the content, the brightness level of the content, the file size of the content, or the processor utilization of the content. For example, content parameters for digital assistant content configured for a home screen on a mobile device may include a lower brightness level threshold than, for example, contextual content that would be provided on a desktop computing device.

[0131] In action 406, the data processing system may generate a quality signal value for the supplemental content item. The data processing system may simulate the provision or submission of the supplemental content item to generate the quality signal. The data processing system may evaluate the supplemental content item to determine the value of the quality signal. The quality signal may include brightness level, file size, duration, audio level, processor utilization, or other quality signals. The data processing system may determine the brightness level, file size, duration, or processor utilization of the supplemental content item provided by the third-party content provider.

[0132] At action 408, the data processing system may compare the generated quality signal value to a threshold established for the content type indicated by the content provider. For example, if the content provider indicates that the supplemental content item is of a digital assistant type, the data processing system may retrieve a brightness level threshold for the digital assistant content type and compare the brightness level quality signal value to the threshold. The data processing system may compare each quality signal value to a threshold for the corresponding content type.

[0133] At decision block 410, the data processing system may determine whether the supplemental content item is valid. Valid may refer to whether the quality signal value of the content item meets the threshold value of the content type. If the data processing system determines at decision block 410 that the content item is valid, the data processing system may proceed to action 414 to verify the content item for the content type. However, if the data processing system determines at decision block 410 that the supplemental content item is invalid, the data processing system may proceed to action 412 to change the content type flag.

[0134] If the generated quality signal value is greater than a threshold, the data processing system may determine to reject the supplemental content item. The data processing system may remove the supplemental content item from the content data repository. In some cases, the data processing system may determine to change the content type flag; for example, if the remaining quality signal value satisfies other content types, then the content type is changed to search content instead of digital assistant content.

[0135] Figure 5is a block diagram of an example computer system 500. The computer system or computing device 500 may include or be used to implement the system 100 or its components, such as the data processing system 102. The data processing system 102 may include an intelligent personal assistant or a voice-based digital assistant. The computing system 500 includes a bus 505 or other communication component for transmitting information, and a processor 510 or processing circuit coupled to the bus 505 for processing information. The computing system 500 may also include one or more processors 510 or processing circuits coupled to the bus for processing information. The computing system 500 also includes a main memory 515, such as a random access memory (RAM) or other dynamic storage device, coupled to the bus 505 for storing information and instructions to be executed by the processor 510. The main memory 515 may be or include a data repository 124. The main memory 515 may also be used to store location information, temporary variables, or other intermediate information during the execution of instructions by the processor 510. The computing system 500 may also include a read-only memory (ROM) 520 or other static storage device coupled to the bus 505 for storing static information and instructions for the processor 510. A storage device 525 , such as a solid-state device, magnetic disk, or optical disk, may be coupled to bus 505 to persistently store information and instructions. Storage device 525 may include or be a part of data repository 124 .

[0136] Computing system 500 may be connected via bus 505 to a display 535, such as a liquid crystal display or an active matrix display, for displaying information to a user. An input device 530, such as a keyboard including alphanumeric and other keys, may be coupled to bus 505 for communicating information and command selections to processor 510. Input device 530 may include a touch screen display 535. Input device 530 may also include a cursor control, such as a mouse, trackball, or cursor direction keys, for communicating direction information and command selections to processor 510, and for controlling cursor movement on display 535. Display 535 may be, for example, a computer program product of data processing system 102, client computing device 140, or the like. Figure 1 part of other components.

[0137] The processes, systems, and methods described herein can be implemented by the computing system 500 in response to the processor 510 executing the instruction arrangement contained in the main memory 515. These instructions can be read into the main memory 515 from another computer-readable medium (such as storage device 525). The execution of the instruction arrangement contained in the main memory 515 causes the computing system 500 to perform the illustrative processes described herein. One or more processors in a multi-processing arrangement can also be used to execute the instructions contained in the main memory 515. Hard-wired circuits can be used in place of software instructions or in combination with software instructions and the systems and methods described herein. The systems and methods described herein are not limited to any specific combination of hardware circuitry and software.

[0138] Although already Figure 5 An example computing system is described in the specification, but the subject matter including the operations described in this specification may be implemented in other types of digital electronic circuits, or in computer software, firmware, or hardware (including the structures disclosed in this specification and their structural equivalents), or in a combination of one or more of them.

[0139] In the case where the systems discussed herein collect user personal information or can utilize personal information, the user may be provided with the opportunity to control programs or functions that can collect personal information (e.g., information about the user's social network, social actions or activities, user preferences, or user location) or to control whether or how to receive content more relevant to the user from a content server or other data processing system. In addition, before storing or using certain data, it may be anonymized in one or more ways to remove identifiable personal information when generating parameters. For example, the identity of the user may be anonymized so that the user's identifiable personal information cannot be determined, or the user's geographic location may be summarized as the location where the location information is obtained (such as to the city, zip code, or state level), so that the user's specific location cannot be determined. Therefore, the user can control how the content server collects and uses information about him or her.

[0140] The subject matter and operations described in this specification may be implemented in digital electronic circuits or in computer software, firmware or hardware (including the structures disclosed in this specification and their structural equivalents) or in one or more combinations thereof. The subject matter described in this specification may be implemented as one or more computer programs (e.g., one or more computer program instruction circuits) encoded on one or more computer storage media for execution by a data processing device or for controlling the operation of the data processing device. Alternatively or additionally, program instructions may be encoded in artificially generated propagation signals (e.g., machine-generated electrical, optical or electromagnetic signals) generated to encode information and thus transmitted to a suitable receiver device for execution by a data processing device. Computer storage media may be or be included in a computer-readable storage device, a computer-readable storage substrate, a random or serial access memory array or device, or a combination of one or more of them. Although computer storage media are not propagation signals, computer storage media may be the source or destination of computer program instructions encoded in artificially generated propagation signals. Computer storage media may also be or be included in one or more separate components or media (e.g., multiple CDs, disks, or other storage devices). The operations described in this specification may be implemented as operations performed by a data processing apparatus on data stored on one or more computer-readable storage devices or received from other sources.

[0141] The terms "data processing system," "computing device," "component," or "data processing apparatus" encompass various apparatuses, devices, and machines for processing data, including, for example, a programmable processor, a computer, a system on a chip, or multiple or combinations of the foregoing. The apparatus may include dedicated logic circuits, such as an FPGA (field programmable gate array) or an ASIC (application-specific integrated circuit). In addition to hardware, the apparatus may also include code that creates an execution environment for the computer program in question, for example, code that constitutes processor firmware, a protocol stack, a database management system, an operating system, a cross-platform runtime environment, a virtual machine, or a combination of one or more of them. The apparatus and execution environment may implement a variety of different computing model infrastructures, such as network services, distributed computing, and grid computing infrastructures. For example, the direct action API 110, the content selector component 120, or the NLP component 108 and other data processing system 102 components may include or share one or more data processing apparatuses, systems, computing devices, or processors.

[0142] A computer program (also referred to as a program, software, software application, application, script, or code) may be written in any form of programming language, including compiled or interpreted, declarative, or procedural languages, and may be deployed in any form, including as a standalone program or module, component, subroutine, object, or other unit suitable for a computing environment. A computer program may correspond to a file in a file system. A computer program may be stored in a portion of a file that holds other programs or data (e.g., one or more scripts stored in a markup language document), in a single file dedicated to the program in question, or in multiple collaborative files (e.g., files storing one or more modules, subroutines, or code portions). A computer program may be deployed to execute on one computer or on multiple computers located in one location or distributed in multiple locations and interconnected by a communication network.

[0143] The processes and logic flows described in this specification may be performed by one or more programmable processors executing one or more computer programs (e.g., components of data processing system 102) to perform actions by operating on input data and generating output. The processes and logic flows may also be performed by, and the apparatus may be implemented as, special purpose logic circuitry, such as an FPGA (field programmable gate array) or an ASIC (application specific integrated circuit). Suitable devices for storing computer program instructions and data include all forms of non-volatile memory, media, and storage devices, including, for example, semiconductor memory devices, such as EPROM, EEPROM, and flash memory devices; magnetic disks, such as internal hard disks or removable disks; magneto-optical disks; and CD ROM and DVD-ROM disks. The processor and memory may be supplemented by or incorporated into special purpose logic circuitry.

[0144] The subject matter described herein can be implemented in a computing system that includes a back-end component (e.g., a data server) or includes a middleware component (e.g., an application server) or includes a front-end component (e.g., a client computer with a graphical user interface or a web browser through which a user can interact with an implementation of the subject matter described in this specification) or a combination of one or more such back-end components, middleware components, or front-end components. The components of the system can be interconnected by digital data communication (e.g., a communication network) in any form or medium. Examples of communication networks include local area networks ("LANs") and wide area networks ("WANs"), internetworks (e.g., the Internet), and peer-to-peer networks (e.g., ad hoc peer-to-peer networks).

[0145] A computing system such as system 100 or system 500 may include a client and a server. The client and server are generally remote from each other and typically interact through a communication network (e.g., network 105). The relationship of client and server arises due to computer programs running on the respective computers and having a client-server relationship to each other. In some implementations, the server transmits data (e.g., data packets representing digital components) to a client device (e.g., in order to display the data to and receive user input from a user interacting with the client device). Data generated at the client device (e.g., the result of the user interaction) can be received from the client device at the server (e.g., received by the data processing system 102 from the local computing device 140 or the supplemental digital content provider device 152 or the service provider device 154).

[0146] Although operations are depicted in the drawings in a particular order, such operations are not required to be performed in the particular order shown or sequentially, and not all of the operations shown are required to be performed. The actions described herein may be performed in a different order.

[0147] The separation of various system components does not require separation in all implementations, and the program components may be included in a single hardware or software product. For example, the NLP component 108 or the direct action API 110 may be a single component, application, or program, or a logical device having one or more processing circuits, or part of one or more servers of the data processing system 102.

[0148] Now that some illustrative implementations have been described, it is clear that the foregoing is illustrative rather than restrictive and has been provided by way of example. Specifically, although the multiple examples presented herein relate to specific combinations of method actions or system elements, those actions and those elements may be combined in other ways to achieve the same goals. The actions, elements, and features discussed in conjunction with one implementation are not intended to be excluded from similar roles or other implementations in other implementations.

[0149] The phraseology and terminology used herein are for descriptive purposes and should not be considered limiting. "Including," "comprising," "having," "involving," "characterized by," "characterized by," and variations thereof as used herein are meant to encompass the items listed thereafter, their equivalents and additional items, and alternative implementations consisting of the items specifically listed thereafter. In one implementation, the systems and methods described herein include one, more than one of each combination, or all of the elements, actions, or components described.

[0150] Any implementation or element or action of the system and method mentioned in the singular form herein may also encompass implementations that include multiple of these elements, and any implementation or element or action mentioned in the plural form herein may also encompass implementations that include only a single element. The mention in the singular or plural form is not intended to limit the currently disclosed system or method, their components, actions or elements to a single or multiple configurations. The mention that any action or element is based on any information, action or element may include an implementation that the action or element is based at least in part on any information, action or element.

[0151] Any implementation disclosed herein may be combined with any other implementation or embodiment, and references to "implementations," "some implementations," "an implementation," etc., are not necessarily mutually exclusive and are intended to mean that a particular feature, structure, or characteristic described in connection with an implementation may be included in at least one implementation or in at least one embodiment. These terms as used herein do not necessarily all refer to the same implementation. Any implementation may be combined with any other implementation, inclusively or exclusively, in any manner consistent with the various aspects and implementations disclosed herein.

[0152] References to "or" may be interpreted as inclusive, so any term described using "or" may mean any of a single, more than one, or all of the terms. References to at least one of a combined list of terms may be interpreted as inclusive, or indicate any of a single, more than one, and all of the terms. For example, references to "at least one of A and B" may include only A, only B, and both A and B. Such references used in conjunction with "including" or other public terms may include additional items.

[0153] Where technical features in the drawings, detailed description or any claims are followed by reference numerals, the reference numerals have been included to increase the intelligibility of the drawings, detailed description and claims. Therefore, reference numerals or their absence have no limiting effect on the scope of any claim element.

[0154] The systems and methods described herein may be embodied in other specific forms without departing from the characteristics of the systems and methods described herein. The foregoing implementations are illustrative rather than limiting of the systems and methods described. Accordingly, the scope of the systems and methods described herein is indicated by the appended claims rather than the foregoing description, and changes within the meaning and range of equivalents of the claims are included within the scope of the systems and methods described herein.

Claims

1. A system for controlling the delivery of supplemental content via a digital assistant, comprising: a data processing system including a memory and one or more processors; A digital assistant component of a data processing system, the digital assistant component being configured to: receiving, from a computing device via a network, a data packet including voice input detected by a microphone of the computing device; processing the data packet to generate an action data structure responsive to the speech input; generating content selection criteria for input into a content selector component to select supplemental content provided by a third-party content provider, the content selection criteria comprising one or more keywords and a digital assistant content type; sending, in a manner overlapping with the generation of the action data structure in response to the speech input, a request to a content selector component to select supplemental content based on the content selection criteria; Prior to generation of the action data structure, receiving, in response to the request, a supplemental content item selected by the content selector component based on content selection criteria generated by the digital assistant component; determining to delay delivery of the supplemental content item until generation of the action data structure is complete; In response to the generation of the action data structure, providing the action data structure responsive to the speech input in the first graphical user interface slot via a display device coupled to the computing device; as well as In response to the generation of the action data structure, the supplemental content item selected by the content selector component is provided in a second graphical user interface slot via a display device coupled to the computing device.

2. The system according to claim 1, comprising: The data processing system generates a second graphical user interface slot in a home screen of the computing device.

3. The system according to claim 1 or 2, comprising: The data processing system generates a second graphical user interface slot having attributes configured for the mobile display device.

4. The system according to claim 1 or 2, comprising the data processing system, wherein the data processing system is used for: receiving an indication that the computing device is receiving power from a battery and is not connected to a charger; and In response to the indication, a second graphical user interface slot is generated having a first attribute that reduces energy consumption relative to a second attribute that consumes more energy.

5. The system according to claim 1 or 2, comprising the data processing system, wherein the data processing system is used for: generating content selection criteria including digital assistant content types, wherein: The digital assistant content type defines the attributes of the supplemental content.

6. The system according to claim 5, wherein: The attribute of the digital assistant content type reduces the amount of energy used to provide supplemental content via the computing device relative to a second attribute of a different type of supplemental content that causes the computing device to use a greater amount of energy.

7. The system according to claim 1 or 2, comprising the data processing system, wherein the data processing system is configured to: sending a request to select supplemental content to a content selector component in a manner overlapping with generation of an action data structure in response to speech input; Prior to generation of the action data structure, receiving a supplemental content item from a content selector component; and In response to providing the action data structure in the first graphical user interface slot, the computing device is instructed to provide the supplemental content item in the second graphical user interface slot.

8. The system according to claim 1 or 2, wherein: The second graphical user interface slot is configured to lock the supplemental content item on a home screen of the computing device.

9. The system according to claim 1 or 2, wherein: The supplemental content items include supplemental digital assistant applications provided by third-party content providers.

10. The system according to claim 1 or 2, comprising the content selector component, wherein the content selector component is configured to: identifying a plurality of candidate supplemental content items corresponding to the digital assistant content type; and Based on the one or more keywords, a supplemental content item is selected from a plurality of candidate supplemental content items.

11. The system according to claim 1 or 2, comprising the content selector component, wherein the content selector component is configured to: identifying, based on the one or more keywords, a first plurality of candidate supplemental content items corresponding to a digital assistant content type; identifying, based on the one or more keywords, a second plurality of candidate supplemental content items having a second content type different from the digital assistant content type; as well as The weights of the first plurality of candidate supplemental content items are increased to increase the likelihood of selecting one of the first plurality of candidate supplemental content items relative to selecting one of the second plurality of candidate supplemental content items.

12. The system according to claim 1 or 2, comprising the content selector component, wherein the content selector component is configured to: receiving one or more supplemental content items from a third-party content provider; for the one or more supplemental content items, receiving an indication from a third-party content provider indicating that the one or more supplemental content items correspond to a digital assistant content type; identifying content parameters for a digital assistant content type; as well as In response to the indication of the digital assistant content type, a validation process is performed on the one or more supplemental content items to identify one or more valid supplemental content items that satisfy content parameters established for the digital assistant content type.

13. The system according to claim 12, comprising: The content selector component selects a supplemental content item from one or more valid supplemental content items.

14. The system according to claim 12, wherein: The content parameters include at least one of image size, brightness level, file size, network bandwidth utilization, video duration, or audio duration.

15. A method for delivering supplemental content via a digital assistant, comprising: a data processing system including a memory and one or more processors; A digital assistant component of a data processing system, the digital assistant component being configured to: Receiving, by a data processing system including a memory and one or more processors, a data packet including voice input detected by a microphone of the computing device from the computing device via a network; processing the data packet by a digital assistant component of the data processing system to generate an action data structure responsive to the speech input; generating, by the digital assistant component, content selection criteria for input into the content selector component to select supplemental content provided by a third-party content provider, the content selection criteria comprising one or more keywords and a digital assistant content type; sending, by the digital assistant component to the content selector component, a request to select supplemental content based on the content selection criteria in a manner overlapping with the generation of the action data structure in response to the speech input; Prior to the generation of the action data structure, receiving, by the digital assistant component in response to the request, a supplemental content item selected by the content selector component based on content selection criteria generated by the digital assistant component; In response to the generation of the action data structure, providing, by the digital assistant component via a display device coupled to the computing device, the action data structure in response to the speech input in the first graphical user interface slot; as well as In response to the generation of the action data structure, the supplemental content item selected by the content selector component is provided by the digital assistant component in the second graphical user interface slot via a display device coupled to the computing device.

16. The method according to claim 15, comprising: A second graphical user interface slot is generated by the data processing system in a home screen of the computing device having attributes configured for the mobile display device.

17. The method according to claim 15 or 16, comprising: Content selection criteria are generated by a data processing system that include a digital assistant content type, wherein the digital assistant content type defines an attribute of supplemental content and reduces an amount of energy used to provide the supplemental content via a computing device relative to a second attribute of a different type of supplemental content that causes the computing device to use a greater amount of energy.

18. The method according to claim 15 or 16, comprising: sending, by the data processing system, a request to select supplemental content to a content selector component in a manner overlapping with generation of an action data structure in response to speech input; receiving, by the data processing system, the supplemental content item from the content selector component prior to generating the action data structure; as well as The computing device is directed by the data processing system to provide a supplemental content item in a second graphical user interface slot in response to providing the action data structure in the first graphical user interface slot.

19. The method according to claim 15 or 16, comprising: receiving, by the content selector component, one or more supplemental content items from a third-party content provider; receiving, by the content selector component, for the one or more supplemental content items, from a third-party content provider an indication that the one or more supplemental content items correspond to a digital assistant content type; identifying, by the content selector component, content parameters for the digital assistant content type, wherein the content parameters include at least one of image size, brightness level, file size, network bandwidth utilization, video duration, or audio duration; as well as A validation process is performed by the content selector component on the one or more supplemental content items in response to an indication of the digital assistant content type to identify one or more valid supplemental content items that satisfy content parameters established for the digital assistant content type.

Citation Information

Patent Citations

  • System and method for selecting and presenting advertisements based on natural language processing of voice-based input

    US20080189110A1