Detection of duplicate packetized data transmission
By identifying request and triggering keywords in a voice-activated computer network environment, generating action data structures and selecting appropriate interfaces to transmit, the problem of network traffic redundancy is solved, and data processing efficiency and resource utilization are improved.
Patent Information
- Application Number
- CN202310910488.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2017-12-08
- Publication Date
- 2025-08-19
- Estimated Expiration
- 2037-12-08
AI Technical Summary
Over-transmission of network traffic data between computing devices leads to overload processing power, reduces response quality and complicates data routing, and redundant content transmission increases bandwidth utilization and resource consumption.
By using natural language processor components to identify requests and trigger keywords in a voice-activated computer network environment, a natural language processor component is used to identify requests and trigger keywords, generate action data structures, and select appropriate interfaces to transmit data through the interface management component, avoid redundant transmission and optimize resource utilization.
It reduces network bandwidth utilization, delay and resource consumption, improves data processing efficiency, saves processing capacity and power, and optimizes the resource use of computing equipment.
Smart Images

Figure CN117059080B_ABST
Abstract
Description
[0001] Description of the case
[0002] This application is a divisional application of Chinese invention patent application No. 201780013430.9, filed on December 8, 2017. Technical Field
[0003] The present application relates generally to the detection of duplicate packetized data transmissions. Background Art
[0004] Excessive network transmission of network traffic data between computing devices, whether packet-based or otherwise, can prevent the computing devices from properly processing the network traffic data, completing operations related to the network traffic data, or responding to the network traffic data in a timely manner. If the responding computing device reaches or exceeds its processing capacity, excessive network transmission of network traffic data can also complicate data routing or reduce the quality of the response, which may result in inefficient bandwidth utilization. A large number of content item objects that can initiate network transmission of network traffic data between computing devices can complicate the control of the network transmission corresponding to the content item objects. Summary of the Invention
[0005] At least one aspect relates to a system for transmitting packetized data in a computer network environment based on voice-activated packets. The system may include a natural language processor component executable by a data processing system. The data processing system may receive, via an interface of the data processing system, a data packet that may include an input audio signal. The input audio signal may be detected by a sensor of a client device. The natural language processor component may parse the input audio signal to identify a request and a trigger keyword that may correspond to the request. The data processing system may include a direct action application programming interface for generating a first action data structure based on at least one of the request and the trigger keyword. The data processing system may include a content selector component for receiving at least one of the request and the trigger keyword identified by the natural language processor. The content selector component may select a digital component to be displayed at the client device based on at least one of the request and the trigger keyword via a real-time content selection process. The data processing system may include an interface management component for identifying a second interface associated with the client device. The interface management component may determine a prior instance in which the second interface associated with the client device previously received the digital component. The interface management component may transmit the first action data structure, rather than the digital component, to the client device for rendering as audio output from the client device.
[0006] At least one aspect relates to a method for transmitting packetized data in a voice-activated packet-based computer network environment. The method may include receiving, via an interface of a data processing system, a data packet that may include an input audio signal. The input audio signal may be detected by a sensor of a client device. The method may include parsing, by a natural language processor component, the input audio signal to identify a request and a trigger keyword that may correspond to the request. The method may include generating, by a direct action application programming interface, a first action data structure based on at least one of the request and the trigger keyword. The method may include receiving, by a content selector component, at least one of the request and the trigger keyword identified by the natural language processor. The method may include selecting, by the content selector component, a digital component to be displayed at the client device based on at least one of the request and the trigger keyword via a real-time content selection process. The method may include identifying, by an interface management component, a second interface associated with the client device. The method may include determining a prior instance in which the second interface associated with the client device previously received the digital component. The method may include transmitting, instead of the digital component, the first action data structure to the client device for rendering as audio output from the client device.
[0007] These and other aspects and embodiments are discussed in detail below. The foregoing information and the following detailed description include illustrative examples of the various aspects and embodiments and provide an overview or framework for understanding the nature and character of the claimed aspects and embodiments. The accompanying drawings provide illustration and further understanding of the various aspects and embodiments and are incorporated into and constitute a part of this specification. BRIEF DESCRIPTION OF THE DRAWINGS
[0008] The drawings are not intended to be drawn to scale. Like reference numbers and names in the various drawings indicate like elements. For clarity, not every component may be labeled in every drawing.
[0009] In the attached figure:
[0010] Figure 1 A system for multimodal transmission of packetized data in a voice-activated computer network environment is described;
[0011] Figure 2 A flow chart depicting multimodal transmission of packetized data in a voice-activated computer network environment;
[0012] Figure 3 A method for multimodal transmission of packetized data in a voice-activated computer network environment is described; and
[0013] Figure 4is a block diagram illustrating the general architecture of a computer system that may be employed to implement elements of the systems and methods described and illustrated herein. DETAILED DESCRIPTION
[0014] The following is a more detailed description of various concepts related to methods, apparatuses, and systems for multimodal transmission of packetized data in a computer network environment based on voice-activated data packets, their implementation, and methods, apparatuses, and systems. The various concepts introduced above and discussed in more detail below can be implemented in any of numerous ways.
[0015] The disclosed systems and methods generally relate to a data processing system for identifying potentially redundant transmissions in a voice-activated computer network environment. The data processing system can improve the efficiency and effectiveness of data packet (or other protocol-based) transmissions over one or more computer networks by, for example, preventing or reducing the number of redundant data packet transmissions. The system can deem transmissions redundant when related or similar transmissions are made to related client computing devices. The system can deem these transmissions redundant because a user may have already viewed the transmission results on the related client device. The system can select from multiple options a transmission modality for routing data packets over a computer network for content items to one or more client computing devices (also referred to as client devices) or to different interfaces (e.g., different applications or programs) of a single client computing device. Signals based on data packets or other protocols corresponding to the selected operation can be routed over the computer network between the multiple computing devices. For example, the data processing system can route the content item to an interface different from the interface from which the request was received. The different interface can be on the same client computing device or on a different client computing device from which the request was received. The data processing system can select at least one candidate interface from a plurality of candidate interfaces for transmitting the content item to the client computing device. Candidate interfaces can be determined based on technical or computing parameters (such as processor power or utilization, memory capacity or availability, battery status, available power, network bandwidth utilization, interface parameters, or other resource utilization values). By selecting an interface from a client computing device for receiving and providing content items for rendering based on the candidate interfaces or utilizations associated with the candidate interfaces, the data processing system can reduce network bandwidth utilization, latency, or processing utilization or power consumption of the client computing device that renders the content items. This saves processing power and other computing resources (such as memory), reduces power consumption of the data processing system, and the reduced data transmission via the computer network reduces bandwidth requirements and usage of the data processing system.
[0016] The systems and methods described herein may include a data processing system that receives an input audio query, which may also be referred to as an input audio signal. Through the input audio query, the data processing system may identify a request and a trigger keyword corresponding to the request. Based on the trigger keyword or request, the data processing system may generate a first action data structure. For example, the first action data structure may include an organic response to the input audio query received from a client computing device, and the data processing system may provide the first action data structure to the same client computing device for rendering as audio output via the same interface from which the request was received.
[0017] The data processing system can also select at least one content item based on a trigger keyword or request. The data processing system can identify or determine multiple candidate interfaces for rendering the content item. The interface can include one or more hardware or software interfaces, such as a display screen, an audio interface, a speaker, an application or program available on the client computing device that initiates the input audio query or on different client computing devices. The interface can include a javascript slot (slot) for inserting the content item and a push notification interface of an online document. The data processing system can determine the utilization value of different candidate interfaces. For example, the utilization value can indicate power capability, processing capability, memory capability, bandwidth capability, or interface parameter capability. Based on the utilization value of the candidate interface, the data processing system can select the candidate interface as the interface selected for presenting or rendering the content item. For example, the data processing system can convert or provide the content item for delivery in a modality compatible with the selected interface. The selected interface can be the interface of the same client computing device or different client computing devices that initiated the input audio signal. By routing data packets through a computing network based on utilization values associated with candidate interfaces, a data processing system selects a destination for a content item from available options in a manner that uses a minimal amount of processing power, memory, or bandwidth, or that conserves power for one or more client computing devices.
[0018] The data processing system can provide content items or the first action data structure to the client computing device via a computer network by transmitting data messages based on packets or other protocols. The content items can also be referred to as digital components. The content items can be included in the digital components. The output signal can cause the audio driver component of the client computing device to generate sound waves that can be output from the client computing device, such as audio output. Audio (or other) output can correspond to the first action data structure or to the content items. For example, the first action data structure can be routed as audio output, and the content items can be routed as text-based messages. By routing the first action data structure and the content items to different interfaces, relative to providing the first action data structure and the content items to the same interface, the data processing system can save the resources utilized by each interface. Compared to the situation where the first action data structure and the content items are not separated and the first action data structure and the content items are independently routed, this causes less data processing operations, less memory usage, or less network bandwidth utilization of the interface (or its corresponding device) selected.
[0019] Figure 1 An example system 100 for multimodal transmission of packetized data in a computer network environment based on voice-activated data packets (or other protocols) is depicted. The system 100 may include at least one data processing system 105. The data processing system 105 may include at least one server having at least one processor. For example, the data processing system 105 may include a plurality of servers located in at least one data center or server farm. The data processing system 105 may determine a request and a trigger keyword associated with the request from an input audio signal. Based on the request and the trigger keyword, the data processing system 105 may determine or select at least one action data structure and may select at least one content item (and initiate other actions as described herein).
[0020] The data processing system 105 can identify the candidate interface for rendering the action data structure or content items. The data processing system 105 can provide the action data structure or content items for one or more candidate interfaces on one or more client computing devices to render. The selection of the interface can be based on the previous transmission or the associated interface to the content items of the interface or the resource utilization value for the candidate interface or the resource utilization value of the candidate interface. The action data structure (or content items) can include one or more audio files that provide audio output or sound waves when rendered. The action data structure or content items can include other content (e.g., text, video, or image content) except the audio content.
[0021] Data processing system 105 may include multiple logically grouped servers and facilitate distributed computing techniques. Logical groups of servers may be referred to as data centers, server farms, or machine farms. Servers may be geographically dispersed. A data center or machine farm may be managed as a single entity, or a machine farm may include multiple machine farms. Servers within each machine farm may be heterogeneous—one or more of the servers or machines may operate according to one or more types of operating system platforms. Data processing system 105 may include servers in a data center, such as an enterprise data center, stored in one or more high-density rack systems along with associated storage systems. Data processing system 105 with consolidated servers in this manner can improve system manageability, data security, physical security of the system, and system performance by locating servers and high-performance storage systems on a localized, high-performance network. Centralizing all or some of the data processing system 105 components (including servers and storage systems) and coupling them with advanced system management tools allows for more efficient use of server resources, saving power and processing requirements, and reducing bandwidth usage.
[0022] The data processing system 105 may include at least one natural language processor (NLP) component 110, at least one interface 115, at least one prediction component 120, at least one content selector component 125, at least one audio signal generator component 130, at least one direct action application programming interface (API) 135, at least one interface management component 140, and at least one data repository 145. The NLP component 110, interface 115, prediction component 120, content selector component 125, audio signal generator component 130, direct action API 135, and interface management component 140 may each include at least one processing unit, server, virtual server, circuit, engine, agent, instrument, or other logic device, such as a programmable logic array configured to communicate with the data repository 145 and other computing devices (e.g., at least one client computing device 150, at least one content provider computing device 155, or at least one service provider computing device 160) via at least one computer network 165. The network 165 may include computer networks such as the Internet, a local area network, a wide area network, a metropolitan area network or other area network, an intranet, a satellite network, other computer networks such as voice or data mobile telephone communication networks, and combinations thereof.
[0023] Network 165 may include or constitute a display network, for example, a subset of information resources available on the Internet that are associated with a content placement or search engine results system, or that are eligible to include third-party content items as part of a content item placement campaign. Network 165 may be used by data processing system 105 to access information resources, such as web pages, websites, domain names, or uniform resource locators that may be presented, output, rendered, or displayed by client computing device 150. For example, via network 165, a user of client computing device 150 may access information or data provided by data processing system 105, content provider computing device 155, or service provider computing device 160.
[0024] Network 165 may include, for example, a point-to-point network, a broadcast network, a wide area network, a local area network, a telecommunications network, a data communications network, a computer network, an ATM (Asynchronous Transfer Mode) network, a SONET (Synchronous Optical Network) network, an SDH (Synchronous Digital Hierarchy) network, a wireless network, or a wired network, and combinations thereof. Network 165 may include wireless links, such as infrared channels or satellite bands. The topology of network 165 may include a bus, star, or ring network topology. Network 165 may include a mobile phone network using any one or more protocols for communication between mobile devices, including Advanced Mobile Phone Protocol ("AMPS"), Time Division Multiple Access ("TDMA"), Code Division Multiple Access ("CDMA"), Global System for Mobile Communications ("GSM"), General Packet Radio Service ("GPRS"), or Universal Mobile Telecommunications System ("UMTS"). Different types of data may be transmitted via different protocols, or the same type of data may be transmitted via different protocols.
[0025] The client computing device 150, the content provider computing device 155, and the service provider computing device 160 may each include at least one logic device (such as a computing device with a processor) to communicate with each other or with the data processing system 105 via the network 165. The client computing device 150, the content provider computing device 155, and the service provider computing device 160 may each include at least one server, processor, or memory, or a plurality of computing resources or servers located in at least one data center. The client computing device 150, the content provider computing device 155, and the service provider computing device 160 may each include at least one computing device, such as a desktop computer, a laptop computer, a tablet computer, a personal digital assistant, a smartphone, a portable computer, a server, a thin client computer, a virtual server, or other computing device.
[0026] Client computing device 150 may include: at least one sensor 151, at least one transducer 152, at least one audio driver 153, and at least one speaker 154. Sensor 151 may include a microphone or an audio input sensor. Transducer 152 may convert audio input into an electronic signal, or vice versa. Audio driver 153 may include a script or program executed by one or more processors of client computing device 150 to control sensor 151, transducer 152 or audio driver 153, as well as other components of client computing device 150 to process audio input or provide audio output. Speaker 154 may transmit audio output signals.
[0027] The client computing device 150 may be associated with an end user who enters a voice query as audio input into the client computing device 150 (via the sensor 151) and receives audio output in the form of computer-generated speech, which may be provided from the data processing system 105 (or the content provider computing device 155 or the service provider computing device 160) to the client computing device 150 and output from the speaker 154. The audio output may correspond to an action data structure received from the direct action API 135 or a content item selected by the content selector component 125. The computer-generated speech may include a recording from a real person or computer-generated speech.
[0028] The content provider computing device 155 (or the data processing system 105 or the service provider computing device 160) can provide an audio-based content item or action data structure for display as an audio output by the client computing device 150. The action data structure or content item can include an organic response or offer for a product or service, such as a voice-based message stating "Today it will be sunny and 80 degrees at the beach" as an organic response to the voice input query "Is today a beach day?" The data processing system 105 (or other system 100 components, such as the content provider computing device 155) can also provide a content item in response, such as a voice-based or text-based content item offering sunscreen.
[0029] The content provider computing device 155 or data repository 145 may include a memory to store a series of audio action data structures or content items that may be provided in response to voice-based queries. Action data structures and content items may include a packet-based data structure for transmission via network 165. The content provider computing device 155 may also provide audio-based or text-based content items (or other content items) to the data processing system 105, in which audio-based or text-based content items (or other content items) may be stored in the data repository 145. The data processing system 105 may select an audio action data structure or text-based content item and provide (or instruct the content provider computing device 155 to provide) the audio action data structure or text-based content item to the same or different client computing devices 150 in response to a query received from one of those client computing devices 150. The audio-based action data structure may be proprietary audio, or may be combined with text, image, or video data. The content item may be proprietary text, or may be combined with audio, image, or video data.
[0030] The service provider computing device 160 may include at least one service provider natural language processor (NLP) component 161 and at least one service provider interface 162. The service provider NLP component 161 (or other components, such as a direct action API of the service provider computing device 160) may engage with the client computing device 150 (via the data processing system 105 or bypassing the data processing system 105) to create a back-and-forth real-time voice or audio-based dialogue (e.g., a session) between the client computing device 150 and the service provider computing device 160. For example, the service provider interface 162 may receive data messages (e.g., action data structures or content items) or provide data messages to the direct action API 135 of the data processing system 105. The direct action API 135 may also generate action data structures independently of or without input from the service provider computing device 160. The service provider computing device 160 and the content provider computing device 155 may be associated with the same entity. For example, content provider computing device 155 may create, store, or make available content items for beach-related services, such as sunscreen, beach towels, or swimsuits, and service provider computing device 160 may establish a session with client computing device 150 in response to a voice-input query about beach weather, beach directions, or recommendations for area beaches, and may provide these content items to an end user of client computing device 150 via an interface of the same client computing device 150 from which the query was received, a different interface of the same client computing device 150, or an interface of a different client computing device. Data processing system 105 may also establish a session with a client computing device (including service provider computing device 160 or bypassing it) via direct action API 135, NLP component 110, or other components to, for example, provide an organic response to a beach-related query.
[0031] Data repository 145 may include one or more local or distributed databases and may include a database management system. Data repository 145 may include computer data storage or memory and may store one or more parameters 146, one or more policies 147, content data 148, or templates 149, among other data. Parameters 146, policies 147, and templates 149 may include information such as rules for a voice-based session between client computing device 150 and data processing system 105 (or service provider computing device 160). Content data 148 may include content items or associated metadata for audio output and input audio messages that may be part of one or more communication sessions with client computing device 150.
[0032] System 100 can optimize the processing of action data structure and content item in voice activated data grouping (or other protocols) environment.For example, data processing system 105 can comprise voice activated assistant service, voice command device, intelligent personal assistant, knowledge navigator, event plan or other assistant program or their part.Data processing system 105 can provide one or more examples of action data structure as the audio output for display to complete the task relevant with input audio signal from client computing device 150.For example, among other things, data processing system can communicate with service provider computing device 160 or other third party computing device to generate the action data structure with the information of relevant beach.For example, end user can be entered into client computing device 150 with input audio signal: "OK, I would like to go to the beach this weekend (OK, I want to go to the beach this weekend)", and action data structure can indicate the weekend weather forecast of regional beach, such as, "it will be sunny and 80 degrees at the beach on Saturday, with high tide at 3pm (Saturday will be sunny, beach temperature 80 degrees, high tide time 3 o'clock in the afternoon)".
[0033] The action data structure may include multiple organic or non-sponsored responses to an input audio signal. For example, the action data structure may include a beach weather forecast or directions to the beach. The action data structure in this example includes organic or non-sponsored content that directly responds to the input audio signal. The content items that respond to the input audio signal may include sponsored or non-organic content, such as an offer to buy sunscreen from a convenience store located near the beach. In this example, the organic action data structure (beach weather forecast) responds to the input audio signal (a query related to the beach), and the content item (a reminder or offer for sunscreen) also responds to the same input audio signal. The data processing system 105 may adjust system 100 parameters (e.g., power usage, available displays, display formats, memory requirements, bandwidth usage, power capacity, or time of input power (e.g., internal battery or external power source, such as power from a wall outlet)) to provide the action data structure and content items to different candidate interfaces on the same client computing device 150 or to different candidate interfaces on different client computing devices 150.
[0034] The data processing system 105 may include an application, script, or program installed at the client computing device 150, such as an application for transmitting an input audio signal (e.g., as data packets via a packetized or other protocol-based transmission) to at least one interface 115 of the data processing system 105 and driving a component of the client computing device 150 to render an output audio signal (e.g., for an action data structure) or other output signal (e.g., a content item). The data processing system 105 may receive a data packet or other signal that includes or identifies an input audio signal. For example, the data processing system 105 may execute or run the NLP component 110 to receive the input audio signal.
[0035] The NLP component 110 can convert an audio input signal into recognized text by comparing the input signal with a set of stored representative audio waveforms (e.g., in the data repository 145) and selecting the closest match. Representative waveforms are generated across a large set of users and can be enhanced with speech samples. After converting the audio signal into recognized text, the NLP component 110 can match the text to words associated with actions that the data processing system 105 can service, for example, through training across users or manually specified.
[0036] The audio input signal may be detected by a sensor 151 (e.g., a microphone) of the client computing device. The sensor 151 may be referred to as an interface of the client computing device 150. Via a transducer 152, an audio driver 153, or other components, the client computing device 150 may provide the input audio signal to the data processing system 105 (e.g., via a network 165), where the audio input signal may be received (e.g., via an interface 115) and provided to the NLP component 110 or stored in the data repository 145 as content data 148.
[0037] The NLP component 110 may receive or otherwise obtain an input audio signal. From the input audio signal, the NLP component 110 may identify at least one request or at least one trigger keyword corresponding to the request. The request may indicate the intent or subject of the input audio signal. The trigger keyword may indicate the type of action that is likely to be taken. For example, the NLP component 110 may parse the input audio signal to identify at least one request to go to the beach on the weekend. The trigger keyword may include: at least one word, phrase, root or partial word, or a derivative word indicating the action to be taken. For example, the trigger keyword "go" or "to go to" in the input audio signal may indicate the need for transportation or a trip away from home. In this example, the input audio signal (or the identified request) does not directly express the intention for transportation, but the trigger keyword indicates that transportation is an auxiliary action for at least one other action indicated by the request.
[0038] The NLP component 110 can identify emotional keywords or emotional states in the input audio signal. The emotional keywords or states can indicate the attitude of the user when the user provides the input audio signal. The content selector component 125 can use the emotional keywords and states to select content items. For example, based on the emotional keywords and states, the content selector component 125 can skip the selection of content items. For example, if the NLP component 110 detects an emotional keyword such as "only" or "just" (for example, "Ok, just give me the results for the movie times"), the content selector component 125 can skip the selection of content items, thereby returning only the action data structure in response to the input audio signal.
[0039] The prediction component 120 (or other mechanism of the data processing system 105) can generate at least one action data structure associated with the input audio signal based on the request or trigger keyword. The action data structure can indicate information related to the subject of the input audio signal. The action data structure can include one or more actions, such as an organic response to the input audio signal. For example, the input audio signal "OK, I would like to go to the beach this weekend" can include at least one request indicating an interest in a beach weather forecast, surf report, or water temperature information and at least one trigger keyword, for example, "go" indicating a trip to the beach, such as the need for items that someone might want to bring to the beach, or the need for transportation to the beach. The prediction component 120 can generate or identify the subject of the at least one action data structure, an indication of a request for a beach weather forecast, and a subject of a content item, such as an indication of a query for sponsored content related to a day at the beach. Based on the request or trigger keyword, the prediction component 120 (or other system 100 components, such as the NLP component 110 or the direct action API 135) predicts, estimates, or otherwise determines the subject of an action data structure or content item. Based on the subject, the direct action API 135 can generate at least one action data structure and can communicate with at least one content provider computing device 155 to obtain at least one content item 155. The prediction component 120 can access parameters 146 or policies 147 in the data repository 145 to determine or otherwise estimate the request for the action data structure or content item. For example, the parameters 146 or policies 147 can indicate a request for a weekend weather forecast action for the beach or a request for content items related to going to the beach, such as a content item for sunscreen.
[0040] The content selector component 125 can obtain an indication of any interest or request for an action data structure or a content item. For example, the prediction component 120 can provide an indication of an action data structure or a content item to the content selector component 125 directly or indirectly (e.g., via a data repository 145). The content selector component 125 can obtain this information from the data repository 145, where it can be stored as part of the content data 148. The indication of the action data structure can notify the content selector component 125 that regional beach information is needed, such as a weather forecast or a product or service that an end user may need for a beach trip. The NLP component 110 can detect keywords associated with a private mode and temporarily place one or more client computing devices 150 associated with the user in a private mode. For example, the input audio signal can include "Ok, private mode" or "Ok, don't save this search." When in private mode, the content selector component 125 may not store the indication of interest, the request for the action data structure, or the selected content item in the data repository 145 as part of the content data 148. In private mode, when the data processing system 105 is not in private mode, the data processing system 105 does not use the input signal (and the data and associations generated by the input signal) to select subsequent content items and action data structures. For example, during a first interaction, when a user wants to order a gift for a significant other from a given store via a speaker-based assistant device, the user can place the speaker-based assistant device in private mode. During a second subsequent interaction with the speaker-based assistant device (by the user or significant other), the content selector component 125 will not select content items associated with the request, action data structure, or content item from the first interaction. For example, the content selector component 125 will not select content items associated with the store where the gift was purchased. The content selector component 125 can use the request, action data store, or content item from the first interaction to select content items to be transmitted to the client computing device 150 that is not placed in private mode. Continuing with the above example, the content selector component 125 may use the request from the first interaction, the action data store, or the content item to select a content item for transmission to the interface of the user's mobile phone.
[0041] Based on the information received by the content selector component 125 (e.g., an indication of an upcoming trip to the beach), the content selector component 125 can identify at least one content item. The content item can be responsive to or related to the subject of the input audio query. For example, the content item can include a data message identifying a store near the beach that sells sunscreen or offering to take a taxi to the beach. The content selector component 125 can query the data repository 145 to select or identify the content item, for example, from the content data 148. The content selector component 125 can also select the content item from the content provider computing device 155. For example, in response to the query received from the data processing system 105, the content provider computing device 155 can provide the content item to the data processing system 105 (or a component thereof) for ultimate output by the client computing device 150 that originated the input audio signal, or for output by the same end user via a different client computing device 150.
[0042] The audio signal generator component 130 can generate or otherwise obtain an output signal comprising a content item (and an action data structure) in response to an input audio signal. For example, the data processing system 105 can execute the audio signal generator component 130 to generate or create an output signal corresponding to the action data structure or content item. The interface component 115 of the data processing system 105 can transmit one or more data packets comprising the output signal to any client computing device 150 via a computer network 165. The interface 115 can be designed, configured, constructed, or operated to receive and transmit information by using, for example, data packets. The interface 115 can receive and transmit information by using one or more protocols (such as, a network protocol). The interface 115 can include: a hardware interface, a software interface, a wired interface, or a wireless interface. The interface 115 can facilitate converting or formatting data from one format to another. For example, the interface 115 can include an application programming interface that includes definitions for communication between various components (such as, software components of the system 100).
[0043] Data processing system 105 may provide output signals including action data structures from data repository 145 or audio signal generator component 130 to client computing device 150. Data processing system 105 may provide output signals including content items from data repository 145 or audio signal generator component 130 to the same or different client computing device 150.
[0044] The data processing system 105 may also instruct the content provider computing device 155 or the service provider computing device 160 to provide an output signal (e.g., an output signal corresponding to an action data structure or a content item) to the client computing device 150 via data packet transmission. The output signal may be obtained, generated, converted into one or more data packets (or other communication protocols), or transmitted from the data processing system 105 (or other computing device) to the client computing device 150 as one or more data packets (or other communication protocols).
[0045] The content selector component 125 can select a content item or an action data structure as part of a real-time content selection process. For example, the action data structure can be provided to the client computing device 150 in a conversational manner in direct response to the input audio signal for transmission as an audio output by an interface of the client computing device 150. The real-time content selection process of identifying the action data structure and providing the content item to the client computing device 150 can occur in one minute or less from the time of the input audio signal and can be considered real-time. The data processing system 105 can also identify the content item and provide the content item to at least one interface of the client computing device 150 that initiated the input audio signal or a different client computing device 150.
[0046] For example, the action data structure (or content item) obtained or generated by the audio signal generator component 130 and transmitted to the client computing device 150 via the interface 115 and the computer network 165 can cause the client computing device 150 to execute the audio driver 153 to drive the speaker 154 to generate a sound wave corresponding to the action data structure or content item. The sound wave can include words of the action data structure or content item or words corresponding to the action data structure or content item.
[0047] A sound wave representing an action data structure can be output from the client computing device 150 separately from the content item. For example, the sound wave can include the audio output: "Today it will be sunny and 80 degrees at the beach." In this example, the data processing system 105 obtains an input audio signal: for example, "OK, I would like to go to the beach this weekend." From this information, the NLP component 110 identifies at least one request or at least one trigger keyword, and the prediction component 120 uses the request or trigger keyword to identify a request for an action data structure or a content item. The content selector component 125 (or other component) can identify, select, or generate a content item for, for example, sunscreen available near the beach. The direct action API 135 (or other component) can identify, select, or generate an action data structure for, for example, a weekend beach weather forecast. The data processing system 105 or a component thereof, such as the audio signal generator component 130, can provide the action data structure for output by the interface of the client computing device 150. For example, a sound wave corresponding to the motion data structure may be output from client computing device 150. Data processing system 105 may provide content items for output by different interfaces of the same client computing device 150 or by interfaces of different client computing devices 150.
[0048] The packet-based transmission of the action data structure by data processing system 105 to client computing device 150 may include a direct or real-time response to the input audio signal "OK, I would like to go to the beach this weekend," such that the packet-based data transmission over computer network 165 as part of a communication session between data processing system 105 and client computing device 150 has the flow and feel of a real-time human-to-human conversation. The packet-based data transmission communication session may also include content provider computing device 155 or service provider computing device 160.
[0049] The content selector component 125 can select content items or action data structures based on at least one request or at least one trigger keyword of the input audio signal. For example, the request of the input audio signal "OK, I would like to go to the beach this weekend" can indicate a beach theme, a trip to the beach, or an item that facilitates a trip to the beach. The NLP component 110 or the prediction component 120 (or other data processing system 105 components executed as part of the direct action API 135) can recognize the trigger keyword "go," "go to," or "to go to" and can determine a transportation request to the beach based at least in part on the trigger keyword. The NLP component 110 (or other system 100 components) can also determine bids for content items related to beach activities, such as bids for sunscreen or beach umbrellas. Thus, the data processing system 105 can infer an action by being a secondary request (e.g., a request for sunscreen) that is not the primary request or theme of the input audio signal (information about the beach this weekend).
[0050] The action data structure and content item can correspond to the subject of the input audio signal. The direct action API 135 can, for example, execute a program or script from the NLP component 110, the prediction component 120, or the content selector component 125 to identify the action data structure or content item of one or more of these actions. The direct action API 135 can perform the specified action to satisfy the end user's intention determined by the data processing system 105. Depending on the action specified in its input, the direct action API 135 can execute code or a dialogue script that identifies the parameters required to fulfill the user's request. Such code can, for example, search for additional information (such as the name of the home automation service) in the data repository 145, or such code can provide an audio output for rendering at the client computing device 150 to ask the end user a question, such as the intended destination of the requested taxi. The direct action API 135 can determine the necessary parameters and can encapsulate the information into an action data structure, which can then be sent to another component (such as the content selector component 125 or the service provider computing device 160) to be fulfilled.
[0051] The direct action API 135 of the data processing system 105 can generate an action data structure based on a request or trigger keyword. The action data structure can be generated in response to a request of an input audio signal. The action data structure can be included in a message transmitted to or received by a service provider computing device 160. Based on the input audio signal parsed by the NLP component 110, the direct action API 135 can determine which service provider computing device 160 (if any) among multiple service provider computing devices 160 the information should be sent to. For example, if the input audio signal includes "OK, I would like to go to the beach this weekend", the NLP component 110 can parse the input audio signal to identify a request or trigger keyword, such as the trigger keyword "to go to" as an indication that a taxi is needed. The direct action API 135 can encapsulate the request into an action data structure for transmission as a message to the service provider computing device 160 of the taxi service. The message can also be passed to the content selector component 125. The action data structure can include information for completing the request. In this example, the information may include a pick-up location (e.g., home) and a destination location (e.g., beach). The direct action API 135 may retrieve a template 149 from the repository 145 to determine which fields to include in the action data structure. The direct action API 135 may retrieve content from the repository 145 to obtain information for the fields in the data structure. The direct action API 135 may use this information to populate the fields in the template to generate the data structure. The direct action API 135 may also use data from the input audio signal to populate the fields. The template 149 may be standardized for a category of service providers or may be standardized for a specific service provider. For example, a ride-sharing service provider may use the following standardized template 149 to create a data structure: {client_device_identifier; authentication_credentials; pickup_location; destination_location; no_passengers; service_level}.
[0052] The content selector component 125 can identify, select, or obtain multiple content items generated from multiple content selection processes. The content selection process can be real-time, for example, part of the same conversation, communication session, or a series of communication sessions between the data processing system 105 and the client computing device 150 involving a common topic. For example, the conversation can include asynchronous communications separated by hours or days. The conversation or communication session can last for a period of time from the receipt of a first input audio signal until the outcome of a final action related to the first input audio signal is estimated or known, or the data processing system 105 receives an indication of the termination or expiration of the conversation. For example, the data processing system 105 can determine that a conversation related to a weekend beach trip begins when the input audio signal is received and expires or terminates at the end of the weekend (e.g., Sunday evening or Monday morning). The data processing system 105 that provides one or more interfaces for the client computing device 150 or another client computing device 150 to render action data structures or content items during the active period of the conversation (e.g., from the receipt of the input audio signal until the determined expiration time) can be considered to operate in real time. In this example, the content selection process and rendering of the content item and action data structures occurs in real time.
[0053] The interface management component 140 can poll, determine, identify, or select an interface for rendering action data structures and content items. For example, the interface management component 140 can identify one or more candidate interfaces of the client computing device 150. The candidate interfaces can be associated with an end user who enters an input audio signal (e.g., "What is the weather at the beach today?") into a client computing device 150 via an audio interface. The interface can include hardware such as a sensor 151 (e.g., a microphone), a speaker 154, or a screen size of a computing device, either alone or in combination with a script or program (e.g., an audio driver 153) and an application, a computer program, an online document (e.g., a web page) interface, and combinations thereof. Each of the candidate interfaces can be within a single computing device or distributed across multiple devices. For example, a first candidate interface can be the speaker 154 of a speaker-based assistant device, while a second candidate interface can be the screen of a mobile device.
[0054] The candidate interface can be a hardware interface, a software interface, or a combination of the two. For example, the software interface can include a social media account (or a component of a social media account), a text messaging application, or an email account associated with the end user of the client computing device 150 that initiated the input audio signal. The software interface can be accessed on multiple computing devices. For example, the candidate software interface can include a web-based email program that the user accesses via a public computer. The user can later access the interface (e.g., a web-based email program) from the user's personal computer.
[0055] The interface may include an audio output of a smartphone (an example of a hardware-based interface), or an application-based messaging device installed on a smartphone or on a wearable computing device, and other client computing devices 150. The interface may also include display screen parameters (e.g., size, resolution), audio parameters, mobile device parameters (e.g., processing power, battery life, presence of installed applications or programs, or sensor 151 or speaker 154 capabilities), a content slot on an online document for text, image, or video rendering of a content item, a chat application, laptop computer parameters, smartwatch or other wearable device parameters (e.g., an indication of its display or processing capabilities), or virtual reality headset parameters.
[0056] The interface management component 140 can poll multiple interfaces to identify candidate interfaces. Candidate interfaces include interfaces that have the ability to render a response to an input audio signal (e.g., an action data structure that is output as audio, or a content item that can be output in various formats including non-audio formats). The interface management component 140 can determine parameters or other capabilities of the interfaces to determine whether they are (or are not) candidate interfaces. For example, the interface management component 140 can determine, based on parameters 146 of the content item or the first client computing device 150 (e.g., a smart watch wearable device), that the smart watch includes an available visual interface of sufficient size or resolution for rendering the content item. The interface management component 140 can also determine that the client computing device 150 that initiates the input audio signal has speaker 154 hardware and installed programs, such as an audio driver or other script for rendering the action data structure.
[0057] The interface management component 140 can determine a utilization value for a candidate interface. The utilization value can indicate whether the candidate interface can (or cannot) render an action data structure or content item provided in response to an input audio signal. The utilization value can also include a parameter indicating the number of previous instances of a content item (or related content item) transmitted or rendered via the interface. For example, each time the interface management component 140 selects an interface from a plurality of candidate interfaces, the interface management component 140 can record the selection in a data repository 145. The utilization value can include parameters 146 obtained from the data repository 145 or other parameters obtained from the client computing device 150, such as bandwidth or processing utilization or requirements, processing capacity, power requirements, battery status, memory utilization or capacity, or other interface parameters indicating whether the interface can be used to render an action data structure or content item. The battery status can indicate the type of power source (e.g., internal battery or external power source, such as via an output), the charging status (e.g., currently charging or not charging), or the remaining battery charge. The interface management component 140 can select an interface based on the battery status or the charging status.
[0058] The interface management component 140 can sort the candidate interfaces into a hierarchy or ranking based on the utilization values. For example, different utilization values (e.g., the number of times a content item is received, processing requirements, display screen size, accessibility to an end user) can be given different weights. The interface management component 140 can rank the one or more utilization values of the candidate interfaces based on the weight of one or more utilization values to determine the best corresponding candidate interface for rendering the content item (or action data structure). Based on the hierarchy, the interface management component 140 can select the highest ranked interface to render the content item.
[0059] Based on the utilization values of the candidate interfaces, the interface management component 140 can select at least one candidate interface as the selected interface for the content item. The selected interface for the content item can be the same interface from which the input audio signal was received (e.g., an audio interface of the client computing device 150) or a different interface (e.g., a text messaging-based application of the same client computing device 150, or an email account accessible from the same client computing device 150).
[0060] The interface management component 140 can sort the candidate interfaces according to a hierarchy or ranking based on the prior content items of the interfaces that are transmitted to the client computing device 150 or associated with the client computing device 150. For example, the interface management component 140 can access the content data 148 to determine which candidate interfaces of the candidate interfaces have been previously selected and what content items or action data structures are transmitted to each of those candidate interfaces. The interface management component 140 can rank the candidate interfaces based on which candidate interfaces of the candidate interfaces have recently received content items or action data structures related to the currently selected content items or action data structures. For example, when ranking multiple candidate interfaces for transmitting content items associated with a music event, the interface management component 140 can determine which candidate interfaces of the candidate interfaces have previously provided content items associated with the music event. The ranking of candidate interfaces that have recently received content items associated with the music event can be made higher than that of candidate interfaces that have not recently received (or have never received) content items associated with the music event.
[0061] The interface management component 140 can rank the candidate interfaces based on which of the candidate interfaces have recently received responses (or other interactions) to content items or action data structures related to the currently selected content item or action data structure. For example, a response to a content item can include clicking a link in the content item or a request for additional information related to the content item. Sorting the candidate interfaces based on the time at which the candidate interfaces received responses to content items or action data structures related to the currently selected content item or action data structure can reduce wasteful network data transmission by transferring the content item or action data structure to the interface where it was previously useful to the user. For example, in response to an input audio signal "Ok, how do I get to the restaurant?", the data processing system 105 can provide instructions for going to the restaurant and then respond with the query "would you like additional details about the restaurant?" If the user responds in the affirmative, the interface management component 140 can select a candidate interface associated with the user's mobile device, speaker-based assistant device, or other client computing device. During the first previous interaction, the user may not have interacted with the content item transmitted to the speaker-based assistant device's interface, but may have interacted with the content item transmitted to the mobile device's interface. In this example, the interface management component 140 can rank the mobile device's interface relatively higher than the speaker-based assistant device's interface among the candidate devices.
[0062] The interface management component 140 can rank the candidate interfaces to throttle the content items or action data structures transmitted to a given interface over time. For example, the interface management component 140 can record in the content data 148 or other portion of the data repository 145 each instance that an interface is selected from a plurality of candidate interfaces. Once the rate at which a candidate interface is selected increases (e.g., the number of times a candidate interface is selected exceeds a predetermined number during a certain interval), the interface management component 140 can lower the ranking of a given candidate interface.
[0063] The interface management component 140 can select an interface for the content item that is an interface of a client computing device 150 that is different from the device that originated the input audio signal. For example, the data processing system 105 can receive an input audio signal from a first client computing device 150 (e.g., a smart phone) and can select an interface, such as a display of a smart watch (or any other client computing device for rendering the content item). Multiple client computing devices 150 can all be associated with the same end user. The data processing system 105 can determine that multiple client computing devices 150 are associated with the same end user based on information received with the end user's consent (such as user access to a public social media or email account across multiple client computing devices 150).
[0064] Interface management component 140 may also determine that an interface is unavailable. An interface may be unavailable if an instance of a content item was previously transferred to the interface. Interface management component 140 may determine that the interface is unavailable if the previously transferred content item is associated with the content item for which interface management component 140 is currently selecting an interface. For example, the content items may originate from the same content provider device 155, or the current and previously transferred content items may relate to the same subject matter. Once interface management component 140 has selected a candidate interface, interface management component 140 may record a timeout period in data repository 145. The selected interface may remain unavailable until the timeout period expires. For example, during the timeout period, data processing system 105 may transfer the selected action data structure but determine not to transfer the digital component to client computing device 150. A user associated with the interface may set the duration of each timeout period. Interface management component 140 may set a timeout period responsive to the length of the previous content item. For example, interface management component 140 may set a proportionally longer timeout period for the interface after a proportionally longer content item has been transferred to the interface.
[0065] A user associated with an interface can indicate to the interface management component 140 the types or categories of content items that can be transmitted to the interface. When the interface management component 140 selects an interface, if the content item is associated with a category that the user indicates should not be transmitted to the corresponding interface, the interface management component 140 can mark the interface as unavailable. Each content item can include one or more tags that indicate what category the content item belongs to. The interface management component 140 can parse the tags to determine whether the content item is associated with a category of content items that should not be transmitted to a given interface. For example, a user can set up a speaker-based assistant device that can be configured in the user's living room so that the assistant device's interface only receives content items suitable for the family. In this example, content items with a category tag of "alcohol" for example will not be sent to the interface of the assistant device in the user's living room.
[0066] The interface management component 140 can determine whether an interface is unavailable based on the state or characteristics of the interface or the associated computing device. For example, the interface management component 140 can poll the interface and determine that the battery state of the client computing device 150 associated with the interface is low or below a threshold level (such as 10%). Alternatively, the interface management component 140 can determine that the client computing device 150 associated with the interface lacks sufficient display screen size or processing power to render the content item, or can determine that the processor utilization is too high because the client computing device is currently executing another application, such as to stream content via the network 165. In these and other examples, the interface management component 140 can determine that the interface is unavailable and can eliminate the interface as a candidate for rendering the content item or action data structure.
[0067] Another characteristic that interface management component 140 can use to determine whether an interface is unavailable is the physical location of the candidate interface. Interface management component 140 can also determine the distance between the candidate interface and the location of the receiving interface (e.g., transducer 152) that received the input audio signal. Interface management component 140 can determine whether the distance between the candidate interface and the receiving interface is above a predetermined threshold at which the candidate interface is unavailable. For example, if a user is away from home and inputs an input audio signal into the microphone of the user's mobile phone, interface management component 140 can determine that the distance between the mobile phone and the user's speaker-based assistant device at home is above a predetermined distance threshold. Interface management component 140 can mark the speaker-based assistant device as an unavailable interface. For example, the threshold can be a distance between approximately 10 yards and approximately 100 yards, between approximately 10 yards and approximately 500 yards, or between approximately 10 yards and approximately 1000 yards. The threshold can be set to enable interface management component 140 to determine whether the candidate interface and the receiving interface are within the same building, room, or general neighborhood.
[0068] The interface management component 140 can determine to link a candidate interface accessible by a first client computing device 150 to an account of the end user, and to also link a second candidate interface accessible by a second client computing device 150 to the same account. For example, both client computing devices 150 can access the same social media account, for example, by installing an application or script at each client computing device 150. The interface management component 140 can also determine that multiple interfaces correspond to the same account, and can provide multiple different content items to the multiple interfaces corresponding to the common account. With the end user's consent, for example, the data processing system 105 can determine that the end user has accessed the account from different client computing devices 150. These multiple interfaces can be separate instances of the same interface (e.g., the same application installed on different client computing devices 150) or different interfaces, such as different applications for different social media accounts that are all linked to a common email address account and accessible from multiple client computing devices 150.
[0069] The interface management component 140 may also determine or estimate the distance between client computing devices 150 associated with candidate interfaces. For example, with the user's consent, the data processing system 105 may obtain an indication that the input audio signal originates from a smartphone or virtual reality headset computing device 150 and that the end user is associated with an active smartwatch client computing device 150. From this information, the interface management component may determine that the smartwatch is active, e.g., worn by the end user when the end user records the input audio signal into the smartphone, such that the two client computing devices 150 are within a threshold distance of each other. In another example, with the end user's consent, the data processing system 105 may determine the location of the smartphone, which is the source of the input audio signal, and may also determine that a laptop account associated with the end user is currently active. For example, the laptop may be logged into a social media account indicating that the user is currently active on the laptop. In this example, the data processing system 105 may determine that the end user is within the threshold distance of both the smartphone and the laptop, making the laptop a suitable choice for rendering the content item via the candidate interface.
[0070] The interface management component 140 can select an interface for the content item based on at least one utilization value indicating that the selected interface is most efficient for the content item. For example, among the candidate interfaces, the interface for rendering the content item at the smartwatch uses the least bandwidth because the content item is smaller and fewer resources can be used to transmit the content item. Alternatively, the interface management component 140 can determine that the candidate interface selected for rendering the content item is currently charging (e.g., connected to a power source) so that rendering the content item by the interface will not deplete the battery power of the corresponding client computing device 150. In another example, the interface management component 140 can select a candidate interface that currently performs fewer processing operations than another, unselected interface of, for example, a different client computing device 150, which is currently streaming video content over the network 165 and is less available for rendering the content item without delay.
[0071] The interface management component 140 (or other data processing system 105 components) can convert the content item for delivery in a modality compatible with the candidate interface. For example, if the candidate interface is a display of a smartwatch, a smart phone, or a tablet computing device, the interface management component 140 can determine the size of the content item for an appropriate visual display given the size of the display screen associated with the interface. The interface management component 140 can also convert the content item into a packet-based or other protocol format (including proprietary formats or industry standard formats) for transmission to the client computing device 150 associated with the selected interface. The interface selected by the interface management component 140 for the content item may include an interface accessible to the end user from multiple client computing devices 150. For example, the interface can be or include a social media account that the end user can access via the client computing device 150 (e.g., a smart phone) that initiates the input audio signal, as well as other client computing devices (such as a desktop computer or desktop computer or other mobile computing device).
[0072] The interface management component 140 can also select at least one candidate interface for the action data structure. The interface can be the same interface from which the input audio signal is obtained, for example, a voice-activated assistant service executed at the client computing device 150. The interface can be the same interface or a different interface than the one selected by the interface management component 140 for the content item. The interface management component 140 (or other data processing system 105 component) can provide the action data structure to the same client computing device 150 that originated the input audio signal for rendering as audio output as part of the assistant service. The interface management component 140 can also transmit or otherwise provide the content item to the selected interface for the content item in any converted modality suitable for rendering by the selected interface.
[0073] Thus, the interface management component 140 can provide an action data structure as audio output in response to an input audio signal received by the same client computing device 150 for interface rendering by the client computing device 150. The interface management component 140 can also provide content items for different interface rendering by the same client computing device 150 or a different client computing device 150 associated with the same end user. For example, an action data structure (e.g., "it will be sunny and 80 degrees at the beach on Saturday") can be provided for audio rendering by a client computing device as part of an assistant interface executed in part at the client computing device 150, and a content item—e.g., text, audio, or a combination of content items indicating "sunscreen is available from the convenience store near the beach"—can be provided for rendering by an interface of the same or a different computing device 150 (such as an email or text message accessible by the same or a different computing device 150 associated with the end user).
[0074] Separating the content items from the action data structure and sending the content items as, for example, text messages rather than audio messages can cause the processing power of the client computing device 150 accessing the content items to decrease, because, for example, text message data transmission is less computationally intensive than audio message data transmission. This separation can also reduce the power usage, memory storage, or transmission bandwidth used to render the content items. This causes the processing, power, and bandwidth efficiency of the system 100 and devices (such as, client computing devices 150 and data processing systems 105) to increase. This improves the efficiency of the computing devices handling these transactions and improves the speed at which the content items can be rendered. The data processing system 105 can process thousands, tens of thousands, or more input audio signals simultaneously, so bandwidth, power, and processing savings can be significant, rather than just incremental or incidental.
[0075] After delivering the action data structure to the client computing device 150, the interface management component 140 may provide or deliver the content item to the same client computing device 150 as the action data structure (or a different device). For example, the content item may be provided for rendering via a selected interface at the end of rendering the action data structure for audio output. The interface management component 140 may also provide the content item to the selected interface while configuring the action data structure to the client computing device 150. The interface management component 140 may provide the content item for delivery via the selected interface within a predetermined time period from the time the NLP component 110 receives the input audio signal. For example, the time period may be any time during the active length of the conversational dialogue. For example, if the input audio signal is "I would like to go to the beach this weekend," the predetermined time period may be any time from the time the input audio signal is received to the end of the weekend, e.g., the active period of the conversation. The predetermined time period may also be a time triggered by the client computing device 150 rendering the action data structure as an audio output, such as within 5 minutes, 1 hour, or a day of the rendering.
[0076] Interface management component 140 can provide an action data structure along with an indication of the presence of a content item to client computing device 150. For example, data processing system 105 can provide an action data structure that is rendered at client computing device 150 to provide an audio output: "it will be sunny and 80 degrees at the beach on Saturday, check your email for more information." The phrase "check your email for more information" can indicate the presence of a content item (e.g., a content item for sunscreen) provided to an interface (e.g., email) by data processing system 105. In this example, sponsored content can be provided as a content item to an email (or other) interface, and organic content (e.g., weather) can be provided as an action data structure for the audio output.
[0077] The data processing system 105 may also provide a prompt to the action data structure to query the user to determine the user's interest in obtaining the content item. For example, the action data structure may indicate "it will be sunny and 80 degrees at the beach on Saturday, would you like to hear about some services to assist with your trip?" In response to the prompt "would you like to hear about some services to assist with your trip?", the data processing system 105 may receive another input audio signal from the client computing device 150, such as "sure". The NLP component 110 may parse the response (e.g., "sure") and interpret it as an approval for the client computing device 150 to perform audio rendering of the content item. In response, the data processing system 105 may provide the content item for audio rendering by the same client computing device 150 from which the response "sure" originated.
[0078] The data processing system 105 can delay transmission of content items associated with the action data structure to optimize processing utilization. For example, the data processing system 105 provides the action data structure for the client computing device to render as audio output in real time, such as in a conversational manner, in response to receiving an input audio signal, and can delay content item transmission until an off-peak or non-peak period of data center usage, which makes data center utilization more efficient by reducing peak bandwidth usage, heat output, or cooling requirements. The data processing system 105 can also initiate conversions or other activities associated with the content item based on data center utilization or bandwidth metrics or requirements of the network 165 or the data center including the data processing system 105, such as booking a car service in response to a response to the action data structure or content item.
[0079] Based on a response to a content item or an action data structure for a subsequent action (such as clicking on a content item rendered via a selected interface), the data processing system 105 can identify a conversion, or initiate a conversion or action. The processor of the data processing system 105 can call the direct action API 135 to execute a script that facilitates the conversion action, such as booking a car with a car sharing service to take the end user to the beach or to take the end user away from the beach. The direct action API 135 can obtain content data 148 (or parameters 146 or policies 147) from the data repository 145 and data received from the client computing device 150, with the end user's consent, for determining the location, time, user account, logistics, or other information, to book a car with the car sharing service. Using the direct action API 135, the data processing system 105 can also communicate with the service provider computing device 160 to complete the conversion by making a car sharing ride reservation in this example.
[0080] Figure 2 A flowchart 200 is depicted for multimodal transmission of packetized data in a voice-activated computer network environment. A data processing system 105 may receive an input audio signal 205, e.g., "OK, I would like to go to the beach this weekend." In response, the data processing system generates at least one action data structure 210 and at least one content item 215. The action data structure 210 may include organic or non-sponsored content, such as an audio rendering of the statement "It will be sunny and 80 degrees at the beach this weekend" or "high tide is at 3 pm." The data processing system 105 may provide the action data structure 210 to the same client computing device 150 that originated the input audio signal 205 for rendering as output by, for example, a candidate interface of the client computing device 150 in real-time or conversational fashion as part of a digital or conversational assistant platform.
[0081] The data processing system 105 may select a candidate interface 220 as the selected interface for the content item 215 and may provide the content item 215 to the selected interface 220. The content item 215 may also include a data structure that is converted by the data processing system 105 into an appropriate modality for rendering by the selected interface 220. The content item 215 may include sponsored content, such as beach chair rentals for the day or offers for sunscreen. The selected interface 220 may be part of or executed by the same client computing device 150, or executed by a different device accessible to the end user of the client computing device 150. The transfer action data structure 210 and the content item 215 may occur at the same time or sequentially. The action data structure 210 may include an indicator that the content item 215 is being or will be transferred separately to the selected interface 220 via a different modality or format, thereby alerting the end user to the presence of the content item 215.
[0082] The action data structure 210 and the content item 215 can be provided separately for rendering to the end user. By separating the sponsored content (content item 215) from the organic response (action data structure 210), the action data structure 210 does not need to be provided with an audio or other alert indicating that the content item 215 is sponsored. This can reduce the bandwidth requirements associated with transmitting the action data structure 210 via the network 165 and can simplify the rendering of the action data structure 210, for example, without the need for an audio disclaimer or warning message.
[0083] The data processing system 105 may receive a response audio signal 225. The response audio signal 225 may include an audio signal such as, "great, please book me a hotel on the beach this weekend." The data processing system 105's receipt of the response audio signal 225 may cause the data processing system to call the direct action API 135 to perform a translation, such as booking a room at a hotel on the beach. The direct action API 135 may also communicate with at least one service provider computing device 160 to provide information to the service provider computing device 160 so that the service provider computing device 160 can complete or confirm the reservation process.
[0084] Figure 3A block diagram depicts an example method 300 for multimodal transmission of packetized data in a voice-activated computer network environment. Method 300 may include receiving a data packet (ACT 305). Method 300 may include identifying a request and a trigger keyword (ACT 310). Method 300 may include generating an action data structure (ACT 315). Method 300 may include selecting a content item (ACT 320). Method 300 may include polling an interface (ACT 325). Method 300 may include determining whether the interface is unavailable (ACT 330). Method 300 may include transmitting the action data structure to a client computing device (ACT 335).
[0085] As described above, method 300 may include receiving data packets (ACT 305). For example, NLP component 110 executed by data processing system 105 may receive data packets including input audio signals from client computing device 150. The data packets may be received via network 165 as data transmission based on packet or other protocols. Input (e.g., input audio signals) may be detected, recorded, or logged at an interface of client computing device 150. For example, the interface may be a microphone.
[0086] Method 300 may include identifying requests and trigger keywords (ACT 310). NLP component 110 may identify requests and trigger keywords in an input audio signal received by data processing system 105 as data packets. For example, NLP component 110 may parse the input audio signal to identify a request related to the subject matter of the input audio signal. NLP component 110 may parse the input audio signal to identify a trigger keyword that may indicate, for example, an action associated with the request.
[0087] Method 300 may include generating at least one action data structure (ACT 315). For example, direct action API 135 may generate an action data structure based on a request or trigger keyword identified in the input audio signal. The action data structure may indicate organic content or non-sponsored content related to the input audio signal.
[0088] Method 300 may include selecting at least one content item (ACT 320). For example, content selector component 125 may receive a request or trigger keyword and, based on this information, may select one or more content items. The content items may include sponsored items having a theme related to the theme of the request or trigger keyword. The content items may be selected by content selector component 125 via a real-time content selection process.
[0089] Method 300 may include polling a plurality of interfaces to determine candidate interfaces (ACT 325). The candidate interfaces may include interfaces associated with a user of a client computing device that transmits an input audio signal to data processing system 105. The candidate interfaces may be interfaces of the client computing device, or may be interfaces of different client computing devices. For example, a first candidate interface may be a speaker of a user's mobile phone, and a second candidate interface may be a speaker of a user's speaker-based assistant device. The candidate interfaces may include interfaces capable of rendering a selected content item (or action data structure).
[0090] Method 300 may include determining whether one or more interfaces are unavailable (ACT 330). Interface management component 140 may determine whether one or more candidate interfaces are unavailable, or whether an interface on a client computing device through which an input audio signal is received is unavailable. Interface management component 140 may determine that an interface is unavailable based on the type or category of content items transmitted to one or more interfaces associated with the client computing device providing the input audio signal. For example, interface management component 140 may determine whether a previous instance of a content item or a content item associated with the content item selected in ACT 320 was previously transmitted to one of the candidate interfaces. If any of the candidate interfaces received a previous instance of a content item or a content item associated with the content item selected in ACT 320, interface management component 140 may designate each of the candidate interfaces as unavailable. In some implementations, only candidate interfaces that previously received a previous instance of a content item or a content item associated with the content item selected in ACT 320 are marked as unavailable.
[0091] The interface management component 140 can determine that an interface is unavailable based on a state or characteristic of the interface (or associated computing device). For example, the interface management component 140 can query the interface to obtain a utilization value, such as parameter information or other characteristics of the interface. If, for example, a battery level associated with the interface is below a predetermined threshold, the interface management component 140 can mark the interface as unavailable. Based on determining that the interface is unavailable, the data processing system can determine not to transmit the digital component to the client computing device.
[0092] Method 300 may include transmitting an action data structure to a client computing device (ACT 335). Data processing system 105 may transmit the action data structure (instead of the digital components) to the client computing device that transmitted the input audio signal to data processing system 105. Data processing system 105 may transmit the action data structure (instead of the digital components) to a second or different client computing device. In some embodiments, the content item may be transmitted to an interface of a second or different client computing device 150. In some embodiments, the action data structure may be transmitted to an interface that detects or otherwise receives input (e.g., an input audio signal). Method 300 may also include converting the content item to a modality for rendering via a selected interface. For example, data processing system 105 or a component thereof (such as interface management component 140) may convert the content item for rendering in a content item slot of an online document (e.g., for display as an email (e.g., via a selected email interface) or as a text message for display in a chat application).
[0093] For example, based on which candidate interfaces among the candidate interfaces are determined by the interface management component to be available, the content item may be transferred to at least one of the candidate interfaces for rendering of the content item (or action data structure). For example, when one or more candidate interfaces among the candidate interfaces are marked as unavailable, the data processing system 105 may transfer the action data structure instead of the selected content item to one of the candidate interfaces (or the client computing device). For example, when deciding not to transfer the content item, the data processing system 105 may discard, exclude, or limit the transmission of the selected content item to the client computing device. In some embodiments, discarding, excluding, or limiting the transmission of the content item to the client computing device 150 may include not including the content item in a transmission with the action data structure or in a transmission within a predetermined time window of a transmission that includes the action data structure. The data processing system 105 may discard the selected content item by returning to ACT 320 and selecting a new content item (e.g., a second content item). The interface management component 140 may re-poll each of the interfaces (e.g., repeating ACT 325), or the method 300 may continue using the originally selected candidate interface. The data processing system 105 may transmit the second content item to the client computing device from which the input audio signal originated or to one of the candidate interfaces.
[0094] Figure 44 is a block diagram of an example computer system 400. Computer system or computing device 400 may include or be used to implement system 100 or components thereof, such as data processing system 105. Computing system 400 includes a bus 405 or other communication component for communicating information and a processor 410 or processing circuit coupled to bus 405 for processing information. Computing system 400 may also include one or more processors 410 or processing circuits coupled to bus 405 for processing information. Computing system 400 also includes a main memory 415 (such as a random access memory (RAM) or other dynamic storage device) coupled to bus 405 for storing information and instructions to be executed by processor 410. Main memory 415 may be or include data repository 145. Main memory 415 may also be used to store location information, temporary variables, or other intermediate information during execution of instructions by processor 410. Computing system 400 may further include a read-only memory (ROM) 420 or other static storage device coupled to bus 405 for storing static information and instructions for processor 410. A storage device 425 , such as a solid-state device, a magnetic disk, or an optical disk, may be coupled to bus 405 for persistent storage of information and instructions. Storage device 425 may include data repository 145 or may be part of data repository 145 .
[0095] The computing system 400 may be coupled to a display 435 (such as a liquid crystal display or an active matrix display) via bus 405 for displaying information to a user. An input device 430 (such as a keyboard including alphanumeric and other keys) may be coupled to bus 405 for communicating information and command selections to the processor 410. The input device 430 may include a touch screen display 435. The input device 430 may also include a cursor controller (such as a mouse, trackball, or cursor direction keys) for communicating direction information and command selections to the processor 410 and for controlling cursor movement on the display 435. For example, the display 435 may be a computer program product of the data processing system 105, the client computing device 150, or the like. Figure 1 part of other components.
[0096] The processes, systems, and methods described herein can be implemented by the computing system 400 in response to the processor 410 executing the arrangement of instructions contained in the main memory 415. Such instructions can be read into the main memory 415 by another computer-readable medium, such as the storage device 425. Execution of the arrangement of instructions contained in the main memory 415 enables the computing system 400 to perform the illustrative processes described herein. One or more processors in a multi-processing arrangement can also be employed to execute the instructions contained in the main memory 415. Hard-wired circuitry can be used in place of software instructions or in combination with software instructions and the systems and methods described herein. The systems and methods described herein are not limited to any specific combination of hardware circuitry and software.
[0097] Although already Figure 4 Although the subject matter including the operations described in this specification is described in terms of an example computing system, the subject matter including the operations described in this specification may be implemented in other types of digital electronic circuitry, or in computer software, firmware, or hardware (including the structures disclosed in this specification and their structural equivalents), or in combinations of one or more of them.
[0098] For situations where the systems discussed herein collect personal information about a user or can utilize personal information, the user can be provided with an opportunity to control whether a program or feature can collect personal information (e.g., information about the user's social network, social actions or activities, the user's preferences, or the user's location), or to control whether and / or how to receive content that may be more relevant to the user from a content server or other data processing system. In addition, before storing or using specific data, the specific data can be anonymized in one or more ways so that personally identifiable information is removed when generating parameters. For example, the user's identity can be anonymized so that the user's personally identifiable information cannot be determined, or the user's geographic location can be generalized (such as to a city, zip code, or state or county level) when location information is obtained so that the user's specific location cannot be determined. Thus, the user can control the way in which the content server collects and uses information about him or her.
[0099] The subject matter and operations described in this specification can be implemented in digital electronic circuit systems, or in computer software, firmware, or hardware (including the structures disclosed in this specification and their structural equivalents), or in a combination of one or more of these. The subject matter described in this specification can be implemented as one or more computer programs, for example, one or more circuits of computer program instructions encoded on one or more computer storage media for execution by a data processing device or to control the operation of the data processing device. Alternatively or in addition, the program instructions can be encoded on an artificially generated propagated signal, for example, a machine-generated electrical, optical, or electromagnetic signal, which is generated to encode information for transmission to a suitable receiver device for execution by the data processing device. The computer storage medium can be or include the following: a computer-readable storage device, a computer-readable storage substrate, a random or serial access memory array or device, or a combination of one or more of them. When the computer storage medium is not a propagated signal, the computer storage medium can be the source or destination of the computer program instructions encoded in the artificially generated propagated signal. The computer storage medium can also be or include the following: one or more separate components or media (for example, multiple CDs, disks, or other storage devices). The operations described in this specification can be implemented as operations performed by a data processing apparatus on data stored in one or more computer-readable storage devices or received from other sources.
[0100] The terms "data processing system," "computing device," "component," or "data processing apparatus" encompass various apparatuses, devices, and machines for processing data, including, for example, a programmable processor, a computer, a system on a chip, or multiple or combinations thereof. The apparatus may include specialized logic circuitry, such as an FPGA (field programmable gate array) or an ASIC (application-specific integrated circuit). In addition to hardware, the apparatus may also include code that creates an execution environment for the computer program in question, such as code constituting processor firmware, a protocol stack, a database management system, an operating system, a cross-platform runtime environment, a virtual machine, or a combination of one or more thereof. The apparatus and execution environment may implement a variety of different computing model infrastructures, such as web services, distributed computing, and grid computing infrastructures. The interface management component 140, the direct action API 135, the content selector component 125, the prediction component 120, or the NLP component 110, as well as other data processing system 105 components, may include or share one or more data processing apparatuses, systems, computing devices, or processors.
[0101] Computer programs (also referred to as programs, software, software applications, applications, scripts, or code) can be written in any form of programming language (including compiled or interpreted languages, declarative or procedural languages) and can be deployed in any form (including as stand-alone programs or modules, components, subroutines, objects, or other units suitable for use in a computing environment). A computer program can correspond to a file in a file system. A computer program can be stored in a portion of a file that holds other programs or data (e.g., one or more scripts stored in a markup language document), or in a single file dedicated to the program in question, or in multiple collaborative files (e.g., files storing one or more modules, subroutines, or partial codes). A computer program can be deployed to execute on one computer or on multiple computers located at one site or distributed across multiple sites and interconnected by a communication network.
[0102] The processes and logic flows described in this specification may be performed by one or more programmable processors executing one or more computer programs (e.g., components of the data processing system 105) to perform actions by operating on input data and generating output. The processes and logic flows may also be performed by, and the apparatus may be implemented as, special-purpose logic circuitry (e.g., an FPGA (field programmable gate array) or an ASIC (application-specific integrated circuit)). Devices suitable for storing computer program instructions and data include all forms of non-volatile memory, media, and storage devices, including, for example, semiconductor memory devices (e.g., EPROM, EEPROM, and flash memory devices), magnetic disks (e.g., internal hard disks or removable disks), magneto-optical disks, CD-ROM disks, and DVD-ROM disks. The processor and memory may be supplemented by or incorporated into special-purpose logic circuitry.
[0103] The subject matter described herein can be implemented in a computing system that includes a back-end component (e.g., as a data server), or a computing system that includes a middleware component (e.g., an application server), or a computing system that includes a front-end component (e.g., a client computer having a graphical user interface or web browser through which a user can interact with an implementation of the subject matter described in this specification), or a computing system that includes a combination of one or more such back-end components, middleware components, or front-end components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include local area networks ("LANs") and wide area networks ("WANs"), internets (e.g., the Internet), and peer-to-peer networks (e.g., point-to-point peer-to-peer networks).
[0104] Computing system (such as, system 100 or system 400) can comprise client and server.Client and server are generally far away from each other and usually interact through communication network (for example, network 165).The relationship of client and server is produced by computer program that runs on corresponding computer and has client-server relationship each other.In some embodiments, server transmits data (for example, the data grouping representing action data structure or content item) to client device (for example, in order to display data to the user that interacts with client device and receive the user input from this user to client device 150, or to service provider computing device 160 or content provider computing device 155).Data (for example, the result of user interaction) generated on client computing device can be received from client computing device at server (for example, by data processing system 105 from computing device 150 or content provider computing device 155 or service provider computing device 160).
[0105] Although operations are described in the drawings in a particular order, such operations need not be performed in the particular order shown or in sequential order, and not all illustrated operations need be performed. The actions described herein may be performed in a different order.
[0106] Separation of the various system components is not required in all embodiments, and the described program components may be included in a single hardware or software product. For example, the NLP component 110, the content selector component 125, the interface management component 140, or the prediction component 120 may be a single component, application, or program, or a logic device having one or more processing circuits, or part of one or more servers of the data processing system 105.
[0107] Now that some illustrative embodiments have been described, it will be apparent that the foregoing is illustrative rather than restrictive and is presented by way of example. Specifically, although many of the examples presented herein relate to specific combinations of method actions or system elements, these actions and these elements may be combined in other ways to achieve the same objectives. Actions, elements, and features discussed in conjunction with one embodiment are not intended to exclude similar effects in other embodiments or embodiments.
[0108] The phraseology and terminology used herein are for descriptive purposes and should not be construed as limiting. The use of "including," "having," "comprising," "involving," "characterized by," and variations thereof herein is intended to encompass the items listed thereafter, their equivalents, and additional items, as well as alternative embodiments consisting of the items listed thereafter exclusively. In one embodiment, the systems and methods described herein consist of one of the described elements, actions, or components, each combination of more than one of the described elements, actions, or components, or all of the described elements, actions, or components.
[0109] Any reference to an embodiment or element or action of a system and method mentioned herein in the singular may also include embodiments having a plurality of such elements, and any reference to any embodiment or element or action herein in the plural may also encompass embodiments including only a single element. Reference in the singular or plural is not intended to limit the presently disclosed systems or methods, their components, actions, or elements to a singular or plural configuration. Reference to any action or element based on any information, action, or element may include embodiments in which the action or element is based at least in part on any information, action, or element.
[0110] Any embodiment disclosed herein may be combined with any other embodiment or example, and references to "an embodiment," "some embodiments," "one embodiment," etc. are not necessarily mutually exclusive and are intended to indicate that a particular feature, structure, or characteristic described in connection with the embodiment may be included in at least one embodiment or example. Such terms as used herein do not necessarily all refer to the same embodiment. Any embodiment may be combined with any other embodiment, inclusively or exclusively, in any manner consistent with the aspects and embodiments disclosed herein.
[0111] References to "or" may be interpreted as inclusive, such that any term described using "or" may refer to any of a single described term, more than one described term, and all described terms. For example, a reference to "at least one of 'A' and 'B'" may include only 'A', only 'B', and both 'A' and 'B'. Such references used in conjunction with "including" or other open-ended terms may include additional items.
[0112] Where a technical feature in the drawings, detailed description, or any claims is followed by a reference numeral, the reference numeral has been included to increase the intelligibility of the drawings, detailed description, and claims. Therefore, neither the reference numeral nor the absence of a reference numeral has any limitation on the scope of any claim element.
[0113] The systems and methods described herein may be embodied in other specific forms without departing from the characteristics of the systems and methods described herein. The above embodiments are illustrative rather than restrictive of the systems and methods described herein. The scope of the systems and methods described herein is therefore indicated by the appended claims rather than the foregoing description, and changes that come within the meaning and range of equivalents of the claims are intended to be embraced therein.
Claims
1. A system for transmitting packetized data in a voice activated packet based computer network environment, comprising: A data processing system having one or more processors coupled to a memory for: identifying a request from a data packet comprising an input audio signal obtained via a first interface of a client device; selecting a digital component based on the request identified from the data packet; identifying a plurality of candidate interfaces associated with the client device; determining a plurality of utilization values corresponding to the plurality of candidate interfaces; selecting a second interface from the plurality of candidate interfaces based on the plurality of utilization values, the second interface being on a second client device different from the client device; determining that the third interface for output is unavailable based on a prior instance of a digital component corresponding to the third interface; Based on the third interface being unavailable, determining not to transmit the digital component to the third interface; as well as The digital component is transmitted to the second interface associated with the client device to output the digital component.
2. The system according to claim 1, comprising the data processing system for: generating an action data structure based on the request identified from the requests, the action data structure including a response to the input audio signal and separated from the digital component; and A third interface is selected from the plurality of candidate interfaces, and the action data structure is transmitted to the third interface based on the plurality of utilization values.
3. The system according to claim 1, comprising the data processing system for: determining availability of the second interface for outputting the digital component based on a utilization value corresponding to the second interface; and Transmitting the digital component to the second interface is determined based on availability of the second interface.
4. The system according to claim 1, comprising the data processing system for: ranking the plurality of candidate interfaces based on the corresponding plurality of utilization values; and The second interface is selected from a plurality of candidate interfaces based on the ranking of the plurality of utilization values.
5. The system of claim 1 , comprising the data processing system for: determining a distance between the first interface and the second interface of the client device obtaining the input audio signal; and Availability of the second interface to be presented is determined based on a comparison between the distance and a threshold value.
6. The system of claim 1 , comprising the data processing system to set a timeout period for the second interface after transmitting the digital component, the timeout period before which transmission of a second digital component to the second interface associated with the client device is permitted.
7. The system of claim 1, comprising the data processing system to transmit the digital component to the second interface associated with the client device for outputting an audio output.
8. The system of claim 1 , comprising the data processing system for determining the plurality of utilization values corresponding to the plurality of candidate interfaces, each utilization value in the plurality of utilization values indicating at least one of: processing capability, power requirements, battery status, memory utilization, or network bandwidth usage on a corresponding candidate interface in the plurality of candidate interfaces.
9. A method for transmitting packet data in a voice activated packet based computer network environment, comprising: identifying, by a data processing system, a request from a data packet comprising an input audio signal obtained via a first interface of a client device; selecting, by the data processing system, a digital component based on the request identified from the data packet; identifying, by the data processing system, a plurality of candidate interfaces associated with the client device; determining, by the data processing system, a plurality of utilization values corresponding to the plurality of candidate interfaces; selecting, by the data processing system, a second interface from the plurality of candidate interfaces based on the plurality of utilization values, the second interface being on a second client device different from the client device; determining, by the data processing system, that the third interface for output is unavailable based on a prior instance of a digital component corresponding to the third interface; determining, by the data processing system, not to transmit the digital component to the third interface based on the third interface being unavailable; as well as The digital component is transmitted by the data processing system to the second interface associated with the client device to output the digital component.
10. The method according to claim 9, comprising: generating, by the data processing system, an action data structure based on the request identified from the requests, the action data structure including a response to the input audio signal and separated from the digital component; and The data processing system selects a third interface from the plurality of candidate interfaces, and transmits the action data structure to the third interface based on the plurality of utilization values.
11. The method according to claim 9, comprising: determining, by the data processing system, availability of the second interface for output based on a utilization value corresponding to the second interface; and The data processing system determines to transmit the digital component to the second interface based on availability of the second interface.
12. The method according to claim 9, comprising: Ranking, by the data processing system, the plurality of candidate interfaces based on the corresponding plurality of utilization values; and The second interface is selected, by the data processing system, from a plurality of candidate interfaces based on the ranking of the plurality of utilization values.
13. The method of claim 9, comprising the data processing system being configured to: determining, by the data processing system, a distance between the first interface and the second interface of the client device obtaining the input audio signal; and Availability of the second interface to be presented is determined by the data processing system based on a comparison between the distance and a threshold value.
14. The method according to claim 9, comprising: After transmitting the digital component, a timeout period is set by the data processing system for the second interface, before which a second digital component is allowed to be transmitted to the second interface associated with the client device.
15. The method according to claim 9, comprising: The digital component is transmitted by the data processing system to the second interface associated with the client device for output as at least one of audio output, image output, or text output.
16. The method according to claim 9, comprising: The data processing system determines the plurality of utilization values corresponding to the plurality of candidate interfaces, each utilization value in the plurality of utilization values indicating at least one of: processing capability, power requirements, battery status, memory utilization, or network bandwidth usage on a corresponding candidate interface in the plurality of candidate interfaces.
17. A system for transmitting packetized data in a voice activated packet based computer network environment, comprising: A data processing system having one or more processors coupled to a memory for: identifying a request from a data packet comprising an input audio signal obtained via a first interface of a client device; selecting a digital component based on the request identified from the data packet; identifying a plurality of candidate interfaces associated with the client device; determining a plurality of utilization values corresponding to the plurality of candidate interfaces; selecting a second interface from the plurality of candidate interfaces based on the plurality of utilization values; determining availability of the second interface for outputting the digital component based on a utilization value corresponding to the second interface; determining that the third interface for output is unavailable based on a prior instance of a digital component corresponding to the third interface; Based on the third interface being unavailable, determining not to transmit the digital component to the third interface; as well as Based on the availability of the second interface, the digital component is transmitted to the second interface associated with the client device to output the digital component.
18. A system for transmitting packetized data in a voice activated packet based computer network environment, comprising: A data processing system having one or more processors coupled to a memory for: identifying a request from a data packet comprising an input audio signal obtained via a first interface of a client device; selecting a digital component based on the request identified from the data packet; identifying a plurality of candidate interfaces associated with the client device; determining a plurality of utilization values corresponding to the plurality of candidate interfaces; selecting a second interface from the plurality of candidate interfaces based on the plurality of utilization values; determining that the third interface for output is unavailable based on a prior instance of a digital component corresponding to the third interface; determining not to transmit the digital component to the third interface based on the third interface being unavailable; and The digital component is transmitted to the second interface associated with the client device to output the digital component.
19. A system for transmitting packetized data in a voice activated packet based computer network environment, comprising: A data processing system having one or more processors coupled to a memory for: identifying a request from a data packet comprising an input audio signal obtained via a first interface of a client device; selecting a digital component based on the request identified from the data packet; identifying a plurality of candidate interfaces associated with the client device; determining a plurality of utilization values corresponding to the plurality of candidate interfaces; selecting a second interface from the plurality of candidate interfaces based on the plurality of utilization values; determining a distance between the first interface and the second interface of the client device obtaining the input audio signal; determining availability of the second interface to present based on a comparison between the distance and a threshold; determining that the third interface for output is unavailable based on a prior instance of a digital component corresponding to the third interface; Based on the third interface being unavailable, determining not to transmit the digital component to the third interface; as well as The digital component is transmitted to the second interface associated with the client device to output the digital component.
Citation Information
Patent Citations
Device configuration-based function delivery
US20170237801A1