Natural Language Processing for Session Establishment with Service Providers

Through the data processing system, the voice commands are analyzed and the action data structure is generated using a model based on aggregated speech training, which solves the problems of low transmission efficiency and accuracy between different computing resources, and realizes efficient voice command processing and resource optimization.

CN113918896BActive Publication Date: 2025-07-29GOOGLE LLC
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202111061959.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2016-12-30
Filing Date
2017-08-31
Publication Date
2025-07-29
Estimated Expiration
2037-08-31

AI Technical Summary

Technical Problem

In the network service data transmission between different computing resources, there are problems of low processing efficiency, untimely response and waste of resources. Especially in voice-based computing environments, the inconsistency of the speech model makes it difficult to guarantee the accuracy and consistency of the analytical instructions.

Method used

Through the data processing system, using a speech model based on aggregate speech training, the speech-based input is parsed and action data structures are generated, and routed to a third-party provider device, reducing resource consumption and processing time, and improving the reliability and efficiency of instructions.

Benefits of technology

It improves the efficiency of information transmission and processing between different computing resources, ensures the accuracy and consistency of voice commands, and reduces resource consumption and bandwidth utilization.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN113918896B_ABST
    Figure CN113918896B_ABST
Patent Text Reader

Abstract

The present disclosure relates to natural language processing for session establishment with a service provider. Packetized actions are routed in a computer network environment based on voice-activated data packets. The system may receive an audio signal detected by a microphone of a device. The system may parse the audio signal to identify a trigger keyword and a request, and generate an action data structure. The system may send the action data structure to a third-party provider device. The system may receive an indication of a communication session established with the device from the third-party provider device.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Division Explanation

[0002] This application is a divisional application of Chinese Patent Application No. 201780001369.6 with a filing date of August 31, 2017.

[0003] Cross - Reference to Related Applications

[0004] This application claims the benefit and priority of U.S. Patent Application No. 15 / 395,689, filed on December 30, 2016, and titled "AUDIO - BASED DATA STRUCTURE GENERATION", which is hereby incorporated by reference in its entirety for all purposes. Background Art

[0005] Over - network transmission of network traffic data between computing devices, whether packet - based or otherwise, can impede a computing device from properly processing network traffic data, completing operations related to the network traffic data, or responding to the network traffic data in a timely manner. If the responding computing device is at or beyond its processing capacity, over - network transmission of network traffic data can also complicate data routing or degrade the quality of the response, which can lead to inefficient bandwidth utilization. Controlling network transmissions corresponding to content item objects can become complex due to the large number of content item objects that can initiate network traffic data transmissions between computing devices. Summary of the Invention

[0006] The present disclosure generally aims to improve the efficiency and effectiveness of information transmission and processing across fundamentally different computing resources. For fundamentally different computing resources, it is challenging to efficiently process and consistently and accurately parse audio - based instructions in a voice - based computing environment. For example, fundamentally different computing resources may not be able to access the same voice model, or may access an outdated or out - of - sync voice model, which can make accurately and consistently parsing audio - based instructions challenging.

[0007] The systems and methods of the present disclosure generally aim at a data processing system for routing packetized action data via a computer network. The data processing system can specifically use a voice model trained based on aggregation speech to process voice - based inputs to parse voice - based instructions and create action data structures. The data processing system can send the action data structures to one or more components of a data processing system or a third - party provider device, thereby allowing the third - party provider device to process the action data structures without having to process voice - based inputs. By processing voice - based inputs for multiple third - party provider devices, the data processing system can improve the reliability, efficiency, and accuracy of processing and executing voice - based instructions.

[0008] At least one aspect is directed to a system for routing packetized actions via a computer network. The system can include a natural language processor (“NLP”) component executed by a data processing system. The NLP component can receive data packets including an input audio signal detected by a sensor of a client device via an interface of the data processing system. The NLP component can parse the input audio signal to identify a request and a trigger keyword corresponding to the request. The data processing system can include a direct action application programming interface (“API”). The direct action API can generate an action data structure in response to the request based on the trigger keyword. The direct action API can send the action data structure to a third-party provider device to cause the third-party provider device to invoke a conversation application programming interface and establish a communication session between the third-party provider device and the client device. The data processing system can receive an indication from the third-party provider device that a communication session has been established between the third-party provider device and the client device.

[0009] At least one aspect is directed to a method for routing packetized actions via a computer network. The method can include a data processing system receiving data packets including an input audio signal detected by a sensor of a client device via an interface of the data processing system. The method can include the data processing system parsing the input audio signal to identify a request and a trigger keyword corresponding to the request. The method can include the data processing system generating an action data structure in response to the request based on the trigger keyword. The method can include the data processing system sending the action data structure to a third-party provider device to cause the third-party provider device to invoke a conversation application programming interface and establish a communication session between the third-party provider device and the client device. The method can include the data processing system receiving an indication from the third-party provider device that a communication session has been established between the third-party provider device and the client device.

[0010] These and other aspects and implementations are discussed in detail below. The above information and the following detailed description include illustrative examples of various aspects and implementations and provide an overview or framework for understanding the nature and characteristics of the claimed aspects and implementations. The drawings provide illustrations and further understanding of the various aspects and implementations and are incorporated into and form a part of this specification. BRIEF DESCRIPTION OF THE DRAWINGS

[0011] The drawings are not intended to be drawn to scale. Like reference numerals and names in the various drawings indicate like elements. For clarity purposes, not every component may be labeled in each drawing. In the drawings:

[0012] Figure 1 is an illustration of a system for routing packetized actions via a computer network.

[0013] Figure 2 is an illustration of the operation of the system for routing packetized actions via a computer network.

[0014] Figure 3 It is an illustration of the operation of the system routing packetized actions via a computer network.

[0015] Figure 4 It is an illustration of a method of routing packetized actions via a computer network.

[0016] Figure 5 It is a block diagram of the general architecture of a computer system illustrating the elements that can be employed to implement the systems and methods described and illustrated herein. DETAILED DESCRIPTION

[0017] The following is a more detailed description of various concepts related to methods, apparatuses, and systems for routing packetized actions via a computer network, as well as implementations of these methods, apparatuses, and systems. The various concepts introduced above and discussed in more detail below can be implemented in any of a number of ways.

[0018] The present disclosure is generally directed to improving the efficiency and effectiveness of information transmission and processing across radically different computing resources. For radically different computing resources, it is challenging to efficiently process and consistently and accurately parse audio-based instructions in a voice-based computing environment. For example, radically different computing resources may not be able to access the same voice model, or may access an outdated or out-of-sync voice model, which can make it challenging to accurately and consistently parse audio-based instructions.

[0019] The systems and methods of the present disclosure are generally directed to a data processing system for routing packetized actions via a computer network. The data processing system can specifically use a voice model trained based on aggregated speech to process voice-based inputs to parse voice-based instructions and create an action data structure. The data processing system can send the action data structure to one or more components of a data processing system or a third-party provider device, thereby allowing the third-party provider device to process the action data structure without having to process voice-based inputs. By processing voice-based inputs for multiple third-party provider devices, the data processing system can improve the reliability, efficiency, and accuracy of processing and executing voice-based instructions.

[0020] The present solution can reduce the amount of resource consumption, processor utilization, battery consumption, bandwidth utilization, the size of audio files, or the time consumed by a speaker by parsing voice-based instructions from an end user, constructing an action data structure using a template, and routing the action data structure to the corresponding third-party provider.

[0021] Figure 1FIG. illustrates an example system 100 for routing packetized traffic over a computer network. System 100 may include a content selection infrastructure. System 100 may include a data processing system 102. The data processing system 102 may communicate with one or more of a content provider computing device 106, a service provider computing device 108, or a client computing device 104 via a network 105. The network 105 may include a computer network such as the Internet, a local area network, a wide area network, a metropolitan area network, or other regional network, an intranet, a satellite network, and other communication networks such as a voice or data mobile telephone network. The network 105 may be used to access information resources such as web pages, websites, domain names, or uniform resource locators that may be presented, output, rendered, or displayed on at least one computing device 104 such as a laptop, desktop, tablet, personal digital assistant, smart phone, portable computer, or speaker. For example, a user of the computing device 104 may access information or data provided by the service provider 108 or the content provider 106 via the network 105. The computing device 104 may or may not include a display; for example, the computing device may include a limited type of user interface such as a microphone and a speaker. In some cases, the primary user interface of the computing device 104 may be a microphone and a speaker.

[0022] The network 105 may include or constitute a display network, e.g., a subset of information resources available on the Internet that are associated with content placement or search engine result systems and are eligible to include third-party content items as part of a content item placement activity. The network 105 may be used by the data processing system 102 to access information resources such as web pages, websites, domain names, or uniform resource locators that may be presented, output, rendered, or displayed by the client computing device 104. For example, a user of the client computing device 104 may access information or data provided by the content provider computing device 106 or the service provider computing device 108 via the network 105.

[0023] Network 105 can be any type or form of network and can include any one of the following: a peer-to-peer network, a broadcast network, a wide area network, a local area network, a telecommunications network, a data communication network, a computer network, an ATM (Asynchronous Transfer Mode) network, a SONET (Synchronous Optical Network) network, an SDH (Synchronous Digital Hierarchy) network, a wireless network, and a wired line network. Network 105 can include wireless links, such as infrared channels or satellite bands. The topology of network 105 can include a bus, star, or ring network topology. The network can include a mobile phone network that uses any one or more protocols for communication between mobile devices, the protocols including Advanced Mobile Phone Protocol (“AMPS”), Time Division Multiple Access (“TDMA”), Code Division Multiple Access (“CDMA”), Global System for Mobile Communications (“GSM”), General Packet Radio Service (“GPRS”), or Universal Mobile Telecommunications System (“UMTS”). Different types of data can be sent via different protocols, or the same type of data can be sent via different protocols.

[0024] System 100 can include at least one data processing system 102. The data processing system 102 can include at least one logic device, such as a computing device having a processor to communicate via network 105 with, for example, a computing device 104, a content provider device 106 (content provider 106), or a service provider device 108 (or service provider 108). The data processing system 102 can include at least one computing resource, server, processor, or memory. For example, the data processing system 102 can include multiple computing resources or servers located in at least one data center. The data processing system 102 can include multiple logically grouped servers and facilitate distributed computing techniques. The logical group of servers can be referred to as a data center, server farm, or machine farm. The servers can also be geographically dispersed. The data center or machine farm can be managed as a single entity, or the machine farm can include multiple machine farms. The servers within each machine farm can be heterogeneous - one or more of these servers or machines can operate according to one or more types of operating system platforms.

[0025] The servers in the machine farm can be stored together with associated storage systems in a high-density rack system and located in an enterprise data center. For example, integrating servers in this way can improve system manageability, data security, physical security of the system, and system performance by locating the servers and high-performance storage systems on a localized high-performance network. Centralization of all or some of the components of the data processing system 102 (including servers and storage systems) and coupling them with advanced system management tools allows for more efficient use of server resources, which saves power and processing requirements and reduces bandwidth usage.

[0026] System 100 may include, access, or otherwise interact with at least one service provider device 108. The service provider device 108 may include at least one logic device, such as a computing device having a processor to communicate, via network 105, for example, with computing device 104, data processing system 102, or content provider 106. The service provider device 108 may include at least one computing resource, server, processor, or memory. For example, the service provider device 108 may include multiple computing resources or servers located in at least one data center. The service provider device 108 may include one or more components or functionality of data processing system 102.

[0027] Content provider computing device 106 may provide audio-based content items for display by client computing device 104 as audio output content items. The content items may include offers of goods or services, such as a voice-based message stating: "Do you want me to book a taxi for you?". For example, content provider computing device 155 may include memory for storing a series of audio content items that may be provided in response to voice-based queries. Content provider computing device 106 may also provide audio-based content items (or other content items) to data processing system 102, where they may be stored in data repository 124. Data processing system 102 may select the audio content items and provide (or instruct content provider computing device 104 to provide) the audio content items to client computing device 104. The audio-based content items may specifically be audio, or may be combined with text, image, or video data.

[0028] The service provider device 108 may include at least one service provider natural language processor component 142 and a service provider interface 144, which docks with or otherwise communicates with at least one service provider natural language processor component 142 and the service provider interface 144. The service provider computing device 108 may include at least one service provider natural language processor (NLP) component 142 and at least one service provider interface 144. The service provider NLP component 142 (or other components such as the direct action API of the service provider computing device 108) may engage with the client computing device 104 (either via the data processing system 102 or bypassing the data processing system 102) to create a conversation (e.g., a session) based on real-time voice or audio exchanges between the client computing device 104 and the service provider computing device 108. The service provider NLP 142 may include one or more functions or features of the NLP component 112 of the data processing system 102. For example, the service provider interface 144 may receive and provide data messages to the direct action API 116 of the data processing system 102. The service provider computing device 108 and the content provider computing device 106 may be associated with the same entity. For example, the content provider computing device 106 may create, store, or produce available content items for a car-sharing service, and the service provider computing device 108 may establish a session with the client computing device 106 to arrange for the delivery of a taxi or car for the car-sharing service to pick up and drop off the end user of the client computing device 104. The data processing system 102 may also establish a session with the client computing device via the direct action API 116, the NLP component 112, or other components, including or bypassing the service provider computing device 104, to arrange for the delivery of a taxi or car for a car-sharing service, for example.

[0029] The computing device 104 may include at least one sensor 134, transducer 136, audio driver 138, or preprocessor 140, interface with at least one sensor 134, transducer 136, audio driver 138, or preprocessor 140, or otherwise communicate with at least one sensor 134, transducer 136, audio driver 138, or preprocessor 140. The sensor 134 may include, for example, an ambient light sensor, proximity sensor, temperature sensor, accelerometer, gyroscope, motion detector, GPS sensor, position sensor, microphone, or touch sensor. The transducer 136 may include a speaker or microphone. The audio driver 138 may provide a software interface to the hardware transducer 136. The audio driver may execute an audio file or other instructions provided by the data processing system 102 to control the transducer 136 to generate a corresponding sound wave or acoustic wave. The preprocessor 140 may be configured to detect keywords and perform actions based on the keywords. The preprocessor 140 may filter out one or more terms or modify terms before sending the terms to the data processing system 102 for further processing. The preprocessor 140 may convert an analog audio signal detected by a microphone into a digital audio signal and send one or more data packets carrying the digital audio signal to the data processing system 102 via the network 105. In some cases, the preprocessor 140 may send data packets carrying some or all of the input audio signal in response to detecting an instruction to perform such a transmission. The instruction may include, for example, a trigger keyword or other keyword or approval for sending a data packet including the input audio signal to the data processing system 102.

[0030] The client computing device 104 may be associated with an end user who enters a voice query as audio input into the client computing device 104 (via the sensor 134) and receives an audio output output from the transducer 136 (e.g., a speaker) in the form of computer-generated speech that may be provided to the client computing device 104 from the data processing system 102 (or the content provider computing device 106 or the service provider computing device 108). The computer-generated speech may include recordings from real people or computer-generated language.

[0031] The data repository 124 may include one or more local or distributed databases and may include a database management system. The data repository 124 may include computer data storage or memory and may store, among other data, one or more parameters 126, one or more policies 128, content data 130, or templates 132. The parameters 126, policies 128, and templates 132 may include information such as rules regarding a voice-based session between the client computing device 104 and the data processing system 102 (or service provider computing device 108). The content data 130 may include content items of audio output or associated metadata, as well as input audio messages that may be part of one or more communication sessions with the client computing device 104.

[0032] The data processing system 102 may include a content placement system having at least one computing resource or server. The data processing system 102 may include at least one interface 110, dock with at least one interface 110, or otherwise communicate with at least one interface 110. The data processing system 102 may include at least one natural language processor component 112, dock with at least one natural language processor component 112, or otherwise communicate with at least one natural language processor component 112. The data processing system 102 may include at least one direct action application programming interface (“API”) 116, dock with at least one direct action application programming interface (“API”) 116, or otherwise communicate with at least one direct action application programming interface (“API”) 116. The data processing system 102 may include at least one session handler 114, dock with at least one session handler 114, or otherwise communicate with at least one session handler 114. The data processing system 102 may include at least one content selector component 118, dock with at least one content selector component 118, or otherwise communicate with at least one content selector component 118. The data processing system 102 may include at least one audio signal generator 122, dock with at least one audio signal generator 122, or otherwise communicate with at least one audio signal generator 122. The data processing system 102 may include at least one data repository 124, dock with at least one data repository 124, or otherwise communicate with at least one data repository 124. At least one data repository 124 may include or store parameters 126, policies 128, content data 130, or templates 132 in one or more data structures or databases. The parameters 126 may include, for example, thresholds, distances, time intervals, durations, scores, or weights. The content data 130 may include, for example, content activity information, content groups, content selection criteria, content item objects, or other information provided by the content provider 106 or obtained or determined by the data processing system to facilitate content selection. The content data 130 may include, for example, the historical performance of content activities.

[0033] Interface 110, natural language processor component 112, session handler 114, direct action API 116, content selector component 118, or audio signal generator component 122 may each include at least one processing unit or other logic device, such as a programmable logic array engine or a module configured to communicate with a database repository or database 124. Interface 110, natural language processor component 112, session handler 114, direct action API 116, content selector component 118, audio signal generator component 122, and data repository 124 may be separate components, a single component, or part of a data processing system 102. System 100 and its components (such as data processing system 102) may include hardware elements, such as one or more processors, logic devices, or circuits.

[0034] Data processing system 102 may obtain anonymous computer network activation information associated with a plurality of computing devices 104. A user of computing device 104 may affirmatively authorize data processing system 102 to obtain network activation information corresponding to the user's computing device 104. For example, data processing system 102 may prompt the user of computing device 104 to consent to obtain one or more types of network activation information. The identity of the user of computing device 104 may remain anonymous and computing device 104 may be associated with a unique identifier (e.g., a unique identifier of the user or computing device provided by the data processing system or the user of the computer). The data processing system may associate each observation with the corresponding unique identifier.

[0035] Content provider 106 may establish an electronic content campaign. The electronic content campaign may be stored as content data 130 in data repository 124. The electronic content campaign may refer to one or more content groups corresponding to a common theme. The content campaign may include a hierarchical data structure that includes content groups, content item data objects, and content selection criteria. To create a content campaign, content provider 106 may specify values for the campaign-level parameters of the content campaign. The campaign-level parameters may include, for example, a campaign name, a preferred content network for placement of content item objects, a value of resources to be used for the content campaign, start and end dates of the content campaign, a duration of the content campaign, a schedule for placement of content item objects, a language, a geographic location, a type of computing device on which the content item objects are to be provided. In some cases, an impression may refer to when a content item object is prefetched from its source (e.g., data processing system 102 or content provider 106), and is countable. In some cases, due to the possibility of click fraud, bot activity may be filtered and excluded as an impression. Thus, in some cases, an impression may refer to a measurement of a response from a web server to a page request from a browser, which is filtered from bot activity and error codes and recorded at the point as close as possible to have an opportunity to render the content item object for display on computing device 104. In some cases, an impression may refer to a visible or audible impression; for example, the content item object is at least partially (e.g., 20%, 30%, 30%, 40%, 50%, 60%, 70% or more) visible on a display device of client computing device 104, or audible via a speaker 136 of computing device 104. A click or selection may refer to a user's interaction with a content item object, such as a voice response to an audible impression, a mouse click, a touch interaction, a gesture, a shake, an audio interaction, or a keyboard click. A conversion may refer to a user taking a desired action with respect to a content item object; for example, purchasing a product or service, completing a survey, visiting a physical store corresponding to the content item, or completing an electronic transaction.

[0036] Content provider 106 may also establish one or more content groups for the content campaign. The content group includes one or more content item objects and corresponding content selection criteria, such as keywords, words, terms, phrases, geographic locations, types of computing devices, times, interests, topics, or verticals. Content groups under the same content campaign may share the same campaign-level parameters, but have customized specifications for specific content group-level parameters (such as keywords, negative keywords (e.g., to block placement of a content item in the case where a negative keyword is present on the main content), a bid for a keyword, or a parameter associated with the bid or content campaign).

[0037] To create a new content group, a content provider may provide values for the content group level parameters of the content group. Content group level parameters include, for example, the content group name or content group topic and bids for different content placement opportunities (e.g., automatic placement or managed placement) or outcomes (e.g., clicks, impressions, or conversions). The content group name or content group topic may be one or more terms that the content provider 106 can use to capture the topic or theme for which content item objects of the content group will be selected for display. For example, an auto dealer may create different content groups for each model of vehicle it carries and may also create different content groups for each brand of vehicle it carries. Examples of content group topics that the auto dealer may use include, for example, "Make A sports cars", "Make B sports cars", "Make C sedans", "Make C trucks", "Make C hybrids", or "Make D plug-in hybrids". For example, an example content campaign topic may be "plug-in hybrids" and include content groups for both "Make C hybrids" and "Make D plug-in hybrids".

[0038] The content provider 106 may provide one or more keywords and content item objects to each content group. The keywords may include terms related to the product or service associated with or identified by the content item object. The keywords may include one or more terms or phrases. For example, an auto dealer may include "sports cars", "V-6 engine", "four-wheel drive", "fuel efficiency" as keywords for a content group or content campaign. In some cases, the content provider may specify negative keywords to avoid, prevent, block, or disable content placement on a particular term or keyword. The content provider may specify the type of match to be used for selecting content item objects, such as exact match, phrase match, or broad match.

[0039] The content provider 106 may provide one or more keywords to be used by the data processing system 102 to select content item objects provided by the content provider 106. The content provider 106 may identify one or more keywords to bid on and may further provide bid amounts for the various keywords. The content provider 106 may provide additional content selection criteria to be used by the data processing system 102 to select content item objects. Multiple content providers 106 may bid on the same or different keywords, and the data processing system 102 may run a content selection process or an ad auction in response to receiving an indication of the keywords in an electronic message.

[0040] Content provider 106 may provide one or more content item objects for selection by data processing system 102. When a content placement opportunity that matches the resource allocation, content schedule, maximum bid, keywords, and other selection criteria specified for a content group becomes available, data processing system 102 (e.g., via content selector component 118) may select a content item object. Different types of content item objects may be included in the content group, such as voice content items, audio content items, text content items, image content items, video content items, multimedia content items, or content item links. When selecting a content item, data processing system 102 may send the content item object for rendering on computing device 104 or a display device of computing device 104. Rendering may include displaying the content item on the display device or playing the content item via a speaker of computing device 104. Data processing system 102 may provide instructions for rendering the content item object to computing device 104. Data processing system 102 may instruct computing device 104 or the audio driver 138 of computing device 104 to generate an audio signal or sound wave.

[0041] Data processing system 102 may include interface component 110 that is designed, configured, constructed, or operable to receive and send information using, for example, data packets. Interface 110 may receive and send information using one or more protocols, such as network protocols. Interface 110 may include a hardware interface, a software interface, a wired interface, or a wireless interface. Interface 110 may assist in converting or formatting data from one format to another. For example, interface 110 may include an application programming interface that includes definitions for communicating between various components, such as software components.

[0042] The data processing system 102 may include an application, script, or program installed at the client computing device 104, such as an app for transmitting an input audio signal to the interface 110 of the data processing system 102 and driving components of the client computing device to render an output audio signal. The data processing system 102 may receive data packets or other signals that include or identify an audio input signal. For example, the data processing system 102 may execute or run the NLP component 112 to receive or obtain the audio signal and parse the audio signal. For example, the NLP component 112 may provide an interaction between humans and computers. The NLP component 112 may be configured with techniques for understanding natural language and allowing the data processing system 102 to derive meaning from human or natural language input. The NLP component 112 may include or be configured with machine learning-based techniques, such as statistical machine learning. The NLP component 112 may utilize decision trees, statistical models, or probability models to parse the input audio signal. The NLP component 112 may perform functions such as named entity recognition (e.g., given a text stream, determining which items in the text map to appropriate names (such as people or places) and what type each such name is, such as person, place, or organization), natural language generation (e.g., converting information from a computer database or semantic intent into understandable human language), natural language understanding (e.g., converting text into a more formal representation, such as a first-order logic structure manipulable by a computer module), machine translation (e.g., automatically translating text from one human language to another), morphological segmentation (e.g., breaking words into individual morphemes and identifying the categories of the morphemes, which can be challenging based on the morphological or structural complexity of the words of the language being considered), question answering (e.g., determining the answer to a human language question, which can be specific or open-ended), semantic processing (e.g., processing that may occur after identifying words and encoding their meanings to associate the identified words with other words having similar meanings).

[0043] The NLP component 112 converts the audio input signal into recognized text by comparing the input signal against a stored set of representative audio waveforms (e.g., in the data repository 124) and selecting the closest match. The set of audio waveforms may be stored in the data repository 124 or other databases accessible to the data processing system 102. The representative waveforms are generated across a large group of users and may then be augmented with voice samples from users. After the audio signal is converted into recognized text, the NLP component 112 matches the text with words associated with actions that the data processing system 102 can service, such as via training across users or by manual specification.

[0044] The audio input signal may be detected by a sensor 134 or transducer 136 (e.g., a microphone) of the client computing device 104. Via the transducer 136 or other component, the client computing device 104 may provide the audio input signal to the data processing system 102 (e.g., via the network 105), where it may be received (e.g., by the interface 110) and provided to the NLP component 112 or stored in the data repository 124.

[0045] The NLP component 112 may obtain an input audio signal. Based on the input audio signal, the NLP component 112 may identify at least one request or at least one trigger keyword corresponding to the request. The request may indicate the intent or subject of the input audio signal. The trigger keyword may indicate the type of action that is likely to be taken. For example, the NLP component 112 may parse the input audio signal to identify at least one request to leave home in the evening to attend dinner and a movie. The trigger keyword may include at least one word, phrase, root or partial word, or derivative indicating the action to be taken. For example, the trigger keyword "go" or "to go" from the input audio signal may indicate the need for transportation. In this example, the input audio signal (or the identified request) does not directly express the intention for transportation, but the trigger keyword indicates that transportation is an auxiliary action for at least one other action indicated by the request.

[0046] The NLP component 112 may parse the input audio signal to identify, determine, retrieve, or otherwise obtain requests and trigger keywords. For example, the NLP component 112 may apply semantic processing techniques to the input audio signal to identify trigger keywords or requests. The NLP component 112 may apply semantic processing techniques to the input audio signal to identify trigger phrases that include one or more trigger keywords (such as a first trigger keyword and a second trigger keyword). For example, the input audio signal may include the sentence "I need someone to do my laundry and my dry cleaning." The NLP component 112 may apply semantic processing techniques or other natural language processing techniques to the data group including the sentence to identify the trigger phrases "do my laundry" and "do my dry cleaning." The NLP component 112 may also identify multiple trigger keywords, such as laundry and dry cleaning. For example, the NLP component 112 may determine that the trigger phrase includes a trigger keyword and a second trigger keyword.

[0047] The NLP component 112 can filter the input audio signal to identify trigger keywords. For example, the data packet carrying the input audio signal may include "It would be great if I could get someone that could help me go to the airport", in which case the NLP component 112 can filter out one or more terms as follows: "it", "would", "be", "great", "if", "I", "could", "get", "someone", "that", "could", or "help". By filtering out these terms, the NLP component 112 can more precisely and reliably identify the trigger keyword, such as "go to the airport", and determine that this is a request for a taxi or ride-sharing service.

[0048] In some cases, the NLP component can determine that the data packet carrying the input audio signal includes one or more requests. For example, the input audio signal may include the sentence "I need someone to do my laundry and my drycleaning". The NLP component 112 can determine that this is a request for a laundry service and a dry cleaning service. The NLP component 112 can determine that this is a single request for a service provider that can provide both the laundry service and the dry cleaning service. The NLP component 112 can determine that these are two requests: a first request for a service provider to perform the laundry service, and a second request for a service provider to provide the dry cleaning service. In some cases, the NLP component 112 can combine the multiple determined requests into a single request and send the single request to the service provider device 108. In some cases, the NLP component 112 can send separate requests to the respective service provider devices 108, or send the two requests separately to the same service provider device 108.

[0049] The data processing system 102 may include a direct action API 116 that is designed and constructed to generate an action data structure in response to a request based on a trigger keyword. A processor of the data processing system 102 may invoke the direct action API 116 to execute a script that generates the data structure for a service provider device 108 to request or order a service or product, such as a car from a car sharing service. The direct action API 116 may obtain data from a data repository 124 and data received from a client computing device 104 with the end user's consent to determine location, time, user account, logic, or other information to allow the service provider device 108 to perform an operation, such as reserving a vehicle from a car sharing service. Using the direct action API 116, the data processing system 102 may also communicate with the service provider device 108 to complete a transaction by, in this example, making a car sharing pick-up reservation.

[0050] The direct action API 116 may perform a specified action to fulfill the end user's intent as determined by the data processing system 102. Depending on the action specified in its input, the direct action API 116 may execute code or a conversation script that identifies the parameters needed to fulfill the user's request. Such code may, for example, look up additional information in the data repository 124, such as the name of a home automation service, or it may provide an audio output for rendering at the client computing device 104 to ask the end user questions such as the intended destination of a requested taxi. The direct action API 116 may determine the necessary parameters and may encapsulate the information into an action data structure, which may then be sent to another component (such as a content selector component 118) or to the service provider computing device 108 to be completed.

[0051] The direct action API 116 may receive instructions or commands from an NLP component 112 or other component of the data processing system 102 to generate or construct an action data structure. The direct action API 116 may determine the type of action in order to select a template from a template repository 132 stored in the data repository 124. The type of action may include, for example, a service, a product, a reservation, or a ticket. The type of action may also include the type of service or product. For example, the type of service may include a car sharing service, a food delivery service, a laundry service, a cleaning service, a repair service, or a housekeeping service. The type of product may include, for example, clothes, shoes, toys, electronic devices, computers, books, or jewelry. The type of reservation may include, for example, a dinner reservation or a hair salon appointment. The type of ticket may include, for example, a movie ticket, a stadium ticket, or an airline ticket. In some cases, the types of services, products, reservations, or tickets may be classified based on price, location, delivery type, availability, or other attributes.

[0052] When identifying the type of the request, the direct action API 116 can access the corresponding template from the template repository 132. The template can include fields in a structured data set that can be filled by the direct action API 116 to further perform the operations requested by the service provider device 108 (such as the operation of dispatching a taxi to pick up and drop off an end user at a pick-up location and transporting the end user to a destination location). The direct action API 116 can perform a lookup in the template repository 132 to select a template that matches the trigger keyword and one or more characteristics of the request. For example, if the request corresponds to a request for a car or a ride to a destination, the data processing system 102 can select a car-sharing service template. The car-sharing service template can include one or more of the following fields: device identifier, pick-up location, destination location, number of passengers, or type of service. The direct action API 116 fills the fields with values. To fill the fields with values, the direct action API 116 can check, poll, or otherwise obtain information from one or more sensors 134 of the computing device 104 or the user interface of the device 104. For example, the direct action API 116 can use a location sensor (such as a GPS sensor) to detect the source location. The direct action API 116 can obtain further information by submitting a survey, prompt, or query to the end user of the computing device 104. The direct action API can submit a survey, prompt, or query via the interface 110 of the data processing system 102 and the user interface of the computing device 104 (such as an audio interface, a voice-based user interface, a display, or a touch screen). Thus, the direct action API 116 can select a template for the action data structure based on the trigger keyword or request, fill one or more fields in the template with information detected by one or more sensors 134 or obtained via the user interface, and generate, create, or otherwise construct an action data structure to facilitate the operation performed by the service provider device 108.

[0053] The data processing system 102 can select a template 132 based on the template data structure based on various factors, including for example one or more of the trigger keyword, the request, the third-party provider device 108, the type of the third-party provider device 108, the category into which the third-party provider device 108 falls (such as taxi service, laundry service, flower service, or food delivery), location, or other sensor information.

[0054] To select a template based on a trigger keyword, the data processing system 102 (e.g., via the direct action API 116) may use the trigger keyword to perform a lookup or other query operation on the template database 132 to identify a template data structure that maps to or otherwise corresponds to the trigger keyword. For example, each template in the template database 132 may be associated with one or more trigger keywords to indicate that the template is configured to generate an action data structure in response to a trigger keyword that the third-party provider device 108 can process to establish a communication session.

[0055] In some cases, the data processing system 102 may identify the third-party provider device 108 based on the trigger keyword. To identify the third-party provider 108 based on the trigger keyword, the data processing system 102 may perform a lookup in the data repository 124 to identify the third-party provider device 108 that maps to the trigger keyword. For example, if the trigger keyword includes "ride" or "to got to", the data processing system 102 (e.g., via the direct action API 116) may identify the third-party provider device 108 as corresponding to taxi service company A. The data processing system 102 may use the identified third-party provider device 108 to select a template from the template database 132. For example, the template database 132 may include a mapping or correlation between the third-party provider device 108 or entity into a template that is configured to generate an action data structure in response to a trigger keyword that the third-party provider device 108 can process to establish a communication session. In some cases, templates may be customized for the third-party provider device 108 or for a category of the third-party provider device 108. The data processing system 102 may generate an action data structure based on the template for the third-party provider 108.

[0056] To construct or generate an action data structure, the data processing system 102 may identify one or more fields in the selected template to be filled with values. The fields may be filled with numerical values, strings, Unicode values, boolean logic, binary values, hexadecimal values, identifiers, location coordinates, geographic regions, timestamps, or other values. The fields or the data structure itself may be encrypted or masked to maintain data security.

[0057] When determining the fields in the template, the data processing system 102 may identify the values of the fields to fill the fields of the template to create the action data structure. The data processing system 102 may obtain, retrieve, determine, or otherwise identify the values of the fields by performing a lookup or other query operation on the data repository 124.

[0058] In some cases, the data processing system 102 may determine that information or values for a field do not exist in the data repository 124. The data processing system 102 may determine that the information or values stored in the data repository 124 are outdated, stale, or otherwise not suitable for the purpose of constructing an action data structure in response to trigger keywords and requests identified by the NLP component 112 (e.g., the location of the client computing device 104 may be an old location rather than the current location; an account may be expired; the destination restaurant may have moved to a new location; physical activity information; or mode of transportation).

[0059] If the data processing system 102 determines that it cannot currently access the values or information for a field of a template in the memory of the data processing system 102, the data processing system 102 may obtain those values or information. The data processing system 102 may obtain or acquire information by querying or polling one or more available sensors of the client computing device 104, prompting the end user of the client computing device 104 for information, or accessing an online web-based resource using the HTTP protocol. For example, the data processing system 102 may determine that it does not have the current location of the client computing device 104, which may be a required field of a template. The data processing system 102 may query the client computing device 104 for location information. The data processing system 102 may request that the client computing device 104 use one or more location sensors 134 (such as a Global Positioning System sensor), WIFI triangulation, cell tower triangulation, Bluetooth beacons, IP address, or other location sensing techniques to provide location information.

[0060] The direct action API 116 may send the action data structure to a third-party provider device (e.g., the service provider device 108) to cause the third-party provider device 108 to invoke a dialogue application programming interface (e.g., the service provider NLP component 142) and establish a communication session between the third-party provider device 108 and the client computing device 104. In response to establishing a communication session between the service provider device 108 and the client computing device 104, the service provider device 108 may directly send data packets to the client computing device 104 via the network 105. In some cases, the service provider device 108 may send data packets to the client computing device 104 via the data processing system 102 and the network 105.

[0061] In some cases, a third-party provider device 108 may execute at least a portion of the dialogue API 142. For example, the third-party provider device 108 may handle certain aspects of a communication session or certain types of queries. The third-party provider device 108 may utilize the NLP component 112 executed by the data processing system 102 to assist in processing an audio signal associated with the communication session and generating a response to the query. In some cases, the data processing system 102 may include a dialogue API 142 configured for the third-party provider 108. In some cases, the data processing system routes data packets between the client computing device and the third-party provider device to establish a communication session. The data processing system 102 may receive an indication from the third-party provider device 108 that the third-party provider device has established a communication session with the client device 104. The indication may include an identifier of the client computing device 104, a timestamp relative to when the communication session was established, or other information associated with the communication session, such as an action data structure associated with the communication session.

[0062] In some cases, the dialogue API may be a second NLP that includes one or more components or functions of the first NLP 112. The second NLP 142 may interact with or utilize the first NLP 112. In some cases, the system 100 may include a single NLP 112 executed by the data processing system 102. The single NLP 112 may support both the data processing system 102 and the third-party service provider device 108. In some cases, the direct action API 116 generates or constructs an action data structure to assist in performing a service, and the dialogue API generates a response or query to further communicate with the end user in a communication session or obtain additional information to improve or enhance the end user's experience of the service or the performance of the service.

[0063] The data processing system 102 may include, execute, access, or otherwise communicate with a session processor component 114 to establish a communication session between the client device 104 and the data processing system 102. The communication session may refer to one or more data transmissions between the client device 104 and the data processing system 102, which include an input audio signal detected by a sensor 134 of the client device 104 and an output signal sent from the data processing system 102 to the client device 104. The data processing system 102 (e.g., via the session processor component 114) may establish a communication session in response to receiving the input audio signal. The data processing system 102 may set a duration for the communication session. The data processing system 102 may set a timer or counter for the duration set for the communication session. In response to the expiration of the timer, the data processing system 102 may terminate the communication session.

[0064] A communication session may refer to a network-based communication session in which the client device 104 provides authentication information or credentials to establish the session. In some cases, a communication session refers to the topic or context of an audio signal carried by data packets during the session. For example, a first communication session may refer to an audio signal related to a taxi service (e.g., including keywords, action data structures, or content item objects) sent between the client device 104 and the data processing system 102; while a second communication session may refer to an audio signal related to laundry and dry cleaning services sent between the client device 104 and the data processing system 102. In this example, the data processing system 102 may determine that the contexts of the audio signals are different (e.g., via the NLP component 112), and separate the two sets of audio signals into different communication sessions. The session processor 114 may terminate the first session related to the ride service in response to identifying one or more audio signals related to dry cleaning and laundry services. Thus, the data processing system 102 may initiate or establish a second session for the audio signals related to dry cleaning and laundry services in response to detecting the context of the audio signals.

[0065] The data processing system 102 may include, execute, or otherwise communicate with the content selector component 118 to receive trigger keywords identified by the natural language processor, and select content items based on the trigger keywords via a real-time content selection process. The content selection process may refer to or include selecting sponsored content item objects provided by a third-party content provider 106. The real-time content selection process may include content items provided by multiple content providers being parsed, processed, weighted, or matched in order to select one or more content items to provide to the computing device 104 as a service. The content selector component 118 may execute the content selection process in real time. Executing the content selection process in real time may refer to executing the content selection process in response to a request for content received via the client computing device 104. The real-time content selection process may be executed (e.g., initiated or completed) within a time interval for receiving the request (e.g., 5 seconds, 10 seconds, 20 seconds, 30 seconds, 1 minute, 2 minutes, 3 minutes, 5 minutes, 10 minutes, or 20 minutes). The real-time content selection process may be executed during a communication session with the client computing device 104 or within a certain time interval after the communication session is terminated.

[0066] For example, the data processing system 102 may include a content selector component 118 that is designed, constructed, configured, or operable to select content item objects. To select content items to be displayed in a voice-based environment, the data processing system 102 (e.g., via the NLP component 112) may parse the input audio signal to identify keywords (e.g., trigger keywords) and use these keywords to select matching content items based on broad match, exact match, or phrase match. For example, the content selector component 118 may analyze, parse, or otherwise process the topics of candidate content items to determine whether the topics of the candidate content items correspond to the topics of the keywords or phrases of the input audio signal detected by the microphone of the client computing device 104. The content selector component 118 may use image processing techniques, character recognition techniques, natural language processing techniques, or database lookups to identify, analyze, or recognize the voice, audio, terms, characters, text, symbols, or images of candidate content items. The candidate content items may include metadata indicating the topics of the candidate content items, in which case the content selector component 118 may process the metadata to determine whether the topics of the candidate content items correspond to the input audio signal.

[0067] The content provider 106 may provide additional indicators when setting up content activities that include content items. By performing a lookup using information about the candidate content items, the content provider 106 may provide information at the content activity or content group level that the content selector component 118 may identify. For example, the candidate content items may include unique identifiers that can be mapped to content groups, content activities, or content providers. The content selector component 118 may determine information about the content provider 106 based on information in the content activity data structure stored in the data repository 124.

[0068] The data processing system 102 may receive a request for content to be presented on the computing device 104 via a computer network. The data processing system 102 may identify the request by processing the input audio signal detected by the microphone of the client computing device 104. The request may include selection criteria for the request, such as device type, location, and keywords associated with the request. The request may include an action data structure or action data structures.

[0069] In response to the request, data processing system 102 may select a content item object from data repository 124 or a database associated with content provider 106 and provide the content item via network 105 for presentation via computing device 104. The content item object may be provided by a content provider device 108 that is different from service provider device 108. The content item may correspond to a type of service that is different from the type of service of the action data structure (e.g., taxi service versus food delivery service). Computing device 104 may interact with the content item object. Computing device 104 may receive an audio response to the content item. Computing device 104 may receive an indication for selecting a hyperlink or other button associated with the content item object, which indication causes or allows computing device 104 to identify service provider 108, request a service from service provider 108, instruct service provider 108 to perform a service, send information to service provider 108, or otherwise query service provider device 108.

[0070] The data processing system 102 may include, execute, or communicate with an audio signal generator component 122 to generate an output signal. The output signal may include one or more parts. For example, the output signal may include a first part and a second part. The first part of the output signal may correspond to an action data structure. The second part of the output signal may correspond to a content item selected by the content selector component 118 during the real-time content selection process.

[0071] The audio signal generator component 122 may generate an output signal having a first portion of sound corresponding to the first data structure. For example, the audio signal generator component 122 may generate the first portion of the output signal based on one or more values populated into fields of the action data structure via the direct action API 116. In the taxi service example, the values of the fields may include, for example, a pick-up location of 123rd Street, a destination location of 1234th Street, a number of passengers of 2, and a service level of economy. The audio signal generator component 122 may generate the first portion of the output signal to confirm that the end user of the computing device 104 wants to proceed with sending the request to the service provider 108. The first portion may include the following output: "Would you like to an economy car from taxi service provider A to pick two people up at 123 Main Street and drop off at 1234 Main Street?"

[0072] In some cases, the first part may include information received from the service provider device 108. The information received from the service provider device 108 may be customized or tailored for the action data structure. For example, the data processing system 102 may send the action data structure to the service provider 108 before instructing the service provider 108 to perform an operation (e.g., via the direct action API 116). Alternatively, the data processing system 102 may instruct the service provider device 108 to perform an initial or preliminary processing on the action data structure to generate preliminary information about the operation. In the example of a taxi service, the preliminary processing of the action data structure may include identifying available taxis that meet the service level requirements located around the pick-up and drop-off locations, estimating the amount of time for the nearest available taxi to reach the pick-up location, estimating the time to reach the destination, and estimating the price of the taxi service. The initial values of the estimates may include fixed values, valuations that are subject to change based on various conditions, or ranges of values. The service provider device 108 may return the preliminary information to the data processing system 102 or directly to the client computing device 104 via the network 104. The data processing system 102 may incorporate the preliminary results from the service provider device 108 into the output signal and send the output signal to the computing device 104. The output signal may include, for example, "Taxi Service Company A can pick you up at 123 Main Street in 10 minutes, and drop you off at 1234 Main Street by 9 AM for $10. Do you want to order this ride?" This may form the first part of the output signal.

[0073] In some cases, data processing system 102 may form a second part of the output signal. The second part of the output signal may include content items selected by content selector component 118 during a real-time content selection process. The first part may be different from the second part. For example, the first part may include information corresponding to an action data structure in direct response to a data packet carrying an input audio signal detected by sensor 134 of client computing device 104, whereas the second part may include content items tangentially related to the action data structure selected by content selector component 104, or may include sponsored content provided by content provider device 106. For example, an end user of computing device 104 may request a taxi from taxi service company A. Data processing system 102 may generate a first part of the output signal to include information about a taxi from taxi service company A. However, data processing system 102 may generate a second part of the output signal to include content items selected based on the keyword "taxiservice" and information in the action data structure that the end user may be interested in. For example, the second part may include content items or information provided by a different taxi service company, such as taxi service company B. Although the user may not have specifically requested taxi service company B, data processing system 102 may still provide a content item from taxi service company B because the user may choose to perform an operation regarding taxi service company B.

[0074] Data processing system 102 may send information from the action data structure to taxi service company B to determine the pick-up time, the time to reach the destination, and the journey price. Data processing system 102 may receive this information and generate a second part of the output signal as follows: "Taxi Service Company B can pick you up at 123 Main Street in 2 minutes, and drop you off at 1234 Main Street by 8:52 AM for $15. Do you want this ride instead?" The end user of computing device 104 may then choose the journey provided by taxi service company A or the journey provided by taxi service company B.

[0075] Before providing a sponsorship content item corresponding to the service provided by taxi service company B in the second part of the output signal, data processing system 102 may notify the end user of the computing device that the second part corresponds to a content item object selected (e.g., by content selector component 118) during a real-time content selection process. However, data processing system 102 may have limited access to different types of interfaces to provide notifications to the end user of computing device 104. For example, computing device 104 may not include a display device, or the display device may be disabled or turned off. The display device of computing device 104 may consume more resources compared to the speaker of computing device 104, so turning on the display device of computing device 104 may be less efficient compared to using the speaker of computing device 104 to convey the notification. Thus, in some cases, data processing system 102 may improve the efficiency and effectiveness of information transmission over one or more interfaces or one or more types of computer networks. For example, data processing system 102 (e.g., via audio signal generator component 122) may modularize a portion of the output audio signal that includes the content item to provide an indication or notification to the end user that this portion of the output signal includes the sponsorship content item.

[0076] Data processing system 102 (e.g., via interface 110 and network 105) may send data packets including the output signal generated by audio signal generator component 122. The output signal may cause the audio driver component 138 of client device 104 or an audio driver component 138 executed by client device 104 to drive the speaker (e.g., transducer 136) of client device 104 to generate sound waves corresponding to the output signal.

[0077] Figure 2 is an illustration of the system 100 routing packetized actions via a computer network. The system may include Figure 1 one or more components of system 100 depicted in. At 205, client computing device 104 may send data packets carrying an input audio signal detected by the microphone or other sensors of computing device 104. Client computing device 104 may send the input audio signal to data processing system 102. Data processing system 102 may parse the input audio signal to identify keywords, requests, or other information to generate an action data structure in response to the request.

[0078] At action 210, data processing system 102 may send the action data structure to service provider device 108 (or third-party provider device 108). Data processing system 102 may send the action data structure via the network. Service provider device 108 may include an interface configured to receive and process the action data structure sent by data processing system 102.

[0079] At action 215, the service provider device 108 (e.g., via a dialogue API) may respond to the action data structure. The response from the service provider device 108 may include an indication of the service to be performed corresponding to the action data structure. The response may include an acknowledgement to continue the operation. The response may include a request for further information to perform the operation corresponding to the action data structure. For example, the action data structure may be for a journey, and the service provider 108 may respond to the request for further information, such as the number of passengers in the journey, the type of car desired by the passengers, the desired facilities in the car, or the preferred pick-up and drop-off locations. The request for additional information may include information that may not be present in the action data structure. For example, the action data structure may include baseline information for performing the operation, such as the pick-up location, the destination location, and the number of passengers. The baseline information may be a standard data set used by multiple service providers 108 in the taxi service category. However, a particular taxi service provider 108 may choose to customize and improve the operation by requesting additional information or preferences from the client computing device 104.

[0080] At action 215, the service provider device 108 may send one or more data packets carrying the response to the data processing system 102. The data processing system 102 may parse the data packets and identify the source and destination of the data packets. At action 220, the data processing system 102 may thus route or forward the data packets to the client computing device 104. The data processing system 102 may route or forward the data packets via the network 105.

[0081] At action 225, the client computing device 220 may send instructions or commands to the data processing system 102 based on the forwarded response. For example, the response forwarded at 225 may be a request for acknowledgement of the number of passengers and to continue arranging the taxi journey. The instructions at 225 may include the number of passengers and instructions to continue arranging the pick-up and drop-off. The client device 104 may send one or more data packets carrying the instructions to the data processing system 102. At action 230, the data processing system 102 may route or forward the data packets carrying the instructions to the service provider device 108.

[0082] In some cases, data processing system 102 may route or forward data packets as is (e.g., without manipulating the data packets) at action 220 or action 230. In some cases, data processing system 102 may process the data packets to filter out information or encapsulate the data packets with information to facilitate processing of the data packets by service provider device 108 or client computing device 104. For example, data processing system 102 may mask, hide, or protect the identity of client computing device 104 from service provider device 108. Thus, data processing system 102 may encrypt the identification information using a hash function such that service provider 108 cannot directly identify the device identifier or username of client computing device 104. Data processing system 102 may maintain a mapping of proxy identifiers provided to service provider device 108 for use during a communication session to the identifier or username of client computing device 104.

[0083] Figure 3 is an illustration of the system 100 routing packetized actions via a computer network. The system may include Figure 1 one or more components of system 100 depicted in. At 305, client computing device 104 may send a data packet carrying an input audio signal detected by a microphone or other sensor of computing device 104. Client computing device 104 may send the input audio signal to data processing system 102. Data processing system 102 may parse the input audio signal to identify keywords, requests, or other information to generate an action data structure in response to the request.

[0084] At action 310, data processing system 102 may send the action data structure to service provider device 108 (or third-party provider device 108). Data processing system 102 may send the action data structure via the network. Service provider device 108 may include an interface configured to receive and process the action data structure sent by data processing system 102.

[0085] At action 315, the service provider device 108 (e.g., via a dialogue API) may respond to the action data structure. The response from the service provider device 108 may include an indication of the service to be performed corresponding to the action data structure. The response may include an acknowledgement to continue the operation. The response may include a request for further information to perform the operation corresponding to the action data structure. For example, the action data structure may be for a journey, and the service provider 108 may respond with further information such as the number of passengers on the journey, the type of car desired by the passengers, the desired facilities in the car, or the preferred pick-up and drop-off locations. The request for additional information may include information that may not be present in the action data structure. For example, the action data structure may include baseline information for performing the operation, such as pick-up location, destination location, and number of passengers. The baseline information may be a standard data set used by multiple service providers 108 in a taxi service category. However, a particular taxi service provider 108 may choose to customize and improve the operation by requesting additional information or preferences from the client computing device 104.

[0086] The service provider device 108 may directly send one or more data packets carrying the response to the client computing device 104 via the network 105. For example, instead of routing the response through the data processing system 102, the service provider device 108 may directly respond to the client computing device 104 via the dialogue API executed by the service provider device 108. This may allow the service provider to customize the communication session.

[0087] At action 320, the client computing device 104 may send instructions or commands to the service provider device 108 based on the response. For example, the response provided at 315 may be a request for confirmation of the number of passengers and to continue arranging the taxi journey. The instructions at 320 may include the number of passengers and an instruction to continue arranging the pick-up and drop-off. The client device 104 may send one or more data packets carrying the instructions to the service provider device 108 instead of routing these data packets through the data processing system 102.

[0088] The data processing system 102 may assist the service provider device 108 and the client computing device 104 in establishing a communication session independent of the data processing system 102 by passing communication identifiers to the respective devices. For example, the data processing system 102 may forward the identifier of device 104 to device 108; and the data processing system 102 may forward the identifier of device 108 to device 104. Thus, device 108 may directly establish a communication session with device 104.

[0089] In some cases, device 108 or device 104 may forward information about a communication session, such as status information, to data processing system 102, respectively. For example, device 108 may provide an indication to the data processing system that a communication session has been successfully established between device 108 and client device 104.

[0090] Figure 4 is an illustration of an example method for performing dynamic modulation of a packetized audio signal. Method 400 may be performed by one or more components, systems, or elements of system 100 or system 500. Method 400 may include a data processing system receiving an input audio signal (act 405). The data processing system may receive the input audio signal from a client computing device. For example, a natural language processor component executed by the data processing system may receive the input audio signal from the client computing device via an interface of the data processing system. The data processing system may receive a data packet carrying or including the input audio signal detected by a sensor of the client computing device (or client device).

[0091] At act 410, method 400 may include the data processing system parsing the input audio signal. The natural language processor component may parse the input audio signal to identify a request and a trigger keyword corresponding to the request. For example, an audio signal detected by the client device may include "Okay device, I need a ride from Taxi Service CompanyA to go to 1234Main Street". In this audio signal, the initial trigger keyword may include "okay device", which may indicate to the client device to send the input audio signal to the data processing system. A preprocessor of the client device may filter out the term "okay device" before sending the remaining audio signal to the data processing system. In some cases, the client device may filter out additional terms or generate keywords to be sent to the data processing system for further processing.

[0092] The data processing system may identify the trigger keyword in the input audio signal. The trigger keyword may include, for example, "togo to" or "ride" or variations of these terms. The trigger keyword may indicate the type of service or product. The data processing system may identify the request in the input audio signal. The request may be determined based on the term "I need". Semantic processing techniques or other natural language processing techniques may be used to determine the trigger keyword and the request.

[0093] At action 415, method 400 may include the data processing system generating an action data structure. The data processing system may generate the action data structure based on a trigger keyword, a request, a third-party provider device, or other information. The action data structure may be in response to a request. For example, if an end user of a client computing device requests a taxi from taxi service company A, the action data structure may include information for requesting a taxi service from taxi service company A. The data processing system may select a template for taxi service company A and populate the fields in the template with values to allow taxi service company A to dispatch a taxi to the user of the client computing device to pick up the user and transport the user to the requested destination.

[0094] At action 420, method 400 may include the data processing system sending the action data structure to a third-party provider device to cause the third-party provider device. The third-party device may parse or process the received action data structure and determine to invoke a conversation API and establish a communication session between the third-party provider device and the client device. Based on the content of the action data structure, service provider device 108 may determine to invoke or otherwise execute or utilize the conversation API. For example, service provider device 108 may determine that additional information may assist in performing the operation corresponding to the action data structure. Service provider device 108 may determine that communicating with client computing device 1042 may improve the service level or reduce resource utilization due to incorrect execution of the operation. Service provider device 108 may determine to customize the operation for client computing device 104 by obtaining additional information.

[0095] At action 425, method 400 may include the data processing system receiving an indication from the third-party provider device that a communication session has been established between the third-party provider device and the client device. The indication may include a timestamp corresponding to when the communication session was established, a unique identifier for the communication session (e.g., a tuple formed by a device identifier, a timestamp of the communication session's time and date, and an identifier of the service provider device).

[0096] Figure 5is a block diagram of an example computer system 500. The computer system or computing device 500 may include or be used to implement system 100 or its components, such as data processing system 102. The data processing system 102 may include an intelligent personal assistant or a voice-based digital assistant. The computing system 500 includes a bus 505 or other communication component for transferring information and a processor 510 or processing circuitry coupled to the bus 505 for processing information. The computing system 500 may also include one or more processors 510 or processing circuitry coupled to the bus for processing information. The computing system 500 also includes a main memory 515 coupled to the bus 505 for storing information and instructions to be executed by the processor 510, such as random access memory (RAM) or other dynamic storage device. The main memory 515 may be or include data repository 145. The main memory 515 may also be used to store location information, temporary variables, or other intermediate information during the execution of instructions by the processor 510. The computing system 500 may further include a read only memory (ROM) 520 or other static storage device coupled to the bus 505 for storing static information and instructions for the processor 510. A storage device 525, such as a solid state device, disk, or optical disc, may be coupled to the bus 505 to persistently store information and instructions. The storage device 525 may include or be part of data repository 145.

[0097] The computing system 500 may be coupled via the bus 505 to a display 535, such as a liquid crystal display or an active matrix display, for displaying information to a user. An input device 530, such as a keyboard including alphanumeric and other keys, may be coupled to the bus 505 for transferring information and commands to the processor 510. The input device 530 may include a touch screen display 535. The input device 530 may also include a cursor control, such as a mouse, trackball, or cursor direction keys, for transferring direction information and command selections to the processor 510 and for controlling the movement of a cursor on the display 535. For example, the display 535 may be Figure 1 part of the data processing system 102, client computing device 150, or other components.

[0098] The processes, systems, and methods described herein can be implemented by a computing system 500 in response to execution by a processor 510 of an arrangement of instructions contained in main memory 515. Such instructions can be read into main memory 515 from another computer-readable medium, such as a storage device 525. Execution of the arrangement of instructions contained in main memory 515 causes the computing system 500 to perform the illustrative processes described herein. One or more processors in a multiprocessing arrangement can also be employed to execute the instructions contained in main memory 515. Hardwired circuitry can be used in place of software instructions or in combination with software instructions with the systems and methods described herein. The systems and methods described herein are not limited to any particular combination of hardware circuitry and software.

[0099] Although example computing systems have been described in Figure 5 , the subject matter including the operations described in this specification can be implemented using other types of digital electronic circuitry, or by computer software, firmware, or hardware, including the structures disclosed in this specification and their structural equivalents, or by a combination of one or more of them.

[0100] For situations in which the systems discussed herein collect personal information about a user or may make use of personal information, the user can be provided with an opportunity to control whether programs or features collect personal information (e.g., information about a user's social network, social actions or activities, a user's preferences, or a user's location) or to control whether or how content may be received from a content server or other data processing system that may be more relevant to the user. Additionally, certain data can be anonymized in one or more ways before it is stored or used, so that personally identifiable information is removed when generating parameters. For example, a user's identity can be anonymized so that personally identifiable information cannot be determined for that user, or a user's geographic location can be generalized (such as to city, zip code, or state level) when location information is obtained so that a user's specific location cannot be determined. Thus, the user can control how information is collected about him or her and used by a content server.

[0101] The subject matter and operations described in this specification can be implemented using digital electronic circuitry, or using computer software, firmware, or hardware (including the structures disclosed in this specification and structural equivalents thereof), or using a combination of one or more of them. The subject matter described in this specification can be implemented as one or more computer programs (e.g., one or more circuits of computer program instructions) encoded on one or more computer storage media for execution by, or to control the operation of, a data processing apparatus. Alternatively or in addition, the program instructions can be encoded on an artificially generated propagated signal (e.g., a machine-generated electrical, optical, or electromagnetic signal generated to encode information for transmission to a suitable receiver apparatus for execution by a data processing apparatus). A computer storage medium can be or be included in a computer-readable storage device, a computer-readable storage substrate, a random or serial access memory array or device, or a combination of one or more of them. Although a computer storage medium is not a propagated signal, a computer storage medium can be the source or destination of computer program instructions encoded in an artificially generated propagated signal. A computer storage medium can also be or be included in one or more separate components or media (e.g., multiple CDs, disks, or other storage devices). The operations described in this specification can be implemented as operations performed by a data processing apparatus on data stored on one or more computer-readable storage devices or received from other sources.

[0102] The terms “data processing system,” “computing device,” “component,” or “data processing apparatus” include a variety of devices, apparatus, and machines for processing data, including, by way of example, programmable processors, computers, system-on-a-chip, or multiple programmable processors, computers, system-on-a-chip, or combinations thereof. The apparatus can include special purpose logic circuitry, such as an FPGA (field programmable gate array) or an ASIC (application specific integrated circuit). The apparatus can also include code that creates an execution environment for the computer program, in addition to hardware, e.g., code that constitutes processor firmware, a protocol stack, a database management system, an operating system, a cross-platform runtime environment, a virtual machine, or a combination of one or more of them. The apparatus and the execution environment can implement various different computing model infrastructures, such as web services, distributed computing, and grid computing infrastructures. For example, the direct action API 116, the content selector component 118, or the NLP component 112 and other components of the data processing system 102 can include or share one or more data processing apparatuses, systems, computing devices, or processors.

[0103] A computer program (also known as a program, software, software application, app, script, or code) can be written in any form of programming language, including compiled or interpreted languages, declarative or procedural languages, and can be deployed in any form, including as a stand-alone program or as a module, component, subroutine, object, or other unit suitable for use in a computing environment. A computer program can correspond to a file in a file system. The computer program can be stored in a part of a file that holds other programs or data (e.g., one or more scripts stored in a markup language document), in a single file dedicated to the program, or in multiple coordinated files (e.g., files that store one or more modules, subroutines, or portions of code). The computer program can be deployed to execute on one computer or on multiple computers distributed at one site or across multiple sites and interconnected by a communication network.

[0104] The processes and logical flows described in this specification can be performed by one or more programmable processors executing one or more computer programs (e.g., components of data processing system 102) to perform actions by operating on input data and generating output. The processes and logical flows can also be performed by, and the apparatus can also be implemented as, special purpose logic circuitry, such as an FPGA (Field Programmable Gate Array) or ASIC (Application Specific Integrated Circuit). Apparatus suitable for storing computer program instructions and data includes all forms of non-volatile memory, media, and memory devices, by way of example including semiconductor memory devices, such as EPROM, EEPROM, and flash memory devices; magnetic disks, such as internal hard disks or removable disks; magneto-optical disks; and CD-ROM and DVD-ROM disks. The processor and the memory can be supplemented by, or incorporated in, special purpose logic circuitry.

[0105] The subject matter described in this document can be implemented in a computing system that includes a backend component (e.g., as a data server), or includes a middleware component (e.g., an application server), or includes a frontend component (e.g., a client computer having a graphical user interface or a web browser through which a user can interact with an implementation of the subject matter described in this specification), or a combination of one or more such backend, middleware, or frontend components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include local area networks (“LANs”) and wide area networks (“WANs”), the Internet (e.g., the Internet), and peer-to-peer networks (e.g., ad hoc peer-to-peer networks).

[0106] A computing system such as system 100 or system 500 may include clients and servers. The clients and servers are typically remote from each other and typically interact via a communication network (e.g., network 165). The relationship between the clients and servers arises by virtue of computer programs running on respective computers and having a client-server relationship with each other. In some embodiments, the server sends data (e.g., data packets representing content items) to the client device (e.g., for the purpose of displaying data to a user interacting with the client device and receiving user input from the user interacting with the client device). Data generated at the client device (e.g., the result of a user interaction) may be received at the server from the client device (e.g., by data processing system 102 from computing device 150 or content provider computing device 155 or service provider computing device 160).

[0107] Although the operations are depicted in the drawings in a particular order, such operations are not required to be performed in the particular order shown or in sequential order, and all illustrated operations are not required to be performed. The actions described herein may be performed in a different order.

[0108] The separation of the various system components is not required in all embodiments, and the described program components may be included in a single hardware or software product. For example, NLP component 112 or content selector component 118 may be a single component, app, or program, or a logic device having one or more processing circuits, or part of one or more servers of data processing system 102.

[0109] After some illustrative embodiments have now been described, it is apparent that the foregoing is illustrative and not restrictive, and has been presented by way of example. In particular, although many of the examples presented herein involve specific combinations of method actions or system elements, those actions and those elements may be combined in other ways to achieve the same objectives. Actions, elements, and features discussed in connection with one embodiment are not intended to be excluded from a similar role in other implementations or embodiments.

[0110] The words and terms used herein are for the purpose of description and should not be regarded as restrictive. The use of "comprising", "including", "having", "containing", "involving", "characterized by", "characterized in that" and variations thereof is intended to include the items listed hereinafter, their equivalents and additional items, as well as alternative embodiments consisting of the items exclusively listed hereinafter. In one embodiment, the systems and methods described herein consist of one, more than one, each combination, or all of all the described elements, actions, or components.

[0111] Any reference in this document to an embodiment, element, or act of a system or method in the singular can also include embodiments that include a plurality of these elements, and any reference in this document to any embodiment, element, or act in the plural can also include embodiments that include only a single element. References in the singular or plural form are not intended to limit the presently disclosed system or method, its components, acts, or elements to a single or multiple configurations. Any reference to an act or element based on any information, act, or element can include embodiments in which the act or element is based at least in part on any information, act, or element.

[0112] Any embodiment disclosed herein can be combined with any other embodiment or example, and references to "an embodiment", "some embodiments", "one embodiment", etc. are not necessarily mutually exclusive and are intended to indicate that the particular features, structures, or characteristics described in connection with that embodiment can be included in at least one embodiment or example. The terms used herein do not necessarily all refer to the same embodiment. Any embodiment can be included or excluded in combination with any other embodiment in any manner consistent with the aspects and embodiments disclosed herein.

[0113] References to "or" can be interpreted as inclusive, such that any term described using "or" can indicate any one of a single, more than one, and all of the terms described. For example, a reference to "at least one of 'A' and 'B'" can include only 'A', only 'B', and both 'A' and 'B'. Such references used in conjunction with "including" or other open-ended terms can include additional items.

[0114] In cases where there are reference numerals following a technical feature in the drawings, the detailed description, or any claim, these reference numerals have been included to enhance the understandability of the drawings, the detailed description, and the claims. Accordingly, neither the presence nor absence of reference numerals has any effect on the scope of any claim element.

[0115] The systems and methods described herein can be embodied in other specific forms without departing from their characteristics. For example, the data processing system 102 can select a content item for a subsequent action (e.g., for the third action 215) based in part on data from a previous action in the action sequence from thread 200 (such as data indicating the completion or impending start of the second action 210 from the second action 210). The above embodiments illustrate but do not limit the described systems and methods. The scope of the systems and methods described herein is thus indicated by the appended claims rather than the above description, and changes that fall within the meaning and scope of the equivalents of the claims are included therein.

Claims

1. A system for routing packetized actions via a computer network to operate a voice-based digital assistant, comprising: A data processing system including a memory and one or more processors, the data processing system performing the following operations: Receiving data packets via an interface of the data processing system, the data packets including input audio signals detected by sensors of a client device remote from the data processing system; Parsing the input audio signals to identify requests and keywords; Based on the keywords in response to the requests, generating an action data structure for a service provided by a third-party provider remote from the data processing system and the client device; Based on the keywords via a real-time content selection process, selecting a content item provided by a second third-party provider different from the third-party provider, wherein the second third-party provider provides content selection criteria including bids for the content item, and the real-time content selection process uses the bids to select the content item; Sending the content item to the client device via an output signal for presentation by the client device; and Sending the action data structure to the third-party provider to cause the third-party provider to execute the action data structure to perform the service or to invoke a conversation application programming interface to establish a communication session with the client device.

2. The system according to claim 1, comprising: The data processing system selects the content item including an indication of a service type or product type provided by the second third-party provider.

3. The system according to claim 1, comprising: The data processing system selects the content item for a second service different from the service of the action data structure provided by the third-party provider.

4. The system according to claim 1, comprising: The data processing system provides the content item including an audio output to cause the client device to present the audio output of the content item via a speaker of the client device.

5. The system according to claim 1, comprising: The data processing system provides the content item to the client device to cause the client device to output the content item via computer-generated voice.

6. The system according to claim 1, comprising: The data processing system provides the content item including a visual output to cause the client device to output the visual output via a display device of the client device.

7. The system according to claim 1, wherein the data processing system performs the following operations: Detecting an interaction with the content item; and Identifying a transformation of the content item in response to the interaction.

8. The system according to claim 1, wherein the data processing system performs the following operations: Identifying the third-party provider based on the keywords; Selecting a template from a database based on the third-party provider; Populating fields in the template with values received from the client device; And Generating the action data structure based on the template and the values of the fields.

9. The system according to claim 1, comprising: The data processing system receives an indication that the third-party provider invokes the dialogue application programming interface to establish the communication session with the client device.

10. The system according to claim 1, comprising: The data processing system sends the action data structure to the third-party provider, so that the third-party provider invokes the dialogue application programming interface executed by the data processing system to establish the communication session with the client device.

11. A method for routing packetized actions via a computer network to operate a voice-based digital assistant, comprising: Receiving, by a data processing system including one or more processors and a memory, via an interface, a data packet, the data packet including an input audio signal detected by a sensor of a client device remote from the data processing system; Parsing, by the data processing system, the input audio signal to identify a request and keywords; Generating, by the data processing system, an action data structure of a service provided by a third-party provider remote from the data processing system and the client device in response to the request based on the keywords; Selecting, by the data processing system, a content item provided by a second third-party provider different from the third-party provider via a real-time content selection process based on the keywords, wherein the second third-party provider provides content selection criteria including a bid for the content item, and the real-time content selection process uses the bid to select the content item; Sending, by the data processing system, the content item to the client device via an output signal for presentation by the client device; and Sending, by the data processing system, the action data structure to the third-party provider so that the third-party provider executes the action data structure to perform the service or invokes a dialogue application programming interface to establish a communication session with the client device.

12. The method according to claim 11, comprising: Selecting, by the data processing system, the content item including an indication of a service type or product type provided by the second third-party provider.

13. The method according to claim 11, comprising: Selecting, by the data processing system, the content item of a second service different from the service of the action data structure provided by the third-party provider.

14. The method according to claim 11, comprising: Providing, by the data processing system, the content item including an audio output so that the client device presents the audio output of the content item via a speaker of the client device.

15. The method according to claim 11, comprising: Providing, by the data processing system, the content item to the client device so that the client device outputs the content item via computer-generated voice.

16. The method according to claim 11, comprising: Providing, by the data processing system, the content item including a visual output so that the client device outputs the visual output via a display device of the client device.

17. The method according to claim 11, comprising: The interaction with the content item is detected by the data processing system; and the transformation of the content item in response to the interaction is identified by the data processing system.

18. The method according to claim 11, comprising: identifying the third-party provider by the data processing system based on the keyword; selecting a template from a database by the data processing system based on the third-party provider; populating fields in the template by the data processing system with values received from the client device; and generating the action data structure by the data processing system based on the template and the values of the fields.

19. The method according to claim 11, comprising: receiving, by the data processing system, an indication that the third-party provider invokes the dialog application programming interface to establish the communication session with the client device.

20. The method according to claim 11, comprising: sending, by the data processing system, the action data structure to the third-party provider, so that the third-party provider invokes the dialog application programming interface executed by the data processing system to establish the communication session with the client device.

Citation Information

Patent Citations

  • Network transaction method, electronic equipment and electronic device

    CN106127556A

  • An apparatus and method for specifying and obtaining services through voice commands

    WO2002037470A2