Feedback control for data transfers
The feedback control system optimizes data transmission by using merged speech models to select content items and adjust selection processes based on communication quality, addressing inefficiencies in existing systems and enhancing resource utilization.
Patent Information
- Application Number
- DE112017000131
- Authority / Receiving Office
- DE · DE
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2016-12-30
- Filing Date
- 2017-08-31
- Publication Date
- 2025-07-31
- Estimated Expiration
- 2037-08-31
AI Technical Summary
Existing systems face challenges in efficiently transmitting data over computer networks due to limited interfaces, resource consumption, and inconsistent language models, leading to inefficient bandwidth utilization and resource waste.
A feedback control system that processes speech-based inputs using merged speech models to select content items and adjusts the content selection process based on communication session quality, utilizing a natural language processor and feedback monitor to optimize resource usage.
This system enhances data transmission efficiency by reducing resource consumption and improving battery usage in client devices, ensuring accurate and consistent parsing of audio-based instructions across diverse computing resources.
Smart Images

Figure 00000000_0000_ABST
Abstract
Description
BACKGROUNDPacket-based or otherwise overheared network communications of network traffic data between computer devices may prevent a computer device from properly processing the network traffic data, checking an operation associated with the network traffic data, or timely responding to the network traffic data. The overheared network transmissions of network traffic data may also complicate data routing or degrade the quality of the response when the responding computer device arrives at or above its processing capacity, which may result in inefficient bandwidth utilization. Control of network communications corresponding to content item objects may be complicated by the groBe number of content item objects that may initiate network communications of network traffic data between computer devices.US 2011 / 0 271 194 A1 discloses a method for content presentation in which a voice input is received from the user and transmitted to a content system, wherein a command is received in response to the voice input and this is executed using one or more processors, including the modification of the content element.US 2008 / 0 147 388 A1 discloses methods and systems for changing the communication quality of a communication session on the basis of the meaning of voice data.US 2016 / 0 274 864 A1 discloses a content management computer device for managing online speech interactive content. The processor of the device is programmed to identify at least one voice interaction.SUMMARYThe present disclosure is directed to feedback control for data transmissions via one or more interfaces or one or more types of computer networks. For example, computer systems may have access to a limited number of interfaces or limited types of interfaces, or there may be a limited number of available interfaces at a particular time. It may be difficult for a system to efficiently transmit information in response to the currently available interfaces because certain types of interfaces may consume more computing resources or battery. Information about dissimilar computing resources may be difficult to efficiently, reliably, and accurately communicate because dissimilar computing resources are difficult to efficiently process and consistently and accurately parse audio-based instructions in a voice-based computing environment. For example, the dissimilar computing resources may not have access to the same language models or may have access to legacy or unsynchronized language models, which may make it difficult to accurately and consistently provide audio-based instructions.Systems and methods of the present disclosure are generally directed to feedback control for data transmissions. The data processing system may process the speech-based input using speech models trained on merged speech to parse the speech-based instructions and select content items via a real-time content selection process performed by a content selection component. The data processing system may transmit the selected content item to the client computing device to initiate a communication session between the client computing device and a third party device associated with the selected content item. The data processing system may monitor or otherwise receive information about the communication session to measure a property of the communication session and generate a quality signal. The data processing system may then adjust or control the content selection component based on the quality signal to affect the real-time content selection process. For example, blocking or preventing the content selection component from selecting content item objects associated with low quality communication sessions may reduce wasted resource consumption as compared to allowing or allowing selection of the content item and establishing a communication session. Further, for client devices that use battery power, the feedback monitor component may result in savings in battery usage.At least one aspect is directed to a feedback control system for communications over a computer network. The system may include a data processing system executing a natural language processor component and a content selection component. The system may include a feedback monitor component. The natural language processor component may receive, via a surface of the data processing system, data packets comprising an input audio signal detected by a sensor of a client device. The natural language processor component may parse the input audio signal to identify a request and a trigger keyword corresponding to the request. The data processing system may include a content selection component to receive the trigger keyword identified by the natural language processor and select a content item based on the trigger keyword via a real-time content selection process. The system may include a feedback monitor component. The feedback monitor component may receive data packets that transmit acoustic signals between the client device and a conversation application programming interface that has established a communication session with the client device in response to interaction with the content item. The feedback monitor may measure a characteristic of the communication session based on the acoustic signals. The feedback monitor component may generate a quality signal based on the measured property. The content selection component may adjust the real-time selection process based on the quality signal.At least one aspect is directed to a method of transmitting data over a computer network using a feedback control system. The method may be performed, at least in part, by a data processing system executing a natural language processor component and a content selection component. The method may be performed at least in part by a feedback monitor component. The method may include the natural language processor component receiving, via a surface of the data processing system, data packets comprising an input audio signal detected by a sensor of a client device. The method may include parsing, by the data processing system, the input audio signal to identify a request and a trigger keyword corresponding to the request. The method may include the content selection component including a trigger keyword identified by the natural language processor. The method may include selecting, based on the trigger keyword, a content item by the content selection component via a real-time content selection process. The method may include receiving, by the feedback monitor component, data packets that transmit acoustic signals between the client device and a conversation application programming interface that has established a communication session with the client device in response to interaction with the content item. The method may include measuring, by the feedback monitor component, the quality of the communication session based on the acoustic signals. The method may include generating, by the feedback monitor component, a quality signal based on the measured property. The method may include adjusting the real-time selection process based on the quality signal by the feedback monitor component.These and other aspects and implementations are discussed in more detail below. The above information and the following detailed description include illustrative examples of various aspects and implementations, and provide an overview or framework for understanding the spirit and character of the claimed aspects and implementations. The drawings provide an illustration and further understanding of the various aspects and implementations and are incorporated in and constitute a part of this specification.BRIEF DESCRIPTION OF THE DRAWINGSThe accompanying drawings are not to scale. Like reference numerals and labels in the various drawings refer to similar elements. For clarity, not every component may be labeled in each drawing. In the drawings: FIG. 1 shows an illustration of a feedback control system for communications over a computer network. FIG. 2 shows an illustration of operation of a feedback control system for data transmissions over a computer network. FIG. 3 shows an illustration of a method for transmitting data over a computer network using the feedback control system. FIG. 4 is a block diagram illustrating a general architecture for a computer system that may be employed to implement elements of the systems and methods described and illustrated herein.DETAILED DESCRIPTIONMore detailed descriptions of various concepts related to methods, devices, and systems of a feedback control system over a computer network and implementations thereof are provided below. The various concepts introduced above and discussed in more detail below may be implemented in any of numerous ways.The present disclosure is directed to feedback control for data transmissions via one or more interfaces or one or more types of computer networks. For example, computer systems may have access to a limited number of interfaces or limited types of interfaces, or there may be a limited number of available interfaces at a particular time. It may be difficult for a system to efficiently transmit information in response to the currently available interfaces because certain types of interfaces may consume more computing resources or battery. Information about dissimilar computing resources may be difficult to communicate efficiently, reliably, and accurately because dissimilar computing resources are difficult to efficiently process and consistently and accurately parse audio-based instructions in a voice-based computing environment. For example, the dissimilar computing resources may not have access to the same language models or may have access to legacy or unsynchronized language models, which may make it difficult to accurately and consistently provide audio-based instructions.Systems and methods of the present disclosure are generally directed to feedback control for data transmissions. The data processing system may process the speech-based input using speech models trained on merged speech to parse the speech-based instructions and select content items via a real-time content selection process performed by a content selection component. The data processing system may transmit the selected content item to the client computing device to initiate a communication session between the client computing device and a third party device associated with the selected content item. The data processing system may monitor or otherwise receive information about the communication session to measure a property of the communication session and generate a quality signal. The data processing system may then adjust or control the content selection component based on the quality signal to affect the real-time content selection process.FIG. 1 illustrates an example feedback control system 100 for data transmissions over a computer network. The system 100 may include a content selection infrastructure. The system 100 may include a data processing system 102. The data processing system 102 may communicate with one or more of a content provider computing device 106, service provider computing device 108, or client computing devices 104 via a network 105. The network 105 may include computer networks such as the Internet, local area networks, wide area networks, regional or other area networks, intranets, satellite networks, or other communication networks such as mobile voice or data cellular telephone networks. The network 105 may be used to access information resources such as web pages, Internet presences, domain names, or uniform resource locators (URLs) that may be presented, output, rendered, or displayed on at least one computing device 104 such as a laptop, desktop, tablet, personal digital assistant, smartphone, portable computers, or speaker. For example, a user of computing device 104 may access information or data provided by a service provider 108 or content provider 106 via network 105.The network 105 may comprise or form a display network, such as a subset of information sources available on the Internet, associated with a content placement or search engine results system, or which may be selected to include third party content items as part of a content item placement campaign. The network 105 may be used by the computing system 102 to access information resources such as web pages, Internet presences, domain names, or URL addresses that may be presented, output, rendered, or displayed by the client computing device 104. For example, a user of the client computing device 104 can access information or data provided by the content provider computing device 106 or the service provider computing device 108 via the network 105.The network 105 may be any type or form of network and may include one of the following: a point-to-point network, a broadcast network, a wide area network, a local area network, a telecommunications network, a data communication network, a computer network, an ATM (Asynchronous Transfer Mode) network, a SONET (Synchronous Optical Network) network, an SDH (Synchronous Digital Hierarchy) network, a wireless network, or a wired network. The network 105 may include a wireless connection, such as an infrared channel or a satellite frequency band. The topology of the network 105 may include a bus, star, or ring network topology. The network may include mobile telephone networks using any protocol or protocols suitable for communication with mobile devices, including advanced mobile phone protocol ("AMPS"), time division multiple access ("TDMA"), code division multiple access ("CDMA"), global system for mobile communication ("GSM"), general packet radio services ("GPRS"), and universal mobile telecommunications system ("UMTS"). Different types of data may be transmitted over different protocols or the same types of data may be transmitted over different protocols.The system 100 may include at least one data processing system 102. The data processing system 102 may include at least one logical device, such as a computing device having a processor for communication over the network 105, e.g., with the computing device 104, the content provider device 106 (content provider 106), or the service provider device 108 (or service provider 108). The computing system 102 may include at least one computing resource, a server, processor, or memory. The data processing system 102 may include, for example, a plurality of computing resources or servers located in at least one data center. The data processing system 102 may include multiple logically grouped servers and facilitate distributed computing techniques. The logical group of servers may be referred to as a data center, a server farm, or a computer farm. The servers may also be distributed to different locations. A data center or computer farm may be managed as a single entity or the computer farm may include a plurality of computer farms. The servers in a computer farm may be heterogeneous - one or more of the servers or computers may operate according to one or more types of operating system platforms.Servers in the computer farm can be maintained in high density rack systems along with associated storage systems and located in an enterprise data center. When servers are consolidated in this manner, system management, data security, physical security of the system, and system performance can be improved, for example, by searching for servers and high performance storage systems in high performance localized networks. Centralizing all or some of the computing system 102 components, including servers and storage systems and coupling them to enhanced system management tools, enables more efficient use of server resources, which saves power and processing requirements and reduces bandwidth usage.The system 100 may include, access, or otherwise interact with at least one service provider device 108. The service provider device 108 may include at least one logical device, such as a computing device having a processor for communication over the network 105, e.g., with the computing device 104, the computing system 102, or the content provider 106. The service provider device 108 may include at least one computing resource, a server, processor, or memory. For example, the service provider device 108 may include multiple computing resources or servers located in at least one data center. The service provider device 108 may include one or more components or functionalities of the data processing system 102.The content provider computing device 106 may provide audio-based content items for display by the client computing device 104 as an audio output content item. The content item may include an offer for an article of merchandise or service, such as a voice-based message, as follows: "Want I to order a taxi for you?"The content provider computing device 155 may include, for example, memory to store a series of audio content items provided in response to a voice-based request. The content provider computing device 106 may also provide audio-based content items (or other content items) to the data processing system 102 where they may be stored in the data repository 124. The data processing system 102 may select the audio content items and provide the audio content items to (or instruct the content provider computing device 104 to provide) the client computing device 104. The audio-based content items may be exclusively audio or may be combined with text, image or video data.The service provider device 108 may include connecting or otherwise communicating with at least one natural language service provider processor component 142 and a service provider interface 144. The service provider computing device 108 may include at least one service provider's natural language processor component (NLP) 142 and at least one service provider interface 144. The service provider NLP component 142 (or other components such as a direct action API of the service provider computing device 108) may drive the client computing device 104 (via the computing system 102 or by bypassing the computing system 102) to generate a reciprocating voice or audio based conversation in real-time (e.g., a session) between the client computing device 104 and the service provider computing device 108. The service provider NLP 142 may include one or more functions or features such as the NLP component 112 of the data processing system 102. For example, the service provider interface 144 may receive or provide data messages to the direct action API 116 of the data processing system 102. The service provider computing device 108 and the content provider computing device 106 may be associated with the same entity. For example, service provider computing device 106 may create, store, or provide content for a ride-sharing service, and service provider computing device 108 may establish a session with client computing device 106 to cause the provision of a taxi or car of the ride-sharing service to pick up the end user of client computer 104. The computing system 102 may also establish the session with the client computing device, including or bypassing the service provider's computing device 104, via the direct action API 116, NLP component 112, or other components, for example, to cause the ride-sharing service to be provided a taxi or car.The computing device 104 may include connecting or otherwise communicating with at least one sensor 134, transducer 136, audio driver 138, or preprocessor 140. The sensor 134 may include, for example, an ambient light sensor, proximity sensor, temperature sensor, accelerometer, gyroscope, motion detector, GPS sensor, location sensor, microphone, or touch sensor. The transducer 136 may include a speaker or a microphone. Audio driver 138 may provide a software interface to hardware converter 136. The audio driver may execute the audio file or other commands provided by the data processing system 102 to control the transducer 136 to generate a corresponding acoustic wave or sound wave. The preprocessor 140 may be configured to recognize a keyword and perform an action based on the keyword. The preprocessor 140 may filter out one or more terms or modify the terms for further processing before sending the terms to the data processing system 102. The preprocessor 140 may convert the analog audio signals detected by the microphone into a digital audio signal and transmit one or more data packets carrying the digital audio signal to the data processing system 102 via the network 105. In some cases, the preprocessor 140 may transmit data packets containing some or all of the input audio signals in response to detecting a command to execute such transmission. The command may include, for example, a trigger keyword or other keyword or approval to send data packets comprising the input audio signal to the data processing system 102.The client computing device 104 may be associated with an end user that inputs voice requests as audio input to the client computing device 104 (via the sensor 134) and receives audio output in the form of a computer-generated voice that may be provided by the computing system 102 (or the content provider computing device 106 or the service provider computing device 108) to the client computing device 104 and output by the transducer 136 (e.g., a speaker). The computer-generated voice may include records from a real person or a computer-generated language.The data repository 124 may include one or more local or distributed databases, as well as a database management system. The data container 124 may include computer storage or memory and store one or more parameters 126, one or more policies 128, interaction modes 130, and templates 132, among other data. The parameters 126, policies 128, and templates 132 may include information such as rules about a voice-based session between the client computing device 104 and the computing system 102 (or the service provider computing device 108). The content data 130 may include content items for audio output or associated metadata, as well as input audio messages that may be part of one or more communication sessions with the client computing device 104.The data processing system 102 may include a content placement system having at least one computing resource or server. The data processing system 102 may include connecting or otherwise communicating with at least one interface 110. The data processing system 102 may include connecting or otherwise communicating with at least one natural language processor component 112. The computing system 102 may include, interface with, or otherwise communicate with at least one application programming interface ("API") 116 for direct action. The data processing system 102 may include, interface with, or otherwise communicate with at least one session handler 114. The data processing system 102 may include connecting or otherwise communicating with at least one content selection component 118. The data processing system 102 may include interfacing or otherwise communicating with at least one feedback monitor component 120. The data processing system 102 may include, interface with, or otherwise communicate with at least one audio signal generator 122. The data processing system 102 may include connecting or otherwise communicating with at least one data container 124. The at least one data container 124 may include or store one or more data structures or databases, parameters 126, policies 128, content data 130, or templates 132. Parameters 126 may include, for example, thresholds, distances, time intervals, periods, scores, or weights. Content data 130 may include, for example, content campagen data, content groups, content selection criteria, content item objects, or other information provided by a content provider 106 or received or determined by the computing system to facilitate content selection. The content data 130 may include, for example, a prior performance of a content campage.The interface 110, the natural language processor component 112, the session handler 114, the direct action API 116, the content selection component 118, the feedback monitor component 120, or the audio signal generator component 122 may each include at least one processing unit or other logical device, such as a programmable logic array engine, or a module configured to communicate with the database container or database 124. The interface 110, the natural language processor component 112, the session handler 114, the direct action API 116, the content selection component 118, the feedback monitor component 120, the audio signal generator component 122, and the data repository 124 may be separate components, a single component, or part of the data processing system 102. The system 100 and its components, such as a data processing system 102, may include hardware elements such as one or more processors, logic devices, or circuits.The data processing system 102 may receive anonymous information about computer network activities associated with multiple computing devices 104. A user of a computing device 104 may selectively authorize the computing system 102 to obtain information about network activities corresponding to the user's computing device 104. For example, the data processing system 102 may prompt the user of the computing device 104 to obtain one or more information about network activity types. The identity of the user of computing device 104 may remain anonymous and computing device 104 may be associated with a unique identifier (e.g., a unique identifier for the user or computing device provided by the computing system or a user of the computing device). The data processing system may assign a corresponding unique identifier to each observation.A content provider 106 may establish an electronic content campaign. The electronic content campaign may be stored in the data container 124 in the form of content data 130. An electronic content campaign may refer to one or more content groups that correspond to a common theme. A content campaign may include a hierarchical data structure that includes content groups, content item data objects, and content selection criteria. To create a content campaign, the content provider 106 can specify specific values for campaign level parameters of the content campaign. The campaign level parameters can include, for example, a campaign name, a preferred content network for placement of content item objects, a value for resources for use for the content campaign, starting and ending data for the content campaign, a duration of the content campaign, a schedule for placement of content item objects, language, geographic locations, and the type of computing devices on which content item objects are to be provided. In some cases, an impression may refer to a content item object being retrieved from its source (e.g., computing system 102 or content provider 106) and being countable. In some cases, given the possibility of fraudulent clicks, computer-controlled activities may be filtered and excluded as an impression. Thus, in some cases, an impression may refer to a measurement of responses from a web server to a page request by a browser filtered from automatic activity and error codes and recorded at a point as close as possible to an opportunity to render the content item object for display on the computing device 104. In some cases, an impression may refer to a visible or audible impression; e.g., the content item object is at least partially visible (e.g., 20%, 30%, 40%, 50%, 60%, 70%, or more) on a display device of the client computing device 104 or audible via a speaker 136 of the computing device 104. A click or selection may refer to a user interaction with the content item object, such as a voice response to an audible feel, a mouse click, a touch interaction, a gesture, shaking, an audio interaction, or a keyboard input. A conversation may refer to a user executing a desired action with respect to the content item object (e.g., purchase a product or service, participation in a survey, visit to a physical business corresponding to the content item, or completion of an electronic transaction).The content provider 106 may also specify one or more content groups for a content campage. A content group includes one or more content item objects and corresponding content selection criteria, such as keywords, words, phrases, phrases, geographic locations, computing device type, time of day, interest, topic, or vertical. Content groups under the same content campaign may share the same campaign level parameters, but may have descriptions tailored for particular content group level parameters, such as keywords, negative keywords (e.g., this arrangement of the block of the content item in the presence of the negative keyword in the main content), offers for keywords, or parameters associated with the offer or content campaign.To create a new content group, the content provider may provide values for parameters at the content group level of the content group. The content group level parameters include, for example, a content group name or topic and offers for different content placement options (e.g., automatic placement or managed placement) or results (e.g., clicks, impressions, or conversions). A content group name or topic may consist of one or more terms that the content provider 106 may use to capture a topic or item for which content item object / objects of the content group are to be selected for display. For example, a car dealer may generate a different content group for each vehicle make that he is carrying and may further generate a different content group for each model of a vehicle that he is carrying. Examples of group content topics that the car dealer may use may include, for example, "Brand A sport car", "Brand B sport car", "Brand C sedan", "Brand C truck", "Brand C hybrid", or "Brand D hybrid". An example content campagne topic may be "hybrid" and may include, for example, content groups for both "Brand C hybrid" and "Brand D hybrid.".The content provider computing device 106 may provide one or more keywords and content item objects for each content group. Keywords may include terms relevant to the product or services associated with or identified by the content item object. A keyword may include one / n or more terms or phrases. The car house may include, for example, "sport car", "six cylinder engine", "all wheel drive", "fuel economy" as keywords for a content group or content campaign. In some cases, negative keywords may be specified by the content provider to avoid, prevent, block, or disable content placement in certain terms or keywords. The content provider may specify a type of match (e.g., exact match, phrase match, or general match) used to select content item objects.The content provider 106 may provide one or more keywords to be used by the computing system 102 to select a content item object provided by the content provider 106. The content provider 106 may identify one or more keywords for offering and may further provide offering amounts for different keywords. Content provider 106 may provide additional content selection criteria to be used by computing system 102 to select content item objects. Multiple content providers 106 may offer the same or different keywords and the computing system 102 may execute a content selection process or an advertisement auction in response to receiving an indication of a keyword of an electronic message.The content provider 106 may provide the data processing system 102 with one or more content item objects for selection. The data processing system 102 may select (e.g., via the content selection component 118) the content item objects when a content placement capability becomes available that matches the resource allocation, content schedule, maximum offers, keywords, and other selection criteria specified for the content group. Different types of content item objects may be included in a content group, such as a language content item, audio content item, text content item, image content item, video content item, multimedia content item, or content item link. After selecting a content item, the data processing system 102 may send the content item object for rendering on a computing device 104 or a display device of the computing device 104. Rendering may include displaying the content item on a display device or playing the content item via a speaker of the computing device 104. The data processing system 102 may provide instructions for a computing device 104 to render the content item object. The data processing system 102 may direct the computing device 104 or an audio driver 138 of the computing device 104 to generate audio signals or sound waves.The data processing system 102 may include an interface component 110 designed, configured, constructed, or operable to receive and transmit information using, for example, data packets. The interface 110 may transmit and receive information using one or more protocols, such as a network protocol. The interface 110 may include a hardware interface, a software interface, a wired interface, or a wireless interface. The interface 110 may facilitate translating or formatting data from one format to another. For example, the interface 110 may include an application programming interface that includes definitions for communicating between various components, such as software components.The data processing system 102 may include an application, script, or program installed on the client computing device 104, such as an application to communicate input audio signals to the interface 110 of the data processing system 102 and to drive components of the client computing device to play back output audio signals. The data processing system 102 may receive data packets or another signal that includes or identifies an audio input signal. For example, the data processing system 102 may execute or execute the NLP component 112 to receive or obtain the audio signal and parse the audio signal. For example, the NLP component 112 may provide interactions between a human and a computer. The NLP component 112 may be configured with natural language understanding techniques and may allow the data processing system 102 to derive meaning from human or natural language input. The NLP component 112 may include or be configured with machine learning based on machine learning, such as statistical machine learning. The NLP component 112 may use decision trees, statistical models, or probability models to parse the input audio signal. For example, the NLP component 112 may perform functions such as self-name recognition (e.g., to determine, for a given text stream, which elements in the text map self-names, such as persons or places, and which type each of these names corresponds, such as person, place, or organization), natural language generation (e.g., convert information from computer databases or semantic intent to comprehensible human language), natural language understanding (e.g., convert text to more granular representations, such as predicate logic structures that a computer module can manipulate), machine translation (e.g., automatically translate text from one human language to another), Morphological segmenting (e.g., separating words into individual polymorphisms and identifying the class of the polymorphisms, which may be difficult based on the complexity of the morphology or the word construction of the viewed speech), answering questions (e.g., determining a response to a human speech question that may be specific or open), semantic processing (e.g., processing that may occur after identifying a word and encoding its meaning to relate the identified word to other words having similar meanings).The NLP component 112 converts the audio input signal to recognized text by comparing the input signal to a stored representative series of audio waveforms (e.g., in the data bin 124) and selecting the greatest matches. The series of audio waveforms may be stored in the data container 124 or other database accessible to the data processing system 102. The representative waveforms are generated over a large set of users and can then be augmented with voice samples from the user. After the audio signal is converted to recognized text, the NLP component 112 matches the text with words associated with actions that the data processing system 102 can provide, such as via training via users or by manual description.The audio input signal may be detected by the sensor 134 or the transducer 136 (e.g., a microphone) from the client computing device 104. Via the transducer 136, audio driver 138, or other components, the client computing device 104 may provide the audio input signal to the data processing system 102 (e.g., via the network 105), where it may be received (e.g., through the surface 110), and provided to the NLP component 112, or stored in the data repository 124.The NLP component 112 may receive the input audio signal. Of the input audio signal, the NLP component 112 may identify at least one request or trigger keyword corresponding to the request. The request may indicate intention or object of the input audio signal. The trigger keyword may indicate a type of action that is expected to be taken. For example, the NLP component 112 may parse the input audio signal to identify at least one request to eat a evening and join a kino. The trigger keyword may include at least one word, phrase, word stem, partial word or derivative indicating an action to be taken. The trigger keyword "go" or "go to" of the input audio signal may indicate a need for transportation, for example. In this example, the input audio signal (or identified request) does not directly express an intent for transport, but the trigger keyword indicates that a transport is an additional action for at least one other action indicated by the request.The NLP component 112 may parse the input audio signal to identify, determine, retrieve, or otherwise obtain the request and the trigger keyword. For example, the NLP component 112 may apply a semantic processing technique to the input audio signal to identify the trigger keyword or the request. The NLP component 112 may apply the semantic processing technique to the input audio signal to identify a trigger phrase that includes one or more trigger keywords, such as a first trigger keyword and a second trigger keyword. For example, the input audio signal may include the phrase "I need someone to wash my laundry and perform my dry cleaning.". The NLP component 112 may apply a semantic processing technique or other natural language processing technique to the data packets that includes the phrase "wash my laundry" and "perform my dry cleaning" to identify the trigger phrases. The NLP component 112 may further identify multiple trigger keywords, such as laundry and dry cleaning. For example, the NLP component 112 may determine that the trigger phrase includes the trigger keyword and a second trigger keyword.The NLP component 112 may filter the input audio signal to identify the trigger keyword. For example, the data packets carrying the audio input signal may include "It would be bulky if I could find someone who could help me to come to airport", in which case the NLP component 112 may filter out one or more terms as follows: "Es", "would", "bulky", "if", "ich", "someone", "find", "could", "the", "could" or "help". By filtering out these terms, the NLP component 112 can more accurately and reliably identify the trigger keywords, such as "airport" and determine that it is a request for a taxi or ride sharing service.In some cases, the NLP component may determine that the data packets transmitting the input audio signal include one or more requests. For example, the input audio signal may include the phrase "I need someone to wash my laundry and perform my dry cleaning.". The NLP component 112 may determine that it is a request to wash the laundry and perform dry cleaning. The NLP component 112 may determine that it is a single request for a service provider who may provide washing of the laundry and performing the dry cleaning. The NLP component 112 may determine that they are two requests: a first request to a service provider that provides laundry services and a second request to a service provider that provides dry cleaning services. In some cases, the NLP component 112 may combine the plurality of determined requests into a single request and transmit the single request to the service provider device 108. In some cases, the NLP component 112 may transmit the individual requests to respective service provider devices 108, or transmit both requests separately to the same service provider device 108.The data processing system 102 may include a direct action API 116 that is configured and configured to generate an action data structure in response to the request based on the trigger keyword. Processors of the data processing system 102 may invoke the direct action API 116 to execute scripts that generate a data structure for a service provider device 108 to request or order a service or product, such as a car from a ride-sharing service. The direct action API 116 may receive data from the data container 124, as well as data received from the client computing device 104 under the approval of the end user to determine location, time, user accounts, logistic, or other information to enable the service provider device 108 to reserve an operation such as a car from the ride sharing service. Using the direct action API 116, the data processing system 102 may also communicate with the service provider device 108 to complete the conversion by making the reservation for ride-sharing pickup, in this example.The direct action API 116 may perform a particular action to meet the end user's intent determined by the data processing system 102. Depending on the action specified in its inputs, the direct action API 116 may execute code or dialog script identifying the parameters required to satisfy a user request. This code may look up additional information, e.g., in the data container 124, such as the name of a home automation service, or may provide audio output for playback on the client computing device 104 to ask the end user for questions, such as the intended destination of a requested taxi. The direct action API 116 may determine the necessary parameters and package the information into an action data structure, which may then be sent to another component, such as the content selection component 118 or to the service provider computing device 108 to be fulfilled.The direct action API 116 may receive an instruction or command from the NLP component 112 or another component of the data processing system 102 to generate or construct the action data structure. The direct action API 116 may determine a type of action to select a template from the template container 132 stored in the data container 124. Types of actions may include services, products, reservations, or tickets, for example. Types of actions may further include types of services or products. For example, types of services may include a ride-sharing service, a food delivery service, a laundry service, a room cleaning service, repair services, or household services. Types of products may include, for example, clothing, shoes, toys, electronics, computers, books, or ornaments. Types of reservations may include, for example, reservations for restaurant or hairdress deadlines. Types of tickets may include, for example, kino cards, entry cards for sporting events, or air tickets. In some cases, the types of services, products, reservations, or tickets may be categorized based on price, location, shipping type, availability, or other attributes.The direct action API 116 may, upon identifying the type of request, access the corresponding template from a template container stored in the data container 132. Templates may include fields in a structured record that may be populated by the direct action API 116 to support the action requested by the service provider device 108 (such as sending a taxi to pick up an end user at a pick-up location and transport the end user to a destination location). The direct action API 116 may perform a lookup operation in the template container 132 to select the template that matches one or more features of the trigger keyword and the request. For example, if the request corresponds to a request for a car or a driveway to a destination location, the data processing system 102 may select a ride-sharing service template. The ride-sharing service template may include one or more of the following fields: device identifier, pickup location, destination location, number of passengers, or type of service. The direct action API 116 may fill the fields with values. To fill the fields with values, the direct action API 116 may append, request, or otherwise obtain information from one or more sensors 134 of the computing device 104 or a user interface of the computing device 104. For example, the direct action API 116 may recognize the source location using a location sensor, such as a GPS sensor. The direct action API 116 may obtain additional information by sending a survey, prompt, or request to the end of the user of the computing device 104. The direct action API may send the survey, prompt, or request via the interface 110 of the data processing system 102 and a user interface of the computing device 104 (e.g., audio interface, voice-based user interface, display, or touch screen). Thus, the direct action API 116 may select a template for the action data structure based on the trigger keyword or request, fill one or more fields in the template with information detected by one or more sensors 134 or obtained via a user interface, and generate, create, or otherwise construct the action data structure to facilitate performance of an operation by the content provider device 108.The data processing system 102 may select the template based on the template data structure 132 based on various factors including, for example, one or more of the trigger keyword, the request, the third party device 108, the type of the third party device 108, a category into which the third party device 108 falls (e.g., taxi service, laundry service, flower service, or food delivery), location, or other sensor information.To select the template based on the trigger keyword, the data processing system 102 may perform (e.g., via the direct action API 116) a lookup or other query operation on the template database 132 using the trigger keyword to identify a template data structure associated with or otherwise corresponding to the trigger keyword. For example, each template may be associated with one or more trigger keywords in the template database 132 to indicate that the template is configured to generate an action data structure in response to the trigger keyword that the third party device 108 can process to establish a communication session.In some cases, the computing system 102 may identify a third party device 108 based on the trigger keyword. To identify the third party provider 108 based on the trigger keyword, the data processing system 102 may perform a lookup in the data container 124 to identify a third party device 108 associated with the trigger keyword. For example, if the trigger keyword includes "drive" or "go to", the computing system 102 (e.g., via the direct action API 116) may identify the third-party device 108 as corresponding to the taxi service company A. The data processing system 102 may select the template from the template database 132 using the identified third party device 108. The template database 132 may include, for example, associating or correlating third party devices 108 or entities with templates configured to generate an action data structure in response to the trigger keyword that the third party device 108 can process to establish a communication session. In some cases, the template may be customized for the third-party device 108 or for a category of third-party devices 108. The data processing system 102 may generate the action data structure based on the template for the third party provider 108.To construct or generate the action data structure, the data processing system 102 may identify one or more fields in the selected template to fill them with values. The fields may be filled with numeric values, strings, unicode values, Boolean logic, binary values, hexadecimal values, identifiers, location coordinates, geographical areas, time stamps, or other values. The fields or data structure itself may be encrypted or masked to maintain data security.After determining the fields in the template, the data processing system 102 may identify the values for the fields to fill the fields of the template to create the action data structure. The data processing system 102 may obtain, retrieve, determine, or otherwise identify the values for the fields by performing a search or other query operation on the data container 124.In some cases, the data processing system 102 may determine that the information or values for the fields are missing in the data container 124. The data processing system 102 may determine that the information or values stored in the data repository 124 are, expired, bad, or otherwise not suitable to construct the action data structure in response to the trigger keyword and the request identified by the NLP component 112 (e.g., the location of the client computing device 104 may be the old location and not the current location; an account may have expired; the target restaurant may be involved in a new location; physical activity information; or means of transport).If the data processing system 102 determines that it is not currently having access to the values or information for the template field in the memory of the data processing system 102, the data processing system 102 may capture the values or information. The data processing system 102 may acquire or obtain the information by querying or polling one or more available sensors of the client computing device 104, querying the end user of the client computing device 104 for the information, or accessing an online web-based resource using an HTTP protocol. For example, the data processing system 102 may determine that it is not at the current location of the client computing device 104, which may be a required field of the template. The data processing system 102 may request location information from client computing device 104. The computing system 102 may request the client computing device 104 to provide the location information using one or more location sensors 134, such as a global positioning system sensor, WLAN triangulation, cell tower triangulation, Bluetooth beacon, IP address, or other location sensing technique.The direct action API 116 may transmit the action data structure to a third party device (e.g., service provider device 108) to cause the third party device 108 to invoke a conversation application programming interface (e.g., service provider NLP component 142) and establish a communication session between the third party device 108 and the client computing device 104. In response to establishing the communication session between the service provider device 108 and the client computing device 1004, the service provider device 108 may transmit data packets directly to the client computing device 104 via the network 105. In some cases, the third party device 108 may transmit data packets to the client computing device 104 via the computing system 102 and the network 105.In some cases, the third party device 108 may execute at least a portion of the conversation API 142. For example, the third party device 108 may handle certain aspects of the communication session or types of requests. The third party device 108 may utilize the NLP component 112 executed by the computing system 102 to facilitate processing the audio signals associated with the communication session and generating responses to requests. In some cases, the computing system 102 may include the conversation API 142 configured for the third party provider 108. In some cases, the data processing system routes data packets between the client computing device and the third party device to establish the communication session. The data processing system 102 may receive an indication from the third party device 108 that the third party device has established the communication session with the client device 104. The indication may include an identifier of the client computing device 104, a time stamp when the communication session was established, or other information associated with the communication session, such as the action data structure associated with the communication session. In some cases, the data processing system 102 may include a session handler component 114 for managing the communication session and a feedback monitor component 120 for measuring the property of the communication session.The computing system 102 may include, execute, access, or otherwise communicate with a session handler component 114 to establish a communication session between the client computing device 104 and the computing system 102. The communication session may relate to one or more data transmissions between the client device 104 and the data processing system 102 that include the input audio signal detected by a sensor 134 of the client device 104 and the output signal transmitted by the data processing system 102 to the client device 104. The data processing system 102 (e.g., via the session handler component 114) may establish the communication session in response to receiving the audio input signal. The data processing system 102 may set a duration for the communication session. The data processing system 102 may set a timer or counter for the duration set for the communication session. In response to the expiration of the timer, the data processing system 102 may end the communication session.The communication session may refer to a network-based communication session in which the client device 104 provides authenticating information or login data to establish the session. In some cases, the communication session refers to a topic or context of audio signals transmitted by data packets during the session. For example, a first communication session may refer to audio signals transmitted between the client device 104 and the data processing system 102 that relate to a taxi service (e.g., include keywords, action data structures, or content item objects); while a second communication session may refer to audio signals transmitted between the client device 104 and the data processing system 102 that relate to a laundry and dry cleaning service. In this example, the data processing system 102 may determine that the context of the audio signals is different (e.g., via the NLP component 112) and separate the two sets of audio signals into different communication sessions. The session handler 114 may terminate the first session associated with the ride service in response to identifying one or more audio signals associated with dry cleaning and laundry service. Thus, the data processing system 102 may initiate or establish the second session for the audio signals associated with the dry cleaning and laundry service in response to detecting the context of the audio signals.The data processing system 102 may include, execute, or otherwise communicate with a content selection component 118 to receive the trigger keyword identified by the natural language processor and select a content item based on the trigger keyword via a real-time content selection process. In some cases, the direct action API 116 may transmit the action data structure to the content selection component 118 to perform the real-time content selection process and establish a communication session between the content provider device 106 (or a third party device 108) and the client computing device 104.The content selection process may refer to or include selecting sponsored content item objects provided by third content providers 106. The content selection process may include a service in which content items provided by multiple content providers are parsed, processed, weighted, or matched to select one or more content items to be provided to the computing device 104. The content selection process may be performed in real-time or offline. Executing the content selection process in real-time may refer to executing the content selection process in response to the content request received via the client computing device 104. The real-time content selection process may be performed (e.g., initiated or completed) within a time interval in which the request (e.g., 5 seconds, 10 seconds, 20 seconds, 30 seconds, 1 minute, 2 minutes, 3 minutes, 5 minutes, 10 minutes, or 20 minutes) is received. The real-time content selection process may be performed during a communication session with the client computing device 104 or within a time interval after the communication session is completed.For example, the data processing system 102 may include a content selection component 118 that is designed, constructed, configured, or operable to select content item objects. To select content items for display in a speech-based environment, the data processing system 102 may parse (e.g., via an NLP component 112) the input audio signal to identify keywords (e.g., a trigger keyword) and use the keywords to select a matching content item based on a general match, exact match, or phrase match. For example, the content selection component 118 may analyze, parse, or otherwise process items of candidate content items to determine whether the item of candidate content items matches the item of keywords or phrases of the input audio signal detected by the microphone of the client computing device 104. The content selection component 118 may identify, analyze, or recognize speech, audio, phrases, characters, text, symbols, or images of the candidate content items using an image processing technique, character recognition technique, natural language processing technique, or database lookups. The candidate content items may include metadata indicative of the subject matter of the candidate content items, in which case the content selection component 118 may process the metadata to determine whether the subject matter of the candidate content item corresponds to the input audio signal.Content providers 106 may provide additional indicators when setting up a content campaign that includes content items. The content provider 106 may provide content campagen or content group level information that the content selection component 118 may identify by performing a lookup operation using information about the candidate content item. For example, the content item candidate may include a unique identifier that may be associated with a content group, a content campagne, or a content provider. The content selection component 118 may determine content provider 106 information based on the data stored in the content campaign data structure in the data container 124.The data processing system 102 may receive, via a computer network, a request for content to be displayed on a computing device 104. The data processing system 102 may identify the request by processing an input audio signal detected by a microphone of the client computing device 104. The request may include selection criteria of the request, such as the device type, location, and a keyword associated with the request. The request may include the action data structure or action data structure.In response to the request, the data processing system 102 may select a content item object from the data container 124 or a database associated with the content provider computing device 106, and provide the content item for presentation via the computing device 104 via the network 105. The content item object may be provided by a content provider device 108 that is different from the service provider device 108. The content item may correspond to a type of service different from a type of service of the action data structure (e.g., taxi service versus food delivery service). The computing device 104 may interact with the content item object. The computing device 104 may receive an audio response related to the content item. The computing device 104 may receive an indication to select a hyperlink or other button associated with the content item object, causing or allowing the computing device 104 to identify the service provider 108, request a service from the service provider 108, instruct the service provider 108 to perform a service, transmit information to the service provider 108, or otherwise query the service provider device 108.The data processing system 102 may include, execute, or communicate with an audio signal generator component 122 to generate an output signal. The output signal may include one or more portions. The output signal may include, for example, a first portion and a second portion. The first portion of the output signal may correspond to the action data structure. The second portion of the output signal may correspond to the content item selected by the content selection component 118 during the real-time content selection process.The audio signal generator component 122 may generate the output signal with a first portion having a sound according to the first data structure. For example, the audio signal generator component 122 may generate the first portion of the output signal based on one or more values entered into the fields of the action data structure by the direct action API 116. In a taxi service example, the values for the fields may include, for example, 123 Main Street for the pickup location, 1234 Main Street for the destination location, 2 for the number of passengers, and economy for the service level. The audio signal generator component 122 may generate the first portion of the output signal to confirm that the end user of the computing device 104 wishes to proceed with transmitting the request to the service provider 108. The first section may include the following edition: "Want to order an economy car at the taxi service A to pick up two persons at the 123 main street and to place them at the 1234 main street?"In some cases, the first portion may include information received from the service provider device 108. The information received from the service provider device 108 may be adjusted or tailored to the action data structure. For example, the data processing system 102 may transmit (e.g., via the direct action API 116) the action data structure to the service provider 108 before instructing the service provider 108 to perform the operation. Instead, the data processing system 102 may instruct the service provider device 108 to perform a first or pre-processing of the action data structure to generate preliminary information about the operation. In the example of the taxi service, preprocessing the action data structure may include identifying available taxis corresponding to the level of service located around the pickup location, estimating a time period for the closest available taxi to reach the pickup location, estimating an arrival time at the destination location, and estimating a price for the taxi service. The estimated provisional values may include a fixed value, an estimate that may be changed due to various conditions, or a range of values. The service provider device 108 may return the preliminary information to the data processing system 102 or directly to the client computing device 104 via the network 104. The data processing system 102 may include the preliminary results from the service provider device 108 in the output signal and transmit the output signal to the computing device 104. The output signal may be, for example: "Taxi Service A may pick up at the 123 main street in 10 minutes and place at the 1234 main street at 9 am for 10 dollars. Do you wish to order this ride?"This can form the first portion of the output signal.In some cases, the data processing system 102 may form a second portion of the output signal. The second portion of the output signal may include a content item selected by the content selection component 118 during a real-time content selection process. The first portion may be different from the second portion. For example, the first portion may include information corresponding to the action data structure that directly responds to the data packets transmitting the input audio signal detected by the sensor 134 of the client computing device 104, while the second portion may include a content item selected by a content selection component 104 that may be tangentially relevant to the action data structure or sponsored content provided by a content provider device 106. The end user of computing device 104 may request a taxi at taxi services company A, for example. The data processing system 102 may generate the first portion of the output signal with information about the taxi from the taxi service company A. However, the data processing system 102 may generate the second portion of the output signal to include a content item selected from the action data structure based on the keywords "taxi service" and information that could be of interest to the end user. The second portion may include, for example, a content item or information provided by another taxi service company, such as taxi service company B. Even if the user did not explicitly request the taxi service company B, the data processing system 102 may still provide content of taxi service company B, as the user may choose to operate with the taxi service company B.The data processing system 102 may transmit information from the action data structure to the taxi service company B to determine a pickup time, destination arrival time, and price for the trip. The data processing system 102 may receive this information and generate the second portion of the output signal as follows: "Taxi Service Company B may pick up at the 123 main street in 2 minutes and place at the 1234 main street at 8:52 a.m. for 15 dollars. If you are interested in this ride with you ?"The end user of computing device 104 can then select the trip offered by taxi service company A or the trip offered by taxi service company B.Before the data processing system 102 provides the sponsored content item corresponding to the taxi utility B service in the second portion of the output signal, it may notify the end user computing device that the second portion corresponds to a content item selected during a real-time content selection process (e.g., by the content selection component 118). However, the data processing system 102 may have limited access to various types of interfaces to enable notification to the end user of the data processing device 104. For example, computing device 104 may not include a display device, or the display device may be disabled or turned off. The display device of computing device 104 may consume more resources than the speaker of computing device 104, such that it may be less efficient to turn on the display device of computing device 104 than when the speaker of computing device 104 is used to communicate the notification. Thus, in some cases, the data processing system 102 may improve the efficiency and effectiveness of information transmission over one or more interfaces or one or more types of computer networks. For example, the data processing system 102 may modulate (e.g., via the audio signal generator component 122) the portion of the output audio signal that includes the content item such that the end user receives the indication or notification that that portion of the output signal includes the sponsored content item.The data processing system 102 (e.g., via interface 110 and network 105) may transmit data packets comprising the output signal generated by the audio signal generator component 122. The output signal may cause the audio driver component 138 of the computing device 104 to drive a speaker (e.g., transducer 136) of the client device 104 to generate an acoustic wave corresponding to the output signal.The data processing system 102 may include a feedback monitor component 120. The feedback monitor component 120 may include hardware or software for measuring the property of the communication session. The feedback monitor component 120 may receive data packets that transmit acoustic signals transmitted between the client device (e.g., computing device 104) and a conversation application programming interface (e.g., NLP component 112 executed by the data processing system or the service provider NLP component 142 executed by the service provider device 108, a third party device, or the content provider device 106) that has established a communication session with the client device in response to interaction with the content item. In some cases, the content provider device 106 may execute an NLP component that includes one or more functions or components of the service provider NLP component 142 or the NLP component 112. The NLP component executed by the service provider device 108 or the content provider device 106 may be adapted for the service provider device 108 or the content provider device 106. By adjusting the NLP component, the NLP component may reduce bandwidth usage and request responses compared to a generic or standard NLP component, as the NLP component may be configured with more precise requests and responses, resulting in reduced front-and-back between the NLP component and the client computing device 104.The feedback monitor component 120 may measure a characteristic of the communication session based on the acoustic signals. The feedback monitor component 120 may generate a quality signal based on the measured property. The quality signal may include or relate to a quality level, a quality metric, a quality evaluation or a quality level. The quality signal may include, for example, a numerical metric (e.g., 0 to 10, where 0 is the lowest quality and 10 is the highest quality, or vice versa), a letter rating (e.g., A to F, where A is the best quality), a binary value (e.g., yes / no; pass / fail; 1 / 0; high / low), a rank, or a percentage. The quality signal may include an average quality signal determined from communication between a plurality of client devices communicating with the same NLP component or provider device 106 or 108.The feedback monitor component 120 may measure the characteristic of the communication session using various measurement techniques, heuristic techniques, policies, conditions, or tests. The feedback monitor component 120 may parse data packets transmitted between the client device 104 and the content provider device, third party devices, service providers, or data processing system to determine a property of the communication session. The quality may refer to the quality of the communication channel used to transmit the data or the quality of the data being communicated. The quality of the communication channel may relate to, for example, a signal-to-noise ratio, an ambient noise level, a delay, a time interval, a latency, a dropout, an echo or broken calls. The quality of the data being communicated may relate to the quality of the responses generated by the NLP component that responds to audio signals detected by the microphone of the computing device. The quality of the data may be based on the response speed of the NLP component, the accuracy of the NLP component, or the latency between the NLP component receiving the audio signal or request from the client device 104 and transmitting a response.The feedback monitor component 120 may determine the quality of the communication channel by measuring the amount of background noise or signal level to determine the signal-to-noise ratio ("SNR"). The feedback monitor component 120 may compare the measured or determined SNR to a threshold to determine the quality level. An SNR of 10 dB may be considered good, for example. The thresholds may be predetermined or determined via a machine learning model (e.g., based on feedback from a plurality of devices).The feedback monitor component 120 may also determine the quality of the communication channel based on the ping time between the client device 104 and the provider device or data processing system. The data processing system may compare the ping time to a threshold to determine the quality level. The ping threshold may be, for example, 20 ms, 30 ms, 50 ms, 100 ms, 200 ms or more. The feedback monitor component 120 may determine the quality of the communication channel based on the audio dropouts (e.g., pauses or interrupts in audio; blanking the audio). The feedback monitor component 120 may identify an echo in the communication channel to determine a low quality level. The feedback monitor component 120 may determine the number of dropped calls for the NLP component during a time interval or a ratio of dropped calls to total calls and compare this to a threshold to determine the quality level. The threshold may be, for example, 2 calls interrupted per hour; or 1 call interrupted every 100 calls.The feedback monitor component 120 may determine the quality of the communication session based on the quality of the responses generated by the NLP component (or conversation API) communicating with the client computing device 104. The quality of the responses may include, for example, the amount of time the NLP component takes to generate a response, the text of the response, the accuracy of the response, the relevance of the response, a semantic analysis of the response, or a network activity of the client device in response to or based on the response provided by the NLP component. The feedback monitor component 120 may determine the amount of time the NLP component takes to generate the response by differentiating a time stamp corresponding to the time of receipt of the audio signal from the client device 104 by the NLP component and a time stamp corresponding to the time of transmission of the response by the NLP. The feedback monitor component 120 may determine the duration of time by differentiating a time stamp corresponding to the time of transmitting the audio signal by the client device and a time stamp corresponding to the time of receiving the response from the NLP component by the client device.The feedback monitor component 120 may determine the quality of the responses by parsing data packets that include the response. For example, the feedback monitor component 120 may parse and analyze the text of the response, the accuracy of the response, or the relevance of the response to the request from the client device. The feedback monitor component 120 may make this assessment by providing the request for another NLP component and comparing the responses from the two NLP components. The feedback monitor component 120 may provide this assessment by providing the request and response by an external verifier. The feedback monitor component 120 may determine the consistency of the response by comparing a plurality of responses to a plurality of similar responses provided by a plurality of client devices. The feedback monitor component 120 may determine the quality of the responses based on how often the client device transmits audio signals that include the same request (e.g., indicating that the responses do not fully respond to the request sent by the client device).The feedback monitor component 120 may determine the quality of the responses generated by the NLP based on the network activity of the client device. For example, the NLP component may receive a voice request from the client device, generate a response to the voice request, and transmit data packets that transmit the response of the client device. The client device may perform a network activity or a change of a network activity upon receiving the response from the NLP component. For example, the client device may end the communication session, which may indicate that the NLP component has fully responded to the client device, or the NLP has not successfully responded to the client device and the client device has abandoned the NLP component. The feedback monitor component may determine that the client device has completed the call for good reasons or bad reasons based on a confidence measure associated with the response generated by the NLP component. The confidence measure may be associated with a probabilistic or statistical semantic analysis used to generate the response.The feedback monitor component 120 may determine that the client device has completed the communication session based on an absence of audio signals transmitted by the client device. The feedback monitor component 120 may determine that the client device has completed the communication session based on a termination or termination command transmitted by the client device. The feedback monitor component 120 may determine a quality level based on an amount of silence from the client device (e.g., absence of audio signals). The absence of audio signals may be identified based on the SNR from the client device being less than a threshold (e.g., 6 dB, 3 dB, or 0 dB). The feedback monitor component may measure the characteristic based on a duration of the communication session. For example, a duration greater than a threshold may indicate that the end user of the client device was satisfied with the communication session. However, a long duration combined with other characteristics such as an increased amplitude of audio signals, repeated request, and decreased tempo may indicate low quality because the user of the client may have expended an unnecessary or undesirable longer time period to participate in communication.The NLP component may perform semantic analysis of the requests sent by the client device to determine that the client device repeatedly transmits the same or similar requests even though the NLP component has generated and provides responses. The feedback monitor component 120 may determine that the quality level is low based on a number of repeated requests within a time interval (or successive repeated requests) that exceed a threshold (e.g., 2, 3, 4, 5, 6, 7 or more).In some cases, the feedback monitor component 120 may determine the quality of the communication session at different portions of the communication session (e.g., beginning, middle, or end; or time intervals). For example, the feedback monitor component 120 may determine the quality of a first portion or first time interval of the communication session; and the quality of a second portion or second time interval in the communication session following the first portion or first time interval. The feedback monitor component 120 may compare the quality at the two sections to determine the quality of the entire communication session. The difference in quality between the two portions that is greater than a threshold may indicate, for example, low quality, inconsistent quality, or unreliable quality.In some cases, the feedback monitor component 120 may determine the quality based on a property of the communication session or at least a portion thereof. The characteristic may include, for example, at least one of the following: amplitude, frequency, tempo, tone, and pitch. For example, the feedback monitor component 120 may use the characteristic to determine a response of the user of the client device or the feel of use of the client. For example, if the amplitude of the audio signals transmitted by the client device increases after each response from the NLP, the feedback monitor may determine that the end user is frustrated with the responses generated by the NLP component. The feedback monitor component 120 may compare the amplitude of the audio signals detected by the client device to a threshold or other audio signals received by the client device during the same communication session or other communication sessions.The feedback monitor component 120 may determine the quality based on a characteristic such as the tempo or pitch of the audio signals detected by the client device and transmitted to the NLP component. For example, the feedback monitor component 120 may determine that a deceleration in tempo (e.g., rate of words spoken per time interval) after each NLP response may indicate that the end user is not satisfied with the response generated by the NLP component and repeats it more slowly to allow the NLP component to better parse the audio signals and improve the response. In some cases, an increased or consistent tempo may indicate that the user of the client device is satisfied with the responses generated by the NLP and has trust in the responses. In some cases, an increase in the pitch of the audio signals detected by the client device may indicate poor quality of the responses from the NLP or lack of confidence in the responses.In some cases, the feedback monitor component 120 may transmit requests to the client device to measure or determine the quality. For example, the feedback monitor component 120 may transmit survey questions to the end user asking for the quality of the communication session, the NLP component, or the provider device. In some cases, the feedback monitor component 120 may generate the request in response to the feedback monitor component 120 determining that the first quality signal is below a threshold. For example, the feedback monitor component 120 may determine a first quality signal based on the measurement of the quality using characteristics such as the increase in the amplitude of the audio signals detected by the client device in combination with the decrease in the tempo of the audio signals detected by the client device. The feedback monitor component 120 may generate a quality signal indicative of a lower quality level based on the combined characteristics of amplitude and tempo. In response to the low quality signals determined based on the combination characteristics, the feedback monitor component 120 may generate and transmit a request to the client device that either implicitly or explicitly asks for the quality of the communication session (e.g., How satisfied they are with the responses generated by the NLP component?; How satisfied they are with the communication session?). In another example, the data processing system may determine quality based on whether the service provider 108 may provide the requested service. For example, the end user may request a product or service, but the service provider 108 responds with the indication that it does not have the product or cannot perform the service, which may cause the end user to express frustration at the service provider 108. The data processing system 102 may identify this frustration and accordingly assign a quality.In some cases, the feedback monitor component 120 may measure the property based on network activity on multiple electronic surfaces and summarize the quality measured by the multiple electronic surfaces to generate a summed quality signal. The summed quality signal may be an average, a weighted average, an absolute sum, or another combined quality signal value. The feedback monitor component 120 may further generate statistics for the combined quality signal value or perform statistical analysis to determine the standard deviation, variance, 3-sigma quality, or 6-sigma qualities.The feedback monitor component 120 may adjust the real-time content selection process performed by the content selection component 118. Adjusting the real-time content selection process may refer to adjusting a weight used to select the content item selected by the content provider device 106 or service provider device 108 or third party device 108 that executed the NLP component used to establish the communication session with the client device 104. For example, if the content item has resulted in a low quality communication session, the feedback monitor component 120 may adjust an attribute or parameter of the content data 130 that includes the content item to reduce the likelihood that that content item will be selected for similar action data structures or similar client devices 104 (or accounts or profiles thereof).In some cases, the feedback monitor component 120 may prevent or block the content selection component 118, in this real-time selection process, from selecting the content item in response to the quality signal being less than a threshold. In some cases, the feedback monitor component 120 may allow or allow the content selection component 118, in this real-time selection process, to select the content item in response to the quality signal being greater than or equal to a threshold.FIG. 2 shows an illustration of operation of a feedback control system for data transmissions over a computer network. The system may include one or more components of the system 100 shown in FIG. 1. The system 100 may include one or more electronic surfaces 202 a- nexecuted or provided by one or more client computing devices 104 a- n. Examples of electronic interfaces 202 a- nmay include audio interfaces, video-based interfaces, display screens, HTML content items, multimedia, images, video, text-based content items, SMS messaging applications, chat applications, or natural language processors.At ACT 204, the client computing device 104 may receive data packets, signals, or other information indicative of feedback from or via an electronic surface 202. At ACT 206, the one or more client computing devices 104 a- n, the one or more service provider devices 108 a- n, or the one or more content provider devices 106 a- nmay transmit data packets to the feedback monitor component 124. The data packets may be associated with the communication session established between the client device 104 and one or more of the service provider devices 108 or the content provider devices 106. The data packets may be transmitted from a respective device to the feedback monitor component 124.In some cases, the feedback monitor component 124 may intercept data packets transmitted from a device 104, 106, or 108 to a respective device. The feedback monitor component 124 may analyze the intercepted data packets or route or forward the data packet to its intended destination. Thus, the feedback monitor component 124 may be interposed between the client device 104 and the service / third party device 108 or the content provider device 106.At ACT 208, the feedback monitor component 124 may transmit the intercepted or received data packets from the communication session to the NLP component 112. At ACT 210, NLP component 112 may perform semantic analysis of the data packets and provide them back to feedback component 124. In some cases, the NLP component 112 may perform natural language processing on the audio signals from the communication session 206 to compare the responses of the NLP component generated by the provider devices 106 or 108. The feedback monitor component 124 may compare the responses generated by a control NLP component 112 to determine whether the third party NLP components are operating at a comparable or satisfactory level.At ACT 212, the feedback monitor component 124 may determine a quality signal for the communication session 206 and adjust the real-time content selection process performed by the content selection component 118 such that the next time the content selection component 118 receives a request for content, the content selection component 118 appropriately weights the content item (or content provider) associated with the communication session 206 to increase or decrease the likelihood that the content item is selected. For example, if provider 108 is associated with a plurality of low quality communication sessions, feedback monitor component 124 may instruct content selection component 118 to prevent selection of content items that result in establishing a communication session with provider 108.FIG. 3 is an illustration of an exemplary method for dynamically modulating packetized audio signals. The method 300 may be performed by one or more components, systems, or elements of system 100 or system 400. The method 300 may include a data processing system that receives an input audio signal (ACT 305). The data processing system may receive the input audio signal from a client computing device. For example, the natural language processor component executed by the data processing system may receive the input audio signal from a client computing device via an interface of the data processing system. The data processing system may receive data packets carrying or including the input audio signal detected by a sensor of the client computer (or client device).At ACT 310, the method 300 may include the data processing system parsing the input audio signal. The natural language processor component may parse the input audio signal to identify a request and a trigger keyword corresponding to the request. The audio signal recognized by the client device may include, for example: "Okay device, I need a ride sharing of taxi services company A to arrive at 1234 Main Street." In this audio signal, the initial trigger keyword may include "OK device," which may indicate to the client device to transmit an input audio signal to the data processing system. A preprocessor of the client device may filter out the terms "OK device" before sending the remaining audio signal to the data processing system. In some cases, the client device may filter out additional terms or generate keywords that are transmitted to the data processing system for further processing.The data processing system may identify a trigger keyword in the input audio signal. The trigger keyword may include, for example, "go to" or "drive", or variations of these terms. The trigger keyword may indicate a type of service or product. The data processing system may identify a request in the input audio signal. The request may be determined based on the terms "I need". The trigger keyword and the request may be determined using a semantic processing technique or other natural language processing technique.In some cases, the data processing system may generate an action data structure. The data processing system may generate the action data structure based on the trigger keyword, the request, a third party device, or other information. The action data structure may be present in response to the request. For example, if the end user of the client computing device requests a taxi from taxi services company A, the action data structure may include information to request a taxi service from taxi services company A. The data processing system may select a template for taxi services company A and fill fields in the template with values that allow taxi services company A to send a taxi to the user of the client computing device to pick up the user and transport it to the desired destination.At ACT 315, the data processing system may select a content item. For example, a content selection component may receive a trigger keyword, request, or action data structure and select a content item via a real-time content selection process. The selected content item may correspond to a content provider, a service provider, or a third party provider. The client device may interact with the content item to establish a communication session with the content item provider or another device associated with the content item. The device associated with the content item may interact with the client device using a conversation API, such as an NLP.At ACT 320, a feedback monitor component may receive data packets that transmit acoustic signals between the client device and a conversation application programming interface that has established a communication session with the client device in response to interaction with the content item. At ACT 325, the feedback monitor component may measure a quality or characteristic of the communication session based on the acoustic signals and generate a quality signal based on the measured characteristic. At ACT 330, the feedback monitor component or data processing system may adjust the real-time selection process based on the quality signal.FIG. 4 shows a block diagram of an example computer system 400. Computer system or computing device 400 may include or be used to implement system 100 or its components, such as data processing system 102. The data processing system 102 may include a smart personal assistant or a voice-based digital assistant. Computer system 400 includes a bus 405 or other communication component for transmitting information, and a processor 410 or processing circuitry coupled to bus 405 for processing information. The computer system 400 may also include one or more processors 410 or processing circuitry coupled to the bus for processing information. Computer system 400 further includes main memory 415, such as random access memory (RAM) or other dynamic storage device, coupled to bus 405 for storing data, as well as instructions to be executed by processor 410. The main memory 415 may be or include the data container 145. Main memory 415, when executing instructions by processor 410, may be further used to store position data, temporary variables, or other medium-term information. Computer system 400 may further include read-only memory (ROM) 420 or other static storage device coupled to bus 405 to store static information and instructions for processor 410. A storage device 425, such as a solid state device, magnetic or optical disk, may be coupled to bus 405 to permanently store information and instructions. The storage device 425 may include or be part of the data container 145.The computer system 400 may be coupled via the bus 405 to a display 435, such as a liquid crystal display (LCD) or active matrix display, such that information may be displayed to a user. An input device 430, such as a keyboard with alphanumeric and other keys, may be coupled to bus 405 to provide selected information and commands to processor 410. The input device 430 may include a touch screen display 435. The input device 430 may also include a cursor control, such as a mouse, trackball, or arrow keys on the keyboard, such that direction data and selected commands may be transmitted to the processor 410 and the movement of the cursor may be controlled on the display 435. The display 435 may be part of, for example, the data processing system 102, the client computing device 150, or other components of FIG. 1.The processes, systems, and methods described herein may be implemented by the computer system 400 in response to the processor 410 executing an instruction set contained in the main memory 415. These instructions may be read from another computer readable medium, such as storage device 425, to main memory 415. Execution of the instruction set contained in main memory 415 causes computer system 400 to execute the processes described and illustrated herein. In a multi-processor arrangement, one or more processors may be used to execute the instructions contained in main memory 415. Hardwired circuits may be used in place of or in combination with software instructions along with the systems and methods described herein. The systems and methods described herein are not limited to a specific combination of hardware circuits and software.Although an example computer system has been described in FIG. 4, the subject matter, including the operations described in this specification, may be implemented in other types of digital electronic circuits or in computer software, firmware, or hardware, including also the structures disclosed in this specification and their structural equivalents, or in combinations of one or more of them.For situations where the systems discussed herein may collect personal information about users or utilize personal information, the users may be provided with a capability to check whether programs or features, collect personal information (e.g., information about a user's social network, social actions or activities, a user preference or location of a user), or to check whether and / or how content is received from a content server or other data processing system that may be more relevant to the user. Additionally, certain data may be anonymized in one or more ways before being stored or used, such that personal data is removed when parameters are generated. For example, a user identity may be anonymous such that no personally identifiable information can be determined for the user, or a geographic location of the user may be generalized, wherein location information (such as city, zip code, or federal country) is extracted such that a particular location of a user cannot be determined. Thus, the user may have control over how information about him or her is collected and used by a content server.The subject matter and the operations described in this specification can be implemented in digital electronic circuitry, or in computer software, firmware, or hardware, including the structures disclosed in this specification and their structural equivalents, or in combinations of one or more of these. The subject matter described in this specification may be implemented as one or more computer programs, e.g., one or more circuits of computer program instructions encoded on one or more computer storage media for execution by, or control the operation of, data processing apparatus. Alternatively or additionally, the program instructions may be encoded in an artificially generated propagating signal, such as a machine generated electrical, optical, or electromagnetic signal, generated to encode information for transmission to a suitable receiver device for execution by a computing device. A computer storage medium may be, or be included in, a computer readable storage device, a computer readable storage substrate, a freely addressable or serial random access memory array or device, or a combination thereof. However, although a computer storage medium is not a propagating signal, a computer storage medium may be a source or destination of computer program instructions encoded in an artificially generated propagating signal. The computer storage medium may also be one or more separate components or media (e.g., multiple CDs, disks, or other storage devices). The operations described in this specification may be implemented as operations performed by a computing device on data stored on or received from one or more computer readable storage device(s).The terms "data processing system", "computing device", "component", or "data processing device" include various devices, devices, and machines for processing data, including, for example, a programmable processor, a computer, one or more systems on a chip, or more of the same, or combinations of the foregoing. The apparatus may include special-purpose logic circuitry such as an FPGA (field programmable gate array) or an ASIC (application specific integrated circuit). The device may also include code, in addition to hardware, that creates an execution environment for the corresponding computer program, such as code representing processor firmware, a protocol stack, a database management system, an operating system, a cross-platform runtime environment, a virtual computer, or a combination thereof. The device and execution environment may implement various computer model infrastructures, such as web services, as well as distributed computing and spatially distributed computing infrastructures. For example, the direct action API 116, content selection component 118, or NLP component 112, and other components of the computing system 102 may include or share one or more computing devices, systems, computing devices, or processors.A computer program (also known as a program, software, software application, app, script, or code) may be written in any form of programming language, including compiled languages or interpreted languages, declarative or procedural languages, and may be deployed in any form, including as a single program or as a module, component, subroutine, object, or other device suitable for use in a computing environment. A computer program may correspond to a file in a file system. A computer program may be stored in a portion of a file containing other programs or data (e.g., one or more scripts stored in a markup language document), a single file specific to the program in question, or multiple coordinated files (e.g., files storing one or more modules, subroutines, or portions of code). A computer program may be provided and executed on one computer or on multiple computers that are distributed at one site or at multiple sites and interconnected via a communication network.The processes and logic flows described in this specification may be performed by one or more programmable processors executing one or more computer programs (e.g., components of data processing system 102) to perform actions by processing input data and generating outputs. The processes and logic flows may also be implemented as, special purpose logic circuits such as an FPGA (field programmable gate array) or an ASIC (application specific integrated circuit), and devices may also be implemented as these. Media suitable for storing computer program instructions and data include any type of read-only memory, media, and storage devices, including semiconductor memory elements, including EPROM, EEPROM, and flash memory devices; magnetic hard drives, such as internal hard drives or removable disks; magneto-optical hard drives; and CD-ROM and DVD-ROM drives. The processor and memory may be supplemented by or incorporated into a special purpose logic circuit.The subject matter described herein may be implemented in a computer system that includes a backend component, such as a data server, or a middleware component, such as an application server, or a front end component, such as a client computer having a graphical user interface or a web browser through which a user may interact with an implementation of the subject matter described in this specification, or a combination of one or more of those backend, middleware, or front end components. The components of the system may be interconnected by any form or medium of digital data communication, such as a communication network. Examples of communication networks include a local area network ("LAN") and a wide area network ("WAN"), an inter-network (e.g., the Internet), and peer-to-peer networks (e.g., ad hoc peer-to-peer networks).The computer system, such as system 100 or system 400, may include clients and servers. A client and a server are generally remote from each other and typically interact with each other via a communication network (e.g., network 165). The relationship between client and server arises from computer programs that are executed on the respective computers and have a client-server relationship with each other. In some implementations, a server sends data (e.g., data packets representing a content item) to a client device (e.g., for purposes of displaying data and receiving user input from a user interacting with the client device). Data generated in the client device (e.g., a result of user interaction) may be received by the client device at the server (e.g., received by the computing system 102 from the computing device 150 or the content provider computing device 155 or the service provider computing device 160).Although the operations are illustrated in the drawings in a particular order, it is not necessary for these operations to be performed in the illustrated particular order or in consecutive order, nor is it necessary for all illustrated operations to be performed. Actions described herein may be performed in a different order.The separation of various system components does not require separation in all implementations, and the described program components may be included in a single hardware or software product. For example, the NLP component 110 or content selection component 125 may be a single component, app or program, or logical device having one or more processing circuits or part of one or more servers of the data processing system 102.Having described some illustrative implementations, it will be apparent that the foregoing is for purposes of illustration and not limitation, and has been presented in an exemplary manner only. In particular, although many of the examples presented herein include specific combinations of method operations or system elements, these operations and elements may be combined in other ways to achieve the same goals. Acts, elements, and features discussed in connection with an implementation are not intended to be excluded from a similar role in other implementations or implementations.The terminology and terminology used herein is for the purpose of description and should not be considered as limiting. The use of the words "including," "comprising," "having," "including," "including," "characterized by," "characterized in that," and variations thereof, is intended to mean herein that the items listed thereafter, equivalents thereof, and additional items, as well as alternative implementations consisting solely of the items listed thereafter are included. In one implementation, the systems and methods described herein consist of one, any combination of more than one, or all of the elements, acts, or components described herein.Any references to implementations or elements or acts of the systems and methods referred to in the singular herein may also include implementations including a plurality of these elements, while any references to an implementation or element or act of any kind referred to in the plural herein may also include implementations including only a single element. References to the singular or plural form are not intended to limit the presently disclosed systems and methods, the components, acts, or elements thereof to single or multiple configurations. References to an action or element of any type, based on information, actions, or elements of any type may include implementations whose action or element is based at least in part on information, actions, or elements of any type.Any of the implementations disclosed herein may be combined with any other implementations or embodiments, where references to "one implementation", "some implementations", "the one implementation", or the like are not necessarily mutually exclusive, and are intended to indicate that a particular feature, structure, or characteristic described in connection with the implementation may be included in at least one implementation or embodiment. Such terms, as used herein, do not necessarily refer to the same implementation. Each implementation may be combined with any other implementation, including or excluding, and in any manner consistent with the aspects and implementations disclosed herein.References to "or" may be construed as inclusive, such that all terms described using "or" may indicate any single, more than one, or all terms described. For example, a reference to "at least one of 'A' and 'B'" may include only 'A', only 'B', and both 'A' and 'B'. These references, used in connection with "comprising" or other open terminology, may include additional elements.Where technical features in the drawings, detailed description, or any claim are followed by reference numerals, the reference numerals have been included to increase the intelligibility of the drawings, detailed description, or claims. Accordingly, neither those reference numerals nor their absence have a limiting effect on the scope of the claim elements.The systems and methods described herein may also be realized by other embodiments without departing from the essential features thereof. For example, the data processing system 102 may select a content item for a subsequent action (e.g., for the third action 215) based in part on data from a previous action in the sequence of actions of the thread 200, such as data from the second action 210 indicating that the second action 210 is complete or about to begin. The foregoing implementations are considered to be illustrative rather than restrictive of the systems and methods described herein. The scope of the systems and methods described herein is, therefore, indicated by the appended claims rather than by the foregoing description, and changes which come within the meaning and range of equivalence of the claims are therefore incorporated herein.
Claims
A feedback control system (100) for communications over a computer network, comprising: a natural language processor component (112) executed by a data processing system (102) to receive, via an interface (110) of the data processing system (102), data packets comprising an input audio signal detected by a sensor (134) of a client device (104); the natural language processor component (112) to parse the input audio signal to identify a request, a content provider, and a trigger keyword corresponding to the request; a content selection component (118) executed by the data processing system (102) to receive the trigger keyword identified by the natural language processor component (112) and, based on the trigger keyword, select a content item via a real-time content selection process; A feedback monitor component (120) for: receiving data packets that transmit acoustic signals between the client device (104) and a conversation application programming interface (116) that established a communication session (206) with the client device (104) in response to interaction with the content item; measuring a property of the communication session (206) based on the acoustic signals; and generating a quality signal based on the measured property; and adjusting the real-time content selection process based on the quality signal by the content selection component (118).The system (100) of claim 1, comprising the conversation application programming interface (116) executed by the data processing system (102).The system (100) of claim 1, comprising the conversation application programming interface (116) executed by a third party device (108).The system (100) of claim 1, comprising the data processing system (102) for: intercepting the data packets transmitted from the client device (104); parsing the data packets to identify a third party device (108); and routing the data packets to the third party device (108).The system (100) of claim 1, comprising the data processing system to: parse the data packets to determine an absence of acoustic signals; and generate the quality signal indicative of a low quality level based on the absence of the acoustic signals.The system (100) of claim 1, comprising: the feedback monitor component (120) to forward the data packets transmitting the acoustic signals to the natural language processor component (112) to determine a first characteristic of the acoustic signals in a first time interval and a second characteristic of the acoustic signals in a second time interval after the first time interval; and measure the characteristic based on a comparison of the first characteristic and the second characteristic.The system (100) of claim 6, wherein the first characteristic and the second characteristic include at least one of amplitude, frequency, tempo, tone, and pitch.The system (100) of claim 1, comprising the data processing system (102) to: transmit a plurality of voice-based requests to the client device (104); and measure the characteristic based on responses to the plurality of voice-based requests.The system (100) of claim 1, comprising the data processing system (102) for: generating a request based on the quality signal that is less than a threshold; receiving a response to the request from the client device (104); and generating a second quality signal based on the response.The system (100) of claim 1, comprising the data processing system (102) to: measure the characteristic based on a duration of the communication session (206).The system (100) of claim 1, comprising the data processing system (102) to: measure the property based on network activity on a plurality of electronic surfaces (202a-n); and aggregate quality signals measured by the plurality of electronic surfaces (202a-n) to generate a summed quality signal.The system (100) of claim 1, comprising the data processing system (102) for: preventing the content selection component (118) from selecting the content item in response to the quality signal being less than a threshold in this real-time content selection process.The system (100) of claim 1, comprising: allowing the content selection component (118), in this real-time content selection process, to select the content item in response to the quality signal being greater than or equal to a threshold.A method of transmitting data over a computer network (105) using a feedback control system (100), comprising: receiving, by a natural language processor component (112) executed by a data processing system (102), data packets over an interface (110) of the data processing system (102) comprising an input audio signal detected by a sensor (134) of a client device (104); parsing, by the data processing system (102), the input audio signal to identify a request, a content provider, and a trigger keyword according to the request; receiving, by a content selection component (118) executed by the data processing system (102), the trigger keyword identified by the natural language processor component (112); selecting, based on the trigger keyword, a content item by the content selection component (118) via a real-time content selection process; receiving, by a feedback monitor component (120), data packets that transmit acoustic signals between the client device (104) and a conversation application programming interface (116) that established a communication session (206) with the client device (104) in response to interaction with the content item; measuring, by the feedback monitor component (120), a characteristic of the communication session (206) based on the acoustic signals by the feedback monitor component (120); generating, by the feedback monitor component (120), a quality signal based on the measured characteristic; and adjusting, by the content selection component (118), the real-time content selection process based on the quality signal by the content selection component (118).The method of claim 14, comprising: executing the conversation application programming interface (116) on the data processing system (102).The method of claim 14, comprising: intercepting, by the data processing system (102), the data packets transmitted from the client device (104); parsing, by the data processing system (102), the data packets to identify a third party device (108); and routing, by the data processing system (102), the data packets to the third party device (108).The method of claim 14, comprising: parsing, by the data processing system (102), the data packets to determine an absence of acoustic signals; and generating, by the data processing system (102), the quality signal indicative of a low quality level based on the absence of the acoustic signals.The method of claim 14, comprising: forwarding, by the feedback monitor component (120), the data packets transmitting the acoustic signals to the natural language processor component (112) to determine a first characteristic of the acoustic signals in a first time interval and a second characteristic of the acoustic signals in a second time interval after the first time interval; and measuring, by the feedback monitor component (120), the quality based on a comparison of the first characteristic and the second characteristic.The method of claim 18, wherein the first characteristic and the second characteristic include at least one of amplitude, frequency, tempo, tone, and pitch.The method of claim 14, comprising: preventing the content selection component (118) from selecting the content item in response to the quality signal being less than a threshold in this real-time content selection process.
Citation Information
Patent Citations
Methods And Systems For Changing A Communication Quality Of A Communication Session Based On A Meaning Of Speech Data
US20080147388A1
Voice ad interactions as ad conversions
US20110271194A1
Systems and methods for enabling user voice interaction with a host computing device
US20160274864A1