Establishing an audio-based network session with a non-registered resource
By using natural language processor components to parse audio signals and generate action data structures in a computer network environment, the scalability limitations of computer systems interacting with network resources are solved, enabling seamless, scalable interaction and navigation, and providing a secure and fast user experience.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2017-06-13
- Publication Date
- 2026-03-24
AI Technical Summary
In existing technologies, computer systems have limited scalability when interacting with network resources, resulting in a waste of computing resources such as bandwidth, events, and power. Furthermore, it is difficult to achieve scalable network resource loading, especially when interacting with various network resources without specific integration.
By using a natural language processor component to parse input audio signals in a computer network environment, identifying request and trigger keywords, generating action data structures, and using a navigation component to establish sessions and select interaction models, seamless interaction with network resources is achieved, employing headless rendering and interaction model navigation technology.
It enables real-time, scalable interaction between computer systems and network resources, avoids specific integration requirements, and provides a secure, fast, and seamless user experience.
Smart Images

Figure CN114491226B_ABST
Abstract
Description
[0001] Case Analysis
[0002] This application is a divisional application of Chinese invention patent application 201780000981.1, filed on June 13, 2017. Background Technology
[0003] Computer systems can interact with network resources configured to interact with them. To interact with network resources, interfaces can be designed for each resource. Because it's not possible to efficiently create custom interfaces for every network resource, the scalability of some computer systems may be limited, and computational resources (such as bandwidth, events, or power) may overlap and be wasted. Therefore, the scalability of some computer systems is not capable of scaling using onboarding techniques for network resources. Summary of the Invention
[0004] At least one aspect relates to a system for retrieving digital components in a computer network environment based on voice-activated data packets, comprising a natural language processor (NLP) component. The NLP component may be executed by a data processing system. The NLP component may receive data packets including input audio signals detected by sensors of a computing device via an interface of the data processing system. The NLP component may parse the input audio signals to identify a request, a content provider, and triggering keywords corresponding to the request. The system may include a direct action application programming interface (API) that can generate an action data structure based on the triggering keywords. The action data structure may be generated in response to the request and the identified content provider. The system may include a navigation component. The navigation component may establish a session with the identified content provider. The navigation component may render digital components received from the content provider via the session. The navigation component may select an interaction model associated with the content provider. The navigation component may generate a data array based on the interaction model and the action data structure. The system may transmit the data array to the content provider via the interface of the data processing system.
[0005] At least one aspect relates to a method for retrieving and interacting with digital components in a computer network environment based on voice-activated data packets. The method may include a natural language processor component, executed by a data processing system, receiving data packets via an interface of the data processing system that may include input audio signals detected by sensors of a computing device. The method may include the natural language processor component parsing the input audio signals to identify a request, a content provider, and triggering keywords corresponding to the request. The method may include generating an action data structure based on the triggering keywords using a direct action application programming interface. The action data structure may be generated in response to the request and the content provider. The method may include establishing a session with the content provider by a navigation component. The method may include rendering digital components received from the remote data processing system via the session by the navigation component. The method may include selecting an interaction model associated with the remote data processing system by the navigation component. The method may include generating a data array by the navigation component based on the interaction model and the action data structure. The method may include transmitting the data array to the remote data processing system via the interface of the data processing system.
[0006] The foregoing general description, the following illustrated description, and the detailed description are exemplary and intended to provide further explanation of the claimed invention. Other objectives, advantages, and novel features will become apparent to those skilled in the art from the following illustrated description and detailed description. Attached Figure Description
[0007] The accompanying drawings are not intended to be drawn to scale. Similar reference numerals and instructions indicate similar elements in various drawings. For clarity, not every component is labeled in every drawing. In the drawings:
[0008] Figure 1 The diagram illustrates a block diagram of an example system used for selecting and interacting with digital components via a computer network.
[0009] Figure 2 The illustration is used for retrieval. Figure 1 The diagram illustrates a block diagram of an example method for a digital component in the system.
[0010] Figure 3 The illustration is provided. Figure 1 The diagram shows a block diagram of an example data flow of the system.
[0011] Figure 4 The illustration is in Figure 1 The diagram shows a block diagram of an example computer system used in the system illustrated in the figure. Detailed Implementation
[0012] The following is a more detailed description of various concepts related to methods, apparatuses, and systems for retrieving and interacting with digital components in an audio-based computer network, as well as implementations of such methods, apparatuses, and systems. The various concepts introduced above and discussed in more detail below can be implemented in any of many ways, as the concepts are not limited to any particular implementation.
[0013] This disclosure generally relates to systems and methods for increasing the scalability of loading network resources, such as digital components, into voice- or image-based networks. The system may include a data processing system that enables navigation and interaction with digital components, such as web pages, portions thereof, or other online documents using voice input and output interfaces on voice, image, or computing devices. The system may receive and process voice input (also referred to herein as input audio signals) to identify the digital component. The system may identify a content provider that provides the digital component. Voice input (or other non-text input, such as image input) may include verb instructions associated with at least one resource of a defined type that can be accessed in the digital component. The resource may be, include, or correspond to an action performed via, contained in, or otherwise identified by the digital component or a specific item or web page. The system may create a session with a content provider hosting the digital component. The system may render the digital component received from the content provider. The system may render the digital component headless without a user interface. The system may use an interaction model to analyze the rendered digital component to navigate to resources using the interaction model. The system may select an interaction model from a first model type and a second model type. The first model type may be or include a general model that incorporates a first training dataset based on a collection of websites with resources of the same type as those recognized in voice input. The second model type may be or include a specific model that incorporates a second training dataset dedicated to resources specific to the digital component. Both model sets may include data for determining access to or navigation of the digital component to the corresponding resource. The system may provide information about the associated resource to the computing device based on received instructions for accessing the corresponding resource in order to perform one or more subsequent operations. The system may also update the interaction model based on one or more determinations made when navigating the digital component to access or attach to a resource.
[0014] The system can input data into the digital component. This data can be provided by the user and stored in a secure storage wallet on the computing device or system. Data can be automatically provided to the system during the creation of a session with the digital component or when required by one or more subsequent operations during the created session.
[0015] The system can generate a general model by acquiring or storing information about one or more of the terms, layouts, categories, and hyperlinks common to accessing corresponding types of resources on multiple different digital components. The system can generate a specific model by acquiring and storing information about one or more of the terms, layouts, menus, categories, and hyperlinks for resources specifically identified in the voice input.
[0016] Loading network resources into a voice-activated network, enabling other network resources to interact with them, can be technically challenging because it may require creating a unique interface for each resource. For example, resource owners might need to create application programming interfaces (APIs) that allow providers of voice-based computing devices to interact with the resources. Providers of voice-based computing devices may also need to generate programs that allow computing devices to interact with resources via the provided APIs. Voice-based computing device providers can offer APIs that resource owners can integrate into their digital components. This loading process can be time-consuming, computationally inefficient, and may require collaboration between the two parties.
[0017] The loading process may be necessary to guide transactions with resources, gain access to secure portions of the resources, or exchange sensitive information. Technologies used to interact with these resources using voice input interfaces to guide transactions and exchange data may require tight integration with the resource interfaces of voice-based computing devices, enabling the devices to navigate through resources using voice commands published to specific user interface modules. Such resource-specific integration may require predefined boundaries between the interface module and the corresponding resource, and this integration can be essential to ensure a secure, fast, and seamless experience for users of the computing device. This is because prior secure configuration of how secure data is available to the interface module (e.g., from a secure wallet on the computing device) and prior integration with specific resources may be required, allowing the interface module to know the layout of resource inputs, components, or other interactive objects.
[0018] Therefore, interaction with resources that have not yet undergone a loading process to establish tight integration between voice-based computing devices and resources may not be possible. Without integration with individual and specific resources, it may be impossible to use voice-based computing devices to perform actions with resources.
[0019] Furthermore, voice (or image or video) based computing devices can use multiple voice recognition-based interface modules. Depending on the specific interface module used, for each resource that a voice-based computing device can interact with, predefined integration with each corresponding interface module may be required.
[0020] Providing such integration with voice-based computing devices for navigating and interacting with resources suffers from the drawbacks described above, namely, the need for very specific and closely predefined integration for each resource. These technologies may not be scalable and cannot be applied to all resources. For example, providing functionality to allow voice-controlled user interface modules to navigate a website or domain might require limiting functionality to one type of interface module, rather than being able to interact with any voice input recognition module.
[0021] Therefore, there is a need for provisioning universal and scalable technologies that enable voice-based computing devices to navigate and interact with all resources using voice, speech, or image input recognition interface modules.
[0022] This disclosure provides the following technical steps to enable real-time and scalable technologies that can operate with all resources providing goods, services, or other actions. A user can initiate commands via a voice-based computing device, such as, “OK, I want to obtain product X on website Y, please let me know the price and availability.” A system (e.g., a data processing system) can use natural language processing to parse and interpret the input audio signal. For example, the system can obtain the user's credentials for website Y via a secure wallet. The system can initiate a session with website Y. A server can render website Y without a user interface. The system can then use one of specific or general interaction models to navigate website Y to obtain the price and availability of product X. The price and availability can be provided to the user's computing device via an output audio file presented to the user via a transducer (e.g., a speaker) on or associated with the computing device. The user can confirm the purchase or provide additional details to the system via a second input audio signal that the system can parse. In this example, the system enables users to interact with website Y without requiring any specific integration between the system and the voice-based computing device. This technology provides a method and system for interacting with websites, web resources, or other digital components in real time using trained interaction models, without requiring specific integration.
[0023] The technology described herein can be general and scalable to all types of digital components, and enables data processing systems to interact with digital components without prior integration or coordination between digital component providers and providers of voice-based computing devices.
[0024] The above technical solution provides a mapping between the server and the website domain by rendering the corresponding digital components on the server side without a user interface. Using at least one trained interaction model, the system can identify structures, elements, input elements, and other components of the digital components. The above steps can occur seamlessly and automatically in real time to further provide an effective end-user experience for processing input audio signals to interact with the digital components.
[0025] Figure 1 The diagram illustrates a block diagram of an example system 100 for selecting and interacting with digital components via a computer network. System 100 may include content selection infrastructure. System 100 may include a data processing system 102. The data processing system 102 may communicate with one or more of a content provider computing device 106 or a client computing device 104 via a network 105. Network 105 may include computer networks such as the Internet, local area networks, wide area networks, metropolitan area networks, or other regional networks, intranets, satellite networks, and other communication networks such as voice or data mobile phone networks and combinations thereof. Network 105 may access information resources such as web pages, websites, domain names, or Uniform Resource Locators, which may be presented, output, rendered, or displayed on at least one computing device 104, such as a laptop computer, desktop computer, tablet, personal digital assistant, smartphone, home assistant device, portable computer, or speaker. For example, via network 105, a user of computing device 104 may access information or data provided by content provider device 106.
[0026] The data processing system 102 may include an interface 110, a natural language processor component 112, and a session controller component 114. The data processing system 102 may also include a direct action application programming interface 116, a navigation component 118, and an audio signal generator component 122. The data processing system 102 may also include a data repository 124 on which parameters 126, policies 128, interaction models 130, and templates 132 are stored.
[0027] Network 105 can be used by data processing system 102 to access information resources, such as web pages, websites, domain names, or Uniform Resource Locators, which can be presented, output, rendered, or displayed by client computing device 104. Web pages, websites, and other digital content stored or otherwise provided by content providing device 106 can be referred to as digital components or content items. Through network 105, users of client computing device 104 can access information or data (e.g., digital components such as content items) provided by content providing computing device 106.
[0028] Digital components may be rendered via a display device of computing device 104 or on data processing system 102. Rendering may include displaying the digital components or other content items on a display device, which may or may not be part of computing device 104. In some embodiments, computing device 104 does not include a display device for rendering digital components. For example, computing device 104 may render digital components simply by playing them through speakers of computing device 104. Data processing system 102 may act as middleware and enable computing device 104 to interact with digital components in an audio-based manner.
[0029] Network 105 can be any type or form of network and may include any of the following: peer-to-peer network, broadcast network, wide area network, local area network, telecommunications network, digital communication network, computer network, ATM (Asynchronous Transfer Model) network, SONET (Synchronous Optical Network) network, SDH (Synchronous Digital Hierarchy) network, wireless network, and wired network. Network 105 may include wireless links, such as infrared channels or satellite strips. The topology of network 105 may include bus, star, or ring network topologies. The network may include mobile phone networks using any one or more protocols for communication between mobile devices, including Advanced Mobile Phone Protocol (“AMPS”), Time Division Multiple Access (“TDMA”), Code Division Multiple Access (“CDMA”), Global System for Mobile Communications (“GSM”), General Packet Radio Service (“GPRS”), or Universal Mobile Telecommunications System (“UMTS”). Different types of data may be transmitted via different protocols, or the same type of data may be transmitted via different protocols.
[0030] System 100 may include at least one data processing system 102. Data processing system 102 may include at least one logical device, such as a computing device having a processor to communicate via network 105 with, for example, computing device 104 or content providing device 106 (content provider 106). Data processing system 102 may include at least one computing resource, server, processor, or memory. For example, data processing system 102 may include multiple computing resources or servers located in at least one data center. Data processing system 102 may include multiple logically grouped servers and technologies that facilitate distributed computing. The logical grouping of servers may be referred to as a data center, server farm, or machine farm. Servers may also be geographically distributed. A data center or machine farm may be managed as a single entity, or a machine farm may include multiple machine farms. Servers within each machine farm may be heterogeneous—one or more of the servers or machines may operate according to one or more types of operating system platforms.
[0031] Servers in a server farm can be stored alongside associated storage systems in high-density rack systems within an enterprise data center. For example, this consolidation of servers improves system manageability, data security, physical security, and system performance by positioning servers and high-performance storage systems on a localized, high-performance network. Centralizing all or some of the components of the data processing system 102, including servers and storage systems, and coupling them with advanced system management tools allows for more efficient use of server resources, saving power and processing requirements and reducing bandwidth usage.
[0032] System 100 may include, access, or otherwise interact with at least one content providing device 106. Content providing device 106 may include at least one logical device, such as a computing device having a processor to communicate via network 105. Content providing device 106 may include at least one computing resource, server, processor, or memory. For example, content providing device 106 may include multiple computing resources or servers located in at least one data center.
[0033] Content providing computing device 106 can provide digital components to data processing system 102 and computing device 104. The digital components can be web pages, including graphics, text, hyperlinks, and machine-executable instructions. The digital components can be visually displayed to end users, such as via a web browser that renders the web page and displays the rendered web page on a monitor. The digital components can be or include web pages, which include offers, goods, services, or information. For example, a digital component could be a website selling clothing.
[0034] The computing device 104 may include or interface with or communicate with at least one sensor 134, transducer 136, audio driver 138, or preprocessor 140. Sensor 134 may include, for example, an ambient light sensor, proximity sensor, temperature sensor, accelerometer, gyroscope, motion detector, GPS sensor, position sensor, microphone, or touch sensor. Transducer 136 may include a speaker or microphone. Audio driver 138 may provide a software interface to hardware transducer 136. Audio driver 138 may execute audio files or other instructions provided by data processing system 102 to control transducer 136 to generate corresponding acoustic waves or sound waves. Preprocessor 140 may be configured to detect keywords and perform actions based on those keywords. Preprocessor 140 may filter out one or more terms and modify terms before sending them to data processing system 102 for further processing. Preprocessor 140 may convert analog audio signals detected by the microphone into digital audio signals and transmit one or more data packets carrying the digital audio signals to data processing system 102 via network 105. In some cases, the preprocessor 140 may transmit data packets (or other protocol-based transmissions) carrying some or all of the input audio signals in response to a detected instruction to perform such a transmission, such as "OK," "Start," or other wake words. The instruction may include, for example, trigger keywords or other keywords, or approval, to transmit data packets including the input audio signals to the data processing system 102. In some cases, the main user interface of the computing device 104 may be a microphone and speaker.
[0035] Types of actions can include, for example, services, goods, reservations, or tickets. The type of action can further include the type of service or goods. For example, types of services can include car-sharing services, food delivery services, laundry services, cleaning services, repair services, or housekeeping services. Types of goods can include, for example, clothing, shoes, toys, electronics, computers, books, or jewelry. Types of reservations can include, for example, dinner reservations or hair salon appointments. Types of tickets can include, for example, movie tickets, sporting event tickets, or airline tickets. In some cases, the type of service, goods, reservation, or ticket can be categorized based on price, location, delivery type, availability, or other attributes.
[0036] Client computing device 104 can be associated with an end user who inputs a voice query as audio input to client computing device 104 (via sensor 134) and receives audio output in the form of computer-generated speech provided from data processing system 102 (or content providing computing device 106) to client computing device 104 and output from transducer 136 (e.g., a speaker). Computer-generated speech can include recordings from a real person or computer-generated language. In addition to voice queries, input can also include one or more image or video clips generated or obtained from client computing device 104 (e.g., via network 105) and parsed by data processing system 102 to obtain the same type of information obtained by parsing the voice query. For example, a user can take a picture of an item they wish to purchase. Data processing system 102 can perform machine vision on the image to identify the content of the image and generate a text string that identifies the image content. The text string can be used as an input query.
[0037] Data processing system 102 may include at least one interface 110, or be connected to or otherwise communicate with other devices, such as via network 105. Data processing system 102 may include at least one natural language processor component 112, or be connected to or otherwise communicate with it. Data processing system 102 may include at least one direct action application programming interface (“API”) 116, or be connected to or otherwise communicate with it. Data processing system 102 may include at least one session controller 114, or be connected to or otherwise communicate with it. Data processing system 102 may include at least one navigation component 118, or be connected to or otherwise communicate with it. Data processing system 102 may include at least one audio signal generator 122, or be connected to or otherwise communicate with it. Data processing system 102 may include at least one data repository 124, or be connected to or otherwise communicate with it.
[0038] Data processing system 102 may include or interface with or otherwise communicate with navigation component 118. Navigation component 118 enables voice-based interaction between computing device 104 and digital components such as websites provided by content provider 106. Digital components provided by content provider 106 may not be configured to allow voice-based interaction. For example, digital components may be web pages, including text, images, videos, input elements, and other non-audio elements. Furthermore, there may be no prior integration between the digital component (or its provider) and data processing system 102. Navigation component 118 may use, for example, a headless browser or a headless web tool renderer to render and recognize input elements, text, and other data in the digital components. When rendered with a headless renderer, the rendered digital components do not require a graphical user interface to function. Navigation component 118 may interact with these elements using an interaction model. For example, navigation component 118 may input data arrays into input fields, select and activate input elements (e.g., navigation or submit buttons), and retrieve data based on the interaction model to perform actions such as those recognized in input audio signals received from computing device 104. As an example, given the input audio signal "buy two shirts", navigation component 118 can generate a data array of "text=2". Navigation component 118 can also recognize input fields and "buy" buttons in a headless rendered webpage. Navigation component 118 can input the text "2" into the input field and then select the "buy" button to complete the transaction.
[0039] Data repository 124 may include one or more local or distributed databases and may include a database management system. Data repository 124 may include computer data storage or storage and may store data such as one or more parameters 126, one or more policies 128, interaction models 130, and templates 132. Parameters 126, policies 128, and templates 132 may include information such as rules regarding voice-based conversations between client computing device 104 and data processing system 102. Data repository 124 may also store content data, which may include content items for audio output or associated metadata, and input audio messages that may be part of one or more communication conversations with client computing device 104. Parameters 126 may include, for example, thresholds, distances, time intervals, durations, scores, or weights.
[0040] Interaction model 130 can be generated and updated by navigation component 118. Data repository 124 can include multiple interaction models. Interaction model 130 can be categorized into general models and content provider-specific models. General models can be further subdivided into different interaction or action categories. For example, interaction model 130 can include general models for different types of commercial websites such as shopping websites, weather provider websites, and booking websites.
[0041] Interaction model 130 may also include specific models. A specific interaction model may be specific to content provider 106 or a specific digital component provided by content provider 106. For example, for a specific website Y, the specific interaction model may know the placement of links and menus, how to navigate the website, and how to store and categorize specific products and data within the website. Navigation component 118 can use this information to navigate throughout the website and provide data arrays to the website when interacting with it to complete actions or other transactions with the website.
[0042] A general interaction model can be used when navigation component 118 does not have a predetermined number of interactions with digital components or content provider 106 to generate a specific interaction model. Navigation component 118 can train models (both general and specific interaction models) by initially collecting data from specific sessions (e.g., user sessions after obtaining their permission, or specific training sessions where the user interacts with digital components). For example, given an input audio signal, a user can complete a task associated with the input audio signal (e.g., “OK, buy a shirt.”). Navigation component 118 can receive user input as the action is completed and build a model that enables navigation component 118 to recognize input elements (such as text fields and buttons) used to complete the action. Specific models can be trained for specific digital components, while general models can be trained using multiple digital components within a given category. Training can enable the model to determine interaction data, such as the steps involved in purchasing a specific item, product catalogs, product sorting, and how sorting is performed for digital components. The model enables navigation component 118 to correctly identify and interact with the items or services identified in the input audio signal.
[0043] Both types of interaction models 130 can be trained and updated during and after a session between the data processing system 102 and the content provider 106. For example, while a general model is used for the digital component, the navigation component 118 can construct a specific model for the digital component. Once the specific interaction model for a particular navigation component 118 is deemed reliable, for example, built using data from a predetermined number of sessions, the navigation component 118 can begin using the specific interaction model for the digital component instead of the general model. The navigation component 118 can update the interaction model 130 using data from additional (or new sessions).
[0044] Interface 110, natural language processor component 112, session controller 114, direct action API 116, navigation component 118, or audio signal generator component 122 may each include at least one processing unit or other logic device such as a programmable logic array engine, or a module configured to communicate with a database repository or database 124. Interface 110, natural language processor component 112, session controller 114, direct action API 116, navigation component 118, audio signal generator component 122, and data repository 124 may be separate components, single components, or part of data processing system 102. System 100 and its components such as data processing system 102 may include hardware elements such as one or more processors, logic devices, or circuitry.
[0045] Data processing system 102 can obtain anonymous computer network activity information associated with multiple computing devices 104. Users of computing devices 104 can authorize data processing system 102 to obtain network activity information corresponding to their computing devices 104. For example, data processing system 102 can prompt users of computing devices 104 to consent to obtaining one or more types of network activity information. The identity of users of computing devices 104 can remain anonymous, and computing devices 104 can be associated with unique identifiers (e.g., unique identifiers provided by data processing system 102 or the users of computing devices for users or computing devices). Data processing system 102 can associate each observation with a corresponding unique identifier.
[0046] Data processing system 102 may include interface component 110, which is designed, configured, constructed, or operated to receive and transmit information using, for example, data packets. Interface 110 may use one or more protocols, such as network protocols, to receive and transmit information. Interface 110 may include a hardware interface, a software interface, a wired interface, or a wireless interface. Interface 110 may facilitate the conversion or formatting of data from one format to another. For example, interface 110 may include an application programming interface that includes definitions for communication between various components, such as software components.
[0047] Data processing system 102 can receive data packets or other signals that include or identify audio input signals. For example, data processing system 102 can execute or run NLP component 112 to receive or acquire audio signals and parse them. For example, NLP component 112 can provide human-computer interaction. NLP component 112 can be configured using techniques for understanding natural language and allowing the data processing system to derive meaning from human or natural language input. NLP component can include or be configured using machine learning-based techniques (such as statistical machine learning). NLP component 112 can utilize decision trees, statistical models, or probabilistic statistical models to parse the input audio signal. NLP component 112 can perform functions such as named entity recognition (e.g., given a stream of text, determining which item in the text maps to an appropriate name, such as a person or place, and what type of each such name is, such as person, location, or organization), natural language generation (e.g., converting information from a computer database or semantic intent into understandable human language), natural language understanding (e.g., converting text into more normal expressions, such as first-order logical structures that a computer module can manipulate), machine translation (e.g., automatically translating text from one human language to another), morpheme segmentation (e.g., separating words into individual morphemes and identifying morpheme classifications based on morpheme complexity or considering the word structure of the language, which can be challenging), question answering (e.g., determining answers to human language questions, which can be specific or open-ended), and semantic processing (e.g., processing that occurs after words are identified and their meanings are encoded in order to correlate identified words with other words using similar meanings).
[0048] NLP component 112 converts the audio input signal into recognized text by comparing the input signal with a corresponding set of stored audio waveforms (e.g., in a data repository 124) and selecting the closest match. The set of audio waveforms may be stored in the data repository 124 or in another database accessible to the data processing system 102. Typical waveforms are generated among a large set of users and can subsequently be enhanced using speech samples from the users. After the audio signal is converted into recognized text, NLP component 112 matches the text to words associated with actions that the data processing system 102 can perform, for example, through inter-user training or through manual specification.
[0049] The audio input signal can be detected by the sensor 134 or transducer 136 (e.g., a microphone) of the client computing device 104. The audio input signal can be provided to the data processing system 102 (e.g., via network 105) via the transducer 136, the audio driver 138, or other components of the client computing device 104, where it can be received (e.g., via interface 110) and provided to the NLP component 112 or stored in the data repository 124.
[0050] NLP component 112 can acquire an input audio signal. Based on the input audio signal, NLP component 112 can identify at least one request or at least one trigger keyword corresponding to the request. The request can indicate the intent or topic of the input audio signal. The trigger keyword can indicate the type of action that may be taken. For example, NLP component 112 can parse the input audio signal to identify at least one request to leave home in the evening to attend dinner and watch a movie. The trigger keyword can include at least one word, phrase, root word, or incomplete word, or a derivation indicating the action to be taken. For example, the trigger keyword "go" or "to go to" from the input audio signal can indicate a need for transportation. In this example, the input audio signal (or the identified request) does not directly express the intent of transportation, but the trigger keyword indicates that transportation is an auxiliary action to at least one other action indicated by the request.
[0051] NLP component 112 can parse the input audio signal to identify, determine, retrieve, or obtain request and trigger keywords. For example, NLP component 112 can apply semantic processing techniques to the input audio signal to identify trigger keywords or requests. NLP component 112 can apply semantic processing techniques to the input audio signal to identify trigger phrases, which include one or more trigger keywords, such as a first trigger keyword and a second trigger keyword. For example, the input audio signal may include the sentence "I needed someone to do my laundry and my dry cleaning." NLP component 112 can apply semantic processing techniques or other natural language processing techniques to data groups including this sentence to identify the trigger phrases "do my laundry" and "do may dry cleaning." NLP component 112 can further identify multiple trigger keywords, such as "laundry" and "dry cleaning." For example, NLP component 112 can determine that the trigger phrases include a trigger keyword and a second trigger keyword.
[0052] NLP component 112 can parse the input audio signal to identify, determine, retrieve, or obtain the identifier of remote content provider 106 or remote data processing system 102 in a method similar to that used by NLP component 112 to obtain request and trigger keywords. For example, an input audio signal including the phrase "ok, buy a red shirt from ABC" can be parsed to identify ABC as the seller of the shirt. Data processing system 102 can then determine the content provider 106 associated with ABC. Content provider 106 may be a server hosting ABC's website. Data processing system 102 can identify ABC's network address, such as "www.ABC.com". Data processing system 102 can transmit a confirmation audio signal to computing device 104, such as "Are you referring to ABC of www.ABC.com?" In response to receiving a confirmation message from computing device 104, data processing system 102 can initiate a session with content provider 106 located at www.ABC.com.
[0053] NLP component 112 can filter the input audio signal to identify triggering keywords. For example, a data packet carrying the input audio signal might include "It would be great if I could get someone that could help me go to the airport." In this case, NLP component 112 can filter out one or more of the following terms: "it," "would," "be," "great," "if," "I," "could," "get," "someone," "that," "could," or "help." By filtering out these terms, NLP component 112 can more accurately and reliably identify triggering keywords, such as "go to the airport," and determine that this is a request for a taxi or ride-sharing service.
[0054] Data processing system 102 may include a Direct Action API 116, designed and configured to generate action data structures based on trigger keywords in response to requests and identification by a remote content provider 106. A processor of data processing system 102 may invoke the Direct Action API 116 to execute scripts that generate data structures for the content provider 106 to request or order services or goods (such as cars in a car-sharing service). The Direct Action API 116 may obtain data from data repository 124 and data received from client computing devices 104 by end users to determine location, time, user account, logistics, and other information to allow data processing system 102 to perform operations, such as reserving cars in a car-sharing service.
[0055] When the data processing system 102 interacts with digital components from the content providing device 106, the Direct Action API 116 can execute specified actions to satisfy the end user's intent. Depending on the action specified in its input, the Direct Action API 116 can execute code or a dialogue script that identifies parameters needed to satisfy the user's request, which can be included in an action data structure. Such code can look up additional information or it can provide audio output for rendering on the client computing device 104 to ask the end user questions such as the user's preferred shirt size, to continue the example above where the input audio signal is "ok, buy a red shirt." The Direct Action API 116 can determine the necessary parameters and can encapsulate the information into an action data structure. For example, when the input audio signal is "ok, buy a red shirt," the action data structure could include the user's preferred shirt size.
[0056] After identifying the request type, the Direct Action API 116 can access the corresponding template from the template repository stored in the data repository 124. Template 132 can include fields populated by the Direct Action API 116 in a structured data set to further satisfy or interact with a digital component provided by the content provider 106. The Direct Action API 116 can perform a lookup in the template repository to select a template that matches one or more characteristics of the triggering keywords and request. For example, if the request corresponds to a request for a car or a ride to a destination, the data processing system 102 can select a car-sharing service template. The car-sharing service template may include one or more of the following fields: device identifier, pick-up location, destination location, number of passengers, or service type. The Direct Action API 116 can populate fields with values. To populate fields with values, the Direct Action API 116 can ping, poll, or otherwise obtain information from one or more sensors 134 of the computing device 104 or the user interface of the computing device 104. For example, the Direct Action API 116 can use a location sensor, such as a GPS sensor, to detect the source location. The Direct Action API 116 can obtain further information by submitting surveys, prompts, or queries to the end user of the computing device 104. The Direct Action API 116 can submit surveys, prompts, or queries via the interface 110 of the data processing system 102 and the user interface of the computing device 104 (e.g., an audio interface, a voice-based user interface, a display, or a touchscreen). Therefore, the Direct Action API 116 can select a template for the action data structure based on trigger keywords or requests, populate one or more fields in the template with information detected by one or more sensors 134 or obtained via the user interface, and generate, create, or otherwise construct the action data structure to facilitate the execution of operations by the content provider 106.
[0057] The data processing system 102 can select a template from the template data structure based on various factors, including one or more of the following: triggering keywords, requests, the type of content provider 106, the category of content provider 106 (e.g., taxi service, laundry service, flower service, retail service, or food delivery), location, or other sensor information.
[0058] To select templates based on trigger keywords, data processing system 102 (e.g., via the Direct Action API 116) can use trigger keywords to perform lookups or other queries against the template database to identify template data structures mapped to or otherwise corresponding to the trigger keywords. For example, each template in the template database may be associated with one or more trigger keywords to indicate that the template is configured to generate an action data structure in response to trigger keywords that data processing system 102 can process to establish a communication session between data processing system 102 and content provider 106.
[0059] To construct or generate action data structures, the data processing system 102 can identify one or more fields in the selected template to populate with values. These fields can be numeric values, strings, Unicode values, Boolean logic, binary values, hexadecimal values, identifiers, location coordinates, geographic regions, timestamps, or other values. These fields or data structures themselves can be encrypted or masked to maintain data security.
[0060] Once the fields in the template are determined, the data processing system 102 can identify the values of these fields in order to populate these fields in the template to create an action data structure. The data processing system 102 can obtain, retrieve, determine, or otherwise identify the values of these fields by performing lookup or other query operations on the data repository 124.
[0061] In some cases, data processing system 102 may determine that information or values for these fields are missing from data repository 124. Data processing system 102 may determine that information or values stored in data repository 124 are outdated, obsolete, or otherwise unsuitable for the purpose of constructing action data structures in response to triggering keywords and requests identified by NLP component 112 (e.g., the location of client computing device 104 may be an old location instead of the current location; the account may have expired; the destination restaurant may have moved to a new location; physical activity information; or transportation patterns).
[0062] If data processing system 102 determines that it cannot currently access the value or information of a field of the template in its memory, it can retrieve the value or information. Data processing system 102 can retrieve or obtain the information by querying or polling one or more available sensors of client computing device 104, prompting the end user of client computing device 104 with the information, or accessing online web-based resources using the HTTP protocol. For example, data processing system 102 may determine that it does not have the current location of client computing device 104, which may be a required field of the template. Data processing system 102 can query location information from client computing device 104. Data processing system 102 can request client computing device 104 to provide location information using one or more location sensors 134, such as GPS sensors, WiFi triangulation, cellular tower triangulation, Bluetooth beacons, IP addresses, or other location sensing technologies.
[0063] In some cases, data processing system 102 can identify a remote content provider 106 based on trigger keywords or requests, thereby establishing a session. To identify content provider 106 based on trigger keywords, data processing system 102 can perform a lookup in data repository 124 to identify content provider 106 mapped to the trigger keywords. For example, if the trigger keywords include "ride" or "to go to," data processing system 102 (e.g., via Direct Action API 116) can identify content provider 106 (or its network address) as corresponding to taxi service company A. Data processing system 102 can select a template from a template database based on the identified content provider 106. Data processing system 102 can also identify content provider 106 by guiding an internet-based search.
[0064] Data processing system 102 may include, execute, access, or otherwise communicate with session controller component 114 to establish a communication session between computing device 104 and data processing system 102. A communication session may also refer to one or more data transfers between data processing system 102 and content provider 106. The communication session between computing device 104 and data processing system 102 may include the transmission of input audio signals detected by sensor 134 of computing device 104, and the transmission of output signals transmitted from data processing system 102 to computing device 104. Data processing system 102 (e.g., via session controller component 114) may establish a communication session in response to receiving an input audio signal. Data processing system 102 may set a duration for the communication session. Data processing system 102 may set a timer or counter for the duration set for the communication session. In response to the expiration of the timer, data processing system 102 may terminate the communication session. The communication session between data processing system 102 and content provider 106 may include the transmission of digital components from content provider 106 to data processing system 102. The communication session between the data processing system 102 and the content provider 106 may also include the transmission of data to the data array of the content provider 106. The communication session may refer to a network-based communication session in which data (e.g., digital components, authentication information, certificates, etc.) is transmitted between the data processing system 102 and the content provider 106, and between the data processing system 102 and the computing device 104.
[0065] The data processing system 102 may include, execute, or communicate with the audio signal generator component 122 to generate an output signal. The output signal may include one or more components. The output signal may include content identified in digital components received from the content provider 106.
[0066] The audio signal generator component 122 can generate an output signal, the first portion of which has sound corresponding to a first data structure. For example, the audio signal generator component 122 can generate the first portion of the output signal based on one or more values in a field of an action data structure populated by the Direct Action API 116. In the example of a taxi service, the field values could include, for example, 123Main Street as the pick-up location, 1234Main Street as the destination location, 2 passengers, and an economy service level.
[0067] Data processing system 102 (e.g., via interface 110 and network 105) can transmit data packets including output signals generated by audio signal generator component 122. The output signals can cause audio driver component 138 of computing device 104, or audio driver component 138 executed by computing device 104, to drive speakers (e.g., transducers 136) of computing device 104 to generate sound waves corresponding to the output.
[0068] Content provider 106 may provide websites, goods, or services (all generally referred to as digital components) to computing device 104 and data processing system 102. Services and goods may be physically provided (e.g., clothing, car services, and other consumables) and associated with digital components. For example, a digital component for car services may be a website through which users dispatch car services. Digital components associated with services and goods may be digital components used for purchasing, initiating, establishing, or other transactions related to goods and services.
[0069] Content provider 106 may include one or more keywords in the digital component. Keywords may be in meta tags, header strings, the body of the digital component, and links. Upon receiving the digital component, navigation component 118 may analyze the keywords to categorize the digital component (or the content provider 106 associated with the digital component) into different categories. For example, the digital component may be categorized into categories such as news, retail, etc., which identify general topics of the digital component. Navigation component 118 may select an interaction model from interaction model 130 based at least in part on the category of the digital component.
[0070] Digital components can be rendered via a display device of computing device 104 or on data processing system 102. Rendering may include displaying content items on a display device. In some embodiments, computing device 104 does not include a display device for rendering digital components. For example, computing device 104 may render digital components simply by playing them via speakers of computing device 104. Data processing system 102 may act as an intermediary and enable computing device 104 to interact with digital components in an audio-based manner. Computing device 104 may include applications, scripts, or programs installed on client computing device 104, such as an app for communicating input audio signals to interface 110 of data processing system 102. The application may also drive components of computing device 104 to render output audio signals.
[0071] Figure 2 The diagram illustrates a block diagram of an example method 200 for retrieving and interacting with digital components in a voice-activated, packet-based computer network. Figure 3 The illustration is in Figure 2During the process of method 200 illustrated in the figure, Figure 1 The diagram illustrates a block diagram of an example data flow for the system. Method 200 includes receiving an input audio signal (ACT 202). Method 200 includes parsing the input audio signal to identify a request, content provider, and triggering keywords (ACT 204). Method 200 includes generating an action data structure (ACT 206). Method 200 includes establishing a session with the content provider (ACT 208). Method 200 includes rendering the received digital components (ACT 210). Method 200 includes selecting an interaction model (ACT 212). Method 200 includes generating a data array based on the interaction model (ACT 214). Method 200 includes transmitting the data array to the content provider (ACT 216).
[0072] As explained above, and see also Figure 2-3 Method 200 includes receiving an input audio signal (ACT 202). Data processing system 102 can receive the input audio signal 320 from computing device 104. The input audio signal 320 can be received by data processing system 102 over a network via NLP component 112. NLP can be performed by data processing system 102. Data processing system 102 can receive the input audio signal 320 as data packets including the input audio signal. The input audio signal can be detected by a sensor of computing device 104, such as a microphone.
[0073] Method 200 includes parsing the input audio signal to identify the request, content provider, and triggering keywords (ACT 204). The input audio signal can be parsed by the natural language processing component 112. For example, the audio signal detected by the computing device 104 may include “Okay device, I want a shirt from ABC Co.”. In this input audio signal, the initial triggering keyword may include “okay device”, which may instruct the computing device 104 to send the input audio signal to the data processing system 102. The preprocessor of the computing device 104 may filter out the term “okay device” before sending the remaining audio signal to the data processing system 102. In some cases, the computing device 104 may filter out additional terms or generate keywords to send to the data processing system 102 for further processing.
[0074] Data processing system 102 can identify triggering keywords in the input audio signal 320. Triggering keywords can be phrases, such as "I want a shirt" in the example above. Triggering keywords can indicate the type of service or commodity (e.g., shirt) and the action to be taken. Data processing system 102 can identify requests in the input audio signal. Requests can be determined based on the term "I want." Triggering keywords and requests can be determined using semantic processing techniques or other natural language processing techniques. Data processing system 102 can identify content provider 106 as ABC Co. Data processing system 102 can identify websites, IP addresses, or other network locations associated with content provider 106, ABC Co.
[0075] Method 200 includes generating an action data structure (ACT 206). A direct action application programming interface can generate the action data structure based on trigger keywords. The action data structure can also be generated in response to a request and an identified content provider 106. The action data structure can be generated from or based on a template. The template can be selected based on the trigger keywords and the identified content provider 106. The generated action data structure can include information and data related to performing the action associated with the trigger keywords. For example, for “I want a shirt from ABC Co.,” the template could indicate the required information related to purchasing the shirt, including size, preferred color, preferred style, and preferred price range. Data processing system 102 can populate fields in the action data structure with values retrieved from memory or based on the user's response to an output signal transmitted from data processing system 102 to computing device 104. Data processing system 102 can populate security fields, such as user credentials from a secure wallet that can be stored on data processing system 102 or computing device 104. Data processing system 102 can request permission from the user to access the secure wallet before obtaining information from it.
[0076] Method 200 includes establishing a session with a content provider (ACT 208). Data processing system 102 can establish a communication session 322 with the content provider in response to identifying the content provider 106 in the input audio signal. Communication session 322 can be established to receive digital components from the content provider 106. The session can be established using the Hypertext Transfer Protocol. The session can be established using a request from data processing system 102 to the content provider 106. Request 324 can be for a webpage transmitted in the response 325 to request 324.
[0077] Method 200 includes rendering the received digital component (ACT 210). The received digital component can be rendered by the navigation component 118 of the data processing system 102. Figure 3 The illustration shows a partial rendering of the digital component 300. Continuing the example above, the digital component can respond to the input audio signal "I want a shirt from ABC Co.". The rendered digital component 300 may include an input field 302, a button 304, a menu, an image field 306, an image 308, and text 310 (generally referred to as components or elements of the digital component). Buttons, links, input fields, and radio buttons may generally be referred to as input elements. The digital component can be rendered without a graphical user interface. For example, the digital component 300 may be an HTML document rendered by a headless browser. The headless browser of the navigation component 118 may include a layout engine that can render the code of the digital component 300, such as HTML and JavaScript within the digital component. When the digital component is rendered in a headless form, the navigation component 118 may render the digital component 300 as an image file that can be analyzed by the machine vision component of the navigation component 118.
[0078] Method 200 includes selecting an interaction model (ACT 212). Navigation component 118 can select an interaction model associated with content provider 106. Navigation component 118 can choose between two general types of interaction models. The first model can be a generic model, which may be the same for each content provider 106 associated with a particular category. For example, data processing system 102 may include a generic model for a shopping website; a generic model for an insurance website; a generic model for a hotel booking website; and a generic model for a food delivery website. The second type of model may be specific to content provider 106 (or digital components received from content provider 106).
[0079] In addition, a second model can be used as a specific data model. For example, the model can be specific to ABC Company. Specific or special characteristics such as the placement of links to access products, the placement of specific menus and how to navigate through them, and how specific products are stored and categorized within the website can be information included in the specific model. Navigation component 118 can use this model to interpret digital component 300. Navigation component 118 can generate a specific interaction model after a predetermined number of sessions have been established between data processing system 102 and content provider 106. For example, data processing system 102 can initially use a general model when interacting with a given content provider 106. Data from the interactions can be used to construct the specific interaction model. Once data processing system 102 has started a predetermined number of sessions and added session data to the specific model for content provider 106, data processing system 102 can begin using the specific interaction model for content provider 106. When the number of previously established sessions is less than the predetermined number, data processing system 102 can continue to use the general interaction model when interacting with content provider 106.
[0080] Using the selected model, navigation component 118 can identify the input field 302, button 304, menu, image field 306, image 308, and text 310 of digital component 300 by performing machine vision analysis of the saved image files of digital component 300. Navigation component 118 can also identify components of digital component 300 by parsing its code. For example, navigation component 118 can identify HTML tags within digital component 300. As an example, navigation component 118 can search for HTML tags. <input> or <form>To identify input field 302.
[0081] When navigation component 118 recognizes an image or button, it can perform machine vision analysis on the image or button to determine one or more characteristics of the image or button. These characteristics may include the determination of colors within the image (e.g., the shirt illustrated in image 308 is a red shirt), the identification of objects in the image (e.g., image 308 illustrates a shirt), or text or icons within the image or button (e.g., button 304 includes an arrow indicating "next" or whether button 304 includes the text "next").
[0082] Method 200 includes generating a data array based on an interaction model (ACT 214). The data array can be generated by navigation component 118 using the interaction model based on information identified in digital component 300. The digital array can be generated using information from an action data structure. For example, using the interaction model, navigation component 118 can determine that text 310 states "size" and is associated with input field 302. The action data structure can include an entry for "medium" in the "size" field. Navigation component 118 can include "medium" in the data array and input the data array into input field 302 to indicate that a medium-sized shirt should be selected.
[0083] Method 200 includes transmitting a data array to a content provider (ACT 216). The data array 330 can be input to input field 302. The data array 330 can be transmitted to content provider 106 in response to navigation component 118 selecting another input field (such as button 304). The data array 330 can be transmitted to content provider 106 in response to an HTTP POST or GET method. The data processing system 102 can continue to interact with digital components to perform actions recognized in the input audio signal. For example, in Figure 3 In the example illustrated, the data processing system 102 can repeat the ACT of method 200 to select a shirt, check out or purchase a shirt, and then send a confirmation to the client computing device 102.
[0084] Data processing system 102 can establish a communication session 322 between data processing system 102 and computing device 104. Communication session 322 can be established via a conventional application programming interface. Communication session 322 can be a real-time, round-trip voice or audio-based conversation. Data processing system 102 can establish communication session 322 with computing device 104 to retrieve additional information for action data structures or data arrays. For example, data processing system 102 can transmit an output audio signal 326 using an instruction that causes the transducer of computing device 104 to generate a sound wave of "What is your preferred color?". A user can provide a second input audio signal 328 in response to the output audio signal 326. Natural language processor component 112 can process the second input audio signal 328 to recognize the user's response, which in this example could be "red". Navigation component 118 can generate a second data array 332 based on the interaction model and the response recognized in the second input audio signal 328. The second data array 332 can be transmitted to content provider 106.
[0085] Data processing system 102 can establish a communication session with a second computing device 104, which is associated with a user of the first computing device 104 that originally transmitted the input audio signal. For example, the first computing device 104 could be a voice-based digital auxiliary speaker system, and the second computing device 104 could be the user's smartphone. Data processing system 102 can request additional information or confirmation from the user via the second computing device 104. For example, in... Figure 3 In the example illustrated, data processing system 102 can provide two images of the selected shirt to the user's smartphone and request the user to select one of the two shirts. Before completing the purchase or making a reservation, data processing system 102 can request verb confirmation via a first computing device 104 or via a second computing device 104 (e.g., selection of the "buy" button).
[0086] Figure 4 This is a block diagram of an example computer system 400. The computer system or computing device 400 may include or be used to implement system 100, or its components, such as data processing system 102. Data processing system 102 may include an intelligent personal assistant or a voice-based digital assistant. Computing system 400 includes a bus 405 or other communication components to communicate information, and also includes a processor 410 or processing circuitry coupled to the bus 405 to process the information. Computing system 400 may also include one or more processors 410 or processing circuitry coupled to the bus to process the information. Computing system 400 also includes a main memory 415 coupled to the bus 405 for storing information, such as random access memory (RAM) or other dynamic storage devices, and instructions to be executed by processor 410. Main memory 415 may be or include a data repository 124. Main memory 415 may also be used to store location information, temporary variables, or other intermediate information during the execution of instructions by processor 410. The computing system 400 may further include a read-only memory (ROM) 420 or other static storage device coupled to the bus 405 for storing static information and instructions for the processor 410. Storage device 42, such as a solid-state device, disk, or optical disk, may be coupled to the bus 405 to permanently store information and instructions. Storage device 425 may include or be part of the data repository 124.
[0087] The computing system 400 can be coupled to a display 435, such as a liquid crystal display or an active matrix display, via a bus 405 for displaying information to a user. An input device 430, such as a keyboard including alphanumeric and other keys, can be coupled to the bus 405 to communicate information and command selection to the processor 410. The input device 430 may include a touchscreen display 435. The input device 430 may also include cursor control, such as a mouse, trackball, or arrow keys, for communicating directional information and command selection to the processor 410 and for controlling cursor movement on the display 435. The display 435 may be, for example, a data processing system 102, a client computing system 104, or... Figure 1 It is part of the other components.
[0088] The processes, systems, and methods described herein can be implemented by computing system 400 in response to processor 410 executing an arrangement of instructions contained in main memory 415. Such instructions can be read into main memory 415 from another computer-readable medium, such as storage device 425. Execution of the arrangement of instructions contained in main memory 415 causes computing system 400 to perform the illustrative processes described herein. One or more processors in a multiprocessor arrangement can also be used to execute the instructions contained in main memory 415. Hardwired circuitry can be used in conjunction with or in combination with software instructions in the systems and methods described herein. The systems and methods described herein are not limited to any particular combination of hardware circuitry and software.
[0089] Despite Figure 4 The example computing system described herein includes the subject matter of the operation described herein, which may be implemented in other types of digital electronic circuits, or in computer software, firmware, or hardware, including the structures disclosed herein and their structural equivalents, or in a combination of one or more of the foregoing.
[0090] In situations where systems discussed here collect or use personal information about users, users can be given the opportunity to control the procedures or features that collect personal information (such as information about a user's social networks, social behavior or activities, user preferences, or user location), or to control whether and / or how content is received from content servers or other data processing systems that may be more relevant to the user. Furthermore, specific data can be anonymized in one or more ways before being stored or used, such that personally identifiable information is removed when parameters are generated. For example, a user's identity information can be anonymized so that no personally identifiable information is determined for the user, or a user's geographic location can be generalized to the place where location information is obtained (such as to the city, zip code, or state level), so that the user's specific location is not determined. In this way, the user can control how information about him or her is collected and how the content server uses it.
[0091] The subject matter and operations described herein can be implemented using digital electronic circuits, or computer software, firmware, or hardware, including the structures and their equivalents disclosed herein, or components of one or more of the foregoing. The subject matter described herein can be implemented as one or more computer programs, such as one or more circuits of computer program instructions, encoded on one or more computer storage media for execution by a data processing apparatus or to control the operation of a data processing apparatus. Alternatively or additionally, the program instructions can be encoded on artificially generated propagating signals, such as machine-generated electrical, optical, or electromagnetic signals, generated to encode information for transmission to a suitable receiving device for execution by the data processing apparatus. The computer storage medium can be a computer-readable storage device, a computer-readable storage substrate, a random or serial access memory array or device, or a combination thereof, or included therein. Although the computer storage medium is not a propagating signal, it can be a source or destination of computer program instructions encoded in artificially generated propagating signals. The computer storage medium can also be one or more discrete components or media (e.g., multiple CDs, discs, or other storage devices), or included therein. The operations described in this specification can be implemented as operations performed by a data processing device on data stored on one or more computer-readable storage devices or received from other sources.
[0092] The terms "data processing system," "computing device 104," "component," or "data processing apparatus" encompass various means, devices, and machines for processing data, including, for example, programmable processors, computers, systems-on-a-chip, or multiple means, or combinations thereof. The means may include special-purpose logic circuitry, such as FPGAs (Field-Programmable Gate Arrays) or ASICs (Application-Specific Integrated Circuits). In addition to hardware, the means may also include code that creates an execution environment for the computer program, such as code constituting processor firmware, protocol stacks, database management systems, operating systems, cross-platform runtime environments, virtual machines, or one or more of the above components. The means and execution environment can implement various different computing model infrastructures, such as web services, distributed computing, and grid computing infrastructures. For example, Direct Action API 116, NLP component 112, and other data processing system 102 components may include or share one or more data processing means, systems, computing devices, or processors.
[0093] Computer programs (also referred to as programs, software, software applications, apps, scripts, or code) can be written in any form of programming language, including compiled or interpreted languages, declarative or procedural languages, and can be deployed in any form, including as a standalone program or as a module, component, subroutine, object, or other unit suitable for use in a computing environment. A computer program may correspond to a file in a file system. A computer program may be stored as part of a file that holds other programs or data (e.g., one or more scripts stored in a markup language document), as a single file dedicated to said program, or as multiple coordinating files (e.g., files storing one or more modules, subroutines, or portions of code). A computer program can be deployed to execute on a single computer, at a single site, or across multiple computers distributed across multiple sites and interconnected by a communication network.
[0094] The processes and logic flows described in this specification can be executed by one or more programmable processors (e.g., components of data processing system 102) that execute one or more computer programs to perform actions by manipulating input data and generating output. The processes and logic flows can also be executed by dedicated logic circuitry, and the apparatus can be implemented as dedicated logic circuitry, such as FPGA (Field-Programmable Gate Array) or ASIC (Application-Specific Integrated Circuit). Suitable devices for storing computer program instructions and data include all forms of non-volatile memory, media, and memory devices, including, for example, semiconductor memory devices such as EPROM, EEPROM, and flash memory devices; magnetic disks, such as internal hard disks or removable disks; magneto-optical disks; and CD-ROMs and DVD-ROMs. Processors and memory can be supplemented by dedicated logic circuitry or incorporated into dedicated logic circuitry.
[0095] The subject matter described herein can be implemented using a computing system comprising back-end components such as data servers, middleware components such as application servers, or front-end components such as client computers having graphical user interfaces and web browsers that allow users to interact with computer system 100 or other components described herein, or combinations of one or more such back-end, middleware, or front-end components. The components of the system can be interconnected by digital data communication of any form or medium, such as communication networks. Examples of communication networks include local area networks ("LANs"), wide area networks ("WANs"), interconnected networks (e.g., the Internet), and peer-to-peer networks (e.g., self-organizing peer-to-peer networks).
[0096] A computing system such as system 100 may include clients and servers. Clients and servers are typically geographically separated and interact via a communication network (e.g., network 105). The client-server relationship is generated by computer programs running on respective computers and having a client-server relationship with each other. The server may transmit data (e.g., data packets representing content items) to computing device 104 (e.g., to display data to a user interacting with computing device 104 or to receive user input from that user). Data generated on computing device 104 (e.g., the result of user interaction) may be received from computing device 104 at the server (e.g., by data processing system 102 from computing device 104 or content provider 106).
[0097] Although the operations are depicted in the accompanying drawings in a specific order, such operations do not need to be performed in the specific order shown or in a sequential order, and it is not necessary to perform all the illustrated operations. The actions described herein can be performed in different orders.
[0098] The separation of various system components is not required in all implementations, and the program components may be included in a single hardware or software product. For example, NLP component 112 may be a single component, an app, a program, a logic device with one or more processing circuits, or part of one or more servers of data processing system 102.
[0099] Some illustrative embodiments have now been described, and it is clear that the foregoing is illustrative and not limiting, presented by way of example. In particular, although many of the examples presented herein involve specific combinations of method actions or system elements, these actions and elements can be combined in other ways to accomplish the same objective. The actions, elements, and features discussed in connection with one embodiment are not intended to exclude similar roles in other embodiments or ways of implementation.
[0100] The wording and terminology used herein are for descriptive purposes and should not be considered limiting. The use of "comprising," "having," "including," "involving," "characterized by," "characterized in," and variations thereof means to include the items listed herein, their equivalents, and additional items, as well as alternative implementations that uniquely constitute the items listed herein. In one implementation, the system and method described herein consist of one or more of each combination, or all of the said elements, actions, or components.
[0101] Any reference to an embodiment, element, or action of a system or method herein expressed in the singular may also include embodiments that include multiple such elements, while any reference to a plural of any embodiment, element, or action herein may also include embodiments that include only a single element. References in singular or plural form are not intended to limit the currently disclosed system or method, its components, actions, or elements to a single or multiple configuration. A reference to any action or element based on any information, action, or element may include an action or element that is at least partially based on an embodiment of that information, action, or element.
[0102] Any implementation disclosed herein may be combined with any other implementation or embodiment, and references to "implementation," "some implementations," "one implementation," etc., are not mutually exclusive but are intended to indicate that a particular feature, structure, or characteristic described in connection with an implementation may be included in at least one implementation or embodiment. Such terms used herein do not necessarily refer to the same implementation. Any implementation may be combined with any other implementation in any manner, inclusively or exclusively, in the same way as the aspects and implementations disclosed herein.
[0103] A reference to "or" can be interpreted as inclusive, such that any term described using "or" can refer to a single term, more than one term, or any of all the stated terms. For example, a reference to "at least one of 'A' or 'B'" can include only 'A', only 'B', or both 'A' and 'B'. Such a reference used in conjunction with "includes" or other open-ended terms can include additional terms.
[0104] Although reference numerals follow technical features in the drawings, detailed descriptions, or any claims, these numerals are included to enhance the comprehensibility of the drawings, detailed descriptions, and claims. Therefore, the presence or absence of reference numerals has no limiting effect on the scope of any claim element.
[0105] The systems and methods described herein may be embodied in other specific forms without departing from their characteristics. The foregoing embodiments are illustrative and not limiting of the systems and methods described herein. The scope of the systems and methods described herein is therefore indicated by the appended claims, rather than the foregoing description, and includes changes falling within the meaning and equivalence of the claims.< / form>
Claims
1. A system for retrieving digital components in a voice-activated data packet-based computer network environment, comprising: A data processing system having one or more processors; A natural language processor component, executed by the data processing system, parses input audio signals acquired via sensors of a first client device to identify a request and a content provider that will fulfill the request, wherein the content provider has not previously created an application programming interface (API) for integrating network resources with the data processing system; and A navigation component executed by the data processing system, the navigation component being used for: The digital component of the content provider is identified based on the request identified from the input audio signal, the digital component having one or more input elements of a graphical user interface; Based on the interpretation and navigation of website features inferred from identified digital components, a specific interaction model is generated for the content provider; Based on the specific interaction model, one or more input elements of the digital component are identified to process the request; Based on the specific interaction model, a data array is generated to include information from at least one of the one or more input elements of the digital component; and The data array is provided to the content provider to satisfy the request to identify from the input audio signal.
2. The system according to claim 1, wherein, The navigation component is further configured to provide an output audio signal to the first client device to retrieve additional information to satisfy the request identified from the input audio signal, and the system further includes: A dialog application programming interface, executed by the data processing system, to receive, via a communication session established with the first client device, a second input audio signal acquired via the sensor of the first client device after the provision of the output audio signal; and The natural language processor component is further configured to parse the second input audio signal to identify a response; and The navigation component is further configured to generate a second data array based on the specific interaction model to include the response to at least one of the one or more input elements of the digital component.
3. The system according to claim 1, wherein, The navigation component is further used for: The graphical user interface uses one or more input elements to render an image corresponding to the digital component on a second client device associated with the first client device; Receive interaction with at least one of the input elements of the digital component of the graphical user interface via the second client device; and A second data array is generated based on the specific interaction model using data corresponding to the interaction with the digital components.
4. The system according to claim 1, wherein, The navigation component is further configured to use training data to build a plurality of interaction models to be selected from, the plurality of interaction models including: The first interaction model defined for the action category. A second interaction model defined for the content providers among multiple content providers, and A third interaction model is defined for a corresponding digital component among multiple digital components.
5. The system according to claim 1, wherein, The navigation component is further used for: Identify the number of previous sessions between the client device and the content provider or the digital component; and An interaction model for selecting the digital component from multiple interaction models based on the number of previous sessions with the content provider, the multiple interaction models including a first interaction model to be selected in response to determining that the number of previous sessions is less than or equal to a threshold number, and a specific interaction model to be selected in response to determining that the number of previous sessions is greater than the threshold number.
6. The system according to claim 1, wherein, The navigation component is further used for: Parsing the content of the digital component identified by the content provider to identify the action category associated with the digital component; and The specific interaction model of the digital component is selected from multiple interaction models based on the action category.
7. The system according to claim 1, wherein, The navigation component is further used for: The request, identified from the input audio signal, is used to identify the digital component, which does not allow voice-based interaction with any of the one or more input elements. and The image corresponding to the digital component is rendered without a user interface using one or more input elements of the graphical user interface.
8. The system according to claim 1, wherein, The navigation component is further used for: In response to the identification of the content provider from the input audio signal, a communication session with the content provider is established; and Using the request identified from the input audio signal, the digital component, including non-audio elements, is received from the content provider via the communication session.
9. The system according to claim 1, wherein, The navigation component is further configured to parse the script corresponding to the digital component to identify one or more input elements of the graphical user interface of the digital component.
10. The system according to claim 1, wherein, The navigation component is further configured to use machine vision analysis to identify one or more input elements of the graphical user interface from the digital component.
11. The system according to claim 1, further comprising: A direct action application programming interface, which is executed by the data processing system to generate an action data structure based on the parsing of the input audio signal; and The navigation component is further configured to generate the data array to include an action data structure, which is provided to the content provider to satisfy the request.
12. The system according to claim 1, wherein, The natural language processor component is further used for: The input audio signal is parsed to identify the triggering keywords that define the request, and The content provider to be communicated is identified based on at least one of the request or the triggering keyword.
13. A method for retrieving digital components in a computer network environment based on voice-activated data packets, comprising: A data processing system having one or more processors parses input audio signals acquired via sensors of a first client device to identify a request and a content provider that will fulfill the request, wherein the content provider has not previously created an application programming interface (API) for integrating network resources with the data processing system; and The data processing system identifies a digital component of the content provider based on the request identified from the input audio signal, the digital component having one or more input elements of a graphical user interface; The data processing system generates a specific interaction model for the content provider based on the interpretation and navigation of website features inferred from identified digital components; The data processing system identifies one or more input elements of the digital component based on the specific interaction model to process the request; The data processing system generates a data array based on the specific interaction model to include information from at least one of the one or more input elements of the digital component; and The data processing system provides the data array to the content provider to satisfy the request identified from the input audio signal.
14. The method of claim 13, further comprising: The data processing system provides an output audio signal to the first client device to retrieve additional information to satisfy the request identified from the input audio signal; The data processing system receives a second input audio signal acquired by the sensor of the first client device after the provision of the output audio signal via a communication session established with the first client device; The data processing system parses the second input audio signal to identify the response; and The data processing system generates a second data array based on the specific interaction model to include the response to at least one of the one or more input elements of the digital component.
15. The method of claim 13, further comprising: The graphical user interface uses one or more input elements to render an image corresponding to the digital component on a second client device associated with the first client device; Receive interaction with at least one of the input elements of the digital component of the graphical user interface via the second client device; and The data processing system generates a second data array based on the specific interaction model using data corresponding to the interaction with the digital component.
16. The method of claim 13, further comprising the data processing system using training data to establish a plurality of interaction models to be selected therefrom, the plurality of interaction models comprising: The first interaction model defined for the action category. A second interaction model defined for the content providers among multiple content providers, and A third interaction model is defined for a corresponding digital component among multiple digital components.
17. The method of claim 13, further comprising: The data processing system identifies the number of previous sessions between the client device and the content provider or the digital component. and The data processing system selects an interaction model for the digital component from a plurality of interaction models based on the number of previous sessions with the content provider. The plurality of interaction models include a first interaction model to be selected in response to determining that the number of previous sessions is less than or equal to a threshold number, and a specific interaction model to be selected in response to determining that the number of previous sessions is greater than the threshold number.
18. The method of claim 13, further comprising: The data processing system parses the content of the digital component into the content identified by the content provider to identify the action category associated with the digital component; and The data processing system selects the specific interaction model of the digital component from multiple interaction models based on the action category.
19. The method of claim 13, further comprising: The data processing system establishes a communication session with the content provider in response to the identification of the content provider from the input audio signal; and The data processing system receives the digital component, including non-audio elements, from the content provider via the communication session using the request identified from the input audio signal.
20. The method of claim 13, further comprising: The data processing system generates the motion data structure based on the analysis of the input audio signal; and The data processing system generates the data array to include action data structures, which are then provided to the content provider to satisfy the request.
Citation Information
Patent Citations
Speaker recognition
US20170092278A1