Real-time micro-profiling using dynamic tree structure
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-07-06
- Publication Date
- 2026-08-11
AI Technical Summary
然而,由于计算设备的输入或输出接口中有限的可用性或功能性,以高效且没有经由可用接口的过多输入/输出过程的方式提供对具有有限信息的查询的有用响应可能是具有挑战性的
Smart Images

Figure CN115867904B_ABST
Abstract
Description
Background Technology
[0001] Computing devices can receive queries from users, process queries, and provide responses to queries. However, due to the limited availability or functionality of the input or output interfaces of computing devices, providing a useful response to queries with limited information in an efficient manner without excessive input / output processes through the available interfaces can be challenging. Summary of the Invention
[0002] This disclosure generally pertains to dynamic tree structures. The technical solution of this invention can use dynamic tree structures to generate microprofiles in real time. For example, in a voice-based computing environment, a computing device executing a digital assistant can receive user input in the form of voice or audio input. The digital assistant can parse the voice input to identify the query. The digital assistant can attempt to generate a response to the query. However, if the query includes limited or insufficient information, or if the user does not have access to enough information, it can be challenging for the digital assistant to generate an accurate or useful response in an efficient manner. For example, the digital assistant may provide a response that may be useless or incorrect. In another example, the user may enter a large or excessive number of voice queries to identify the desired result. In yet another example, the digital assistant may output a list of candidate responses, which may be cumbersome or excessive for the user to listen to or process. Providing incorrect or faulty results or responses to queries, or generating a large list of candidate responses, can be processor-intensive, computationally intensive, utilize additional network bandwidth, or prolong the duration of communication, while introducing latency or delay into the amount of time spent generating an accurate, desired result for the query.
[0003] This technical solution provides a system and method for real-time microprofile generation using a dynamic tree structure. It can receive general voice queries from users and generate pivot points. It can leverage historical search queries executed by multiple computing devices to identify disjoint sets of results. It can determine the similarity distances between each item in the disjoint sets and then select two or more items from the disjoint sets that have the largest distance between them. It can then generate an output audio prompt that asks the user to choose between two items at the pivot point, or ask something else. It can continue generating pivot points in this manner, progressing through child nodes in the tree structure until the final node is identified.
[0004] To reduce the number of input / output requests, this technical solution can generate checkpoints during the process to ask the user if they are satisfied with the current node, or to suggest a final node for the user to choose from. In some cases, checkpoints may need to attempt to decompose the query without generating additional pivot points and child nodes. Using this technical solution, the digital assistant can dynamically or automatically regroup the tree structure and generate one or more child nodes. This technical solution can then generate a micro-profile based on the tree structure that leads to the selection of the final node. This technical solution can associate the micro-profile with the current context of the query (such as time of day, day of the week, location, or other contextual information associated with the electronic identifier that made the request). In response to a subsequent new voice query, the digital assistant of this technical solution can determine whether the context matches the context of the previously generated micro-profile, and if so, determine to load the micro-profile to leverage the previously dynamically generated tree structure and more efficiently select a response to the query.
[0005] At least one aspect relates to a system for generating dynamic tree structures. The system may include a data processing system comprising a memory and one or more processors. The data processing system may receive a first voice query detected via a microphone associated with an electronic account identifier. The data processing system may generate a first pivot point in a tree structure of the first voice query, including multiple child nodes, from a historical search performed by multiple computing devices related to the first voice query. The data processing system may output an audio prompt requesting selection of one of the multiple child nodes. In response to the audio prompt, the data processing system may receive voice input including selection of the first child node among the multiple child nodes. The data processing system may generate a second pivot point in a tree structure, including multiple grandchild nodes, from a historical search associated with the first child node. The data processing system may determine generation checkpoints based on a resource reduction strategy to reduce the generation of additional child nodes. The data processing system may construct a micro-profile of the electronic account identifier using a tree structure based on the response to the checkpoint.
[0006] At least one aspect relates to a method for generating a dynamic tree structure. The method is performed by a data processing system including memory and one or more processors. The method may include the data processing system receiving a first voice query detected via a microphone associated with an electronic account identifier. The method may include the data processing system generating a first pivot point in a tree structure of the first voice query, including multiple child nodes, from a historical search performed by multiple computing devices related to the first voice query. The method may include the data processing system outputting an audio prompt requesting selection of one of the multiple child nodes. The method may include the data processing system receiving voice input, including selection of a first child node among the multiple child nodes, in response to the audio prompt. The method may include the data processing system generating a second pivot point in the tree structure, including multiple grandchild nodes, from a historical search associated with the first child node. The method may include the data processing system determining generation checkpoints based on a resource reduction strategy to reduce the generation of additional child nodes. The method may include the data processing system constructing a micro-profile of the electronic account identifier using a tree structure based on responses to checkpoints.
[0007] These and other aspects and implementations are discussed in detail below. The foregoing information and the following detailed description include illustrative examples of various aspects and implementations, and provide an overview or framework for understanding the nature and characteristics of the claimed aspects and implementations. The accompanying drawings provide illustrations and further understanding of the various aspects and implementations, and are incorporated in and constitute a part of this specification. Attached Figure Description
[0008] The accompanying drawings are not intended to be drawn to scale. The same reference numerals and names in different drawings indicate the same elements. For clarity, not every component may be labeled in every drawing. In the drawings:
[0009] Figure 1 This is an illustration of an example system for generating dynamic tree structures according to an implementation method;
[0010] Figure 2 This is a diagram illustrating an example query process for generating a dynamic tree structure according to an implementation method;
[0011] Figure 3 This is a diagram illustrating an example query process for generating a dynamic tree structure according to an implementation method;
[0012] Figure 4 This is an illustration of an example method for generating dynamic tree structures according to an implementation method;
[0013] Figure 5 This is a block diagram illustrating the architecture of a computer system whose elements can be used to implement the systems and methods described and illustrated herein, including, for example... Figure 1The system described in Figure 2 and Figure 3 The query process described in the text and Figure 4 The method described in the text. Detailed Implementation
[0014] The following is a more detailed description of various concepts and implementations related to methods, apparatuses, and systems for generating dynamic tree structures. The various concepts introduced above and discussed in more detail below can be implemented in any of a variety of ways.
[0015] This technical solution can use a dynamic tree structure to generate microprofiles in real time. In a voice-based computing environment, the computing device executing a digital assistant can receive user input in the form of voice or audio input. For example, the digital assistant device may not include a display device or monitor, or it may not include keyboard or mouse input. In some cases where the digital assistant device includes both a display device and an input device, the user may prefer to operate the digital assistant in a pure voice mode where the input / output interface is a microphone for voice input and a speaker for audio output. When using a voice-based interface, the digital assistant can receive input audio signals with voice input and recognize the query. The digital assistant can attempt to generate a response to the query. However, if the query includes limited or insufficient information, or if the user does not have access to enough information, it can be challenging for the digital assistant to generate an accurate or useful response in an efficient manner.
[0016] For example, using natural language processing, a digital assistant can receive voice commands and output various information, or order products and services. However, a voice query may have to be specific for the digital assistant to provide the correct information or order the correct product or service. The voice query may be general, or the user may not have sufficient information or ability to generate a specific voice query. For example, a general and non-specific voice query could be "Why did the operating system on my laptop crash?". Based on this general query, the data processing system of this technical solution can generate one or more pivot points in one or more layers of a tree structure from a disjoint result set based on historical search results, and obtain the user's selection of pivot points until the final node is reached. Upon reaching the final node, the data processing system can identify the action and then execute it. For example, the action for the final node could be installing an application on the crashed computing device, reinstalling the application, or upgrading the application version.
[0017] Another example of a general and non-specific voice query could be "Help me find a gift for my mother-in-law." The user may have no ideas and therefore cannot provide additional information. If the voice query is not specific enough, the information, products, or services recognized by the digital assistant may be incorrect. In some cases, the digital assistant may proceed with ordering the wrong product or service, thus performing a wasteful or unwanted transaction that wastes computing resources, network usage, or battery power. In other cases, narrowing down the suggestions may be too cumbersome for the digital assistant, resulting in excessive back-and-forth between the user and the digital assistant. Without a screen or other display device, the digital assistant may not be able to display a list of results for the user to scroll through.
[0018] Therefore, the system and method of this technical solution can receive a general voice query, then generate a pivot point with two suggested results for the voice query, and ask the user to choose one of the suggested results or ask something else. This technical solution can provide a digital assistant configured to quickly and efficiently narrow down the results in a way that minimizes or reduces the number of back-and-forth interactions between the digital assistant and the user, while providing accurate and useful results. For example, in response to the general voice query “Help me find a gift for my mother-in-law,” the digital assistant can generate a pivot point query “Are you interested in jewelry, show tickets, or something else?” The pivot point query can be formed by two child nodes: a first child node “jewelry” and a second child node “show tickets.” The digital assistant can identify the two child nodes using historical search queries or search results aggregated from multiple computing devices. The digital assistant of this technical solution can select two child nodes from multiple result sets that have the greatest similarity between them.
[0019] Digital assistants can build upon past searches from other users to clarify questions and results. Over time, digital assistants can build microprofiles based on different contexts. Context can be topic-based, such as buying shoes, purchasing gifts, or movie recommendations. Context can be based on geolocation, time of year (e.g., approaching holidays). Microprofile context can be more or less granular. For example, "mother-in-law" could be the context of a specific microprofile generated in real time using a dynamic tree structure. Digital assistants can use past searches to build microprofiles, allowing past searches to influence future searches, thus allowing digital assistants to reduce back-and-forth interactions and efficiently and accurately narrow down recommendations to the final node.
[0020] To this end, the digital assistant can receive an initial query from the user. In response to receiving the initial query, the digital assistant can begin generating a tree structure. The initial query can form the top node in the tree structure. The digital assistant can then generate and ask the user clarifying questions to proceed to the child nodes in the tree structure. This tree can be binary (e.g., the digital assistant can ask "yes" or "no" questions) or more general. The digital assistant can then generate pivot points. The digital assistant can generate questions to divide the total number of possible outcomes into disjoint sets of equal or nearly equal size (e.g., with 1%, 2%, 3%, 4%, 5%, 6%, 10%, or other quantities that facilitate the generation of micro-profiles in real time using a dynamic tree structure). In a similarity distance metric, disjoint sets can be far apart from each other. Based on the user's response, the digital assistant can proceed to a child node, which can act as the next pivot point. The digital assistant can repeat this process until it reaches a result or the final node.
[0021] Digital assistants can periodically and dynamically generate checkpoints. A checkpoint can refer to presenting possible final results as options to the user. For example, the final result could be a specific product or item that responds to the user's original query. This allows the user to skip intermediate child nodes and proceed directly to the final node. However, because the digital assistant skips intermediate nodes when suggesting the final result, a final result with a higher likelihood of success might be incorrect. Therefore, digital assistants can use strategies or rules to determine when the tree structure progresses to the optimal or satisfactory node level to suggest a final result with sufficient success likelihood or confidence. Digital assistants can be optimized for different objectives, such as helping users find results as quickly as possible on average, or avoiding the possibility of numerous back-and-forth attempts.
[0022] Using this technical solution, child nodes may not be predetermined. The digital assistant can generate open-ended questions, such as "What color shoes are you looking for?". The spectrum of responses is broad, and the digital assistant can proceed to generate the next pivot point or child node based on the response. If the user provides more information than the digital assistant requests, the digital assistant can be configured to skip a level in the tree structure, thereby improving efficiency and reducing latency in tree structure or micro-profile generation, allowing the user to reach the final result as quickly as possible.
[0023] To generate pivot points, the digital assistant in this technical solution can use relevant searches. The digital assistant can use relevant searches associated with electronic account identifiers or other electronic account identifiers with similar profiles. The digital assistant can divide the results into groups of approximately the same size with minimal or no overlap. Searches can be considered relevant based on various factors such as time, geographic location, and language. The digital assistant can generate child nodes that are equally likely to contain the user's expected final result, even if some child nodes are "larger" in the sense of having more final nodes. These probabilities can be determined based on past queries associated with other electronic account identifiers or using previously constructed micro-profiles. The digital assistant can traverse from the original nodes to the final nodes using various paths, in which case the result sets grouped under a particular node may not be completely disjoint.
[0024] Once a tree structure has been generated for the e-account identifier in response to the original query, the digital assistant can generate micro-profiles based on interactions related to the tree structure. The digital assistant can use micro-profiles to improve the efficiency and accuracy of responses to subsequent queries from the e-account identifier. Which micro-profile the digital assistant uses and modifies can depend on the original user query and the preceding questions. Micro-profiles can cover contexts such as buying shoes, holiday shopping, movie searches, or buying a gift for a specific person (e.g., the mother-in-law).
[0025] These microprofiles differ from general user profiles. For example, very similar queries can have very different results due to small changes in context. There might be specific microprofiles based on the time of day and / or the user's mood. For instance, "I want to watch a movie" could be in different contexts: weekday evening (e.g., action movie) versus Saturday morning (e.g., family movie). Buying a gift would depend on who the user is buying it for. Therefore, a digital assistant can build microprofiles for specific contexts and then load those microprofiles when a context match is found.
[0026] The final result can significantly influence the shaping of microprofiles and affect future outcomes. However, the digital assistant can refine the microprofile using each answer along the way. The digital assistant can create a microprofile locally for each user session. The digital assistant can store the microprofile on the server side for future use, across multiple devices, or erase it at the end of the user session. At the end of microprofile generation, the digital assistant can ask the user if they wish to store the microprofile for future use. If the microprofile has already been stored and the user agrees, a new, enhanced microprofile (from the current search) can overwrite a previously generated microprofile with the same context. If the user indicates they do not wish to save the microprofile, the digital assistant can erase it at the end of the current session.
[0027] By asking questions to guide the search and using micro-profiles, digital assistants can effectively help users find answers to open-ended questions and searches. In some cases, users can terminate the search early or move to a visual surface to view the results corresponding to a node. Digital assistants can integrate this process with third-party applications upon completion of the search to perform actions such as, for example, ordering products or streaming movies.
[0028] Digital assistants can perform actions to make purchases, such as using linked payment options, once they identify the final node and the actions associated with it. Seamless integration allows digital assistants to authenticate the chosen payment method based on a combination of voice signatures or voice authentication and other security credentials based on biometrics or cryptography.
[0029] Figure 1 An example system 100 for dynamic tree structure generation according to an embodiment is illustrated. System 100 may include content selection infrastructure. System 100 may include a data processing system 102. The data processing system 102 may communicate with one or more of client computing devices 140 or supplemental digital content provider devices 130 via a network 105. Network 105 may include computer networks (such as the Internet, local area networks, wide area networks, metropolitan area networks or other regional networks, intranets, satellite networks) and other communication networks (such as voice or data mobile phone networks). Network 105 may be used to access information resources, such as web pages, websites, domain names, or Uniform Resource Locators that can be provided, output, displayed, or shown on client computing devices 140.
[0030] Network 105 may include or constitute a display network, such as a subset of information resources available on the Internet or eligible to include third-party digital components as part of a digital component placement campaign, associated with content placement or search engine results systems. Network 105 may be used by data processing system 102 to access information resources, such as web pages, websites, domain names, or Uniform Resource Locators that can be provided, output, displayed, or shown by client computing device 140. For example, via network 105, a user of client computing device 140 may access information or data provided by supplemental digital content provider device 130.
[0031] Network 105 can be any type or form of network and can include any of the following: point-to-point network, broadcast network, wide area network, local area network, telecommunications network, data communication network, computer network, ATM (Asynchronous Transfer Mode) network, SONET (Synchronous Optical Network) network, SDH (Synchronous Digital Hierarchy) network, wireless network, and wired network. Network 105 may include wireless links, such as infrared channels or satellite bands. The topology of network 105 may include bus, star, or ring network topologies. The network may include mobile phone networks using any one or more protocols for communication between mobile devices, including Advanced Mobile Telephone Protocol (“AMPS”), Time Division Multiple Access (“TDMA”), Code Division Multiple Access (“CDMA”), Global System for Mobile Communications (“GSM”), General Packet Radio Service (“GPRS”), or Universal Mobile Telecommunications System (“UMTS”). Different types of data may be sent via different protocols, or the same type of data may be sent via different protocols.
[0032] Client computing device 140 may include, for example, a laptop, desktop computer, tablet computer, digital assistant device, smartphone, mobile telecommunications device, portable computer, smartwatch, wearable device, headset, speaker, television, smart display, or automotive unit. For example, via network 105, a user of client computing device 140 may access information or data provided by supplemental digital content provider device 130. In some cases, client computing device 140 may or may not include a display; for example, the computing device may include limited types of user interfaces, such as microphones and speakers. In some cases, the primary user interface of client computing device 140 may be a microphone and speaker, or a voice interface. In some cases, client computing device 140 includes a display device coupled to client computing device 140, and the primary user interface of client computing device 140 may utilize the display device.
[0033] Client computing device 140 may include microphone 142. Microphone 142 may include a transducer or other hardware configured to detect sound waves (such as voice input from a user) and convert the sound waves into another format that can be processed by client computing device 140. For example, microphone 142 may detect sound waves and convert them into analog or digital signals. Client computing device 140 may use hardware or software to convert the analog or digital signals into data packets corresponding to voice input or other detected audio input. Client computing device 140 may send the data packets with voice input to data processing system 102 for further processing.
[0034] Client computing device 140 may include speaker 144. Speaker 144 may output audio or sound. Speaker 144 may be driven by an audio driver to generate audio output. Speaker 144 may output speech or other audio generated by data processing system 102 and provided to client computing device 140 for output. For example, a user may converse with digital assistant 108 via microphone 142 and speaker 144 of client computing device 140.
[0035] In some cases, client computing device 140 may include one or more components or functions of data processing system 102, such as NLP 106, digital assistant 108, interface 104, or data store 118. For example, client computing device 140 may include a local digital assistant or digital assistant agent having one or more components or functions of server digital assistant 108 or NLP 106. Client computing device 140 may include a data store for storing microprofiles. Client computing device 140 may include... Figure 5 One or more components or functions of the computing system 500 described herein.
[0036] System 100 may include at least one data processing system 102. Data processing system 102 may include at least one logical device, such as a computing device with a processor, to communicate via network 105, for example, with client computing device 140 or supplemental digital content provider device 130 (or third-party content provider device, content provider equipment). Data processing system 102 may include at least one computing resource, server, processor, or memory. For example, data processing system 102 may include multiple computing resources or servers located in at least one data center. Data processing system 102 may include multiple logically grouped servers and facilitates distributed computing technologies. The logical group of servers may be referred to as a data center, server cluster, or machine cluster. Servers may also be geographically distributed. A data center or machine cluster may be managed as a single entity, or a machine cluster may include multiple machine clusters. Servers within each machine cluster may be heterogeneous—one or more of the servers or machines may operate according to one or more types of operating system platforms.
[0037] Servers in a cluster can be stored alongside associated storage systems in a high-density rack system within an enterprise data center. For example, consolidating servers in this way by positioning them and high-performance storage systems on a local high-performance network can improve system manageability, data security, physical security, and system performance. Centralizing all or some of the data processing system components 102, including servers and storage systems, and coupling them with advanced system management tools, allows for more efficient use of server resources, saving power and processing requirements, and reducing bandwidth usage.
[0038] System 100 may include, access, or otherwise interact with at least one third-party device, such as supplemental digital content provider device 130 or service provider device. Supplemental digital content provider device 130 or other service provider device may include at least one logical device, such as a computing device with a processor, to communicate via network 105, for example, with client computing device 140 or data processing system 102.
[0039] Supplemental digital content provider device 130 may provide audio-based digital components for display as audio output digital components by client computing device 140. These digital components may be referred to as sponsored digital components because they are provided by a third-party sponsor. Digital components may include offers for goods or services, such as voice-based messages stating, “Would you like me to book a taxi for you?”. For example, supplemental digital content provider device 130 may include memory to store a range of audio digital components that may be provided in response to a voice-based query. Supplemental digital content provider device 130 may also provide audio-based digital components (or other digital components) to data processing system 102, where they may be stored in a data repository of data processing system 102. Data processing system 102 may select audio digital components and provide (or instruct supplemental digital content provider device 130 to provide) the audio digital components to client computing device 140. Audio-based digital components may be exclusively audio or may be combined with text, image, or video data.
[0040] Data processing system 102 may include a content placement system having at least one computing resource or server. Data processing system 102 may include at least one interface 104, interfacing with or otherwise communicating with it. Data processing system 102 may include at least one natural language processor 106 (or natural language processor component), interfacing with or otherwise communicating with it. Interface 104 or natural language processor 106 may form or be referred to as server digital assistant 108. Data processing system 102 may include at least one server digital assistant 108 (or server digital assistant component), interfacing with or otherwise communicating with it. Server digital assistant 108 may communicate or interface with one or more voice-based interfaces or various digital assistant devices or surfaces to provide or receive data or perform other functions. Data processing system 102 may include at least one content selector 110 (or content selector component). Data processing system 102 may include at least one tree generator 112. Data processing system 102 may include at least one checkpoint generator 114. Data processing system 102 may include at least one profile manager 116. The data processing system 102 may include at least one data store 118. The data store 118 may store micro-profiles 120 including keywords and contextual information associated with a dynamically generated tree structure. The data store 118 may store historical searches 122 performed by other computing devices 140 during previous time periods or intervals. The data store 118 may include strategies 124 that can be used to generate checkpoints to reduce additional resource consumption and node layer generation.
[0041] Data processing system 102, interface 104, NLP 106, content selector 110, tree generator 112, checkpoint generator 114, and microprofile manager 116 may each include at least one processing unit or other logic device (such as a programmable logic array engine), or a module configured to communicate with a data store or database of data processing system 102. Interface 104, NLP 106, or content selector 110, tree generator 112, checkpoint generator 114, and microprofile manager 116 may be separate components, individual components, or part of data processing system 102. System 100 and its components (such as data processing system 102) may include hardware elements such as one or more processors, logic devices, or circuitry.
[0042] Data processing system 102 can obtain anonymous computer network activity information associated with multiple computing devices 140 (or computing devices or digital assistant devices). Users of client computing devices 140 or mobile computing devices can authorize data processing system 102 to obtain network activity information corresponding to client computing devices 140 or mobile computing devices. For example, data processing system 102 can prompt users of client computing devices 140 to consent to obtaining one or more types of network activity information. The identity of users of client computing devices 140 can remain anonymous, and client computing devices 140 can be associated with unique identifiers (e.g., unique identifiers of users or computing devices provided by users of the data processing system or computing devices). Data processing system 102 can associate each observation with a corresponding unique identifier.
[0043] Data processing system 102 may include an interface 104 (or interface component) designed, configured, constructed, or operated to receive and transmit information using, for example, data packets. Interface 104 may use one or more protocols (such as network protocols) to receive and transmit information. Interface 104 may include a hardware interface, a software interface, a wired interface, or a wireless interface. Interface 104 may facilitate the transformation or formatting of data from one format to another. For example, interface 104 may include an application programming interface (API) that includes definitions for communication between various components, such as software components. Interface 104 may communicate via network 105 with one or more of client computing devices 140 or supplemental digital content provider devices 130.
[0044] The data processing system 102 may interface with applications, scripts, or programs installed on the client computing device 140, such as an interface 104 that communicates input audio signals to the data processing system 102 and drives components of the client computing device 140 to display, present, or otherwise output visual or audio signals. The data processing system 102 may receive data packets or other signals that include or identify audio input signals.
[0045] Data processing system 102 may include a natural language processor (“NLP”) 106. For example, data processing system 102 may execute or run NLP 106 to parse received input audio signals or queries. For example, NLP 106 may provide human-computer interaction. NLP 106 may be configured with techniques for understanding natural language and allowing data processing system 102 to derive meaning from human or natural language input. NLP 106 may include or be configured with machine learning-based techniques, such as statistical machine learning. NLP 106 may utilize decision trees, statistical models, or probabilistic models to parse input audio signals. NLP 106 can perform functions such as named entity recognition (e.g., given a stream of text, determining which items in the text map to proper names (such as people or places) and what type each such name is, such as person, location, or organization), natural language generation (e.g., translating information from computer databases or semantic intents into understandable human language), natural language understanding (e.g., translating text into more formal representations, such as first-order logical structures that computer modules can manipulate), machine translation (e.g., automatically translating text from one human language to another), morpheme segmentation (e.g., separating words into individual morphemes and identifying the category of morphemes, which can be challenging based on the complexity of the morphemes or structures of words in the language under consideration), question answering (e.g., determining the answer to a human language question, which can be specific or open-ended), and semantic processing (e.g., processing that can occur after words are identified and their meanings are encoded in order to associate the identified words with other words that have similar meanings).
[0046] NLP 106 converts an audio input signal into recognized text by comparing it to a stored set of representative audio waveforms and selecting the closest match. The set of audio waveforms can be stored in a data store or other database accessible to data processing system 102. Representative waveforms are generated across a large set of users and can then be augmented with language samples from users. After the audio signal is converted into recognized text, NLP 106 matches the text with words associated with actions that data processing system 102 can serve (e.g., through cross-user training or manual specification). Aspects or functions of NLP 106 can be performed by data processing system 102 or client computing device 140. For example, NLP components can be executed on client computing device 140 to perform aspects of converting the input audio signal into text and sending the text via data packets to data processing system 102 for further natural language processing.
[0047] The audio input signal can be detected by a sensor or transducer (e.g., a microphone) of the client computing device 140. The client computing device 140 can provide the audio input signal to the data processing system 102 (e.g., via network 105) via a transducer, audio driver, or other components, whereby the audio input signal can be received (e.g., via interface 104) and provided to NLP 106 or stored in a data repository.
[0048] Data processing system 102 can receive data packets including input audio signals detected by the microphone of client computing device 140 via interface 104. Data processing system 102 can receive data packets generated based on the input audio signals detected by the microphone. Data packets may or may not be filtered. Data packets may be a digital version of the detected input audio signals. Data packets may include text generated by client computing device 140 based on the detected input audio signals. For example, a local digital assistant of client computing device 140 can process the detected input audio signals and send data packets to server digital assistant 108 based on the processed input audio signals for further processing or action.
[0049] Data processing system 102 may include server digital assistant 108. Server digital assistant 108 and NLP 106 may be a single component, or server digital assistant 108 may include one or more components or functions of NLP 106. Server digital assistant 108 may interface with NLP 106. Data processing system 102 (e.g., server digital assistant 108) may process data packets to perform actions or otherwise respond to voice input. In some cases, data processing system 102 may identify acoustic signatures from input audio signals. Data processing system 102 may identify electronic accounts corresponding to acoustic signatures based on lookups in a data repository (e.g., querying a database). In response to the identification of electronic accounts, data processing system 102 may establish a session and an account for that session. The account may include a profile with one or more policies. Data processing system 102 may parse input audio signals to identify requests and trigger keywords corresponding to those requests.
[0050] NLP 106 can acquire input audio signals. In response to the local digital assistant on client computing device 140 detecting a trigger keyword, NLP 106 of data processing system 102 can receive data packets with voice input or input audio signals. The trigger keyword can be a wake-up signal or a hotword, which instructs client computing device 140 to convert subsequent audio input into text and send the text to data processing system 102 for further processing.
[0051] Upon receiving an input audio signal, NLP 106 can identify at least one query or request, or at least one keyword corresponding to that request. A request may indicate the intent or topic of the input audio signal. A keyword may indicate the type of action that might be taken. For example, NLP 106 can parse the input audio signal to identify at least one request to leave home for dinner and a movie in the evening. Trigger keywords may include at least one word, phrase, root word, or part of a word, or a derived word indicating the action to be taken. For example, the trigger keyword “go” or “to go to” from the input audio signal may indicate that a transmission is required. In this example, the input audio signal (or the identified request) does not directly express the intent to transmit; however, the trigger keyword indicates that the transmission is an auxiliary action to at least one other action indicated by the request. In another example, voice input may include a search query, such as “find jobs near me.” However, if the original query is general or lacks sufficient information to generate accurate results, NLP 106 or the server digital assistant 108 may forward the original query to tree generator 112, checkpoint generator 114, or microprofile manager 116 for further processing.
[0052] NLP 106 can parse input audio signals to identify, determine, retrieve, or otherwise obtain a request and one or more keywords associated with that request. For example, NLP 106 can apply semantic processing techniques to the input audio signal to identify keywords or requests. NLP 106 can apply semantic processing techniques to the input audio signal to identify keywords or phrases including one or more keywords, such as a first keyword and a second keyword. For example, the input audio signal may include the sentence "I want to buy an audiobook." NLP 106 can apply semantic processing techniques or other natural language processing techniques to data groups including sentences to identify the keywords or phrases "want to buy" and "audiobook." NLP 106 can further identify multiple keywords, such as "buy" and "audiobook." For example, NLP 106 can determine that the phrase includes a first keyword and a second keyword.
[0053] NLP 106 can filter input audio signals to identify trigger keywords. For example, a data packet carrying the input audio signal might include "It would be great if I could get someone that could help me go to the airport." In this case, NLP 106 can filter out one or more of the following terms: "it," "would," "be," "great," "if," "I," "could," "get," "someone," "that," "could," or "help." By filtering out these terms, NLP 106 can more accurately and reliably identify trigger keywords, such as "go to the airport," and determine that this is a request for a taxi or ride-sharing service.
[0054] However, in some cases, the original voice query may be general or lack sufficient information. The digital assistant 108 may not be able to determine an accurate result with a high level of confidence. The digital assistant 108 may decide to invoke the tree generator 112 instead of outputting a list of candidate results or engaging in multiple inefficient back-and-forth exchanges with the user. The data processing system 102 may include a tree generator 112 designed, constructed, and operated to dynamically generate tree structures in real time. The tree generator 112 may receive a first voice query detected via a microphone and associated with an electronic account identifier. For example, the tree generator 112 may receive an original voice query or other voice query detected by the microphone 142 of the client computing device 140. The client computing device 140 may be associated with or linked to an electronic account identifier. The electronic account identifier may refer to or include a username, identifier, or other unique identifier of a user on a network or platform associated with the data processing system 102. The user can use the electronic account identifier to log in to an operating system or application on the client computing device 140. The client computing device 140 may provide the data processing system 102 with an indication of the electronic account identifier, either in conjunction with, before, or after the transmission of a voice query. In some cases, the client computing device 140 may use the user's electronic account identifier to establish a communication session with the data processing system 102.
[0055] Tree generator 112 can parse the first speech query or the original speech query. Tree generator 112 can identify queries, keywords, or semantic information associated with queries. In some cases, tree generator 112 can utilize one or more components or functions of NLP 106 to parse or process the query. For example, NLP 106 can parse or process the query to identify one or more keywords or semantic meanings and provide the parsed information to tree generator 112. In some cases, tree generator 112 can utilize or use historical searches 122 stored in data store 118. Tree generator 112 can use historical searches to generate pivot points in the tree structure of the first speech query. Pivot points can indicate multiple child nodes. In this tree structure, the first speech query or the original speech query can form the top node, and the pivot point can provide options or candidate child nodes from which the user can select child nodes at the next level in the tree structure.
[0056] Tree generator 112 can identify related searches. Related searches can be stored in search history 122. Tree generator 112 can use various factors to determine related searches. Based on geographic location, time or date, language, or electronic account identifier, a search can be related to a voice query. If performed through another electronic account identifier with a profile or similar characteristics to the electronic account identifier that provided the original voice query, the search can be related to a voice query. For example, by inputting a voice query into a search engine, tree generator 112 can identify search results. Tree generator 112 can also obtain search results from related search queries, which can be similar to voice queries received from client computing device 140. Tree generator 112 can identify search results that are more likely to be related to electronic account identifiers based on various factors.
[0057] Tree generator 112 can divide the totality of possible results into disjoint sets of equal or substantially equal size (e.g., with 1%, 2%, 3%, 4%, 5%, 6%, 10%, or other quantities that facilitate the generation of microprofiles in real time using a dynamic tree structure). Sets can be disjoint if they do not contain identical results or results with the same concept. Sets can be disjoint if they do not overlap on results or topics. Sets can be substantially disjoint, but may not be completely disjoint. For example, different sets may contain less than 10% overlap, 8% overlap, 6% overlap, 5% overlap, 3% overlap, or other amounts of overlap that facilitate the generation of microprofiles in real time using a dynamic tree structure.
[0058] Tree generator 112 can identify different sets and then determine the similarity between the different sets. Tree generator 112 can divide the results into groups of approximately the same size with minimal or no overlap. Tree generator 112 can generate child nodes that are equally likely to contain the user's expected final result, even if some child nodes are "larger" in the sense of having more final nodes. These probabilities can be determined based on past queries associated with other electronic account identifiers or using previously constructed microprofiles.
[0059] Tree generator 112 can use semantic analysis techniques to determine the similarity between disjoint sets. Tree generator 112 can use similarity metrics to determine or quantify the similarity between disjoint sets. Tree generator 112 can be configured with one or more similarity functions. Tree generator 112 can determine a distance metric or its reciprocal, which takes a large value for similar objects and a zero or negative value for dissimilar objects. Tree generator 112 can apply cosine similarity to real-valued vectors. Tree generator 112 can score the similarity of disjoint sets in a vector space model. Tree generator 112 can use clustering functions to determine the similarity between disjoint sets. Tree generator 112 can determine the semantic similarity between terms in disjoint sets to determine the distance between disjoint sets based on their semantic content.
[0060] In determining the distance between each disjoint set identified using historical searches related to the voice query, tree generator 112 can identify two disjoint sets that have the maximum distance relative to other disjoint sets. For example, tree generator 112 can identify two disjoint sets that are at opposite ends of the spectrum or have the fewest commonalities.
[0061] Using these two disjoint sets, tree generator 112 can generate pivot points. A pivot point can comprise two nodes corresponding to the two disjoint sets with the largest possible distance between them. Child nodes can be characterized by terms or keywords. Tree generator 112 can use these two child nodes to generate clarification questions to ask the user. The digital assistant can then generate and ask the user clarification questions to proceed to the child nodes in the tree structure. The tree can be binary (e.g., the digital assistant can present it as a "yes" or "no" question) or more general. The digital assistant can then generate pivot points.
[0062] like Figure 2As shown, data processing system 102 can receive the query "Find a gift for my mother-in-law". Data processing system 102 can identify two disjoint sets separated by the largest distance as: jewelry and tickets. Data processing system 102 can form a first child node for jewelry and a second child node for tickets. Data processing system 102 can generate a pivot point with the first and second child nodes. Data processing system 102 can add a third option, "something else", to provide options along a third child node not provided in the initial list. In some cases, data processing system 102 can provide three child node options. For example, if there are three disjoint sets that are equally separated by the same distance (or substantially the same distance, such as within 1%, 2%, 3%, 5%, 6%, etc.), data processing system 102 can ask the user to select one of the three child nodes. Data processing system 102 can still include a fourth option, "something else".
[0063] If the user responds by selecting "something else," the data processing system 102 can return to the relevant search to identify additional disjoint sets and the distances between them, thereby selecting the next set of child nodes for suggestion. The data processing system 102 can repeat this process until the user selects a child node.
[0064] Therefore, the data processing system 102 can generate a first pivot point with a first child node and a second child node, wherein the first child node and the second child node are separated by a distance greater than a threshold. The threshold can be a relative threshold or a dynamic threshold. The threshold can refer to a ranking, such as a maximum distance. The threshold can be an absolute distance threshold or a normalized distance threshold. For example, the data processing system 102 can rank the distances between each node and then select the highest-ranked distance, or the top two highest-ranked distances, or the top three highest-ranked distances. Therefore, the data processing system 102 can use the threshold to determine how many child nodes to select to form the pivot point based on the distance. The data processing system 102 can generate a first pivot point with a first child node and a second child node, wherein the distance between the first child node and the second child node is the maximum distance between a plurality of candidate nodes identified in response to a first voice query, and the first child node and the second child node are formed from a disjoint set of sizes within a threshold size (e.g., within 1%, 2%, 3%, 5%, 10%, or other amounts of size).
[0065] Tree generator 112 (e.g., via NLP 106 or digital assistant 108) can generate audio cues or provide instructions for generating audio cues. Audio cues can be formed from selected child nodes. Templates can be used to construct output cues, including placeholders for child nodes and adding "something else" options. For example, an output cue could be "jewelry, tickets, or something else." Data processing system 102 can provide audio cues to client computing device 140, causing client computing device 140 to output audio cues via speaker 144.
[0066] Data processing system 102 receives voice input, including a selection of one of the child nodes, in response to an audio prompt. The user can select either a first or a second child node. In some cases, the user may not select one of the suggested child nodes but instead indicate "something else," in which case tree generator 112 can identify one or more other child nodes to suggest generating a new pivot point. The user can select a child node using voice input detected by microphone 142. The voice input can be, for example, the name of the child node (e.g., "jewelry"), or it can indicate a number or order, such as "first" or "second." In some cases, the user can ask data processing system 102 to select either a first or second child node by saying "either is fine," and data processing system 102 can select either the first or the second (e.g., defaulting to the first, or using a random number generator to select either the first or the second child node).
[0067] Based on the selected child nodes, the data processing system 102 can generate a second pivot point. The data processing system can repeat this process to generate second-level child nodes, which can be referred to as grandchild nodes. The data processing system 102 can generate a second pivot point that includes grandchild nodes from the historical search associated with the first child nodes. Figure 2 If the choice is the second node "Ticket", then the example grandchild node could be "Opera", "Musical", "Ballet", or "Sports".
[0068] In some cases, tree generator 112 may determine to select two or more grandchild nodes to generate the second pivot point. Tree generator 112 may determine that providing additional candidate grandchild nodes may be more efficient and accurate because the tree structure has progressed to a more granular or deeper layer in the tree (e.g., a second layer). In some cases, data processing system 102 may determine that there are a larger number of disjoint sets of equal size and equally separated by a distance, thus justifying providing additional grandchild nodes in the second audio cue.
[0069] The data processing system 102 can construct a second output audio cue using grandchild nodes. The second audio cue may include two or more grandchild nodes. Depending on the number of grandchild nodes suggested in the second audio cue, the second audio cue may include an "something else" option, or it may not include an "something else" option. For example, if there are four grandchild nodes in the second audio cue, the data processing system 102 may discard the "something else" option. However, if there are two grandchild options, the data processing system 102 may determine to include the "something else" option.
[0070] In some cases, the data processing system 102 can output a second audio prompt to request the selection of one of multiple grandchild nodes, and based on the response to the second audio prompt, determine to skip a layer in the tree structure to generate a third pivot point with multiple great-great-grandchild nodes. For example, the response to the second audio prompt could be "Lakers sports tickets," in which case the data processing system 102 can skip layers (e.g., Figure 2 The tree structure is generated by skipping layers (as depicted in layer 224) and proceeding directly to the great-great-grandchild layer 226 with Lakers tickets or the final node. When skipping layers, intermediate layers can be omitted or not generated, making the tree structure more efficient and dynamic.
[0071] Data processing system 102 (e.g., tree generator 112) can continue generating pivot points with nodes and receiving selections of nodes until the final node is identified. The user can provide responses to audio prompts, and data processing system 102 can continue to narrow down the options for information, products, or services of interest to the user. However, to reduce computational resource utilization and the duration of communication sessions (e.g., to reduce network bandwidth and battery consumption), data processing system 102 may include a checkpoint generator 114 designed, constructed, and operated to generate checkpoints. Data processing system 102 can invoke checkpoint generator 114 based on how far tree generator 112 has progressed in the tree.
[0072] Data processing system 102 can determine the invocation of checkpoint generator 114 based on strategy 124, such as a resource reduction strategy. The resource reduction strategy can be configured with thresholds, such as the number of layers in the generated tree, the duration of the communication session from the timestamp associated with the reception of the original voice query, or a metric or measure associated with child nodes. For example, if the distance between nodes resulting from the formation is below a threshold, indicating a high degree of similarity between nodes, data processing system 102 can determine to generate a checkpoint. By setting the threshold to fewer layers in the tree, data processing system 102 can reduce the amount of bandwidth consumed by repeated round-trip requests for information from the user, and reduce processor and memory utilization by reaching checkpoints and final nodes at earlier layers in the tree.
[0073] Checkpoints can include or refer to selecting a final or narrower node. A checkpoint can refer to asking the user if they want to continue along the tree. For example, checkpoint generator 114 can determine whether to ask the user if they want to continue along the tree or get a final suggestion based on how many levels the user has progressed in the tree. The user can instruct to continue along the tree, in which case data processing system 102 can generate another pivot point with a node. However, if the user instructs for a final suggestion, or checkpoint generator 114 determines to generate a final suggestion, tree generator 112 can determine to skip one or more levels in the tree and generate a final node with a narrow suggestion, such as "Lakers Ticket". Data processing system 102 can determine the generation of checkpoints based on resource reduction strategies to reduce the generation of additional child nodes associated with generating more pivot points, outputting audio prompts, and receiving responses. By generating checkpoints after a certain number of levels in the tree (e.g., 3, 4, 5, etc.), data processing system 102 can reduce the amount of bandwidth consumed by repeated round-trip requests for information from the user, and reduce processor and memory utilization by reaching checkpoints and final nodes at earlier levels in the tree.
[0074] Tree generator 112 can receive a response to a checkpoint (which may be to continue along the tree) or an indication to select a suggested final node. If tree generator 112 receives a selection for a final node, it can identify the final node based on the response to the checkpoint. Data processing system 102 can construct an action based on the final node. Data processing system 102 can execute the action on a remote device configured to perform the action indicated in the final node.
[0075] For example, tree generator 112 can forward the final node to server digital assistant 108. Server digital assistant 108 can receive the final node and process it to identify actions. Server digital assistant 108 can continue to generate action data structures based on the actions in the final node. Server digital assistant 108 can be configured to perform actions, such as accessing information, making electronic transactions, purchasing goods, or ordering services. Server digital assistant 108 can access one or more accounts associated with or linked to electronic account identifiers, which can be used to perform actions such as ordering ride-sharing services, food, or purchasing other goods.
[0076] To improve the efficiency of subsequent tree generation and final node selection, data processing system 102 can generate microprofiles. For example, if tree generator 112 receives a selection for a final node, tree generator 112 can invoke microprofile manager 116 to generate or construct microprofile 120 based on tree completion and successful selection of a final node. Data processing system 102 may include microprofile manager 116, which is designed, constructed, and operated to maintain, construct, generate, or otherwise provide electronic account identifiers. Microprofile manager 116 can construct or generate microprofiles using the tree structure generated by tree generator 112. Microprofile manager 116 can generate microprofiles using information about one or more pivot points and selected child or final nodes. Due to the specific context associated with the microprofile, microprofiles may differ from conventionally used profiles.
[0077] Context can be associated with semantic information, location information, the user's acoustic characteristics, language, time of day, day of the week, season, holiday, or any other contextual information associated with the original query. In some cases, context may include the user's emotion, which can be determined based on acoustic characteristics associated with the user's voice input that issued the voice query. Acoustic characteristics can indicate tone based on amplitude, frequency, intonation, or other voice characteristics. Data processing system 102 can compare the user's previous voice characteristics to determine context. Data processing system 102 can also identify context based on comparisons of the user's emotions when issuing similar voice queries.
[0078] For example, if a query for "buy a gift for my mother-in-law" is issued closer to Mother's Day than another holiday (such as Christmas), the query could have a different final node. In another example, if a query for "I want to see a movie" is issued on Friday night, it could have a different final node compared to Saturday morning. Therefore, a microprofile generator can recognize and store contextual information, along with the resulting tree structure, to generate microprofiles. Microprofiles can be generated or customized for specific e-account identifiers and contexts.
[0079] Microprofile manager 116 can store microprofiles in data repository 118, such as microprofile data structure 120. Microprofile manager 116 can restrict access to microprofiles to the device associated with or linked to the electronic account identifier. In some cases, microprofiles can be stored on local client computing device 140. Microprofiles may not be stored on data processing system 102 or other cloud computing environments or servers. In some cases, microprofiles can be shared between one or more client computing devices 140 owned, used, or authorized by the same user associated with the electronic account identifier. In some cases, microprofiles can be erased and not stored anywhere after the session is completed. For example, microprofiles may be used during a communication session established between client computing device 140 and data processing system 102, and after reaching the final node and completing one or more electronic transactions, microprofile manager 116 can ask the user whether they want to store the microprofile, where they want to store it, or whether they want to erase it. Microprofile manager 116 can store or not store the microprofile based on the user's response. Therefore, the data processing system 102 can erase the microfile in response to the termination of the communication session (e.g., completion of a transaction, session timeout based on an idle timeout value, user exiting the application or computing device, or computing device entering standby mode).
[0080] Data processing system 102 can load a previously generated microprofile in response to receiving a first voice query forming a top node in a tree structure. For example, in response to receiving a first voice query, data processing system 102 can determine whether a microprofile exists for an electronic account identifier. Data processing system 102 can identify the context associated with the current reception of the first voice query and then perform a lookup in microprofile data structure 120 using the electronic account identifier and context information. Context information may include, for example, concepts or topics associated with the first query, such as time of day, time of year, season, holiday, buying shoes, buying gifts, watching movies, geographic location, language, etc. In some cases, context may include the user's mood, which can be determined based on acoustic characteristics associated with the user's voice input that issued the voice query. Acoustic characteristics may indicate mood based on amplitude, frequency, tone, or other voice characteristics. Data processing system 102 can compare the user's previous voice characteristics to determine context. Data processing system 102 can identify context based on a comparison of the user's mood when issuing similar voice queries. If the data processing system 102 identifies a micro-profile with an electronic account identifier that has a similar context, the data processing system 102 can load the micro-profile.
[0081] Tree generator 112 can load microprofiles and use them to generate pivot points. Tree generator 112 can use microprofiles and related search results to generate pivot points. For example, tree generator 112 can more heavily weight certain nodes in the pivot points based on the microprofiles to select those nodes for the pivot points. In another example, tree generator 112 can skip one or more layers in the tree based on previous node selections or interactions stored in the microprofiles. For example, if a user responds to the query "Show me a movie" and selects to watch a movie of the same genre on Friday night, data processing system 102 can jump to movies in that genre and provide options within that genre, thereby improving efficiency while reducing the number of back-and-forth interactions with the user.
[0082] Therefore, in response to receiving the first query, the data processing system 102 can determine whether a microprofile exists for the user, and whether a microprofile with a similar context exists. If a matching microprofile exists, the data processing system 102 can load the microprofile. Otherwise, the data processing system 102 can continue generating a new dynamic tree structure.
[0083] For example, data processing system 102 may receive a third voice query and determine the current context associated with the electronic account identifier. Data processing system 102 may determine that the current context matches the current context. In response to determining that the current context matches the current context, data processing system 102 may load a microprofile to generate a second audio prompt in response to the third voice query.
[0084] In some cases, data processing system 102 can determine when to update the microprofile. Data processing system 102 can load the microprofile, but the user can select different child nodes or different final nodes. Data processing system 102 can update the microprofile based on different interactions in a new instance, thereby adjusting or continuously improving the microprofile for a specific context.
[0085] Upon identifying the final node, the server digital assistant 108 can perform an action associated with that final node. For example, using natural language processing, the digital assistant can receive voice commands and output various information, or order products and services, or install an application on a computing device or perform another action. For example, the initial voice query could be “Why did the operating system on my laptop crash?”. Based on this general query, the data processing system 102 of this technical solution can generate one or more pivot points in one or more layers of a tree structure from disjoint result sets based on historical search results, and obtain the user's selection of pivot points until the final node is reached. Upon reaching the final node, the data processing system 102 can identify an action and then perform it on the computing device 140. For example, the action for the final node could be installing an application, reinstalling an application, or upgrading an application version on the crashed computing device 140 (e.g., the user's laptop). Therefore, the data processing system 102 can identify the final node based on the response to the checkpoint. The data processing system 102 constructs an action to install an application on the computing device 140 associated with the electronic account identifier based on the final node. Furthermore, the data processing system 102 can perform actions on the computing device 140 to install applications.
[0086] In some cases, the server digital assistant 108 can determine that, in addition to performing an action, supplementary digital content items are also provided. The server digital assistant 108 can send a request for content to the content selector 110 based on an input audio signal. The server digital assistant 108 can send a request for supplementary or sponsored content from a third-party content provider. The content selector 110 can perform a content selection process to select supplementary or sponsored content items based on the action in the voice query. Content items can be sponsored or supplementary digital component objects. Content items can be provided by a third-party content provider (such as supplementary digital content provider device 130). Supplementary content items may include advertisements for goods or services. The content selector 110 can select content items using content selection criteria in response to receiving a request for content from the server digital assistant 108.
[0087] Server digital assistant 108 can receive supplementary or sponsored content items from content selector 110. Server digital assistant 108 can receive content items in response to requests. Server digital assistant 108 can receive content items from content selector 110 and present the content items via audio output or visual output. Server digital assistant 108 can present content items via client computing device 140.
[0088] Data processing system 102 may include a content selector 110 designed, constructed, or operated to select supplementary content items (or sponsored content items or digital component objects). To select a sponsored content item or digital component, content selector 110 may use generated content selection criteria to select matching sponsored content items based on broad matching, exact matching, or phrase matching. For example, content selector 110 may analyze, parse, or otherwise process the subject matter of candidate sponsored content items to determine whether the subject matter of the candidate sponsored content items corresponds to the subject matter of keywords or phrases in the content selection criteria (e.g., actions or intentions). Content selector 110 may use image processing techniques, character recognition techniques, natural language processing techniques, or database lookups to identify, analyze, or recognize the speech, audio, terminology, characters, text, symbols, or images of candidate digital components. Candidate sponsored content items may include metadata indicating the subject matter of the candidate digital component; in this case, content selector 110 may process the metadata to determine whether the subject matter of the candidate digital component corresponds to an input audio signal. The content campaign provided by the supplemental digital content provider device 130 may include content selection criteria, which the data processing system 102 may match with criteria indicated in the second profile layer or the first profile layer.
[0089] When establishing a content movement that includes digital components, the supplemental digital content provider can provide additional indicators. The supplemental digital content provider device 130 can provide content selector 110 that can identify a content movement or content group layer by performing a lookup using information about candidate digital components. For example, candidate digital components may include unique identifiers that can be mapped to content groups, content movements, or content providers.
[0090] In response to a request, content selector 110 can select a digital component object associated with supplemental digital content provider device 130. The supplemental digital content may be provided by a supplemental digital content provider device that is different from a service provider device (e.g., a ride-sharing service provider). The supplemental digital content may correspond to a service type different from an action data structure (e.g., taxi service versus food delivery service). Client computing device 140 can interact with the supplemental digital content. Client computing device 140 can receive audio responses to the digital components. Client computing device 140 can receive indications to select a hyperlink or other button associated with the digital component object, which enables or allows client computing device 140 to identify supplemental digital content provider device 130, request services from supplemental digital content provider device 130, instruct supplemental digital content provider device 130 to perform services, send information to supplemental digital content provider device 130 or a service provider device, or otherwise query supplemental digital content provider device 130.
[0091] The supplementary digital content provider device 130 can establish electronic content campaigns. An electronic content campaign can refer to one or more content groups corresponding to a common theme. A content campaign can include a hierarchical data structure comprising content groups, digital component data objects, and content selection criteria provided by the content provider. The content selection criteria provided by the content provider device 130 can include content types, such as digital assistant content types, search content types, streaming video content types, streaming audio content types, or contextual content types. To create a content campaign, the supplementary digital content provider device 130 can specify values for campaign layer parameters. Campaign layer parameters can include, for example, a campaign name, a preferred content network for placing digital component objects, resource values to be used for the content campaign, start and end dates of the content campaign, duration of the content campaign, a schedule for placing digital component objects, language, geographic location, and the type of computing device on which the digital component objects are provided. In some cases, an impression can refer to when digital component objects are obtained from a source of digital component objects (e.g., data processing system 102 or supplementary digital content provider device 130) and is countable. In some cases, bot activity can be filtered and excluded as an impression due to the possibility of click fraud. Therefore, in some cases, an impression can refer to a measurement of the response from a web server to a page request from a browser, filtered from bot activity and error codes, and recorded at a point as close as possible to the opportunity to display the digital component object for display on the client computing device 140. In some cases, an impression can refer to a visible or audible impression; for example, the digital component object is at least partially (e.g., 20%, 30%, 30%, 40%, 50%, 60%, 70% or more) visible on the display device of the client computing device 140, or audible via the speakers of the client computing device 140. A click or selection can refer to user interaction with the digital component object, such as a voice response to an audible impression, a mouse click, a touch interaction, a gesture, a shake, an audio interaction, or a keyboard click. A conversion can refer to a user taking a desired action regarding the digital component object; for example, purchasing a product or service, completing a survey, visiting a physical store corresponding to the digital component, or completing an electronic transaction.
[0092] The supplemental digital content provider device 130 can further establish one or more content groups for content campaigns. A content group includes one or more digital component objects and corresponding content selection criteria, such as keywords, words, terms, phrases, geographic location, computing device type, time of day, interests, topics, or verticals. Content groups under the same content campaign can share the same campaign-level parameters, but can have specific specifications for particular content group-level parameters, such as keywords, negative keywords (e.g., preventing the placement of digital components if negative keywords are present in the main content), bids for keywords, or parameters associated with bids or content campaigns.
[0093] To create a new content group, a content provider can provide values for content group layer parameters. Content group layer parameters include, for example, a content group name or content group subject, and bids for different content placement opportunities (e.g., automatic or managed placement) or results (e.g., clicks, impressions, or conversions). The content group name or content group subject can be one or more terms that the supplementary digital content provider device 130 can use to capture the topic or theme of the digital component objects of the content group to be selected for display. For example, a car dealership can create different content groups for each brand of vehicles it ships, and can further create different content groups for each model of vehicle it ships. Examples of content group subjects that a car dealership can use include, for example, "Manufacturing Model A Sports Cars," "Manufacturing Model B Sports Cars," "Manufacturing Model C Sedan," "Manufacturing Model C Trucks," "Manufacturing Model C Hybrids," or "Manufacturing Model D Hybrids." For example, an example content movement subject could be "hybrid," and include content groups for both "Manufacturing Model C Hybrids" and "Manufacturing Model D Hybrids."
[0094] The supplemental digital content provider device 130 can provide one or more keywords and digital component objects to each content group. Keywords can include terms related to or identified by the product or service associated with or identified by the digital component object. Keywords can include one or more terms or phrases. For example, a car dealership could include "sports car," "V-6 engine," "four-wheel drive," and "fuel efficiency" as keywords for a content group or content campaign. In some cases, negative keywords can be specified by the content provider to avoid, prevent, block, or disable content placement on certain terms or keywords. The content provider can specify the matching type used to select digital component objects, such as exact match, phrase match, or broad match.
[0095] The supplemental digital content provider device 130 can provide one or more keywords for the data processing system 102 to select digital component objects provided by the supplemental digital content provider device 130. The supplemental digital content provider device 130 can identify one or more bidding keywords and further provide bid amounts for various keywords. The supplemental digital content provider device 130 can provide additional content selection criteria to be used by the data processing system 102 to select digital component objects. Multiple supplemental digital content provider devices 130 can bid on the same or different keywords, and the data processing system 102 can run a content selection process or advertising auction in response to keyword instructions received via electronic messaging.
[0096] The supplemental digital content provider device 130 can provide one or more digital component objects for selection by the data processing system 102. The data processing system 102 (e.g., via content selector 110) can select digital component objects when content placement opportunities become available, matching resource allocation, content scheduling, maximum bid, keywords, and other selection criteria specified for the content group. Different types of digital component objects can be included in the content group, such as voice digital components, audio digital components, text digital components, image digital components, video digital components, multimedia digital components, digital component links, or auxiliary application components. Digital component objects (or digital components, supplemental content items, or sponsored content items) can include, for example, content items, online documents, audio, images, video, multimedia content, sponsored content, or auxiliary applications. When selecting a digital component, the data processing system 102 can send the digital component object for display on the client computing device 140 or its display device. Display can include showing the digital component on the display device, performing an application such as a chatbot or conversational bot, or playing the digital component via the speaker of the client computing device 140. The data processing system 102 can provide instructions to the client computing device 140 to display the digital component object. The data processing system 102 can instruct the client computing device 140 to generate audio signals or sound waves.
[0097] Content selector 110 can perform a real-time content selection process in response to a request. Real-time content selection can refer to or include performing content selection in response to a request. Real-time can refer to or include selecting content within 0.2 seconds, 0.3 seconds, 0.4 seconds, 0.5 seconds, 0.6 seconds, or 1 second after receiving the request. Real-time can refer to selecting content in response to receiving an input audio signal from the client computing device 140.
[0098] Content selector 110 can identify multiple candidate supplementary content items. Content selector 110 can determine the score or ranking of each of the multiple candidate supplementary content items in order to select the supplementary content item with the highest ranking to provide to client computing device 140.
[0099] Figure 2 This is an illustration of an example query flow for dynamic tree structure generation according to an implementation method. Flow 200 can be generated by... Figure 1 One or more systems or components described herein perform this action, including, for example, a data processing system. The data processing system may receive an original or first query 202, such as "find a gift for my mother-in-law." The data processing system may process the original query and generate a pivot point 204 having two or more child nodes selected from disjoint sets that are distinct from each other. The pivot point 204 may be "jewelry, tickets, or something else." The pivot point 204 may be formed from child nodes jewelry 206, tickets 208, and something else node 210 that allows the user to receive different sets of child nodes.
[0100] The data processing system can receive responses from users to pivot points, including indications of one of the nodes within the pivot point, such as node 208 for tickets. The data processing system can proceed down the tree hierarchy to child node 208. The data processing system can generate a second pivot point 214 for child node 208 with grandchild node options "opera, musical, ballet, or sports". The suggested grandchild nodes 216 (opera), 218 (musical), 220 (ballet), and 222 (sports) in pivot point 214 can be associated with node 208 for tickets. If the user has selected a jewelry child node, the data processing system can generate a second pivot point 212 for jewelry child nodes with grandchild node options such as "necklace, bracelet, ring, or earrings". The suggested grandchild nodes in pivot point 212 can be associated with node 206 for jewelry.
[0101] The data processing system can receive a selection for the grandchild node Sports 222. The data processing system can generate additional pivot points. For example, the data processing system can generate a third pivot point "Basketball, Football, or Hockey" 224 corresponding to the Sports 222 node and the Ticket 208 node. In some cases, the data processing system can generate checkpoints and suggest final nodes, such as Lakers Ticket 226 corresponding to the Basketball node within pivot point 224.
[0102] If the user selects final node 226, the data processing system can then perform actions such as purchasing tickets to a Lakers basketball game. The data processing system can generate or update a micro-profile of this context. The data processing system can recognize the context as buying a gift or buying a gift for one's mother-in-law. The data processing system can also specify the context based on time, year, holiday, location, etc.
[0103] Figure 3 This is an illustration of an example query flow for dynamic tree structure generation according to an implementation method. Flow 300 can be generated by... Figure 1The description refers to one or more systems or components that perform this action, including, for example, a data processing system. The data processing system may receive an initial query or a first query, "I want to buy a hat for my mother-in-law" 302. This initial query may be more specific or less general than query 202. Therefore, the data processing system can generate a new dynamic tree structure that skips one or more layers. In some cases, the data processing system can determine whether a micro-profile exists for that electronic account identifier and context. For example, the data processing system can perform a lookup in a micro-profile data structure. The data processing system can identify... Figure 2 The process described in section 200 involves generating a micro-profile for the purchase of a gift. The data processing system can identify the final node 226 corresponding to the Lakers ticket.
[0104] By loading the microprofile, the data processing system can generate a first pivot point with child nodes "sun hat, baseball cap, or Lakers cap" 304. As indicated by this first pivot point, the child nodes can be based on general relevant search results and the specific final node "Lakers ticket" from the microprofile. The user can select "Lakers cap" as the final node 306. Therefore, by loading the microprofile generated for this user and context, a query to buy a "sun hat" can be resolved with a single pivot point, thereby reducing computational resource utilization and the duration of network transactions and communication sessions, while improving accuracy.
[0105] Figure 4 This is an illustration of an example method for generating dynamic tree structures according to an implementation method. Method 400 can be derived from... Figure 1 One or more components or systems described herein perform the action, including, for example, a data processing system. At 402, the data processing system may receive a voice query. The voice query may be a raw or first voice query. The voice query may be general or have some degree of specificity. At 404, the data processing system may determine the context associated with the voice query. The context may include the time of day, the time of year, semantic concepts (buying a gift, traveling, requesting information, requesting a movie), language, and acoustic characteristics of the voice input (e.g., emotion, tone).
[0106] At 406, the data processing system can determine whether a microprofile exists for the specific e-account identifier associated with the device that detected the voice query. The data processing system can determine whether a matching context exists for the microprofile. If a microprofile exists for the user and context, the data processing system can load the microprofile at 408 and then proceed to box 410. If the microprofile does not exist for the user or does not exist in the same context, the data processing system can proceed directly to box 410 to generate a pivot point and output a prompt. A tree generator can be used to generate the pivot point based on relevant searches for similar users and the microprofile (if it exists).
[0107] The data processing system can generate and output pivot points, and then receive responses to the pivot points at 412. The responses to the pivot points can indicate the selection of child nodes within the pivot point. The data processing system can proceed to decision box 414 to determine whether to generate a checkpoint. A checkpoint can refer to asking the user whether to continue the tree process, suggesting a final node, or suggesting a narrower pivot point to skip a layer. The data processing system can determine whether to generate a checkpoint based on one or more factors, including, for example, the number of layers the user has descended from the top node in the tree structure, the distance between disjoint sets of previous or subsequent available pivot points, the tone of the user's voice (e.g., dissatisfaction with the tree process), the duration of the communication session, or the duration since the data processing system received the original query. If the data processing system determines at 414 not to generate a checkpoint, the data processing system can return to box 410 to generate a second pivot point and obtain another response to the second pivot point. The data processing system can iterate through the generation of pivot points and receiving selections of child nodes until the data processing system determines to generate a checkpoint.
[0108] If the data processing system determines that a checkpoint has been generated, it can proceed to decision box 416 to determine whether the user has reached the final node. For example, if the checkpoint question is "Do you want to continue?" and the user says "No," then the final node has been reached. If the user says they do want to continue tree generation, the data processing system can determine that the final node has not been reached and then return from decision box 416 to box 410. In another example, the data processing system can output a suggested final node, such as "Lakers tickets." In decision box 416, the data processing system can ask the user if the final result is acceptable. If the user indicates that the final node is satisfactory, the data processing system can proceed to box 418.
[0109] At step 418, the data processing system can generate, build, or update microprofiles. The data processing system can build microprofiles based on keywords or interactions associated with the spanning tree structure and reaching the final node. If the data processing system has previously loaded a microprofile with the same context, it can update the microprofile. The data processing system can proceed to step 420 to perform or execute actions associated with the final node, such as purchasing Lakers tickets, Lakers hats, or performing any other action to order products, services, or access requested information.
[0110] Figure 5This is a block diagram of an example computer system 500. The computer system or computing device 500 may include or be used to implement system 100 or components thereof, such as data processing system 102 or client computing device 140. Data processing system 102 or client computing device 140 may include an intelligent personal assistant or a voice-based digital assistant. The computing system 500 includes a bus 505 or other communication components for communicating information, and a processor 510 or processing circuitry coupled to the bus 505 for processing information. The computing system 500 may also include one or more processors 510 or processing circuitry coupled to the bus for processing information. The computing system 500 also includes a main memory 515, such as random access memory (RAM) or other dynamic storage device, coupled to the bus 505 for storing information and instructions to be executed by the processor 510. The main memory 515 may be or include a data store 118. The main memory 515 may also be used to store location information, temporary variables, or other intermediate information during the execution of instructions by the processor 510. The computing system 500 may also include a read-only memory (ROM) 520 or other static storage device coupled to the bus 505 for storing static information and instructions for the processor 510. A storage device 425, such as a solid-state device, disk, or optical disk, may be coupled to the bus 505 to persistently store information and instructions. Storage device 425 may include a data storage library 118 or a portion thereof.
[0111] The computing system 500 can be coupled to a display 535 (such as a liquid crystal display or an active matrix display) via a bus 505 for displaying information to a user. An input device 530 (such as a keyboard including alphanumeric keys and other keys) can be coupled to the bus 505 for communicating information and command selections to the processor 510. The input device 530 may include a touchscreen display 535. The input device 530 may also include cursor control, such as a mouse, trackball, or arrow keys, for communicating directional information and command selections to the processor 510 and for controlling cursor movement on the display 535. For example, the display 535 may be... Figure 1 It is part of the data processing system 102 or the client computing device 140 or other components.
[0112] The processes, systems, and methods described herein can be implemented by a computing system 500 in response to a processor 510 executing an instruction arrangement contained in main memory 515. Such instructions may be read into main memory 515 from another computer-readable medium, such as storage device 425. Execution of the instruction arrangement contained in main memory 515 causes the computing system 500 to perform the illustrative processes described herein. One or more processors in a multiprocessor configuration may also be used to execute the instructions contained in main memory 515. Hardwired circuitry may be used in place of or in combination with software instructions in conjunction with the systems and methods described herein. The systems and methods described herein are not limited to any particular combination of hardware circuitry and software.
[0113] Although already Figure 5 An example computing system is described herein, but the subject matter including the operations described herein may be implemented in other types of digital electronic circuits, or in computer software, firmware, or hardware that includes the structures disclosed herein and their equivalents, or in a combination of one or more of them.
[0114] In the context of the systems discussed in this paper that collect or use personal information about users, users can be given the opportunity to control whether a program or feature can collect personal information (e.g., information about a user's social networks, social actions or activities, user preferences, or user location), or to control whether or how content that may be more relevant to the user is received from content servers or other data processing systems. Furthermore, some data can be anonymized in one or more ways before storage or use, so that personally identifiable information is removed when parameters are generated. For example, a user's identity can be anonymized so that personally identifiable information cannot be determined for the user, or the geographic location of the user's location information can be generalized (such as by city, zip code, or state) so that the user's specific location cannot be determined. Therefore, users can control how information is collected about them and how it is used by content servers.
[0115] The subject matter and operations described in this specification can be implemented in digital electronic circuits, or in computer software, firmware, or hardware (including the structures disclosed in this specification and their equivalents), or in a combination of one or more of these. The subject matter described in this specification can be implemented as one or more computer programs, such as one or more circuits of computer program instructions, encoded on one or more computer storage media for execution by or control of the operation of a data processing device. Alternatively or additionally, the program instructions can be encoded on artificially generated propagated signals, such as machine-generated electrical, optical, or electromagnetic signals, generated to encode information for transmission to a suitable receiver device for execution by the data processing device. The computer storage medium can be a computer-readable storage device, a computer-readable storage substrate, a random or serial access memory array or device, or a combination of one or more of these, or be included in a computer-readable storage device, a computer-readable storage substrate, a random or serial access memory array or device, or a combination of one or more of these. Although the computer storage medium is not a propagated signal, it can be a source or destination of computer program instructions encoded in an artificially generated propagated signal. Computer storage media may also be one or more separate components or media (e.g., multiple CDs, disks, or other storage devices) or included in one or more separate components or media (e.g., multiple CDs, disks, or other storage devices). The operations described in this specification can be implemented as operations performed by a data processing apparatus on data stored on one or more computer-readable storage devices or received from other sources.
[0116] The terms “data processing system,” “computing device,” “component,” or “data processing apparatus” encompass a variety of means, devices, and machines for processing data, including, for example, programmable processors, computers, systems-on-a-chip, or a combination thereof. Apparatus may include special-purpose logic circuitry, such as FPGAs (Field-Programmable Gate Arrays) or ASICs (Application-Specific Integrated Circuits). In addition to hardware, apparatus may include code that creates an execution environment for the computer program in question, such as code that constitutes processor firmware, protocol stacks, database management systems, operating systems, cross-platform runtime environments, virtual machines, or combinations thereof. Apparatus and execution environments can implement a variety of different computing model infrastructures, such as network services, distributed computing, and grid computing infrastructures. For example, tree generator 112, digital assistant 108, checkpoint generator 114, or other components may include or share one or more data processing apparatuses, systems, computing devices, or processors.
[0117] Computer programs (also known as programs, software, software applications, apps, scripts, or code) can be written in any form of programming language (including compiled or interpreted languages, declarative or procedural languages) and can be deployed in any form (including as standalone programs or as modules, components, subroutines, objects, or other units suitable for a computing environment). Computer programs can correspond to files in a file system. A computer program can be stored as a part of a file that holds other programs or data (e.g., one or more scripts stored in a markup language document), in a single file dedicated to the program in question, or in multiple coordinating files (e.g., files storing one or more modules, subroutines, or code sections). A computer program can be deployed to execute on a single computer, or on multiple computers located at a single site or distributed across multiple sites and interconnected through a communications network.
[0118] The processes and logic flows described in this specification can be executed by one or more programmable processors (e.g., components of data processing system 102) that execute one or more computer programs to perform actions by manipulating input data and generating output. The processes and logic flows can also be executed by special-purpose logic circuitry (e.g., FPGA (Field-Programmable Gate Array) or ASIC (Application-Specific Integrated Circuit)), and the apparatus can also be implemented as special-purpose logic circuitry. Suitable devices for storing computer program instructions and data include all forms of non-volatile memory, media, and memory devices, including, for example, semiconductor memory devices (e.g., EPROM, EEPROM, and flash memory devices); magnetic disks (e.g., internal hard disks or removable disks); magneto-optical disks; and CD-ROM and DVD-ROM disks. Processors and memory can be supplemented by or incorporated into special-purpose logic circuitry.
[0119] The subject matter described herein can be implemented in a computing system that includes backend components (e.g., as a data server), or middleware components (e.g., an application server), or frontend components (e.g., a client computer having a graphical user interface or web browser through which a user can interact with the embodiments described herein), or a combination of one or more such backend, middleware, or frontend components. The components of the system can be interconnected via digital data communication (e.g., a communication network) of any form or medium. Examples of communication networks include local area networks (“LANs”) and wide area networks (“WANs”), interconnected networks (e.g., the Internet), and peer-to-peer networks (e.g., ad hoc peer-to-peer networks).
[0120] Computing systems (such as system 100 or system 500) may include clients and servers. Clients and servers are typically geographically separated and generally interact via a communication network (e.g., network 105). The client-server relationship arises by means of computer programs running on separate computers and having a client-server relationship with each other. In some implementations, the server sends data (e.g., data packets representing digital components) to the client device (e.g., for the purpose of displaying data to a user interacting with the client device and receiving user input from the user interacting with the client device). Data generated at the client device (e.g., the result of user interaction) may be received at the server (e.g., received by data processing system 102 from client computing device 140 or supplemental digital content provider device 130).
[0121] Although the operations are depicted in a specific order in the diagram, it is not required that such operations be performed in the specific order or sequence shown, nor is it required that all the operations shown be performed. The actions described herein can be performed in different orders.
[0122] The separation of various system components is not required in all implementations, and the described program components may be included in a single hardware or software product. For example, the natural language processor 106 and interface 104 may be a single component, app, or program, or a logic device with one or more processing circuits, or part of one or more servers of the data processing system 102.
[0123] Some illustrative embodiments have been described, and it is clear that the foregoing is illustrative and not limiting, and has been provided by way of example. Specifically, although many of the examples presented herein involve specific combinations of method actions or system elements, those actions and those elements can be combined in other ways to achieve the same objective. Actions, elements, and features discussed in conjunction with one embodiment are not intended to exclude similar roles or embodiments in other embodiments.
[0124] The wording and technical terms used herein are for descriptive purposes and should not be considered restrictive. The terms “comprising,” “including,” “having,” “containing,” “involving,” “characterized by,” “characterized by,” and variations thereof are intended to cover the items listed thereafter, their equivalents and additions, and alternative embodiments consisting of the items exclusively listed thereafter. In one embodiment, the systems and methods described herein consist of one, more than one, or all of the described elements, actions, or components.
[0125] Any reference to an embodiment, element, or action of a system or method mentioned herein in the singular may also include embodiments that include multiple such elements, and any reference to any embodiment, element, or action mentioned herein in the plural may also include embodiments that include only a single element. References in the singular or plural form are not intended to limit the currently disclosed systems or methods, their components, actions, or elements to a single or multiple configuration. A reference to any action or element based on any information, action, or element may include an embodiment in which such action or element is at least partially based on any information, action, or element.
[0126] Any implementation disclosed herein may be combined with any other implementation or embodiment, and references to “implementation,” “some implementations,” “one implementation,” etc., are not necessarily mutually exclusive and are intended to indicate that a particular feature, structure, or characteristic described in connection with an implementation may be included in at least one implementation or embodiment. Such terms as used herein do not necessarily refer to the same implementation. Any implementation may be combined inclusively or exclusively with any other implementation in any manner consistent with the aspects and implementations disclosed herein.
[0127] A reference to "or" can be interpreted as inclusive, such that any term described using "or" can refer to any one, more than one, or all of the terms described. A reference to at least one of a list of combinations of terms can be interpreted as inclusive "or" to refer to any one, more than one, or all of the terms described. A reference to "at least one of 'A' and 'B'" can include only 'A', only 'B', and both 'A' and 'B'. Such references used in conjunction with "including" or other open-ended technical terms can include additional terms.
[0128] Where a reference numeral follows a technical feature in the drawings, detailed description, or any claim, the reference numeral is included to enhance the comprehensibility of the drawings, detailed description, and claims. Therefore, the presence or absence of reference numerals has no limiting effect on the scope of any claim element.
[0129] The systems and methods described herein may be embodied in other specific forms without departing from their characteristics. The foregoing embodiments are illustrative and not intended to limit the systems and methods described. The scope of the systems and methods described herein is therefore indicated by the appended claims rather than the foregoing description, and changes in the meaning and scope of equivalents of the claims are included therein.
Claims
1. A system for generating dynamic tree structures, comprising: A data processing system, including memory and one or more processors, is used for: Receive a first voice query detected via a microphone associated with an electronic account identifier; A first pivot point, comprising multiple child nodes, is generated from historical searches related to the first voice query performed by multiple computing devices in a tree structure for the first voice query. Output an audio prompt to request the selection of one of several child nodes; In response to audio prompts, it receives voice input, including the selection of the first child node among multiple child nodes; In response to the user's selection of the first child node, a second pivot point is generated in the tree structure, which includes multiple grandchild nodes of the first child node, from the historical search related to the first child node. Output a second audio prompt to request the selection of one of several grandchild nodes; Based on the response to the second audio cue, skip one layer in the tree structure to generate a third pivot point or final node with multiple great-great-grandchild nodes. In the event of generating the third pivot point, the one or more processors are configured to: Based on a resource reduction strategy, checkpoints are identified to reduce the generation of additional child nodes. A checkpoint refers to presenting possible final results as options to the user. Identify the final node based on the user's response to the checkpoint; Using the previous search, construct a micro-profile of the electronic account identifier using a tree structure based on the context; Based on the final node, construct the action of installing the application on the computing device associated with the electronic account identifier; and Perform the action on the computing device to install the application.
2. The system according to claim 1, comprising: The data processing system generates a first pivot point with a first child node and a second child node, wherein the first child node and the second child node are separated by a distance greater than a threshold.
3. The system according to claim 1 or 2, comprising: The data processing system generates a first pivot point with a first child node and a second child node, wherein the distance between the first child node and the second child node is the maximum distance between multiple candidate nodes identified in response to a first voice query.
4. The system according to claim 1 or 2, comprising: The data processing system generates a first pivot point with a first child node and a second child node, wherein the first child node and the second child node are formed from separate disjoint sets of size within a threshold size.
5. The system according to claim 1 or 2, further comprising a data processing system for: In response to receiving the first voice query, determine whether a micro-profile exists for the electronic account identifier; and In response to determining that a microfile does not exist for the electronic account identifier, a first pivot point is generated.
6. The system according to claim 1 or 2, further comprising a data processing system for: In response to receiving a first voice query, a context associated with the electronic account identifier is determined, the context being based on at least one of the time of day, date, or acoustic characteristics of the first voice query; and Associate the context with the tree structure in the microfile.
7. The system according to claim 6, further comprising a data processing system for: Receive a third voice query and determine the current context associated with the electronic account identifier; Determine if the current context matches the context; as well as In response to determining that the current context matches the context, a microprofile is loaded to generate a second audio prompt in response to a third voice query.
8. The system according to claim 1 or 2, comprising: In response to the termination of the session with the electronic account identifier, the data processing system erases the microfile from the data processing system.
9. A method for generating dynamic tree structures, comprising: A data processing system, including memory and one or more processors, receives a first voice query detected via a microphone associated with an electronic account identifier. The data processing system generates a first pivot point, which includes multiple child nodes, in a tree structure for the first voice query from historical searches related to the first voice query performed by multiple computing devices. The data processing system outputs an audio prompt to request the selection of one of several child nodes; In response to audio prompts, the data processing system receives voice input, including the selection of the first child node among multiple child nodes; In response to the user's selection of the first child node, the data processing system generates a second pivot point in the tree structure, which includes multiple grandchild nodes of the first child node, from the historical search related to the first child node. The data processing system outputs a second audio prompt to request the selection of one of several grandchild nodes; Based on the response to the second audio cue, the data processing system skips one layer in the tree structure to generate a third pivot point or final node with multiple great-great-grandchild nodes. In the event of generating the third pivot point, the method further includes: Based on a resource reduction strategy, the data processing system determines checkpoints to reduce the generation of additional child nodes. A checkpoint refers to presenting possible final results as options to the user. Identify the final node based on the user's response to the checkpoint; The data processing system uses previous searches to construct micro-profiles of electronic account identifiers based on context using a tree structure; Based on the final node, construct the action of installing the application on the computing device associated with the electronic account identifier; and Perform the action on the computing device to install the application.
10. The method of claim 9, comprising: A first pivot point is generated by the data processing system, having a first child node and a second child node, wherein the first child node and the second child node are separated by a distance greater than a threshold.
11. The method according to claim 9 or 10, comprising: A first pivot point with a first child node and a second child node is generated by the data processing system, wherein the distance between the first child node and the second child node is the maximum distance between multiple candidate nodes identified in response to the first voice query.
12. The method according to claim 9 or 10, comprising: A first pivot point with a first child node and a second child node is generated by the data processing system, wherein the first child node and the second child node are formed from separate disjoint sets of size within a threshold size.
13. The method according to claim 9 or 10, comprising: The data processing system responds to the first voice query and determines whether a micro-shortcut exists for the electronic account identifier; as well as In response to determining that a micro-profile does not exist for the electronic account identifier, the data processing system generates a first pivot point.
14. The method according to claim 9 or 10, comprising: In response to receiving a first voice query, the data processing system determines a context associated with the electronic account identifier, the context being based on at least one of the time of day, date, or acoustic characteristics of the first voice query; as well as The data processing system associates the context with the tree structure in the microfile.
15. The method of claim 14, comprising: The data processing system receives a third voice query and determines the current context associated with the electronic account identifier; The data processing system determines that the current context matches the context. as well as In response to determining that the current context matches the context, the data processing system loads a microprofile to generate a second audio prompt in response to a third voice query.
Citation Information
Patent Citations
Intelligent automated assistant
US20200327895A1