Aggregating information from different data feed services
Machine learning-based aggregation of data from multiple data feed services addresses the inefficiency of manual navigation by processing natural language queries to retrieve and present aggregated information across various platforms.
Patent Information
- Application Number
- JP2024571331
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2022-06-16
- Filing Date
- 2023-06-15
- Publication Date
- 2025-08-13
- Estimated Expiration
- 2043-06-15
AI Technical Summary
Individuals face the tedious and time-consuming task of sifting through multiple data feed services to obtain specific information, especially when using voice user interfaces, due to the habitual seeking of semantically related information across different sources and frequent changes in data presentation.
Implementing machine learning to aggregate information from multiple data feed services using data feed-agnostic aggregator embeddings, allowing users to issue free-form natural language queries, which are processed to select and execute actions across various data feed services, presenting aggregated responses.
Enables efficient retrieval of diverse perspectives on a topic from multiple data feed services, reducing the need for manual navigation and adapting to changes in data presentation.
Smart Images

Figure 2025526228000001_ABST
Abstract
Description
[Technical Field]
[0001] Individuals often habitually seek semantically related information from multiple different sources. For example, a particular individual may tend to check the same data feed service, such as a news website and / or a social networking data feed, to obtain different perspectives on a particular topic, such as developing events, new movies, sporting events, etc. However, in addition to the specific information sought by the individual, various data feed services may communicate myriad other information that is not relevant to the individual at the time. As a result, it can be tedious and / or time-consuming for individuals to sift through multiple different data feed services to obtain the specific information they seek. This problem can be amplified when interacting with a voice user interface (VUI) provided by a computing device, such as a smart speaker or an in-car voice command system. Additionally, individual data feed services routinely change how data is presented and / or made accessible, which can further hinder an individual's efforts. Summary of the Invention
[0002]
[0003] Implementations are described herein for aggregating information in response to queries from multiple different data feed services using machine learning. More particularly, but not exclusively, implementations are described herein for leveraging domain-specific machine learning models to query information from heterogeneous data feed services using data feed-agnostic aggregator embeddings. The embodiments described herein provide various technical advantages. Rather than navigating multiple different data feed services to obtain information about a topic, an individual can simply issue a free-form natural language query that describes what information the individual is seeking and, if applicable, where the individual is seeking this information from. Response information from the multiple different data feed services can then be obtained and presented to the individual in a aggregate, e.g., as part of an aggregated data feed. Thus, the individual can obtain different perspectives on a topic and / or multiple perceptions of the topic from different people across multiple different data feed services.
[0003] In some implementations, a method can be executed using one or more processors and includes obtaining a natural language input including a query for information; performing natural language processing (NLP) on the natural language input to generate a data feed-independent aggregator embedding; selecting a plurality of data feed services, each of the plurality of data feed services to be selected having its own data feed service action space, each data feed service action space including actions that can be performed to access data communicated via the respective data feed service; processing the feed-independent aggregator embedding based on a plurality of domain-specific machine learning models corresponding to the plurality of data feed services, each domain-specific machine learning model trained to transform between its respective data feed service action space and a data feed-independent semantic embedding space including the data feed-independent aggregator embedding; selecting and executing one or more actions from each of the data feed service action spaces based on the processing to aggregate data from the plurality of data feed services in response to the query; and presenting the aggregated response data as output.
[0004] In various implementations, the multiple data feed services may be selected based on an entity identifier included in the query for information. In various implementations, the query for information may include a request for social media posts from a particular individual, and the selected multiple data feed services may include two or more social media services, and the two or more social media services are selected based on the particular individual's membership with the two or more social media services. In various implementations, for a given social media service of the two or more social media services, the one or more actions selected from the data feed service action space of the given social media service may include accessing one or more posts by the particular individual from the particular individual's posting history. In various implementations, for a given social media service of the two or more social media services, the one or more actions selected from the data feed service action space of the given social media service may include filtering one or more posts by the particular individual from a general data feed on the given social media service provided to the user who issued the natural language input.
[0005] In various implementations, the multiple data feed services may be selected based on a lookup table controlled by the user who issued the natural language input. In various implementations, the lookup table may include the user's contact list. In various implementations, the multiple data feed services may be selected based on the user who issued the natural language input having previously provided the aggregator agent with permission to access the multiple data feed services. In various implementations, at least one of the data feed services may include a virtual space that forms part of a larger metaverse comprising multiple virtual spaces.
[0006] In various implementations, the aggregated response data may be presented to the user who issued the natural language input as part of a metaverse graphical user interface. In various implementations, the multiple data feed services may be selected based on the browsing history of the user who issued the natural language input.
[0007] Additionally, some implementations include one or more processors of one or more computing devices, the one or more processors operable to execute instructions stored in associated memory, the instructions configured to cause any of the aforementioned methods to be performed. Some implementations include at least one non-transitory computer-readable storage medium storing computer instructions executable by the one or more processors to perform any of the aforementioned methods.
[0008] It should be understood that all combinations of the foregoing concepts and additional concepts described in more detail herein are contemplated as being part of the subject matter disclosed herein, for example, all combinations of claimed subject matter appearing at the end of this disclosure are contemplated as being part of the subject matter disclosed herein. [Brief explanation of the drawings]
[0009] [Figure 1] FIG. 1 is a schematic diagram of an example environment in which implementations disclosed herein may be practiced. [Figure 2] 1A-1C illustrate generally examples of how data may be exchanged and / or processed to perform selected aspects of the present disclosure, according to various implementations. [Figure 3] 10 illustrates schematically another example of how data may be processed for querying multiple domains using data feed-agnostic aggregator embedding, according to various implementations. [Figure 4] 1 is a flowchart illustrating an exemplary method for practicing selected aspects of the present disclosure, according to implementations disclosed herein. [Figure 5]1 illustrates an exemplary architecture of a computing device. DETAILED DESCRIPTION OF THE INVENTION
[0010]
[0003] Implementations are described herein for aggregating information in response to queries from multiple different data feed services using machine learning. More particularly, but not exclusively, implementations are described herein for leveraging domain-specific machine learning models to query information from heterogeneous data feed services using data feed-agnostic aggregator embeddings. The embodiments described herein provide various technical advantages. Rather than navigating multiple different data feed services to obtain information about a topic, an individual can simply issue a free-form natural language query that describes what information the individual is seeking and, if applicable, where the individual is seeking this information from. Response information from the multiple different data feed services can then be obtained and presented to the individual in a aggregate, e.g., as part of an aggregated data feed. Thus, the individual can obtain different perspectives on a topic and / or multiple perceptions of the topic from different people across multiple different data feed services.
[0011] As used herein, a "data feed service" may be any computer-implemented service, such as a web service, that can be used by individuals to publish, broadcast, push, and / or multicast content to others. Data feed services are typically accessible by multiple individuals and are generally updated frequently, or at least periodically, with new content, often presented in reverse chronological order. One common example of a data feed service is a social networking service, where individual users can publish information, e.g., facts, opinions, photos, statuses, to their "friends" or "followers," or anyone who accesses the social networking service. Another example of a data feed service is a news website (or, more generally, a news service, which may publish to both a website and its own application) that continuously publishes news articles, opinions, columns, and the like.
[0012] These examples of data feed services are not intended to be limiting. Other examples of data feed services include, but are not limited to, weblogs (more commonly called "blogs"), web feeds (e.g., really simply syndication or "RSS" feeds), review aggregation websites, social news websites featuring user-contributed content, etc. Additionally, data feed services are not limited to services provided to websites. For example, a data feed service could take the form of a virtual space (e.g., a massively multiplayer online role-playing game ("MMORPG") or portion thereof, a virtual coffee shop or forum, etc.) that forms part of a larger metaverse that includes multiple virtual spaces.
[0013] To retrieve information responsive to a natural language query from multiple data feed services at once, natural language processing (NLP) may be performed on the natural language query to generate a semantic representation, referred to herein as a "data feed-agnostic aggregator embedding." The data feed-agnostic aggregator embedding may abstractly represent the meaning contained in the natural language query. In some cases, the data feed-agnostic aggregator embedding may be a dense numerical representation, such as a vector of real numbers, that serves as an embedding in a continuous vector space.
[0014] Leveraging data feed-agnostic aggregator embedding to obtain information from multiple sources at once, multiple data feed services may be selected that potentially contain information responsive to the individual's query. Data feed services may be selected on a query-by-query basis and / or across multiple different queries. Potentially responsive data feed services may be selected in various ways based on various signals.
[0015] In some implementations, a data feed service may be selected based on a lookup table associated with the individual issuing the query. As one example, an individual may have a contact list (e.g., a phone book, friends on social media, etc.) that specifies which data feed services are used by the individual's contacts and which data feed service is selected accordingly. As another example, an individual may provide explicit permission to what is referred to herein as an "aggregator agent" to obtain responsive information from an enumerated or compiled list of data feed services. As yet another example, an individual's browsing history may be examined to determine which data feed services the individual tends to explore generally and / or in a particular context, with or without input from the individual. For example, if an individual issues a sports-related query, a sports-related data feed service that the individual has previously visited may be selected. In some implementations, a data feed service may be selected based on an entity (person, place, or thing) identified in the query. For example, an individual may request information about ongoing events communicated (e.g., published, posted, composed) by a particular media commentator. The data feed service used by that particular media commentator may be selected.
[0016] Each data feed service may include its own data feed service action space, which includes actions that can be performed to access data communicated through that data feed service. Actions in such an action space may include, for example, actions that can be performed using input devices, such as a keyboard and pointer device, to navigate a graphical user interface (GUI) provided by the data feed service. For example, if the GUI takes the form of an interactive web page, the action space may include actions that can be performed using graphical input, such as fields, pull-down menus, buttons, and the like, presented as part of the interactive web page. Additionally or alternatively, actions may include commands, queries, and / or parameters that can be used to navigate the GUI or VUI provided by the data feed service to obtain particular information. For example, the interface (GUI or VUI) of a particular data feed service may include a search field that facilitates the submission of natural language queries to quickly retrieve responsive content and / or a filter field for narrowing search results.
[0017] Domain-specific machine learning models configured with selected aspects of the present disclosure may be trained to transform between these action spaces and a data feed-independent semantic embedding space that includes the aforementioned data feed-independent aggregator embeddings. These domain-specific machine learning models (or simply "domain models" elsewhere herein) may take various forms, such as various types of neural networks, transformers, RNNs, graph-based neural networks, etc. In various implementations, the domain-specific machine learning models may be used to process the data feed-independent aggregator embeddings to generate one or more probability distributions over the actions in the data feed service's action space. Based on these probability distribution(s), an action may be selected and executed to retrieve content from the data feed service that is responsive to the query semantically represented by the data feed-independent aggregator embeddings. The response content from the data feed service may be aggregated with response content similarly retrieved from other data feed services. The aggregated response content may then be presented to the individual who issued the original natural language input.
[0018] As used herein, a "domain" may refer to a target subject area in which a computing component is intended to operate, e.g., the scope of knowledge, influence, and / or activity around which the computing component's logic revolves. In some implementations, the domain to which a query is submitted may be identified by heuristically matching keywords in a user-provided input with domain keywords. In other implementations, the user-provided input may be processed using NLP techniques, such as word2vec, bidirectional encoder representation with transformers (BERT), various types of recurrent neural networks ("RNNs," e.g., long short-term memories or "LSTMs," gated recurrent units or "GRUs"), etc., to generate semantic embeddings representing the natural language input. In some implementations, this natural language input semantic embedding (which may be referred to as a "data feed-independent aggregator embedding," as described above) may be used to identify one or more domains based on the distance(s) in the embedding space (or vector space) between the data feed-independent aggregator embedding and other embeddings associated with various domains. These distances in the embedding space can be computed using techniques such as Euclidean distance, dot product, cosine similarity, etc.
[0019] In various implementations, one or more domain models may be pre-generated for each domain. For example, one or more machine learning models, such as RNNs (e.g., LSTM, GRU), BERT transformers, various types of neural networks, reinforcement learning policies, etc., may be trained based on a corpus of documentation associated with the domain. As a result of this training, one or more of the domain model(s) may be at least bootstrapped so that they are usable to process what is referred to herein as a "domain-independent aggregator embedding," and may generate one or more probability distributions over an action space associated with the target domain. Based on these probability distribution(s), multiple actions can be selected and executed to execute queries submitted by users in the target domain.
[0020] FIG. 1 schematically illustrates an exemplary environment in which selected aspects of the present disclosure may be implemented, according to various implementations. Any computing device illustrated in FIG. 1 or elsewhere in the figure may include logic such as one or more microprocessors (e.g., central processing units or “CPUs,” graphical processing units or “GPUs,” tensor processing units (“TPUs”)) that execute computer-readable instructions stored in memory, or other types of logic such as application-specific integrated circuits (“ASICs”), field-programmable gate arrays (“FPGAs”), etc. Some of the systems illustrated in FIG. 1 , such as the cross-domain knowledge system 102, may be implemented using one or more server computing devices forming what is sometimes referred to as a “cloud infrastructure,” although this is not required. In other implementations, aspects of the cross-domain knowledge system 102 may be implemented on a client device 120, for purposes such as, for example, protecting privacy, reducing latency, etc.
[0021] The cross-domain knowledge system 102 may include several different components configured in accordance with selected aspects of the present disclosure, such as, for example, a domain module 104, an interface module 106, and a machine learning ("ML" in FIG. 1) module 108, to name a few. The cross-domain knowledge system 102 may also include any number of databases for storing machine learning model weights and / or other data used to perform selected aspects of the present disclosure. In FIG. 1, for example, the cross-domain knowledge system 102 includes a database 110 that stores a global domain model and another database 112 that stores data indicative of global action embeddings.
[0022] The cross-domain knowledge system 102 may be operatively coupled to any number of client computing devices operated by any number of users via one or more computer networks (114). In Figure 1, for example, a first user 118-1 operates a coordinated ecosystem of one or more client devices 120-1 (e.g., client devices controlled by user 118-1 and / or associated with user 118-1's online profile). A pth user 118-P operates one or more client device(s) 120-P. As used herein, client device(s) 120 may include, for example, one or more of a desktop computing device, a laptop computing device, a tablet computing device, a mobile phone computing device, a computing device in a user's vehicle (e.g., an in-vehicle communication system, an in-vehicle entertainment system, an in-vehicle navigation system), a standalone interactive speaker (which may in some cases include a visual sensor and / or a touchscreen display), a smart appliance such as a smart television (or a standard television with a networked dongle having automated assistant functionality), and / or a user wearable device that includes a computing device (e.g., a user's watch with a computing device, a user's glasses with a computing device, a virtual or augmented reality computing device). Additional and / or alternative client computing devices may be provided.
[0023] The domain module 104 may be configured to determine a variety of different information about domains (e.g., data feed services) associated with a given user 118 at a given time (e.g., the data feed service the user 118 is currently engaged with, the data feed service(s) the user would like to query for information, etc.) To this end, the domain module 104 may collect contextual information about, for example, foreground and / or background applications running on the client device(s) 120 operated by the user 118, web pages currently / recently visited by the user 118, the domain(s) to which the user 118 has access and / or frequently accessed domains, etc.
[0024] Using this collected context information, in some implementations, the domain module 104 may be configured to identify one or more domains (e.g., data feed services) associated with the natural language input provided by the user. For example, a user-composed query seeking response information from multiple different data feed services may be processed by the domain module 104 to identify the data feed service(s) that the user 118 intends to query.
[0025] In some implementations, the domain module 104 may also be configured to retrieve domain knowledge from various different sources related to the identified domain. In some such implementations, this retrieved domain knowledge (and / or embeddings generated therefrom) may be provided to downstream component(s), for example, in addition to the natural language input or contextual information described above. This additional domain knowledge may enable downstream component(s), particularly machine learning models, to be used to make predictions that are more likely to be satisfactory (e.g., aggregate response information across multiple different domains).
[0026] In some implementations, the domain module 104 can apply the collected context information (e.g., the current state) across one or more “domain selection” machine learning model(s) 105 that differ from the domain models described herein. These domain selection machine learning model(s) 105 can take various forms, such as various types of neural networks, support vector machines, random forests, BERT transformers, etc. In various implementations, the domain selection machine learning model(s) 105 may be trained to select applicable domains based on attributes (or “context signals”) of the current context or state of the user 118 and / or client device 120. For example, if the user 118 is interacting with an input form on a particular website to purchase a product or service, the website's uniform resource locator (URL), or attributes of the underlying webpage(s), such as keywords, tags, document object model (DOM) element(s), etc., in either their native form or a dimensionality-reduced embedding, can be applied as input across the model. Other context signals that may be considered include, but are not limited to, the user's IP address (e.g., work vs. home vs. mobile IP address), time of day, social media status, calendar, email / text messaging content, etc.
[0027] Interface module 106 may provide one or more GUIs or VUIs that can be operated by various individuals, such as users 118-1 through 118-P, to perform various actions made available by the semantic task automation system. In various implementations, user 118 may operate a GUI (e.g., a standalone application or a web page) provided by interface module 106 to select or utilize various techniques described herein. For example, users 118-1 through 118-P may be required to provide explicit permission for each data feed service (or, more generally, domain) they wish to query before their search requests can be used to retrieve response information from those data feed services.
[0028] The ML module 108 can access data representing various global domain / machine learning models / policies in the database 110. These trained global domain / machine learning models / policies can take a variety of forms, including, but not limited to, graph-based networks such as graph neural networks (GNNs), graph attention neural networks (GANNs), or graph convolutional neural networks (GCNs), sequence-to-sequence models such as encoder-decoders, various recurrent neural networks (e.g., LSTMs, GRUs, etc.), BERT transformer networks, reinforcement learning policies, and any other type of machine learning model that can be applied to facilitate selected aspects of the present disclosure. The ML module 108 may process various data based on these machine learning models at the request or command of other components, such as the domain module 104 and / or the interface module 106.
[0029] Each client device 120 may operate at least a portion of what is referred to herein as an “aggregator agent” 122. The aggregator agent 122 may be a computer application operable by the user 118 to perform selected aspects of the present disclosure to facilitate cross-domain data aggregation as described herein. For example, the aggregator agent 122 may receive a request and / or permission from the user 118 to aggregate query response data from multiple different domains. In some implementations, the aggregator agent 122 may be operable to grant access to the domain controlled by the user to others, e.g., other aggregator agents 122. For example, user 118-1 may interact with aggregator agent 122-1 to allow certain other individuals to access portions of the domain controlled by user 118-1, such as the user's own social networking profile feed. Without such explicit permission, other aggregator agents 122 may not be able to retrieve response information from the user's social networking profile feed.
[0030] In some implementations, the aggregator agent 122 may take the form of what is often referred to as a “virtual assistant” or “automated assistant” configured to engage in human-computer natural language interactions with the user 118. For example, the aggregator agent 122 may be configured to semantically process natural language input(s) provided by the user 118 to identify one or more intent(s). Based on these intent(s), the aggregator agent 122 may perform various tasks, such as operating a smart device, retrieving information, or performing a task. In some implementations, the interactions between the user 118 and the aggregator agent 122 (or a separate automated assistant accessible to / by the aggregator agent 122) may constitute a series of actions that can be captured and abstracted into domain-independent embeddings, as described herein, and then extended to other domains.
[0031] 1 , each of client device(s) 120-1 may include an aggregator agent 122-1 that provides services to a first user 118-1. The first user 118-1 and its aggregator agent 122-1 may access and / or associate with a “profile” for the first user 118-1 that includes various data relevant to performing selected aspects of the present disclosure. For example, the aggregator agent 122 may access one or more edge databases or data stores associated with the first user 118-1, including an edge database 124-1 that stores local domain model(s) and / or another edge database 126-1 that stores recorded actions. Other users 118-1 may have similar configurations. Any of the data stored in the edge databases 124-1 and 126-1 may be stored partially or entirely on the client device 120-1, for example, to protect the privacy of the first user 118-1. For example, the recorded actions 126-1 may include confidential and / or personal user information of the first user 118-1, such as payment information, address, phone number, etc., and may be stored locally in its raw form on the client device 120-1.
[0032] The local domain model(s) stored in edge database 124-1 may include, for example, local versions of the global model(s) stored in global domain model(s) database 110. For example, in some implementations, global models may be propagated to edges for the purpose of bootstrapping aggregator agents 122 to extend tasks to new domains associated with those propagated models, after which local models at the edges may or may not be trained locally based on activity and / or feedback of users 118. In some such implementations, local models (alternatively referred to as “local gradients” in edge database 124) may be periodically used to train the global model (in database 110), for example, as part of a federated learning framework. Because the global model is trained based on the local models, the global model may, in some cases, be propagated back to other edge databases (124), thereby keeping the local models up to date.
[0033] However, employing federated learning is not a requirement in all implementations. In some implementations, the aggregator agent 122 can provide scrubbed data to the cross-domain knowledge system 102, and the ML module 108 can remotely apply models to the scrubbed data. In some implementations, the "scrubbed" data can be data from which sensitive and / or personal information has been removed and / or obfuscated. In some implementations, personal information can be scrubbed at the edge, for example, by the aggregator agent 122, based on various rules. In other implementations, the scrubbed data provided by the aggregator agent 122 to the cross-domain knowledge system 102 can be in the form of dimensionality-reduced embeddings generated from raw data at the client device 120.
[0034] As previously mentioned, edge database 126-1 can store actions recorded by aggregator agent 122-1. Aggregator agent 122-1 can record actions in a variety of different manners depending on the access level of aggregator agent 122-1 to computer applications executing on client device 120-1 and the permissions granted by user 118. For example, most smartphones include operating system (OS) interfaces for providing or revoking permissions to various computer applications (e.g., location, camera access, etc.). In various implementations, such OS interfaces can be operable to provide / revoke access to aggregator agent 122 and / or to select the particular level of access aggregator agent 122 has to particular computer applications and / or domains to which those applications provide access.
[0035] Aggregator agent 122-1 may have varying levels of access to the workings of a computer application, depending on the permissions granted by user 118 and cooperation from the software developer providing the computer application. Some computer applications may provide aggregator agent 122 with “veiled” access to the application's API or to scripts written using a programming language (e.g., macros) embedded in the computer application, for example, with user 118's permission. Other computer applications may not provide as much access. In such cases, aggregator agent 122 may record actions in other ways, such as by capturing screenshots, performing optical character recognition (OCR) on those screenshots to identify menu items, and / or monitoring user input (e.g., interrupts captured by the OS) to determine which graphical elements were manipulated by user 118 and in what order. In some implementations, aggregator agent 122 may intercept actions performed using the computer application from data exchanged between the computer application and the underlying OS (e.g., via system calls). In some implementations, the aggregator agent 122 can intercept and / or access data exchanged between or used by the window manager and / or window system.
[0036] 2 is a schematic diagram illustrating an example of how data may be processed by and / or using various components across domains. Starting at the top left, user 118 operates client device 120 to provide a typed or spoken natural language query, NL QUERY1. In the latter case, the spoken utterance may first be processed using a speech-to-text (STT) engine (not shown) to generate a speech recognition output. In either case, NL QUERY1 may be provided to aggregator agent 122.
[0037] The aggregator agent 122 may process, or have the ML module 108 process, the data representing NL QUERY 1 to generate a data feed independent aggregator embedding (DAAE) Q1′. In some implementations, the aggregator agent 122 and / or the ML module 108 may use a machine learning model such as a transformer network or a recurrent neural network, e.g., an LSTM, a GRU, etc., to generate the DAAE Q1′. The DAAE Q1′ may then be processed, for example, by the aggregator agent 122 or the ML module 108 using a domain model B associated with the first data feed service and a domain model C associated with the second data feed service. Processing of the DAAE Q1′ using domain model B generates probability distribution(s), which may be used to select multiple actions {B1, B2, ...} from a first data feed service action space associated with the first data feed service. Similarly, processing DAAE Q1' using domain model C results in multiple actions {C1, C2, ...} being selected from a second data feed service action space associated with the second data feed service.
[0038] These selected actions {B1, B2, ...} and {C1, C2, ...} may be executed by various components to retrieve information responsive to the original query NL QUERY 1 from the respective domains of the first and second data feed services. In Figure 2, for example, the selected actions {B1, B2, ...} and {C1, C2, ...} are provided to client device 120 by aggregator agent 122. Client device 120 can then execute the selected actions, for example, by a corresponding application installed and running on client device 120.
[0039] For example, if a first data feed service associated with domain model B is a social networking service, a compatible social networking application (e.g., a client) running on client device 120 can automatically perform actions {B1, B2, ...} to retrieve response data from the social networking service. Similarly, if a second data feed service associated with domain model C is a blog curated by a second individual, a compatible client application (e.g., a web browser) running on client device 120 can automatically perform actions {C1, C2, ...} to retrieve response data from the blog. User 118 may or may not be able to see these automatically performed actions in a GUI rendered by the client application.
[0040] The quality and / or responsiveness of the aggregated information returned in response to a user's initial query may depend, at least in part, on the specificity of the query. Ambiguous queries may result in actions that do not truly fulfill the user's intent, for example, because the resulting aggregated information may have limited value and / or because the actions selected and executed for different data feed services may not match the user's intent, and the response information retrieved from the different data feed services may also not match the user's intent. In FIG. 2 , for example, if NL QUERY1 is vague and / or ambiguous, the selected actions {B1, B2, ...} and {C1, C2, ...} may obtain substantially different results from each data feed service. On the other hand, queries that are more detailed and clearly indicate the user's intent are likely to yield better results. However, if the burden on the user to be specific and clear is too great, the user may prefer to manually aggregate information from multiple different data feed services.
[0041] Thus, in some implementations, a user may be able to record actions that the user performs on a particular data feed service to obtain response information and associate those actions with natural language statements that the user provides that are customary to the user. In other words, the recorded actions, rather than the statements, contain contextual information that can be used to perform similar actions with other data feed services. In this way, natural language statements that otherwise lack detail (e.g., "long-tail" natural language statements) can be associated with specific actions that can be used to aggregate information from multiple different data feed services.
[0042] An example of this is illustrated in Figure 2. Below the dashed line, user 118 operates client device 120 to request and / or authorize recording of actions performed by user 118 using client device 120 in connection with another natural language query, NL QUERY2. In various implementations, aggregator agent 122 cannot record actions without receiving this authorization. In some implementations, this authorization may be granted per individual data feed service and / or per application, much like an application is granted permission to access GPS coordinates, local files, use an onboard camera, etc. In other implementations, this authorization may only be granted unless user 118 indicates otherwise, for example, by pressing a "stop recording" button or by providing a voice input such as "stop recording" or "end," similar to recording a macro.
[0043] Once the request / permission is received, in some implementations, aggregator agent 122 may acknowledge (ACK) the request / permission. Then, a sequence of actions {B3, B1, ...} and a sequence of actions {C5, C2, ...} performed by user 118 using client device 120 may be captured and stored in edge database 126. These actions {B3, B1, ...} and {C5, C2, ...} may take various forms or combinations of forms, such as, for example, command line input, and interactions with one or more graphical elements of a VUI or GUI using various types of input, such as pointer device (e.g., mouse) input, keyboard input, voice input, gaze input, speech input, and any other type of input capable of interacting with graphical elements of a VUI or GUI.
[0044] In various implementations, the domain(s) in which these actions are performed may be identified, for example, by the domain module 104 using any combination of, for example, NL QUERY 2, the computer application(s) operated by the user 118 to perform these actions, the remote data feed services (e.g., email, text messaging, social media) accessed by the user, the projects the user is working on, etc. In some implementations, the domain(s) may be identified at least in part by areas of a simulated digital world, sometimes referred to as the “metaverse,” that the user 118 virtually operates or visits. For example, the user 118 may record actions {B3, B1, ...} from a first metaverse game associated with domain model B that obtain scores for certain other users (e.g., their online gaming friends), brief video replays of those other users' performances, etc. Similarly, the user 118 may record actions {C5, C2, ...} from a second metaverse game associated with domain model C that obtains scores for the same other users, brief video replays of those other users' performances, etc.
[0045] Aggregator agent 122 may process actions {B3, B1, ...} using domain model B (or another domain model associated with the same domain) to generate action embedding B'. Similarly, aggregator agent 122 may process actions {C5, C2, ...} using domain model C (or another domain model associated with the same domain) to generate action embedding C'. As before, aggregator agent 122 (or another component, such as ML module 108) may process NL QUERY2 to generate another data feed-agnostic aggregator that embeds Q2'.
[0046] The aggregator agent 122 (or another component, such as the ML module 108) may then associate the embeddings Q2′, B′, and C′ within and / or across one or more embedding spaces using a variety of different techniques, such as triplet loss. In some implementations, the embeddings Q2′, B′, and C′ may be combined into a single data-feed-independent aggregator embedding via concatenation, averaging, etc. Regardless of how the embeddings Q2′, B′, and C′ are associated with each other or combined into a unified embedding, the user 118 or other users may be able to issue semantically similar natural language queries in the future. In some cases, those queries may be mapped to action embeddings B′ and / or C′, which can be utilized to aggregate data from multiple different data feed services, including data feed services other than those associated with domain models B and C. Notably, NL QUERY2 need not include details, and other semantically similar queries issued later need not include details.
[0047] In various implementations, simulations can be performed, for example, by aggregator agent 122 and / or components of cross-domain knowledge system 102, to further train the domain model. More specifically, various permutations of actions can be simulated to determine composite results. These composite results can be compared, for example, with the natural language input associated with the set of actions from which the simulated permutations were selected. The success or failure of these composite results can be used as positive and / or negative training examples for the domain model. In this way, it is possible to train a domain model based on much more than user-recorded actions and accompanying natural language input.
[0048] Figure 3, from a different perspective than Figure 2, schematically illustrates another example of how the techniques described herein may be used to aggregate information responsive to a user-issued query from multiple different data feed services. Starting at the bottom left, a user 118 operates a client device 120 (in this example, a standalone interactive speaker) and speaks the natural language command, "What are the commentators and my work colleagues saying about last night's playoff game?"
[0049] The STT module 330 can perform STT processing to generate speech recognition output. The speech recognition output can be processed by a natural language processing (NLP) module 332, for example, using machine learning model(s) (e.g., transformer(s), RNN(s), etc.) to generate data feed-agnostic aggregator embeddings 334. An embedding finder ("EF" in FIG. 3 ) module 336 can map or project the data feed-agnostic aggregator embeddings 334 to existing data feed-agnostic aggregator embeddings (white stars) in an embedding space 338. In various implementations, the STT module 330, the NLP module 332, and / or the EF module 336 can be realized as part of the aggregator agent 122, as part of the cross-domain knowledge system 102, or any combination thereof.
[0050] The embedding space 338 may be a continuous space containing multiple data feed-agnostic aggregator embeddings, each represented by a black dot in FIG. 3. These data feed-agnostic aggregator embeddings may be abstractions of previous natural language queries and, where applicable, domain-specific actions recorded in various domain action spaces. The embedding space 338 is shown as two-dimensional for purposes of illustration and understanding only. It should be understood that the embedding space 338 may actually have as many dimensions as the individual embeddings, and may be hundreds or thousands of dimensions.
[0051] The white star represents the coordinate in the action embedding space 338 associated with the data feed-agnostic aggregator embedding 334. As can be seen in FIG. 3 , this white star is actually between two data feed-agnostic aggregator embeddings surrounded by oval 340. In some implementations, multiple embeddings may match a single natural language input, for example, because the multiple embeddings are semantically similar to one another. In some implementations, multiple matching data feed-agnostic aggregator embeddings, such as the two in oval 340, may be combined into a unified representation, for example, via concatenation or averaging, and the unified data feed-agnostic aggregator embedding may be processed by downstream components.
[0052] The aggregator agent 122 may then process, or already processes, data feed-agnostic aggregator embedding(s) using multiple domain models A-C, each associated with a different domain from which the user 118 wishes to aggregate information. Domain A may represent, for example, a social media data feed service provided by one or more social media servers 342A. Domain B may represent, for example, a sports data feed service (e.g., a sports-centric website) served by one or more servers 342B. Domain C may represent, for example, a microblogging data feed service served by one or more servers 342C. One or more of the servers 342A-C may or may not be part of a cloud infrastructure and thus may not necessarily be tied to a particular server instance.
[0053] Processing the selected action embedding(s) based on domain model A may generate probability distribution(s) over the actions in the applicable action space. Based on those probability distribution(s), actions {A1, A2, ...} may be selected in the same manner as described above. Similarly, processing the selected action embedding(s) based on domain models B and C may result in the selection of actions {B1, B2, ...} and {C1, C2, ...}, respectively. These actions may be executed in their respective domains, for example, by servers 342A-C and / or by compatible client application(s) executing on client device 120.
[0054] As a result, the social media server 342A can retrieve and return (e.g., via aggregator agent 122) to, for example, client device 120, the most recent social media posts from anyone included in the list of media commentators and / or work colleagues compiled by user 118. In some implementations, the user's own default social media feed, in which posts from the user's friends and / or people the user follows appear in reverse chronological order, may be searched for responsive content posted by commentators and / or the user's work colleagues. Alternatively, the personal feeds of each of the commentators and / or the user's work colleagues may be searched for responsive content. In various implementations, the manner in which a social media data feed (or any other data feed) is searched may depend on permissions granted by the entity providing the data feed. For example, a social media data feed service may offer an API that can be tapped into to obtain responsive information. Alternatively, if the social media data feed service so desires, the aggregator agent 122 may prevent retrieval of the response content by using terms of use and / or technological measures such as a completely automated public Turing test to tell computers and humans apart (CAPTCHA).
[0055] The sports data feed server(s) 342B may, for example, retrieve and return to the client device 120 (e.g., via the aggregator agent 122) the most recent content provided by any of the commentators. While none of the coworkers may publish content to such a sports website, they may be able to post comments to the sports website, for example, at the bottom of articles, in a comment section. In some such cases, the comments of those coworkers may be searched for response data using the techniques described herein. The microblog server(s) 342C may, for example, retrieve and return to the client device 120 (e.g., via the aggregator agent 122) the most recent microblogging posts by any of the identified commentators or coworkers.
[0056] In some implementations, all of these returned messages may be collated and / or aggregated and presented audibly or visually to user 118. In other implementations, these returned messages may be compared to identify the most recent one, and only that message may be presented to user 118. For example, if client device 120 is a standalone interactive speaker without a display capability, as in FIG. 3 (or, for example, as in a vehicle), it may be advantageous to minimize the amount of output to avoid overwhelm or distracting user 118, in which case the most recent content of any of the domains may be read aloud. Additionally or alternatively, the returned content may be processed, for example, by ML module 108, using a sequence-to-sequence machine learning model trained to paraphrase and / or summarize longer text content, so that the content ultimately presented to user 118 is shorter and / or more concise.
[0057] 4 is a flowchart illustrating an example method 400 for implementing selected aspects of the present disclosure, according to implementations disclosed herein. For convenience, the operations of the flowchart are described with reference to a system that performs the operations. This system may include various components of various computer systems, such as one or more components of the cross-domain knowledge system 102. Furthermore, although the operations of the method 400 are shown in a particular order, this is not meant to be limiting. One or more operations may be rearranged, omitted, or added.
[0058] In block 402, the system may obtain natural language input containing a query for information. The user 118 may type such natural language and / or provide spoken terms, which may be processed, for example, by the STT module 330, to generate a speech recognition output. In some cases, the natural language input may identify one or more entities relevant to the search. These entities may include, for example, individuals who post content to data feed services. These individuals may be friends, commentators, academics, journalists, politicians, celebrities, or other public or private figures. Additionally or alternatively, the natural language input may identify one or more data feed services to be searched. For example, an individual may request content from all of their subscribed social media feeds.
[0059] Whatever the form of the natural language input, in block 404, natural language processing (NLP) can be performed on the natural language input, for example, by the ML module 108, to generate data feed-agnostic aggregator embeddings. This NLP can be performed using a machine learning model, such as an RNN or a transformer. As previously explained, when the natural language input contains sufficient detail, the data feed-agnostic aggregator embeddings may themselves be sufficient to be processed using a domain model to generate probability distributions. Alternatively, the data feed-agnostic aggregator embeddings may be relatively ambiguous, as long as they are semantically similar to previous data feed-agnostic aggregator embeddings generated from “long tail” or other low-detail natural language queries that are mapped to (or combined with) action embeddings generated by the same or different users.
[0060] In block 406, the system may select, by, for example, the aggregator agent 122 or the domain module 104, multiple data feed services for which content should be aggregated. The data feed services may be selected based on various signals or factors, such as the content of the natural language input, synonyms of tokens in the natural language input, the user's context (e.g., at work, driving a vehicle, time of day), permissions granted by the user (e.g., the user may have already selected the aggregator agent 122 for various selected data feed services), the user's contact list (which may indicate which data feed services the user's friends push content to), etc. In various implementations, each selected data feed service may include its own data feed service action space of actions that can be performed to access data communicated through the respective data feed service. These actions may include, for example, actions that can be performed in a client application (e.g., actions that can be performed using voice, a keyboard, a pointing device, etc.) or actions that can be performed on the server side, e.g., actions that can be performed via an API call.
[0061] In block 408, the system may process the feed-agnostic aggregator embeddings, for example by the ML module 108 or the aggregator agent 122, based on multiple domain-specific machine learning models corresponding to the multiple data feed services selected in block 406. Each domain-specific machine learning model may be trained to transform between the respective data feed service action space and the data feed-agnostic semantic embedding space (e.g., 338 in FIG. 3 ), which includes the data feed-agnostic aggregator embeddings.
[0062] Based on the processing of block 408, in block 410, the system may select and execute one or more actions from each of the data feed service action spaces to aggregate data responsive to the query from multiple data feed services. For example, the processing of block 408 may generate, for each applicable action space, a probability distribution over the actions in that space. The action with the highest probability may be executed first, and the action with a lower probability (but still above some minimum threshold) may be executed next. In some cases, the same domain model may be iteratively applied to a sequence of states. This sequence of states may represent, for example, the evolving state of a client application communicating with the data feed service, the evolving state of an interaction or exchange with the data feed service, the changing context of a user or the client device on which they operate, etc. The iteratively applied domain model may be trained using reinforcement learning in some implementations, although this is not required.
[0063] In block 412, the system may cause the aggregated response data to be presented as an output, such as by the interface module 106, audibly, on a screen, etc.
[0064] 5 is a block diagram of an exemplary computing device 510 that may optionally be utilized to perform one or more aspects of the techniques described herein. In some implementations, one or more of client computing devices 120-1 through 120-P, cross-domain knowledge system 102, and / or other component(s) may comprise one or more components of exemplary computing device 510.
[0065] Computing device 510 typically includes at least one processor 514 that communicates with several peripheral devices via a bus subsystem 512. These peripheral devices may include, for example, a storage subsystem 524 including a memory subsystem 525 and a file storage subsystem 526, a user interface output device 520, a user interface input device 522, and a network interface subsystem 516. The input and output devices enable a user to interact with computing device 510. Network interface subsystem 516 provides an interface to external networks and is coupled to corresponding interface devices in other computing devices.
[0066] The user interface input devices 522 may include pointing devices such as a keyboard, a mouse, a trackball, a touchpad, or a graphics tablet, a scanner, a touchscreen integrated into a display, a voice input device such as a voice recognition system, a microphone, and / or other types of input devices. In general, use of the term "input device" is intended to include all possible types of devices and methods for inputting information into the computing device 510 or a communications network.
[0067] The user interface output devices 520 may include a display subsystem, a printer, a fax machine, or a non-visual display such as an audio output device. The display subsystem may include a cathode ray tube (CRT), a flat panel device such as a liquid crystal display (LCD), a projection device, or some other mechanism for creating a visible image. The display subsystem may also provide a non-visual display, such as via an audio output device. In general, use of the term "output device" is intended to include all possible types of devices and methods for outputting information from the computing device 510 to a user or to another machine or computing device.
[0068] Storage subsystem 524 stores programming and data structures that provide the functionality of some or all of the modules described herein. For example, storage subsystem 524 may include logic for performing selected aspects of method 400 of FIG. 4.
[0069] These software modules are generally executed by the processor 514 alone or in combination with other processors. The memory 525 used within the storage subsystem 524 may include several memories, including a main random access memory (RAM) 530 for storing instructions and data during program execution, and a read-only memory (ROM) 532 in which fixed instructions are stored. The file storage subsystem 526 may provide persistent storage for program and data files and may include a hard disk drive, a floppy disk drive with associated removable media, a CD-ROM drive, an optical drive, or a removable media cartridge. Modules that perform the functions of a particular implementation may be stored by the file storage subsystem 526, within the storage subsystem 524, or on another machine accessible by the processor(s) 514.
[0070] The bus subsystem 512 provides a mechanism for allowing the various components and subsystems of the computing device 510 to communicate with each other as intended. Although the bus subsystem 512 is shown schematically as a single bus, alternative implementations of the bus subsystem may use multiple buses.
[0071] Computing device 510 can be of various types, including a workstation, a server, a computing cluster, a blade server, a server farm, or any other data processing system or computing device. Due to the ever-changing nature of computers and networks, the description of computing device 510 shown in Figure 5 is intended only as a specific example to illustrate some implementations. Many other configurations of computing device 510 are possible, having more or fewer components than the computing device shown in Figure 5.
[0072] While several implementations have been described and illustrated herein, various other means and / or structures can be utilized to perform the functions and / or obtain one or more of the results and / or advantages described herein, and each such variation and / or modification is considered to be within the scope of the implementations described herein. More generally, all parameters, dimensions, materials, and configurations described herein are intended to be exemplary, and the actual parameters, dimensions, materials, and / or configurations will depend on the particular application or applications in which the present teachings are used. Those skilled in the art will recognize, or be able to ascertain using no more than routine experimentation, many equivalents to the specific implementations described herein. It should therefore be understood that the above-described implementations are presented by way of example only, and that, within the scope of the appended claims and their equivalents, implementations may be practiced otherwise than as specifically described and claimed. Implementations of the present disclosure are directed to each individual feature, system, article, material, kit, and / or method described herein. Furthermore, any combination of two or more such features, systems, articles, materials, kits, and / or methods is included within the scope of the present disclosure, if such features, systems, articles, materials, kits, and / or methods are not mutually inconsistent.
Claims
1. A method implemented using one or more processors, comprising: obtaining a natural language input comprising an information query; performing natural language processing (NLP) on the natural language input to generate a data feed agnostic aggregator embedding; selecting a plurality of data feed services, each of the plurality of data feed services to be selected including its own data feed service action space, each data feed service action space including actions that can be performed to access data communicated via the respective data feed service; processing the feed-agnostic aggregator embeddings based on a plurality of domain-specific machine learning models corresponding to the plurality of data feed services, each domain-specific machine learning model trained to transform between a respective data feed service action space and a data feed-agnostic semantic embedding space that includes the data feed-agnostic aggregator embedding; selecting and executing one or more actions from each of the data feed service action spaces based on the processing to aggregate data responsive to the query from the plurality of data feed services; presenting the aggregated response data as an output; and A method comprising:
2. The method of claim 1 , wherein the plurality of data feed services are selected based on an entity identifier included in the query for information.
3. 3. The method of claim 1 or 2, wherein the query for information includes a request for social media posts from a particular individual, and the selected plurality of data feed services includes two or more social media services, and the two or more social media services are selected based on the particular individual's membership with the two or more social media services.
4. 4. The method of claim 3, wherein for a given social media service among the two or more social media services, the one or more actions selected from the data feed service action space of the given social media service include accessing one or more posts by the particular individual from the particular individual's posting history.
5. 5. The method of claim 3 or 4, wherein, for a given social media service among the two or more social media services, the one or more actions selected from the data feed service action space of the given social media service include filtering one or more posts by the particular individual from a general data feed on the given social media service provided to the user who issued the natural language input.
6. The method of any one of claims 1 to 5, wherein the plurality of data feed services are selected based on a look-up table controlled by a user who issued the natural language input.
7. The method of claim 6 , wherein the lookup table includes a contact list of the user.
8. 8. The method of claim 1, wherein the plurality of data feed services are selected based on a user who issued the natural language input having previously provided an aggregator agent with permission to access the plurality of data feed services.
9. The method of any preceding claim, wherein at least one of the data feed services comprises a virtual space that forms part of a larger metaverse comprising multiple virtual spaces.
10. The method of any preceding claim, wherein the aggregated response data is presented to the user who issued the natural language input as part of a metaverse graphical user interface.
11. The method of any one of claims 1 to 10, wherein the plurality of data feed services are selected based on a browsing history of a user who submitted the natural language input.
12. 1. A system comprising one or more processors and a memory storing instructions, the instructions causing the one or more processors, in response to execution of the instructions, to: receiving natural language input containing an information query; performing natural language processing (NLP) on the natural language input to generate a data feed agnostic aggregator embedding; selecting a plurality of data feed services, each of the plurality of data feed services to be selected including its own data feed service action space, each data feed service action space including actions that can be performed to access data communicated via the respective data feed service; processing the feed-agnostic aggregator embeddings based on a plurality of domain-specific machine learning models corresponding to the plurality of data feed services, each domain-specific machine learning model trained to transform between a respective data feed service action space and a data feed-agnostic semantic embedding space that includes the data feed-agnostic aggregator embedding; selecting and executing one or more actions from each of the data feed service action spaces to aggregate data responsive to the query from the plurality of data feed services; presenting the aggregated response data as an output; system.
13. The system of claim 12 , wherein the plurality of data feed services are selected based on an entity identifier included in the query for information.
14. 14. The system of claim 12 or 13, wherein the query for information includes a request for social media posts from a particular individual, and the selected plurality of data feed services includes two or more social media services, and the two or more social media services are selected based on the particular individual's membership with the two or more social media services.
15. 15. The system of claim 14, wherein for a given social media service of the two or more social media services, the one or more actions selected from the data feed service action space of the given social media service include accessing one or more posts by the particular individual from the particular individual's posting history.
16. 16. The system of claim 14 or 15, wherein for a given social media service among the two or more social media services, the one or more actions selected from the data feed service action space of the given social media service include filtering one or more posts by the particular individual from a general data feed on the given social media service provided to the user who issued the natural language input.
17. The system of any one of claims 12 to 16, wherein the plurality of data feed services are selected based on a lookup table controlled by a user who issued the natural language input.
18. 20. The system of claim 17, wherein the lookup table includes the user's contact list.
19. The system of any one of claims 12 to 16, wherein the plurality of data feed services are selected based on a user who issued the natural language input having previously provided an aggregator agent with permission to access the plurality of data feed services.
20. A non-transitory computer-readable medium containing instructions that, upon execution of the instructions by a processor, cause the processor to: receiving natural language input containing an information query; performing natural language processing (NLP) on the natural language input to generate a data feed agnostic aggregator embedding; selecting a plurality of data feed services, each of the plurality of data feed services to be selected including its own data feed service action space, each data feed service action space including actions that can be performed to access data communicated via the respective data feed service; processing the feed-agnostic aggregator embeddings based on a plurality of domain-specific machine learning models corresponding to the plurality of data feed services, each domain-specific machine learning model trained to transform between a respective data feed service action space and a data feed-agnostic semantic embedding space that includes the data feed-agnostic aggregator embedding; selecting and executing one or more actions from each of the data feed service action spaces to aggregate data responsive to the query from the plurality of data feed services; presenting the aggregated response data as an output; Non-transitory computer-readable medium.
Citation Information
Patent Citations
Generating recommendations using deep learning models
JP2020503590A
Apparatus and method for performing auxiliary functions for natural language queries
JP2022079728A
System for creating and method for providing a news feed website and application
US20130041893A1