Natural language understanding based meta-voice system using assistant system improves voice recognition accuracy

By using multiple automatic speech recognition engines and meta-speech engine selection strategies, the problem of inaccurate transcription caused by insufficient training data was solved, achieving high-precision transcription and fast response of audio input.

CN114600099BActive Publication Date: 2025-12-16CTRL-LABS CORP
View PDF 18 Cites 0 Cited by

Patent Information

Application Number
CN202080072956.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2020-01-13
Filing Date
2020-09-26
Publication Date
2025-12-16
Estimated Expiration
2040-09-26

AI Technical Summary

Technical Problem

Existing automatic speech recognition systems struggle to achieve high-precision transcription of audio input when training data is insufficient, leading to inaccurate processing of user requests.

Method used

Multiple automatic speech recognition engines are employed, and specific training data for each engine is used to select the most suitable combination of intent and slot through the meta-speech engine to generate a response to improve transcription accuracy.

Benefits of technology

By working in tandem with multiple engines, the transcription accuracy of audio input is improved, the time required to train the speech model is reduced, and the operability and accuracy of the system are enhanced.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114600099B_ABST
    Figure CN114600099B_ABST
Patent Text Reader

Abstract

In one embodiment, a method includes receiving, from a client system associated with a first user, a first audio input. The method includes generating, based on a plurality of automatic speech recognition (ASR) engines, a plurality of transcriptions corresponding to the first audio input. Each ASR engine is associated with a respective domain of a plurality of domains. The method includes determining, for each transcription, a combination of one or more intents and one or more slots associated with the transcription. The method includes selecting, by a meta speech engine, one or more combinations of intents and slots associated with the first user input from the plurality of combinations. The method includes generating a response to the first audio input based on the selected combination and sending, to the client system, instructions for presenting the response to the first audio input.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention generally relates to database and file management in a network environment, and more specifically to hardware and software for smart assistant systems.

[0002] background

[0003] Assistant systems can provide information or services to users based on a combination of user input, location awareness, and the ability to access information from various online sources, such as weather conditions, traffic congestion, news, stock prices, user schedules, retail prices, etc. User input can include text (e.g., online chat) (especially in instant messaging applications or other applications), voice, images, motion, or a combination thereof. Assistant systems can perform concierge-type services (e.g., booking dinner, purchasing event tickets, arranging trips) or provide information based on user input. Assistant systems can also perform administrative or data processing tasks based on online information and events without user initiation or interaction. Examples of tasks that can be performed by an assistant system include schedule management (e.g., sending an alert to a user about being late for a dinner appointment due to traffic, updating both parties' schedules, and changing restaurant reservation times). Assistant systems can be implemented through a combination of computing devices, application programming interfaces (APIs), and application proliferation on user devices.

[0004] Social networking systems, including social networking websites, enable their users (such as individuals or organizations) to interact with them and with each other. Social networking systems can utilize user input to create and store user profiles associated with each user. User profiles can include the user's demographic information, communication channels, and information about their personal interests. Social networking systems can also use user input to create and store records of a user's relationships with other users on the social network, and provide services (such as profile / news feed posts, photo sharing, event organization, messaging, games, or advertising) to facilitate social interaction between or among users.

[0005] Social networking systems can send content or messages related to their services to users' mobile devices or other computing devices via one or more networks. Users can also install software applications on their mobile devices or other computing devices to access their user profiles and other data within the social networking system. Social networking systems can generate a set of personalized content objects to display to users, such as a newsfeed that aggregates stories of other users connected to that user. Invention Overview

[0007] In certain embodiments, the assistant system of the present invention can assist users in obtaining information or services. The assistant system enables users to interact with it in stateful and multi-turn conversations using multimodal user input (e.g., voice, text, images, video, motion). By way of example and not limitation, the assistant system may support audio (verbal) input and non-verbal input, such as visual, location, gesture, motion, or mixed / multimodal input. The assistant system can create and store user profiles that include personal and contextual information associated with the user. In certain embodiments, the assistant system can use natural language understanding to analyze user input. The analysis can be based on the user's user profile to obtain a more personalized and context-aware understanding. The assistant system can use the analysis to resolve entities associated with the user input. In certain embodiments, the assistant system can interact with different agents to obtain information or services associated with the resolved entities. The assistant system can generate responses for the user regarding information or services using natural language generation. Through interaction with the user, the assistant system can use dialogue management techniques to manage and advance the conversation flow with the user. In certain embodiments, the assistant system can also assist users in effectively and efficiently processing the information they receive by summarizing it. The assistant system can also help users better participate in online social networks by providing tools to help them interact with these networks (e.g., creating posts, comments, and messages). Additionally, the assistant system can help users manage different tasks, such as continuously tracking events. In certain embodiments, the assistant system can proactively perform tasks related to the user's interests and preferences based on the user's profile at user-relevant times without user input. In certain embodiments, the assistant system can check privacy settings to ensure that access to the user's profile or other user information and the execution of different tasks are permitted according to the user's privacy settings.

[0008] In a particular embodiment, the assistant system can assist the user through a hybrid architecture built on client-side and server-side processes. The client-side and server-side processes can be two parallel workflows for processing user input and providing assistance to the user. In a particular embodiment, the client-side process can execute locally on a client system associated with the user. In contrast, the server-side process can execute remotely on one or more computing systems. In a particular embodiment, an arbitrator on the client system can coordinate the reception of user input (e.g., audio signals), determine whether to respond to the user input using a client-side process, a server-side process, or both, and analyze the processing results from each process. Based on the above analysis, the arbitrator can instruct a client-side or server-side agent to perform the task associated with the user input. The execution result can then be further rendered as output to the client system. By utilizing client-side and server-side processes, the assistant system can effectively help users optimize the use of computing resources while protecting user privacy and enhancing security.

[0009] In a particular embodiment, the assistant system may utilize multiple Automatic Speech Recognition (ASR) engines to analyze audio input via a meta-speech engine. For an ASR engine to operate with sufficient accuracy, it may require a large amount of training data to build the basis of a speech model corresponding to the ASR engine. As an example, and not a limitation, a large amount of training data could include 100,000 audio inputs and their respective transcriptions. However, when initially training the speech model, there may not be enough training data to build a sufficiently operable speech model. That is, as an example, and not a limitation, the speech model may not have enough training data to accurately generate transcriptions for at least 95% of the audio inputs. This situation can be identified when the user needs to repeat requests and / or when errors occur in the generated transcriptions. Therefore, the speech model may require an even larger amount of training data to accurately generate transcriptions for a threshold number of audio inputs. On the other hand, there may exist ASR engines that have already been trained on a finite dataset associated with a finite task set, operating with sufficient accuracy for that finite task set (e.g., 95% of the audio inputs are accurately transcribed). As an example, and not a limitation, there can be an ASR engine for messaging / calling, an ASR engine for music-related functions, and an ASR engine for default system operations. For example, an ASR engine for messaging / calling can accurately transcribe a threshold number of audio inputs associated with a messaging / call request (e.g., 95%). Therefore, the assistant system can utilize individual ASR engines to improve the accuracy of ASR results. To do this, the assistant system can receive audio input and send it to multiple ASR engines. By sending the audio input to multiple ASR engines, each ASR engine can generate a transcription based on its corresponding speech model. This improves the accuracy of ASR results by increasing the probability of accurate transcription of the audio input. As an example, and not a limitation, if a user requests to play music using the assistant system, the audio input is sent to all available ASR engines, one of which is an ASR engine for music-related functions. The ASR engine for music-related functions can accurately transcribe the audio input into a request to play music. By sending the audio input corresponding to a request to play music to an ASR engine designed for music-related functions, the assistant system can improve the accuracy of audio input transcription because a particular ASR engine can have a large amount of training data corresponding to audio inputs associated with music-related functions. By using multiple ASR engines, the assistant system can have a robust foundation of speech models to handle and transcribe different requests, such as music-related requests or messaging / call-related requests. Using multiple ASR engines—each associated with its own speech model—can help avoid the need to extensively train a single speech model on large datasets, such as training a single speech model to transcribe requests to play music as well as requests to transcribe weather information.These individual ASR engines may have already been trained for their respective functions and may not require any further training to achieve sufficient accuracy for their respective functions to operate. Therefore, using individual ASR engines can reduce or eliminate the time required to train speech models to achieve sufficient operability when processing / transcription of audio inputs associated with a wide range of functions (e.g., the ability to accurately transcribe a threshold number of audio inputs). The assistant system can acquire the output of the ASR engines and send it to a Natural Language Understanding (NLU) module, which identifies one or more intents and one or more slots associated with each output of the respective ASR engine. The output of the NLU module can be sent to a meta-speech engine, which can select one or more intents and one or more slots associated with the received audio input from possible choices based on a selection strategy. The selection strategy can be, for example, a simple selection strategy, a system combination of strategies with machine learning models, or a ranking strategy. After determining the intents and slots, the meta-speech engine can send its output for processing by one or more agents of the assistant system. In a particular embodiment, the meta-speech engine can be a component of the assistant system that processes the audio input to generate combinations of intents and slots. In a particular embodiment, the meta-speech engine may include multiple ASR engines and an NLU module. Although this disclosure describes the use of multiple ASR engines to analyze audio input in a specific manner, this disclosure contemplates the use of multiple ASR engines to analyze audio input in any suitable manner.

[0010] In a particular embodiment, the assistant system may receive a first audio input from a client system associated with a first user. In a particular embodiment, the assistant system may generate multiple transcriptions corresponding to the first audio input based on multiple Automatic Speech Recognition (ASR) engines. Each ASR engine may be associated with a corresponding domain among multiple domains. In a particular embodiment, the assistant system may determine a combination of one or more intents and one or more slots associated with each transcription. In a particular embodiment, the assistant system may select one or more combinations of intents and slots associated with the first user input from the multiple combinations using a meta-speech engine. In a particular embodiment, the assistant system may generate a response to the first audio input based on the selected combination. In a particular embodiment, the assistant system may send instructions to the client system for presenting a response to the first audio input.

[0011] In one aspect, the present invention provides a method comprising one or more computing systems:

[0012] Receive the first audio input from the client system associated with the first user;

[0013] Multiple transcriptions corresponding to the first audio input are generated based on multiple automatic speech recognition (ASR) engines, wherein each ASR engine is associated with a corresponding domain in multiple domains;

[0014] For each transcription, determine one or more intentions and one or more slot combinations associated with the transcription;

[0015] The meta-speech engine selects one or more combinations of intents and slots associated with the first audio input from multiple combinations;

[0016] Generate a response to the first audio input based on the selected combination; and

[0017] Send instructions to the client system to present a response to the first audio input.

[0018] Each ASR engine can be associated with one or more agents that are specific to that ASR engine.

[0019] Each of the multiple domains may include one or more agents specific to that domain. Agents may include one or more first-party agents or third-party agents.

[0020] Each of the multiple domains can include a set of tasks specific to that domain.

[0021] Multiple domains can be associated with multiple agents, and each agent can operate to perform one or more tasks specific to one or more domains.

[0022] The method may also include: for each combination of intent and slot, identifying a domain of a plurality of domains, wherein selecting one or more combinations of intent and slot includes mapping the domain of each combination of intent and slot to a domain associated with one of a plurality of ASR engines.

[0023] When the domain of the corresponding combination of intents and slots matches the domain of one of multiple ASR engines, one or more combinations of intents and slots can be selected.

[0024] Generating multiple transcriptions may include: sending a first audio input to each of the multiple ASR engines; and receiving multiple transcriptions from the multiple ASR engines.

[0025] One or more of the ASR engines in a plurality of ASR engines may be third-party ASR engines associated with a third-party system that is separate from and outside one or more computing systems. The method further includes: sending a first audio input to one of the third-party ASR engines to generate one or more transcripts; and receiving one or more transcripts generated by one of the third-party ASR engines from one of the third-party ASR engines, wherein generating the plurality of transcripts includes selecting one or more transcripts generated by the third-party ASR engines to determine a combination of intents and slots associated with each respective transcript.

[0026] The method may further include: identifying one or more features for each combination of intent and slot, wherein the one or more features indicate whether the combination of intent and slot has an attribute; and ranking the multiple combinations based on their respective identified features, wherein selecting one or more combinations of intent and slot includes selecting one or more combinations of intent and slot based on the ranking of the multiple combinations.

[0027] The method may further include: identifying one or more identical combinations of intents and slots from a plurality of combinations; and sorting one or more identical combinations of intents and slots based on the number of identical combinations of intents and slots, wherein selecting one or more combinations of intents and slots includes using the sorting of one or more identical combinations of intents and slots.

[0028] Generating a response to a first audio input may include: sending a selected combination to multiple agents; receiving multiple responses corresponding to the selected combination from multiple agents; sorting the multiple responses received from multiple agents; and selecting a response from the multiple responses based on the sorting of the multiple responses.

[0029] One of the multiple ASR engines can be a combination of two or more discrete ASR engines, and each of the two or more discrete ASR engines is associated with an independent domain in multiple domains.

[0030] The response may include the action to be performed or one or more of the results generated from the query.

[0031] Instructions used to present a response may include a notification of an action to be performed or a list of one or more results.

[0032] In another aspect, the invention provides one or more computer-readable non-transitory storage media comprising software that, when executed, is operable to:

[0033] Receive the first audio input from the client system associated with the first user;

[0034] Multiple transcriptions corresponding to the first audio input are generated based on multiple automatic speech recognition (ASR) engines, wherein each ASR engine is associated with a corresponding domain in multiple domains;

[0035] For each transcription, determine one or more intentions and one or more slot combinations associated with the transcription;

[0036] The meta-speech engine selects one or more combinations of intents and slots associated with the first user input from multiple combinations;

[0037] Generate a response to the first audio input based on the selected combination; and

[0038] Send instructions to the client system to present a response to the first audio input.

[0039] In another aspect, the present invention provides a system comprising: one or more processors; and a non-transitory memory coupled to the processors, comprising instructions executable by the processors, wherein, when the instructions are executed, the processors are operable to:

[0040] Receive the first audio input from the client system associated with the first user;

[0041] Multiple transcriptions corresponding to the first audio input are generated based on multiple automatic speech recognition (ASR) engines, wherein each ASR engine is associated with a corresponding domain in multiple domains;

[0042] For each transcription, determine one or more intentions and one or more slot combinations associated with the transcription;

[0043] The meta-speech engine selects one or more combinations of intents and slots associated with the first user input from multiple combinations;

[0044] Generate a response to the first audio input based on the selected combination; and

[0045] Send instructions to the client system to present a response to the first audio input.

[0046] The embodiments disclosed herein are merely examples, and the scope of the invention is not limited thereto. Specific embodiments may include all, some, or none of the components, elements, features, functions, operations, or steps of the embodiments disclosed herein. Embodiments of the invention are specifically disclosed in the appended claims and relate to methods, storage media, systems, and computer program products, wherein any feature mentioned in one claim class (e.g., methods) may also be claimed in another claim class (e.g., systems). Dependencies or backreferences in the appended claims are chosen only for formal reasons. However, any subject matter arising from intentional backreferences to any prior claim (particularly multiple dependencies) may also be claimed, thereby disclosing and claiming any combination of claims and their features, regardless of the dependencies chosen in the appended claims. Claimable subject matter includes not only combinations of features set forth in the appended claims but also any other combination of features in the claims, wherein each feature mentioned in the claims may be combined with any other feature or combination of features in the claims. Furthermore, any of the embodiments and features described or depicted herein may be claimed in a separate claim and / or in any combination with any embodiment or feature described or depicted herein or in any combination with any feature of the appended claims. Brief description of the attached diagram

[0048] Figure 1 An example network environment associated with the assistant system is shown.

[0049] Figure 2 An example architecture of the assistant system is shown.

[0050] Figure 3 An example flowchart of the server-side process of the assistant system is shown.

[0051] Figure 4 An exemplary flowchart is shown, illustrating how an assistant system processes user input.

[0052] Figure 5 An exemplary flowchart is shown that uses multiple automatic speech recognition engines to generate transcriptions of audio input.

[0053] Figure 6 An exemplary flowchart is shown, illustrating the use of multiple selection strategies to choose a combination of intents and slots to generate a response.

[0054] Figure 7 An example mapping of an automatic speech recognition engine to a domain and its corresponding agent and available tasks is shown.

[0055] Figure 8An example process for generating transcriptions for audio input using multiple automatic speech recognition engines is shown.

[0056] Figure 9 An example method for generating transcriptions for audio input using multiple automatic speech recognition engines is shown.

[0057] Figure 10 An example social graph is shown.

[0058] Figure 11 An example view of the embedded space is shown.

[0059] Figure 12 An example artificial neural network is shown.

[0060] Figure 13 An example computer system is shown.

[0061] Invention Description

[0062] System Overview

[0063] Figure 1 An example network environment 100 associated with an assistant system is shown. Network environment 100 includes client systems 130, assistant systems 140, social networking systems 160, and third-party systems 170 connected to each other via network 110. Although Figure 1 A specific arrangement of client system 130, assistant system 140, social networking system 160, third-party system 170, and network 110 is shown, but this disclosure contemplates any suitable arrangement of client system 130, assistant system 140, social networking system 160, third-party system 170, and network 110. By way of example and not limitation, two or more of client system 130, social networking system 160, assistant system 140, and third-party system 170 may bypass network 110 and be directly connected to each other. As another example, two or more of client system 130, assistant system 140, social networking system 160, and third-party system 170 may be physically or logically located in the same place as each other, wholly or partially. Furthermore, although... Figure 1 A specific number of client systems 130, assistant systems 140, social networking systems 160, third-party systems 170, and networks 110 are shown, but this disclosure contemplates any suitable number of client systems 130, assistant systems 140, social networking systems 160, third-party systems 170, and networks 110. As an example and not as a limitation, network environment 100 may include multiple client systems 130, assistant systems 140, social networking systems 160, third-party systems 170, and networks 110.

[0064] This disclosure contemplates any suitable network 110. By way of example and not limitation, one or more portions of network 110 may include an ad hoc network, intranet, extranet, virtual private network (VPN), local area network (LAN), wireless LAN (WLAN), wide area network (WAN), wireless WAN (WWAN), metropolitan area network (MAN), a portion of the Internet, a portion of the public switched telephone network (PSTN), a cellular telephone network, or a combination of two or more of these. Network 110 may include one or more networks 110.

[0065] Link 150 can connect client system 130, assistant system 140, social networking system 160, and third-party system 170 to communication network 110 or to each other. This disclosure contemplates any suitable link 150. In a particular embodiment, one or more links 150 include one or more wired (such as, for example, Digital Subscriber Line (DSL) or Cable-Based Data Service Interface Specification (DOCSIS)) links, wireless (such as, for example, Wi-Fi or Global Interoperability Microwave Access (WiMAX)) links, or optical (such as, for example, Synchronous Optical Network (SONET) or Synchronous Digital Hierarchy (SDH)) links. In a particular embodiment, one or more links 150 each include an ad hoc network, intranet, extranet, VPN, LAN, WLAN, WAN, WWAN, MAN, a portion of the Internet, a portion of the PSTN, a cellular-based network, a satellite-based network, another link 150, or a combination of two or more such links 150. Links 150 do not need to be identical throughout the network environment 100. One or more first links 150 may differ from one or more second links 150 in one or more respects.

[0066] In a particular embodiment, client system 130 may be an electronic device that includes hardware, software, or embedded logic components, or a combination of two or more such components, and is capable of performing appropriate functions implemented or supported by client system 130. By way of example and not limitation, client system 130 may include computer systems such as desktop computers, laptop or notebook computers, netbooks, tablet computers, e-book readers, GPS devices, cameras, personal digital assistants (PDAs), handheld electronic devices, cellular phones, smartphones, smart speakers, virtual reality (VR) headsets, augmented reality (AR) smart glasses, other suitable electronic devices, or any suitable combination thereof. In a particular embodiment, client system 130 may be a smart assistant device. More information about intelligent assistant devices can be found in U.S. Patent Application No. 15 / 949011, filed April 9, 2018; U.S. Patent Application No. 16 / 153574, filed October 5, 2018; U.S. Design Patent Application No. 29 / 631910, filed January 3, 2018; U.S. Design Patent Application No. 29 / 631747, filed January 2, 2018; U.S. Design Patent Application No. 29 / 631913, filed January 3, 2018; and U.S. Design Patent Application No. 29 / 631914, filed January 3, 2018, each of which is incorporated herein by reference. This disclosure contemplates any suitable client system 130. Client system 130 enables network users at client system 130 to access network 110. Client system 130 enables its users to communicate with other users at other client systems 130.

[0067] In a particular embodiment, client system 130 may include web browser 132 and may have one or more add-ons, plugins, or other extensions. A user at client system 130 may enter a Uniform Resource Locator (URL) or another address that directs web browser 132 to a specific server (e.g., server 162 or a server associated with third-party system 170), and web browser 132 may generate a Hypertext Transfer Protocol (HTTP) request and pass the HTTP request to the server. The server may accept the HTTP request and, in response, pass one or more Hypertext Markup Language (HTML) files to client system 130. Client system 130 may display a web interface (e.g., a webpage) based on the HTML files from the server for presentation to the user. This disclosure contemplates any suitable source file. By way of example and not limitation, a web interface may be displayed based on HTML files, Extensible Hypertext Markup Language (XHTML) files, or Extensible Markup Language (XML) files, depending on specific needs. Such an interface may also execute scripts, markup languages, and combinations thereof. In this article, references to a web interface include one or more corresponding source files (which the browser can use to display the web interface), and vice versa, where appropriate.

[0068] In a particular embodiment, client system 130 may include a social networking application 134 installed on client system 130. Users at client system 130 can use social networking application 134 to access online social networks. Users at client system 130 can use social networking application 134 to communicate with their social connections (e.g., friends, followers, followed accounts, contacts, etc.). Users at client system 130 can also use social networking application 134 to interact with multiple content objects (e.g., posts, news articles, temporary content, etc.) on the online social network. By way of example and not limitation, users can use social networking application 134 to browse trending topics and breaking news.

[0069] In a particular embodiment, client system 130 may include assistant application 136. A user of client system 130 may use assistant application 136 to interact with assistant system 140. In a particular embodiment, assistant application 136 may include a standalone application. In a particular embodiment, assistant application 136 may be integrated into social networking application 134 or another suitable application (e.g., messaging application). In a particular embodiment, assistant application 136 may also be integrated into client system 130, assistant hardware device, or any other suitable hardware device. In a particular embodiment, assistant application 136 may be accessed via web browser 132. In a particular embodiment, a user may provide input via different modalities. By way of example and not limitation, modalities may include audio, text, images, video, motion, orientation, etc. Assistant application 136 may transmit user input to assistant system 140. Based on the user input, assistant system 140 may generate a response. Assistant system 140 may send the generated response to assistant application 136. Assistant application 136 may then present the response to the user of client system 130. The presented response may be based on different modalities, such as audio, text, images, and video. As an example, and not a limitation, a user can verbally inquire about traffic information from the assistant application 136 by speaking into the microphone of the client system 130 (i.e., via audio modality). The assistant application 136 can then transmit the request to the assistant system 140. The assistant system 140 can generate a response accordingly and send it back to the assistant application 136. The assistant application 136 can also present the response to the user in text and / or images on the display of the client system 130.

[0070] In a particular embodiment, assistant system 140 can assist a user in retrieving information from different sources. Assistant system 140 can also assist a user in requesting services from different service providers. In a particular embodiment, assistant system 140 can receive a user's request for information or services via assistant application 136 in client system 130. Assistant system 140 can use natural language understanding to analyze the user's request based on the user profile and other relevant information. The results of the analysis may include different entities associated with online social networks. Assistant system 140 can then retrieve information or request services associated with these entities. In a particular embodiment, when retrieving information or requesting services for a user, assistant system 140 can interact with social network system 160 and / or third-party system 170. In a particular embodiment, assistant system 140 can use natural language generation technology to generate personalized communication content for the user. Personalized communication content may include, for example, the status of retrieved information or requested services. In a particular embodiment, assistant system 140 enables the user to interact with it regarding information or services in stateful and multi-turn conversations using dialogue management technology. (See below) Figure 2The discussion describes the functions of the assistant system 140 in more detail.

[0071] In a particular embodiment, the social networking system 160 may be a network-addressable computing system capable of hosting an online social network. The social networking system 160 may generate, store, receive, and transmit social network data (such as, for example, user profile data, concept profile data, social graph information, or other suitable data related to the online social network). The social networking system 160 may be accessed directly or via network 110 by other components of the network environment 100. By way of example and not limitation, the client system 130 may access the social networking system 160 directly or via network 110 using a web browser 132 or a native application associated with the social networking system 160 (e.g., a mobile social networking application, a messaging application, another suitable application, or any combination thereof). In a particular embodiment, the social networking system 160 may include one or more servers 162. Each server 162 may be a unitary server or a distributed server spanning multiple computers or multiple data centers. Server 162 can be of various types, for example and without limitation, web server, news server, mail server, message server, advertising server, file server, application server, exchange server, database server, proxy server, another server suitable for performing the functions or processes described herein, or any combination thereof. In a particular embodiment, each server 162 may include hardware, software, or embedded logic components, or a combination of two or more such components for performing suitable functions implemented or supported by server 162. In a particular embodiment, social networking system 160 may include one or more data stores 164. Data stores 164 can be used to store various types of information. In a particular embodiment, the information stored in data stores 164 may be organized according to a particular data structure. In a particular embodiment, each data store 164 may be a relational database, a columnar database, a relevance database, or other suitable database. Although this disclosure describes or illustrates specific types of databases, this disclosure contemplates any suitable type of database. Specific embodiments may provide an interface that enables client system 130, social network system 160, assistant system 140, or third-party system 170 to manage, retrieve, modify, add, or delete information stored in data storage 164.

[0072] In a particular embodiment, the social network system 160 may store one or more social graphs in one or more data stores 164. In a particular embodiment, the social graph may include multiple nodes—which may include multiple user nodes (each corresponding to a specific user) or multiple concept nodes (each corresponding to a specific concept)—and multiple edges connecting the nodes. The social network system 160 may provide users of the online social network with the ability to communicate and interact with other users. In a particular embodiment, a user may join an online social network via the social network system 160 and then add connections (e.g., relationships) with multiple other users in the social network system 160 that they wish to associate with. In this document, the term "friend" may refer to any other user of the social network system 160 with whom the user has formed an association or relationship via the social network system 160.

[0073] In a particular embodiment, social networking system 160 may provide users with the ability to take action on various types of items or objects supported by social networking system 160. By way of example and not limitation, items and objects may include groups or social networks to which a user of social networking system 160 can belong, events or calendar entries that a user may be interested in, computer-based applications that a user can use, transactions that allow a user to buy or sell goods via a service, interactions with advertisements that a user can perform, or other suitable items or objects. Users may interact with anything that can be represented in social networking system 160 or by an external system, third-party system 170, which is decoupled from social networking system 160 and coupled to social networking system 160 via network 110.

[0074] In a particular embodiment, the social networking system 160 is capable of linking various entities. By way of example and not limitation, the social networking system 160 may enable users to interact with each other and receive content from third-party systems 170 or other entities, or allow users to interact with these entities through application programming interfaces (APIs) or other communication channels.

[0075] In certain embodiments, third-party system 170 may include one or more types of servers, one or more data storage devices, one or more interfaces (including but not limited to APIs), one or more web services, one or more content sources, one or more networks, or any other suitable components (e.g., servers may communicate with these components). Third-party system 170 may be operated by an entity different from the entity operating social networking system 160. However, in certain embodiments, social networking system 160 and third-party system 170 may operate in combination to provide social networking services to users of either social networking system 160 or third-party system 170. In this sense, social networking system 160 may provide a platform or backbone that other systems (e.g., third-party system 170) can use to provide social networking services and functionality to users across the Internet.

[0076] In a particular embodiment, third-party system 170 may include a third-party content object provider. The third-party content object provider may include one or more sources of content objects that can be delivered to client system 130. As an example, and not a limitation, the content object may include information about things or activities that a user is interested in, such as movie showtimes, movie reviews, restaurant reviews, restaurant menus, product information and reviews, or other suitable information. As another example, and not a limitation, the content object may include incentivized content objects (e.g., coupons, discount vouchers, gift certificates, or other suitable incentives). In a particular embodiment, the third-party content provider may use one or more third-party proxies to deliver content objects and / or services. The third-party proxies may be implementations hosted and executed on third-party system 170.

[0077] In a particular embodiment, the social networking system 160 also includes user-generated content objects that can enhance user interaction with the social networking system 160. User-generated content can include any content that a user can add, upload, send, or "post" to the social networking system 160. As an example, and not as a limitation, a user transmits a post from client system 130 to social networking system 160. A post can include data such as status updates or other text data, location information, photos, videos, links, music, or other similar data or media. Content can also be added to the social networking system 160 by third parties via "communication channels" (such as news feeds or streams).

[0078] In certain embodiments, the social networking system 160 may include various servers, subsystems, programs, modules, logs, and data stores. In certain embodiments, the social networking system 160 may include one or more of the following: a web server, an action logger, an API request server, a relevance and ranking engine, a content object classifier, a notification controller, an action log, a third-party content object exposure log, an inference module, an authorization / privacy server, a search module, an advertising-targeting module, a user interface module, a user profile store, a relationship store, a third-party content store, or a location store. The social networking system 160 may also include suitable components such as network interfaces, security mechanisms, load balancers, failover servers, management and network operations consoles, other suitable components, or any suitable combination thereof. In certain embodiments, the social networking system 160 may include one or more user profile stores for storing user profiles. User profiles may include, for example, biographical information, demographic information, behavioral information, social information, or other types of descriptive information (e.g., work experience, educational history, hobbies or preferences, interests, affinity, or location). Interest information may include interests associated with one or more categories. Categories can be general or specific. As an example, and not a limitation, if a user “likes” an article about a brand of shoes, the category could be the brand, or the general category of “shoes” or “clothing.” A relationship store can be used to store relationship information about a user. Relationship information can indicate users who have similar or shared work experience, group memberships, hobbies, educational history, or are related or share common attributes in any way. Relationship information can also include user-defined relationships between different users and content (internal and external). A web server can be used to link the social networking system 160 to one or more client systems 130 or one or more third-party systems 170 via network 110. The web server can include a mail server or other messaging functionality for receiving and routing messages between the social networking system 160 and one or more client systems 130. An API request server can allow, for example, an assistant system 140 or a third-party system 170 to access information from the social networking system 160 by calling one or more APIs. An action recorder can be used to receive communications from the web server regarding user actions on or outside the social networking system 160. By combining action logs, a log of third-party content objects exposed to users can be maintained. The notification controller can provide information about the content objects to the client system 130. The information can be pushed to the client system 130 as a notification, or it can be retrieved from the client system 130 in response to a request received from the client system 130.An authorization server can be used to implement one or more privacy settings for users of the social networking system 160. A user's privacy settings determine how specific information associated with that user can be shared. The authorization server can, for example, allow users to opt in or out by setting appropriate privacy settings, enabling their actions to be recorded by the social networking system 160 or shared with other systems (e.g., third-party system 170). A third-party content object store can be used to store content objects received from third parties (e.g., third-party system 170). A location store can be used to store location information received from the client system 130 associated with the user. An advertising pricing module can combine social information, current time, location information, or other suitable information to provide relevant advertisements to users in the form of notifications.

[0079] Assistant System

[0080] Figure 2An example architecture of assistant system 140 is shown. In a particular embodiment, assistant system 140 can assist a user in obtaining information or services. Assistant system 140 enables users to interact with it in stateful and multi-turn sessions using multimodal user input (such as voice, text, images, video, motion) to obtain assistance. By way of example and not limitation, assistant system 140 may support audio input (verbal) and non-verbal input, such as visual, location, gesture, motion, or mixed / multimodal input. Assistant system 140 may create and store user profiles that include personal and contextual information associated with the user. In a particular embodiment, assistant system 140 may use natural language understanding to analyze user input. The analysis may be based on the user's user profile to obtain a more personalized and context-aware understanding. Assistant system 140 may use the analysis to resolve entities associated with the user input. In a particular embodiment, assistant system 140 may interact with different agents to obtain information or services associated with the resolved entities. Assistant system 140 may generate responses for the user regarding information or services by using natural language generation. Through interaction with the user, the assistant system 140 can use conversation management techniques to manage and forward the conversation flow with the user. In a particular embodiment, the assistant system 140 can also help the user effectively and efficiently process the information received by aggregating information. The assistant system 140 can also help the user participate more in online social networks by providing tools to help the user interact with online social networks (e.g., create posts, comments, messages). The assistant system 140 can also help the user manage different tasks, such as continuously tracking events. In a particular embodiment, the assistant system 140 can proactively perform pre-authorized tasks related to the user's interests and preferences based on the user's profile at user-related times, without user input. In a particular embodiment, the assistant system 140 can check privacy settings to ensure that access to the user's profile or other user information and the performance of different tasks are permitted according to the user's privacy settings. More information on assisting users subject to privacy settings can be found in U.S. Patent Application No. 16 / 182542, filed November 6, 2018.

[0081] In a particular embodiment, the assistant system 140 can assist the user through a hybrid architecture built on client-side and server-side processes. The client-side and server-side processes can be two parallel workflows for processing user input and providing assistance to the user. In a particular embodiment, the client-side process can execute locally on the client system 130 associated with the user. In contrast, the server-side process can execute remotely on one or more computing systems. In a particular embodiment, an assistant orchestrator on the client system 130 can coordinate the reception of user input (e.g., audio signals) and determine whether to respond to the user input using the client-side process, the server-side process, or both. A dialogue arbitrator can analyze the processing results from each process. Based on the above analysis, the dialogue arbitrator can instruct the client-side or server-side agent to perform the task associated with the user input. The execution result can then be further rendered as output to the client system 130. By utilizing both client-side and server-side processes, the assistant system 140 can effectively help the user optimize the use of computing resources while protecting user privacy and enhancing security.

[0082] In a particular embodiment, assistant system 140 may receive user input from client system 130 associated with a user. In a particular embodiment, user input may be user-generated input sent to assistant system 140 in a single round. User input may be verbal, nonverbal, or a combination thereof. By way of example, and not limitation, nonverbal user input may be based on the user's voice, vision, location, activity, gestures, actions, or a combination thereof. If user input is based on the user's voice (e.g., the user may speak to client system 130), such user input may first be processed by system audio API 202 (application programming interface). System audio API 202 may perform echo cancellation, noise removal, beamforming and self-activated user voice, speaker recognition, voice activity detection (VAD), and any other acoustic techniques to generate audio data that can be easily processed by assistant system 140. In a particular embodiment, system audio API 202 may perform wake word detection 204 based on user input. By way of example, and not limitation, a wake word may be "Hey, assistant." If such a wake word is detected, assistant system 140 may be activated accordingly. In an alternative embodiment, the user can activate the assistant system 140 via a visual signal without a wake word. The visual signal can be received at a low-power sensor (e.g., a camera) capable of detecting various visual signals. As an example, and not a limitation, the visual signal could be a barcode, QR code, or Universal Product Code (UPC) detected by the client system 130. As another example, and not a limitation, the visual signal could be the user's gaze at an object. As yet another example, and not a limitation, the visual signal could be a user gesture, such as the user pointing at an object.

[0083] In a particular embodiment, audio data from the system audio API 202 can be sent to the assistant orchestrator 206. The assistant orchestrator 206 can execute on the client system 130. In a particular embodiment, the assistant orchestrator 206 can determine whether to respond to user input using a client process, a server process, or both. Figure 2 As shown, the client process is shown below dashed line 207, while the server process is shown above dashed line 207. The assistant orchestrator 206 can also determine whether to respond to user input by using both client and server processes simultaneously. Although Figure 2 Assistant orchestrator 206 is shown as a client process, but assistant orchestrator 206 can be a server process, or it can be a hybrid process that is separate between client and server processes.

[0084] In a particular embodiment, after audio data is generated from the system audio API 202, the server-side process may proceed as follows. The assistant orchestrator 206 may send the audio data to a remote computing system that hosts different modules of the assistant system 140 in response to user input. In a particular embodiment, the audio data may be received at a remote automatic speech recognition (ASR) module 208. The ASR module 208 may allow a user to dictate and have speech transcribed into written text, synthesize a document into an audio stream, or issue a command that is recognized by the system as such. The ASR module 208 may use a statistical model to determine the most probable sequence of words corresponding to a given portion of speech received by the assistant system 140 as audio input. The model may include one or more of a hidden Markov model, a neural network, a deep learning model, or any combination thereof. The received audio input may be encoded into digital data at a specific sampling rate (e.g., 16, 44.1, or 96 kHz) and have a specific number of bits representing each sample (e.g., 8 or 16 bits out of 24 bits).

[0085] In a particular embodiment, ASR module 208 may include different components. ASR module 208 may include one or more of a grapheme-to-phoneme (G2P) model, a pronunciation learning model, a personalized acoustic model, a personalized language model (PLM), or an endpoint model. In a particular embodiment, a G2P model can be used to determine a user's grapheme-to-phoneme style, for example, what it sounds like when a particular user says a particular word. A personalized acoustic model can be a model of the relationship between an audio signal and the sounds of speech units in a language. Thus, such a personalized acoustic model can identify how a user's voice sounds. A personalized acoustic model can be generated using training data, such as training speech received as audio input and corresponding speech units. A personalized acoustic model can be trained or refined using a particular user's voice to recognize that user's speech. In a particular embodiment, a personalized language model can then determine the most likely phrase corresponding to the identified speech units for a particular audio input. A personalized language model can be a probabilistic model of various word sequences that may occur in a language. Personalized language models can be used to match the sounds of speech units in audio input with word sequences, assigning greater weight to word sequences that are more likely to be phrases in that language. The word sequence with the highest weight can then be selected as the text corresponding to the audio input. In a particular embodiment, the personalized language model can also be used to predict what words a user is most likely to say given a context. In a particular embodiment, an endpoint model can detect when the endpoint of a utterance is reached.

[0086] In a particular embodiment, the output of ASR module 208 can be sent to remote natural language understanding (NLU) module 210. NLU module 210 can perform named entity resolution (NER). When analyzing user input, NLU module 210 can additionally consider contextual information. In a particular embodiment, intent and / or slots can be the output of NLU module 210. An intent can be an element in a predefined category of semantic intents that can indicate the purpose of the user's interaction with assistant system 140. NLU module 210 can classify user input into a member of a predefined category; for example, for the input "play Beethoven's Fifth Symphony," NLU module 210 can classify the input as having the intent [IN: play_music]. In a particular embodiment, a domain can represent the social context of the interaction, such as education, or a namespace of a set of intents, such as music. A slot can be a named substring corresponding to a string in the user input, representing a basic semantic entity. For example, the slot for "pizza" could be [SL: dish]. In a particular embodiment, the set of valid or expected named slots can be based on the classified intents. As an example, and not a limitation, for the intent [IN: play_music], a valid slot could be [SL: song_name]. In a particular embodiment, NLU module 210 may additionally extract information from one or more social graphs, knowledge graphs, or concept graphs, and retrieve user profiles from one or more remote data stores 212. NLU module 210 may also process information from these diverse sources by determining what information to aggregate, annotating the user input n-grams, ranking the n-grams with confidence scores based on the aggregated information, and formulating the ranked n-grams into features that can be used by NLU module 210 to understand the user input.

[0087] In a particular embodiment, NLU module 210 can identify one or more domains, intents, or slots from user input in a personalized and context-aware manner. As an example, and not a limitation, user input may include “show me how to get to the coffee shop”. NLU module 210 can identify the specific coffee shop the user wants to go to based on the user's personal information and associated contextual information. In a particular embodiment, NLU module 210 may include a language-specific lexicon, a parser, and grammar rules to segment sentences into internal representations. NLU module 210 may also include one or more programs that use pragmatics to perform naive or stochastic semantic analysis to understand user input. In a particular embodiment, the parser may be based on a deep learning architecture including multiple Long Short-Term Memory (LSTM) networks. As an example, and not a limitation, the parser may be based on a Recurrent Neural Network Grammar (RNNG) model, a type of recurrent and recursive LSTM algorithm. More information on natural language understanding can be found in U.S. Patent Application No. 16 / 011062, filed June 18, 2018; U.S. Patent Application No. 16 / 025317, filed July 2, 2018; and U.S. Patent Application No. 16 / 038120, filed July 17, 2018.

[0088] In a particular embodiment, the output of NLU module 210 can be sent to remote inference module 214. Inference module 214 may include a dialogue manager and an entity resolution component. In a particular embodiment, the dialogue manager may have complex dialogue logic and product-related business logic. The dialogue manager can manage the dialogue state and session flow between the user and assistant system 140. The dialogue manager may additionally store previous sessions between the user and assistant system 140. In a particular embodiment, the dialogue manager can communicate with the entity resolution component to resolve entities associated with one or more slots, which supports the dialogue manager in advancing the session flow between the user and assistant system 140. In a particular embodiment, when resolving entities, the entity resolution component can access one or more of a social graph, knowledge graph, or concept graph. Entities may include, for example, unique users or concepts, each of which may have a unique identifier (ID). By way of example and not limitation, a knowledge graph may include multiple entities. Each entity may include a single record associated with one or more attribute values. A specific record may be associated with a unique entity identifier. Each record may have a different value for one attribute of an entity. Each attribute value may be associated with a confidence probability. The confidence probability of an attribute value represents the probability that the value is accurate for a given attribute. Each attribute value can also be associated with a semantic weight. The semantic weight of an attribute value can represent the degree to which the value is semantically appropriate for a given attribute, taking into account all available information. For example, a knowledge graph can include entities from the book "The Adventures of Alice," which include information extracted from multiple content sources (e.g., online social networks, online encyclopedias, book review sources, media databases, and entertainment content sources), then deduplicated, parsed, and fused to generate a single unique record for the knowledge graph. Entities can be associated with the "fantasy" attribute value, which indicates the genre of the book "The Adventures of Alice." More information about knowledge graphs can be found in U.S. Patent Application No. 16 / 048049, filed July 27, 2018, and U.S. Patent Application No. 16 / 048101, filed July 27, 2018.

[0089] In certain embodiments, the entity parsing component may check privacy constraints to ensure that the parsing of entities does not violate a privacy policy. By way of example, and not limitation, the entity to be parsed could be another user whose privacy settings specify that their identity should not be searchable on the online social network; therefore, the entity parsing component may not return that user's identifier in response to a request. Based on information obtained from social graphs, knowledge graphs, concept graphs, and user profiles, and in accordance with applicable privacy policies, the entity parsing component can therefore parse entities associated with user input in a personalized, context-aware, and privacy-aware manner. In certain embodiments, each parsed entity may be associated with one or more identifiers hosted by the social network system 160. By way of example, and not limitation, identifiers may include a unique user identifier (ID) corresponding to a specific user (e.g., a unique username or user ID number). In certain embodiments, each parsed entity may also be associated with a confidence score. More information about parsed entities can be found in U.S. Patent Application No. 16 / 048049, filed July 27, 2018, and U.S. Patent Application No. 16 / 048072, filed July 27, 2018.

[0090] In certain embodiments, the dialogue manager may perform dialogue optimization and assistant state tracking. Dialogue optimization is the problem of using data to understand what the most likely branch of the dialogue should be. As an example, and not a limitation, using dialogue optimization, the assistant system 140 may not need to confirm who the user wants to call, because the assistant system 140 has a high confidence that the person inferred based on dialogue optimization is likely the person the user wants to call. In certain embodiments, the dialogue manager may use reinforcement learning for dialogue optimization. Assistant state tracking aims to track the state that changes over time as the user interacts with the world and the assistant system 140 interacts with the user. As an example, and not a limitation, depending on an applicable privacy policy, assistant state tracking may track what the user is talking about, who the user is with, where the user is, what task is currently being performed, and where the user is gazing, etc. In certain embodiments, the dialogue manager may use a set of operators to track the dialogue state. Operators may include the necessary data and logic to update the dialogue state. After processing an incoming request, each operator may act as an increment (delta) of the dialogue state. In certain embodiments, the dialogue manager may further include a dialogue state tracker and an action selector. In an alternative embodiment, a dialogue state tracker can replace the entity parsing component and parse references / mentions, and track the state.

[0091] In a particular embodiment, the inference module 214 may further perform error trigger mitigation. The goal of error trigger mitigation is to detect erroneous triggers of the assistant request (e.g., wake words) and avoid generating error logs when the user does not actually intend to invoke the assistant system 140. As an example, and not a limitation, the inference module 214 may implement error trigger mitigation based on a meaninglessness detector. If the meaninglessness detector determines that the wake word is meaningless at this point in the user interaction, the inference module 214 may determine that inferring the user's intention to invoke the assistant system 140 is likely incorrect. In a particular embodiment, the output of the inference module 214 may be sent to a remote dialogue arbitrator 216.

[0092] In a particular embodiment, each of the ASR module 208, NLU module 210, and inference module 214 may access a remote data storage 212 that includes user episode memories to determine how to more effectively assist the user. More information about episode memories can be found in U.S. Patent Application No. 16 / 552559, filed August 27, 2019. Data storage 212 may additionally store a user profile. A user profile may include profile data that includes demographic, social, and contextual information associated with the user. The profile data may also include the user's interests and preferences on multiple topics aggregated through conversations on news feeds, search logs, messaging platforms, etc. Use of user profiles may be subject to privacy restrictions to ensure that a user's information is used only for his / her benefit and not shared with any other person. More information about user profiles can be found in U.S. Patent Application No. 15 / 967239, filed April 30, 2018.

[0093] In a particular embodiment, the client process can run in parallel with the aforementioned server-side processes involving ASR module 208, NLU module 210, and inference module 214. In a particular embodiment, the output of assistant orchestrator 206 can be sent to a local ASR module 216 on client system 130. ASR module 216 may include a personalized language model (PLM), a G2P model, and an endpoint model. Due to the limited computing power of client system 130, assistant system 140 can optimize the personalized language model at runtime during the client process. By way of example and not limitation, assistant system 140 can pre-compute multiple personalized language models for multiple possible topics the user might discuss. When the user requests help, assistant system 140 can then quickly exchange these pre-computed language models, allowing the personalized language model to be locally optimized by assistant system 140 at runtime based on user activity. As a result, assistant system 140 can have the technical advantage of saving computing resources while efficiently determining what the user might be talking about. In a particular embodiment, assistant system 140 can also quickly relearn the user's pronunciation at runtime.

[0094] In a particular embodiment, the output of ASR module 216 may be sent to local NLU module 218. In a particular embodiment, NLU module 218 may be more compact than a server-supported remote NLU module 210. When ASR module 216 and NLU module 218 process user input, they may access local assistant memory 220. For user privacy purposes, local assistant memory 220 may differ from user memory stored on data storage 212. In a particular embodiment, local assistant memory 220 may be synchronized with user memory stored on data storage 212 via network 110. By way of example and not limitation, local assistant memory 220 may synchronize the calendar on user client system 130 with the server-side calendar associated with the user. In a particular embodiment, any secure data in local assistant memory 220 may only be accessed by modules of assistant system 140 that execute locally on client system 130.

[0095] In a particular embodiment, the output of NLU module 218 can be sent to local inference module 222. Inference module 222 may include a dialogue manager and entity parsing components. Due to limited computing power, inference module 222 can perform on-device learning based on learning algorithms specifically tailored for client system 130. By way of example and not limitation, inference module 222 may use federated learning. Federated learning is a specific category of distributed machine learning methods that uses distributed data residing on terminal devices such as mobile phones to train machine learning models. In a particular embodiment, inference module 222 may use a specific federated learning model, namely federated user representation learning, to extend existing neural network personalization techniques to federated learning. Federated user representation learning can personalize the model in federated learning by learning task-specific user representations (i.e., embeddings) or by personalizing model weights. Federated user representation learning is simple, scalable, privacy-preserving, and resource-efficient. Federated user representation learning can categorize model parameters into federated parameters and private parameters. Private parameters, such as private user embeddings, can be trained locally on client system 130 instead of being transmitted to a remote server or averaged on a remote server. In contrast, federated parameters can be trained remotely on a server. In a particular embodiment, inference module 222 may use another specific federated learning model, namely active federated learning, to transfer a global model trained on a remote server to client systems 130 and compute gradients locally on these client systems 130. Active federated learning allows the inference module to minimize the transmission costs associated with downloading the model and uploading gradients. For active federated learning, in each round, client systems are not uniformly randomized but probabilistically selected conditioned on the current model and the data on the client systems to maximize efficiency. In a particular embodiment, inference module 222 may use another specific federated learning model, namely federated Adam. Traditional federated learning models may use stochastic gradient descent (SGD) optimizers. In contrast, the federated Adam model may use a moment-based optimizer. The federated Adam model may use an averaged model to compute approximate gradients instead of directly using an averaged model as in traditional work. These gradients can then be fed into the federated Adam model, which can remove noise from stochastic gradients and use an adaptive learning rate for each parameter. Gradients generated by federated learning may be noisier than those generated by stochastic gradient descent (because the data may not be independent and identically distributed), so the federated Adam model may help to handle noise better. The federated Adam model can use gradients to take smarter steps to minimize the objective function. Experiments show that traditional federated learning on the baseline has a 1.6% decrease in the ROC (Receiver Operating Characteristic) curve, while the federated Adam model only has a 0.4% decrease. Furthermore, the federated Adam model does not add computation to the communication or on-device.In a particular embodiment, the inference module 222 may also perform error trigger mitigation. This error trigger mitigation can help detect erroneous activation requests, such as wake words, on the client system 130 when the user's voice input includes privacy-constrained data. By way of example and not limitation, when a user is in the middle of a voice call, the user's session is private, and error trigger detection based on this session can only occur locally on the user's client system 130.

[0096] In a particular embodiment, assistant system 140 may include a local context engine 224. Context engine 224 may process all other available signals to provide more information prompts to inference module 222. By way of example, and not limitation, context engine 224 may have human-related information, sensing data from sensors (e.g., microphones, cameras) of client system 130 (which are further analyzed using computer vision techniques), geometry, activity data, inertial data (e.g., data collected by a VR headset), location, etc. In a particular embodiment, computer vision techniques may include human skeleton reconstruction, face detection, face recognition, hand tracking, eye tracking, etc. In a particular embodiment, geometry may include constructing objects around the user using data collected by client system 130. By way of example, and not limitation, the user may be wearing AR glasses, and the geometry may be designed to determine where the floor is, where the walls are, where the user's hands are, etc. In a particular embodiment, inertial data may be data associated with linear and angular motion. By way of example, and not limitation, inertial data may be captured by AR glasses that measure how parts of the user's body move.

[0097] In a particular embodiment, the output of the local inference module 222 can be sent to the dialogue arbitrator 216. The dialogue arbitrator 216 can operate differently in three scenarios. In the first scenario, the assistant orchestrator 206 determines to use the server-side process; for this purpose, the dialogue arbitrator 216 can transmit the output of the inference module 214 to the remote action execution module 226. In the second scenario, the assistant orchestrator 206 determines to use both the server-side and client-side processes; for this purpose, the dialogue arbitrator 216 can aggregate and analyze the outputs of the two inference modules (i.e., the remote inference module 214 and the local inference module 222) from the two processes. By way of example and not limitation, the dialogue arbitrator 216 can perform sorting and select the best inference result to respond to user input. In a particular embodiment, the dialogue arbitrator 216 can also determine, based on analysis, whether to use an agent on the server or the client to perform the relevant task. In the third scenario, the assistant orchestrator 206 determines to use the client-side process, and the dialogue arbitrator 216 needs to evaluate the output of the local inference module 222 to determine whether the client-side process is capable of handling the user input task.

[0098] In a particular embodiment, for the first and second scenarios described above, the dialogue arbitrator 216 can determine that a server-side agent is necessary to perform a task in response to user input. Therefore, the dialogue arbitrator 216 can send the necessary information about the user input to the action execution module 226. The action execution module 226 can invoke one or more agents to perform the task. In an alternative embodiment, the dialogue manager's action selector can determine the action to be performed and accordingly instruct the action execution module 226. In a particular embodiment, an agent can be an implementation that acts as a broker between multiple content providers in a domain. A content provider can be an entity responsible for performing an action associated with an intent or completing a task associated with an intent. In a particular embodiment, agents can include first-party agents and third-party agents. In a particular embodiment, a first-party agent can include an internal agent accessible and controllable by the assistant system 140 (e.g., an agent associated with a service provided by an online social network, such as a messaging service or a photo-sharing service). In a particular embodiment, a third-party agent can include an external agent that the assistant system 140 cannot control (e.g., a third-party online music application agent, a ticket sales agent). First-party agents may be associated with first-party providers that offer content objects and / or services hosted by social networking system 160. Third-party agents may be associated with third-party providers that offer content objects and / or services hosted by third-party system 170. In a particular embodiment, each of the first-party or third-party agents may be designated for a specific domain. By way of example, and not limitation, the domain may include weather, traffic, music, etc. In a particular embodiment, assistant system 140 may use multiple agents in concert to respond to user input. By way of example, and not limitation, user input may include “direct me to my next meeting”. Assistant system 140 may use a calendar agent to retrieve the location of the next meeting. Assistant system 140 may then use a navigation agent to guide the user to the next meeting.

[0099] In a specific embodiment, for the second and third scenarios described above, the dialogue arbitrator 216 can determine that the agent on the client side can perform a task in response to user input, but requires additional information (e.g., a response template), or that the task can only be handled by the agent on the server side. If the dialogue arbitrator 216 determines that the task can only be handled by the agent on the server side, then the dialogue arbitrator 216 can send the necessary information about the user input to the action execution module 226. If the dialogue arbitrator 216 determines that the agent on the client side can perform the task, but requires a response template, then the dialogue arbitrator 216 can send the necessary information about the user input to the remote response template generation module 228. The output of the response template generation module 228 can also be sent to the local action execution module 230 executed on the client system 130.

[0100] In a particular embodiment, the action execution module 230 may invoke a local agent to perform a task. The local agent on the client system 130 is capable of performing simpler tasks compared to a server-side agent. As an example, and not a limitation, multiple device-specific implementations (e.g., real-time invocations of client system 130 or messaging applications on client system 130) may be handled internally by a single agent. Alternatively, these device-specific implementations may be handled by multiple agents associated with multiple domains. In a particular embodiment, the action execution module 230 may additionally execute a set of generic executable dialogue actions. This set of executable dialogue actions may interact with the agent, user, and assistant system 140 itself. These dialogue actions may include dialogue actions for slot requests, acknowledgments, ambiguity resolution, agent execution, etc. Dialogue actions may be independent of the underlying implementation of the action selector or dialogue policy. Both tree-based and model-based policies can generate the same basic dialogue actions, and callback functions hide any implementation details specific to the action selector.

[0101] In a particular embodiment, the output from the remote action execution module 226 on the server side can be sent to the remote response execution module 232. In a particular embodiment, the action execution module 226 can send back more information to the dialogue arbitrator 216. The response execution module 232 can be based on a remote session understanding (CU) writer. In a particular embodiment, the output from the action execution module 226 can be formulated as follows:<k,c,u,d> The tuples, where k indicates the knowledge source, c indicates the communication goal, u indicates the user model, and d indicates the discourse model. In a particular embodiment, the CU writer may include a Natural Language Generation (NLG) module and a User Interface (UI) payload generator. The NLG generator may use different language models and / or language templates to generate communication content based on the output of the action execution module 226. In a particular embodiment, the generation of communication content may be application-specific and personalized for each user. The CU writer may also use the UI payload generator to determine the modality of the generated communication content. In a particular embodiment, the NLG module may include a content determination component, a sentence planner, and a surface realization component. The content determination component may determine the communication content based on the knowledge source, the communication goal, and the user's expectations. By way of example and not limitation, determination may be based on description logic. Description logic may include, for example, three basic notions: an individual (representing an object in a domain), a concept (describing a set of individuals), and a role (representing a binary relationship between individuals or concepts). The description logic can be characterized by a set of constructors that allow the natural language generator to construct complex concepts / roles from atomic concepts / roles. In a particular embodiment, the content determination component can perform the following tasks to determine the communication content. A first task can include a translation task, where input to the natural language generator can be translated into concepts. A second task can include a selection task, where relevant concepts can be selected from the concepts generated by the translation task based on a user model. A third task can include a verification task, where the consistency of the selected concepts can be verified. A fourth task can include an instantiation task, where the verified concepts can be instantiated into an executable file that can be processed by the natural language generator. A sentence planner can determine the organization of the communication content to make it human-readable. A surface implementation component can determine the specific words to use, the sentence order, and the style of the communication content. A UI payload generator can determine the preferred modality of the communication content to be presented to the user. In a particular embodiment, the CU writer can check privacy constraints associated with the user to ensure that the generation of the communication content follows a privacy policy. More information on natural language generation can be found in U.S. Patent Application No. 15 / 967279, filed April 30, 2018, and U.S. Patent Application No. 15 / 966455, filed April 30, 2018.

[0102] In a particular embodiment, the output from the local action execution module 230 on the client system 130 can be sent to the local response execution module 234. The response execution module 234 can be based on a local conversation understanding (CU) writer. The CU writer may include a natural language generation (NLG) module. Since the computing power of the client system 130 may be limited, the NLG module can be simple for computational efficiency. Because the NLG module can be simple, the output of the response execution module 234 can be sent to the local response extension module 236. The response extension module 236 can further extend the results of the response execution module 234 to make the response more natural and contain richer semantic information.

[0103] In a particular embodiment, if the user input is based on an audio signal, the output of the server-side response execution module 232 can be sent to the remote text-to-speech (TTS) module 238. Similarly, the output of the client-side response extension module 236 can be sent to the local TTS module 240. Both TTS modules can convert the response into an audio signal. In a particular embodiment, the output from the response execution module 232, the response extension module 236, or the TTS modules at both ends can ultimately be sent to the local rendering output module 242. The rendering output module 242 can generate a response suitable for the client system 130. By way of example and not limitation, the output of the response execution module 232 or the response extension module 236 may include one or more of a natural language string, speech, parameterized actions, or a rendered image or video that can be displayed in a VR headset or AR smart glasses. As a result, the rendering output module 242 can determine what task to perform based on the output of the CU writer to appropriately render the response for display on a VR headset or AR smart glasses. For example, the response could be a vision-based modality (e.g., an image or video clip) that can be displayed via a VR headset or AR smart glasses. As another example, the response could be an audio signal that a user can play through a VR headset or AR smart glasses. Yet another example, the response could be augmented reality data that can be rendered by a VR headset or AR smart glasses to enhance the user experience.

[0104] In certain embodiments, the assistant system 140 may possess various capabilities, including audio recognition, visual recognition, signal intelligence, reasoning, and memory. In certain embodiments, the audio recognition capability enables the assistant system 140 to understand user input associated with various domains of different languages, understand and summarize conversations, perform audio recognition on complex command execution devices, recognize users through speech, extract topics from conversations and automatically tag parts of conversations, enable audio interaction without wake words, filter and amplify user speech from ambient noise and conversations, and understand which client system 130 the user is talking to (if multiple client systems 130 are nearby).

[0105] In a particular embodiment, the visual recognition capability enables the assistant system 140 to perform face detection and tracking, identify users, identify most people of interest in major metropolitan areas from different angles, identify most objects of interest in the world through a combination of existing machine learning models and one-time learning, identify moments of interest and automatically capture them, achieve semantic understanding of multiple visual frames across different time periods, provide platform support for additional capabilities for people, places, and objects recognition, identify full-set settings and micro-locations including personalized locations, identify complex activities, recognize complex gestures to control the client system 130, process images / videos from egocentric cameras (e.g., with motion, capture angle, resolution, etc.), achieve accuracy and speed regarding similarity levels for images with lower resolution, perform one-time registration and recognition of people, places, and objects, and perform visual recognition on the client system 130.

[0106] In certain embodiments, assistant system 140 may utilize computer vision techniques to achieve visual cognition. In addition to computer vision techniques, assistant system 140 may explore options that complement these techniques to expand object recognition. In certain embodiments, assistant system 140 may use supplementary signals, such as optical character recognition (OCR) for object tags, GPS signals for location recognition, and signals from user client system 130, to identify the user. In certain embodiments, assistant system 140 may perform general scene recognition (home, work, public space, etc.) to set context for the user and narrow the computer vision search space to identify the most likely objects or people. In certain embodiments, assistant system 140 may guide the user to train assistant system 140. For example, crowdsourcing may be used to allow users to tag and help assistant system 140 recognize more objects over time. As another example, when using assistant system 140, users can register their personal objects as part of the initial setup. Assistant system 140 may also allow users to provide positive / negative signals for the objects they interact with in order to train and improve personalized models for them.

[0107] In a particular embodiment, the signal intelligence capability enables the assistant system 140 to determine the user's location, understand the date / time, determine the home location, understand the user's calendar and expected future location, integrate richer voice understanding to identify settings / context by voice alone, and build a signal intelligence model at runtime that can be personalized for the user's personal routine.

[0108] In a particular embodiment, reasoning capabilities enable the assistant system 140 to pick up any previous conversation thread at any point in the future, synthesize all signals to understand micro and personalized context, learn interaction patterns and preferences from the user’s historical behavior and accurately suggest interactions they may value, generate highly predictive proactive suggestions based on micro context understanding, understand what content the user may want to see at what time of day, and understand changes in the scene and how this may affect the content the user expects.

[0109] In a particular embodiment, the memory capability enables the assistant system 140 to remember which social connections the user previously invoked or interacted with, write to and query memory at will (i.e., open dictation and automatic tagging), extract richer preferences based on previous interactions and long-term learning, remember the user's life history, extract rich information from egocentric data streams and automatic catalogs, and write to memory in a structured form to form rich short-term, episodic, and long-term memories.

[0110] Figure 3An exemplary flowchart of the server-side process of the assistant system 140 is shown. In a particular embodiment, the server assistant service module 301 may access the request manager 302 upon receiving a user request. In an alternative embodiment, if the user request is based on an audio signal, the user request may first be processed by the remote ASR module 208. In a particular embodiment, the request manager 302 may include a context extractor 303 and a session understanding object generator (CU object generator) 304. The context extractor 303 may extract context information associated with the user request. The context extractor 303 may also update the context information based on the assistant application 136 executing on the client system 130. By way of example, and not limitation, updating the context information may include displaying content items on the client system 130. By way of another example, and not limitation, updating the context information may include whether an alarm clock is set on the client system 130. By way of another example, and not limitation, updating the context information may include whether a song is playing on the client system 130. The CU object generator 304 may generate a specific content object associated with the user request. The content object may include dialogue session data and features associated with the user request, which can be shared with all modules of the assistant system 140. In a particular embodiment, the request manager 302 may store context information and the generated content object in a data storage 212, which is a specific data storage implemented in the assistant system 140.

[0111] In a particular embodiment, the request manager 302 may send the generated content object to a remote NLU module 210. The NLU module 210 may perform multiple steps to process the content object. In step 305, the NLU module 210 may generate a whitelist of content objects. In a particular embodiment, the whitelist may include explanatory data matching the user request. In step 306, the NLU module 210 may perform feature identification based on the whitelist. In step 307, the NLU module 210 may perform domain classification / selection on the user request based on the features generated from the feature identification to classify the user request into a predefined domain. The domain classification / selection results may also be further processed based on two related processes. In step 308a, the NLU module 210 may use an intent classifier to process the domain classification / selection results. The intent classifier may determine the user intent associated with the user request. In a particular embodiment, each domain may have one intent classifier to determine the most probable intent in a given domain. As an example, and not a limitation, the intent classifier may be based on a machine learning model that takes the domain classification / selection result as input and calculates the probability that the input is associated with a specific predefined intent. In step 308b, NLU module 210 can use a meta-intent classifier to process the domain classification / selection results. The meta-intent classifier can determine the category describing the user's intent. In a particular embodiment, intents common to multiple domains can be processed by the meta-intent classifier. By way of example and not limitation, the meta-intent classifier can be based on a machine learning model that takes the domain classification / selection results as input and calculates the probability that the input is associated with a specific predefined meta-intent. In step 309a, NLU module 210 can use a slot tagger to annotate one or more slots associated with the user request. In a particular embodiment, the slot tagger can annotate one or more slots for the n-grams of the user request. In step 309b, NLU module 210 can use a meta-slot tagger to annotate one or more slots for the classification results from the meta-intent classifier. In a particular embodiment, the meta-slot tagger can tag generic slots, such as references to items (e.g., the first item), slot types, slot values, etc. By way of example and not limitation, a user request could include "convert $500 in my account into Japanese Yen". An intent classifier takes a user request as input and formulates it as a vector. The intent classifier can then calculate the probability that the user request is associated with different predefined intents based on vector comparisons between the vector representing the user request and vectors representing different predefined intents. Similarly, a slot tagger takes a user request as input and formulates each word as a vector. The intent classifier can then calculate the probability that each word is associated with different predefined slots based on vector comparisons between the vector representing the word and vectors representing different predefined slots.A user's intent can be categorized as "changing money." The slots for user requests can include "500," "dollars," "account," and "Japanese yen." A user's meta intent can be categorized as "financial service." The meta slot can include "finance."

[0112] In a particular embodiment, NLU module 210 may include semantic information aggregator 310. Semantic information aggregator 310 can help NLU module 210 improve domain classification / selection of content objects by providing semantic information. In a particular embodiment, semantic information aggregator 310 may aggregate semantic information in the following manner: Semantic information aggregator 310 may first retrieve information from user context engine 315. In a particular embodiment, user context engine 315 may include an offline aggregator and an online inference service. The offline aggregator may process multiple data associated with the user collected from previous time windows. By way of example and not limitation, the data may include news feed posts / comments collected within a predetermined time range (e.g., from a previous 90-day window), interactions with news feed posts / comments, search history, etc. The processing results may be stored in user context engine 315 as part of the user profile. The online inference service may analyze session data associated with the user received by assistant system 140 at the current time. The analysis results may also be stored in user context engine 315 as part of the user profile. In a particular embodiment, both the offline aggregator and the online inference service may extract personalized features from multiple data. The extracted personalized features can be used by other modules of the assistant system 140 to better understand user input. In a particular embodiment, the semantic information aggregator 310 can then process the information retrieved from the user context engine 315, i.e., the user profile, in the following steps. In step 311, the semantic information aggregator 310 can process the information retrieved from the user context engine 315 based on natural language processing (NLP). In a particular embodiment, the semantic information aggregator 310 can: tokenize the text through text normalization, extract syntactic features from the text, and extract semantic features from the text based on NLP. The semantic information aggregator 310 can also extract features from contextual information accessed from the dialogue history between the user and the assistant system 140. The semantic information aggregator 310 can also perform global word embedding, domain-specific embedding, and / or dynamic embedding based on the contextual information. In step 312, the processing results can be annotated with entities by an entity tokenizer. In step 313, based on the annotations, the semantic information aggregator 310 can generate a dictionary for the retrieved information. In a particular embodiment, the dictionary may include global dictionary features that can be dynamically updated offline. In step 314, the semantic information aggregator 310 may sort the entities tagged by the entity tagger. In a particular embodiment, the semantic information aggregator 310 may communicate with one or more different graphs 320, including social graphs, knowledge graphs, or concept graphs, to extract ontology data related to the retrieved information from the user context engine 315. In a particular embodiment, the semantic information aggregator 310 may aggregate user profiles, sorted entities, and information from graph 320.The semantic information aggregator 310 can then provide the aggregated information to the NLU module 210 to facilitate domain classification / selection.

[0113] In a particular embodiment, the output of NLU module 210 may be sent to remote inference module 214. Inference module 214 may include coreference component 325, entity resolution component 330, and dialogue manager 335. The output of NLU module 210 may first be received at coreference component 325 to interpret the reference of the content object associated with the user request. In a particular embodiment, coreference component 325 may be used to identify the item referred to in the user request. Coreference component 325 may include reference creation 326 and reference resolution 327. In a particular embodiment, reference creation 326 may create references for entities determined by NLU module 210. Reference resolution 327 may accurately resolve these references. By way of example and not limitation, a user request may include “Find me the nearest grocery store and direct me there.” Coreference component 325 may interpret “there” as “the nearest grocery store.” In a particular embodiment, coreference component 325 may access user context engine 315 and dialogue engine 335 as needed to interpret references with improved accuracy.

[0114] In a particular embodiment, the identified domains, intents, meta-intents, slots and meta-slots, and the resolved referents can be sent to entity resolution component 330 to resolve related entities. Entity resolution component 330 can perform general and domain-specific entity resolution. In a particular embodiment, entity resolution component 330 may include domain entity resolution 331 and general entity resolution 332. Domain entity resolution 331 can resolve entities by classifying slots and meta-slots into different domains. In a particular embodiment, entities can be resolved based on ontology data extracted from FIG. 320. The ontology data may include structural relationships between different slots / meta-slots and domains. The ontology may also include information on how slots / meta-slots can be grouped, related, and subdivided according to similarity and difference within a higher-level hierarchy including domains. General entity resolution 332 can resolve entities by classifying slots and meta-slots into different general topics. In a particular embodiment, resolution may also be based on ontology data extracted from FIG. 320. The ontology data may include structural relationships between different slots / meta-slots and general topics. An ontology can also include information about how slots / meta-slots can be grouped, related, and subdivided based on similarity and difference within a higher-level hierarchy that includes topics. As an example, and not a limitation, in response to an input query for the advantages of a specific brand of electric vehicle, generic entity resolution 332 can resolve the brand name of electric vehicle to vehicle, and domain entity resolution 331 can resolve the brand name of electric vehicle to electric vehicle.

[0115] In a particular embodiment, the output of entity resolution component 330 can be sent to dialogue manager 335 to advance the conversation flow with the user. Dialogue manager 335 can be an asynchronous state machine that repeatedly updates its state and selects actions based on the new state. Dialogue manager 335 can include dialogue intent parser 336 and dialogue state tracker 337. In a particular embodiment, dialogue manager 335 can execute the selected action and then call dialogue state tracker 337 again until the selected action requires a user response or there are no more actions to execute. Each selected action may depend on the result of the execution of a previous action. In a particular embodiment, dialogue intent parser 336 can parse user intents associated with the current dialogue session based on the dialogue history between the user and assistant system 140. Dialogue intent parser 336 can map the intents determined by NLU module 210 to different dialogue intents. Dialogue intent parser 336 can also sort dialogue intents based on signals from NLU module 210, entity resolution component 330, and dialogue history between user and assistant system 140. In a particular embodiment, the dialogue state tracker 337 may be a component without side effects, generating n best candidates for dialogue state update operators that suggest dialogue state updates, rather than directly changing the dialogue state. The dialogue state tracker 337 may include an intent resolver containing logic for processing different types of NLU intents based on the dialogue state and generating operators. In a particular embodiment, the logic may be organized by an intent processor, such as an ambiguity resolution intent processor that processes intents when the assistant system 140 requests ambiguity resolution, an acknowledgment intent processor that includes logic for handling acknowledgments, etc. The intent resolver may combine turn intents with the dialogue state to generate a context update for the user session. The slot resolution component may then recursively resolve slots in the update operators using a resolution provider that includes a knowledge graph and domain proxies. In a particular embodiment, the dialogue state tracker 337 may update / sort the dialogue state of the current dialogue session. By way of example, and not limitation, if the dialogue session ends, the dialogue state tracker 337 may update the dialogue state to "complete". By way of another example, and not limitation, the dialogue state tracker 337 may sort the dialogue states based on their associated priorities.

[0116] In a particular embodiment, the inference module 214 may communicate with the remote action execution module 226 and the dialogue arbitrator 216, respectively. In a particular embodiment, the dialogue manager 335 of the inference module 214 may communicate with the task completion component 340 of the action execution module 226 regarding dialogue intents and associated content objects. In a particular embodiment, the task completion module 340 may sort different dialogue hypotheses for different dialogue intents. The task completion module 340 may include an action selector 341. In an alternative embodiment, the action selector 341 may be included in the dialogue manager 335. In a particular embodiment, the dialogue manager 335 may also check against a dialogue policy 345 regarding the dialogue state included in the dialogue arbitrator 216. In a particular embodiment, the dialogue policy 345 may include a data structure describing the action execution plan of the agent 350. The dialogue policy 345 may include a general policy 346 and a task policy 347. In a particular embodiment, the general policy 346 may be used for actions not specific to a single task. The general policy 346 may include handling low-confidence intents, internal errors, unacceptable user retry responses, skipping or inserting acknowledgments based on ASR or NLU confidence scores, etc. The general policy 346 may also include logic for sorting dialogue state update candidates from the output of the dialogue state tracker 337 and selecting one to update (e.g., selecting the highest-ranked task intent). In a particular embodiment, the assistant system 140 may have a specific interface for the general policy 346 that allows the incorporation of disparate cross-domain policies / business rules, particularly those found in the dialogue state tracker 337, into the functionality of the action selector 341. The interface of the general policy 346 may also allow the creation of independent sub-policy units that can be bound to specific situations or clients, such as policy functions that can be easily turned on or off based on client, situation, etc. The interface of the general policy 346 may also allow for a policy hierarchy with backoff, i.e., multiple policy units, where highly specialized policy units handling specific situations are supported by a more general policy 346 applicable in a broader context. In this case, the general policy 346 may alternatively include intent- or task-specific policies. In a particular embodiment, task policy 347 may include logic for an action selector 341 based on the task and the current state. In a particular embodiment, there may be four types of task policies 347: 1) a hand-crafted tree-based dialogue plan; 2) an coded policy that directly implements an interface for generating actions; 3) a slot-filling task specified by a configurator; and 4) a machine learning model-based policy learned from data. In a particular embodiment, assistant system 140 may guide a new domain with rule-based logic and later refine the task policy 347 using a machine learning model. In a particular embodiment, dialogue policy 345 may be a tree-based policy, which is a pre-built dialogue plan.Based on the current dialogue state, dialogue policy 345 can select nodes to execute and generate corresponding actions. As an example and not a limitation, tree-based policies can include topic grouping nodes and dialogue action (leaf) nodes.

[0117] In a particular embodiment, the action selector 341 may employ candidate operators from the dialogue state and consult the dialogue policy 345 to determine which action should be performed. The assistant system 140 may use a hierarchical dialogue policy, where a general policy 346 handles cross-domain business logic, and a task policy 347 handles task / domain-specific logic. In a particular embodiment, the general policy 346 may select an operator from the candidate operators to update the dialogue state, and then select a user-facing action via the task policy 347. Once a task is active in the dialogue state, the appropriate task policy 347 can be consulted to select the correct action. In a particular embodiment, both the dialogue state tracker 337 and the action selector 341 may remain unchanged in the dialogue state until the selected action is executed. This allows the assistant system 140 to execute the dialogue state tracker 337 and the action selector 341 to process speculative ASR results and utilize dry runs for n-optimal ranking. In a particular embodiment, the action selector 341 may select a dialogue action by taking the dialogue state update operator as part of its input. The execution of a dialogue action can generate a set of expectations to instruct the dialogue state tracker 337 to process future rounds. In a particular embodiment, these expectations can be used to provide context to the dialogue state tracker 337 when processing user input from the next round. As an example, and not a limitation, a slot request dialogue action may expect to prove the value of the requested slot.

[0118] In a particular embodiment, the dialogue manager 335 may support multi-turn composition parsing of slot references. For composition parsing from NLU 210, the parser may recursively parse nested slots. The dialogue manager 335 may additionally support disambiguation of nested slots. As an example, and not a limitation, a user request could be “Remind me to call Alex.” The parser may need to know which Alex to call before creating an actionable to-do reminder entity. When a particular slot requires further user clarification, the parser may pause parsing and set a parsing state. A general policy 346 may check the parsing state and create a corresponding dialogue action for user clarification. In the dialogue state tracker 337, the dialogue manager may update nested slots based on the user request and the last dialogue action. This capability allows the assistant system 140 to interact with the user, not only collecting missing slot values ​​but also reducing the ambiguity of more complex / vague utterances to complete tasks. In a particular embodiment, the dialogue manager may further support requesting nested intents and missing slots in multi-intent user requests (e.g., “Take this picture and send it to Dad”). In a particular embodiment, the dialogue manager 335 may support machine learning models for a more robust dialogue experience. As an example, and not a limitation, the dialogue state tracker 337 may use a neural network-based model (or any other suitable machine learning model) to model beliefs on task assumptions. As another example, and not a limitation, for the action selector 341, the highest priority policy unit may include whitelist / blacklist overriding, which may have to be designed in; medium priority units may include a machine learning model designed for action selection; and lower priority units may include rule-based backoff when the machine learning model chooses not to handle a situation. In a particular embodiment, a general policy unit based on a machine learning model can help the assistant system 140 reduce redundant disambiguation or confirmation steps, thereby reducing the number of rounds required to execute a user request.

[0119] In a particular embodiment, the action execution module 226 may invoke different agents 350 to perform a task. Agent 350 may be selected from registered content providers to complete the action. The data structure may be constructed by the dialogue manager 335 based on an intent and one or more slots associated with that intent. The dialogue policy 345 may also include multiple targets that are interrelated through logical operators. In a particular embodiment, a target may be the output of a dialogue policy, and it may be constructed by the dialogue manager 335. A target may be represented by an identifier (e.g., a string) with one or more named parameters that parameterize the target. By way of example, and not limitation, a target and its associated target parameters may be represented as {Confirmation_Artist, Parameter: {Artist: "Madonna"}}. In a particular embodiment, the dialogue policy may be represented based on a tree structure, where targets are mapped to leaves. In a particular embodiment, the dialogue manager 335 may execute the dialogue policy 345 to determine the next action to be performed. The dialogue policy 345 may include a general policy 346 and a domain-specific policy 347, both of which may guide how to select the next system action based on the dialogue state. In a particular embodiment, the task completion component 340 of the action execution module 226 can communicate with the dialogue policy 345 included in the dialogue arbitrator 216 to obtain guidance for the next system action. In a particular embodiment, the action selection component 341 can therefore select an action based on the dialogue intent, the associated content object, and the guidance from the dialogue policy 345.

[0120] In a particular embodiment, the output of the action execution module 226 can be sent to the remote response execution module 232. Specifically, the output of the task completion component 340 of the action execution module 226 can be sent to the CU writer 355 of the response execution module 226. In an alternative embodiment, the selected action may require the participation of one or more agents 350. Therefore, the task completion module 340 can notify the agents 350 of the selected action. Simultaneously, the dialogue manager 335 can receive instructions to update the dialogue state. By way of example and not limitation, the update may include waiting for a response from the agents 350. In a particular embodiment, the CU writer 355 can generate communication content for the user using the natural language generation (NLG) module 356 based on the output of the task completion module 340. In a particular embodiment, the NLG module 356 can use different language models and / or language templates to generate natural language output. The generation of natural language output can be application-specific. The generation of natural language output can also be personalized for each user. The CU writer 355 can also use the UI payload generator 357 to determine the modality of the generated communication content. Since the generated communication content can be considered a response to a user request, the CU writer 355 can additionally use a response sorter 358 to sort the generated communication content. As an example, and not a limitation, the sorting can indicate the priority of the responses.

[0121] In a particular embodiment, the response execution module 232 may perform different tasks based on the output of the CU writer 355. These tasks may include writing (i.e., storing / updating) the dialogue state 361 retrieved from the data storage 212 and generating a response 362. In a particular embodiment, the output of the CU writer 355 may include one or more of a natural language string, speech, parameterized actions, or a rendered image or video that can be displayed in a VR headset or AR smart glasses. As a result, the response execution module 232 may determine what task to perform based on the output of the CU writer 355. In a particular embodiment, the generated response and communication content may be sent by the response execution module 232 to the local rendering output module 242. In an alternative embodiment, if the modality of the determined communication content is audio, the output of the CU writer 355 may be additionally sent to a remote TTS module 238. The speech generated by the TTS module 238 and the response generated by the response execution module 232 may then be sent to the rendering output module 242.

[0122] Figure 4An exemplary flowchart illustrating the processing of user input by assistant system 140 is shown. As an example and not a limitation, user input may be based on an audio signal. In a particular embodiment, microphone array 402 of client system 130 may receive audio signals (e.g., voice). The audio signal may be transmitted to processing loop 404 in the format of audio frames. In a particular embodiment, processing loop 404 may send audio frames for Voice Activity Detection (VAD) 406 and Voice Wake-up Detection (WoV) 408. The detection results may be returned to processing loop 404. If WoV detection 408 indicates that the user wants to invoke assistant system 140, the audio frame, along with the VAD 406 result, may be sent to encoding unit 410 to generate encoded audio data. After encoding, for privacy and security purposes, the encoded audio data may be sent to encryption unit 412, followed by linking unit 414 and decryption unit 416. After decryption, the audio data may be sent to microphone driver 418, which may further transmit the audio data to audio service module 420. In an alternative embodiment, user input can be received at a wireless device (e.g., a Bluetooth device) paired with client system 130. Accordingly, audio data can be sent from wireless device driver 422 (e.g., a Bluetooth driver) to audio service module 420. In a particular embodiment, audio service module 420 can determine that the user input can be implemented by an application executing on client system 130. Therefore, audio service module 420 can send the user input to real-time communication (RTC) module 424. RTC module 424 can transmit audio packets to a video or audio communication system (e.g., VoIP or video call). RTC module 424 can invoke a related application (App) 426 to perform tasks related to the user input.

[0123] In a particular embodiment, the audio service module 420 can determine that the user is requesting assistance that requires a response from the assistant system 140. Therefore, the audio service module 420 can notify the client assistant service module 426. In a particular embodiment, the client assistant service module 426 can communicate with the assistant orchestrator 206. The assistant orchestrator 206 can determine whether to use a client process or a server process to respond to the user input. In a particular embodiment, the assistant orchestrator 206 can determine to use a client process and notify the client assistant service module 426 of this decision. As a result, the client assistant service module 426 can invoke the relevant module to respond to the user input.

[0124] In a particular embodiment, the client assistant service module 426 may use a local ASR module 216 to analyze user input. The ASR module 216 may include a glyph-to-phoneme (G2P) model, a pronunciation learning model, a personalized language model (PLM), an endpoint model, and a personalized acoustic model. In a particular embodiment, the client assistant service module 426 may further use a local NLU module 218 to understand user input. The NLU module 218 may include a named entity resolution (NER) component and a context-based session NLU component. In a particular embodiment, the client assistant service module 426 may use an intent mediator 428 to analyze the user's intent. To accurately understand the user's intent, the intent mediator 428 may access an entity store 430 that includes entities associated with the user and the world. In an alternative embodiment, user input may be submitted via an application 432 executing on the client system 130. In this case, the input manager 434 may receive the user input and analyze it by the application environment (AppEnv) module 436. The analysis results can be sent to application 432, which can further send the analysis results to ASR module 216 and NLU module 218. In an alternative embodiment, user input can be submitted directly to client assistant service module 426 via assistant application 438 running on client system 130. Client assistant service module 426 can then perform a similar process based on the modules described above, namely ASR module 216, NLU module 218, and intent mediator 428.

[0125] In a particular embodiment, the assistant orchestrator 206 may determine the server-side process for the user. Therefore, the assistant orchestrator 206 may send user input to one or more computing systems of different modules of the managed assistant system 140. In a particular embodiment, the server assistant service module 301 may receive user input from the assistant orchestrator 206. The server assistant service module 301 may instruct a remote ASR module 208 to analyze the audio data of the user input. The ASR module 208 may include a glyph-to-phoneme (G2P) model, a pronunciation learning model, a personalized language model (PLM), an endpoint model, and a personalized acoustic model. In a particular embodiment, the server assistant service module 301 may further instruct a remote NLU module 210 to understand the user input. In a particular embodiment, the server assistant service module 301 may invoke a remote inference model 214 to process the output from the ASR module 208 and the NLU module 210. In a particular embodiment, the inference model 214 may perform entity resolution and dialogue optimization. In a particular embodiment, the output of the inference model 314 may be sent to an agent 350 to perform one or more related tasks.

[0126] In a particular embodiment, agent 350 can access ontology module 440 to accurately understand the results from entity parsing and dialogue optimization, thereby enabling it to accurately perform relevant tasks. Ontology module 440 can provide ontology data associated with multiple predefined domains, intents, and slots. The ontology data may also include structural relationships between different slots and domains. The ontology data may also include information on how slots can be grouped, related, and subdivided based on similarity and difference within a higher-level hierarchy including domains. The ontology data may also include information on how slots can be grouped, related, and subdivided based on similarity and difference within a higher-level hierarchy including topics. Once a task is executed, agent 350 can return the execution result along with a task completion indication to inference module 214.

[0127] The embodiments disclosed herein may include or be implemented in conjunction with artificial reality systems. Artificial reality is a form of reality that has been adapted in some way before being presented to a user, and may include, for example, virtual reality (VR), augmented reality (AR), mixed reality (MR), hybrid reality, or some combination and / or derivative thereof. Artificial reality content may include fully generated content or generated content combined with captured content (e.g., real-world photographs). Artificial reality content may include video, audio, haptic feedback, or some combination thereof, and any of these may be presented in a single channel or in multiple channels (e.g., generating stereoscopic video with a three-dimensional effect for the viewer). Additionally, in some embodiments, artificial reality may be associated with applications, products, accessories, services, or some combination thereof, which are used, for example, to create content in artificial reality and / or be used in artificial reality (e.g., performing activities in artificial reality). Artificial reality systems that deliver artificial reality content can be implemented on a variety of platforms, including head-mounted displays (HMDs) connected to a host computer system, stand-alone HMDs, mobile devices or computing systems, or any other hardware platform capable of delivering artificial reality content to one or more viewers.

[0128] Metaspeech System Based on Natural Language Understanding

[0129] In a particular embodiment, the assistant system 140 may utilize multiple Automatic Speech Recognition (ASR) engines to analyze audio input via a metaspeech engine. In a particular embodiment, the ASR engine may be part of ASR modules 208, 216. For an ASR engine to operate with sufficient accuracy, it may require a large amount of training data to build the basis of a speech model corresponding to the ASR engine. As an example, and not a limitation, a large amount of training data may include 100,000 audio inputs and their respective transcriptions. However, when initially training the speech model, there may not be a sufficient amount of training data to build a sufficiently operable speech model. That is, as an example, and not a limitation, the speech model may not have enough training data to accurately generate transcriptions for at least 95% of the audio inputs. This situation can be identified when the user needs to repeat requests and / or when errors occur in the generated transcriptions. Therefore, the speech model may require an even larger amount of training data to accurately generate transcriptions for a threshold number of audio inputs. On the other hand, there may exist ASR engines that have already been trained on a finite dataset associated with a finite task set, operating with sufficient accuracy for that finite task set (e.g., 95% of the audio inputs are accurately transcribed). As an example, and not a limitation, there can be an ASR engine for messaging / calling, an ASR engine for music-related functions, and an ASR engine for default system operations. For example, the ASR engine for messaging / calling can accurately transcribe a threshold number of audio inputs associated with a messaging / call request (e.g., 95%). Therefore, the assistant system 140 can utilize individual ASR engines to improve the accuracy of the ASR results. To this end, the assistant system 140 can receive audio input and send it to multiple ASR engines. By sending the audio input to multiple ASR engines, each ASR engine can generate a transcription based on its corresponding speech model. This improves the accuracy of the ASR results by increasing the probability of accurate transcription of the audio input. As an example, and not a limitation, if a user requests to play music using the assistant system 140, the audio input is sent to all available ASR engines, one of which is the ASR engine for music-related functions. The ASR engine for music-related functions can accurately transcribe the audio input into a request to play music. By sending the audio input corresponding to a request to play music to an ASR engine used for music-related functions, the assistant system 140 can improve the accuracy of audio input transcription because a particular ASR engine can have a large amount of training data corresponding to audio input associated with music-related functions. By using multiple ASR engines, the assistant system 140 can have a strong foundation of speech models to process and transcribe different requests, such as music-related requests or messaging / call-related requests.Using multiple ASR engines, each associated with its own speech model, can help avoid the need to extensively train a single speech model on large datasets, such as training a speech model to transcribe requests to play music as well as requests to transcribe weather information. These individual ASR engines may already be trained for their respective functions and may not require any further training to achieve sufficient accuracy for operability of their respective functions. Therefore, using individual ASR engines can reduce or eliminate the time required to train the speech model to achieve sufficient operability when processing / transing audio inputs associated with a wide range of functions (e.g., being able to accurately transcribe a threshold number of audio inputs). The assistant system 140 can acquire the output of the ASR engines and send it to a Natural Language Understanding (NLU) module, which identifies one or more intents and one or more slots associated with each output of the respective ASR engine. The output of the NLU module can be sent to a meta-speech engine, which can select one or more intents and one or more slots associated with the received audio input from possible choices based on a selection strategy. The selection strategy can be, for example, a simple selection strategy, a system combination of strategies with machine learning models, or a ranking strategy. After determining the intents and slots, the meta-speech engine can send its output for processing by one or more agents of the assistant system 140. In a particular embodiment, the meta-speech engine may be a component of the assistant system 140, which processes audio input to generate combinations of intents and slots. In a particular embodiment, the meta-speech engine may include multiple ASR engines and an NLU module. Although this disclosure describes the use of multiple ASR engines to analyze audio input in a particular manner, this disclosure contemplates the use of multiple ASR engines to analyze audio input in any suitable manner.

[0130] In a particular embodiment, assistant system 140 may receive audio input from client system 130. In a particular embodiment, client system 130 may be associated with a user. By way of example, and not limitation, audio input may be received from a user's smartphone, a user's assistant system, or another computing device of the user. In a particular embodiment, the user may speak to client system 130, which then passes the audio input to assistant system 140. In a particular embodiment, assistant system 140 may be client system 130 receiving audio input from a user. After receiving the audio input, assistant system 140 may pass the audio input to an assistant orchestrator, which passes the audio input to be processed in a server-side process. In a particular embodiment, assistant system 140 may have one or more stored ASR engines capable of processing the audio input. These ASR engines may be updated periodically based on remotely located speech models. By way of example, and not limitation, ASR engines may be updated monthly, weekly, daily, etc., based on speech models from remote computing systems. ASR engines may also be updated based on when a new update is detected. There may be locally stored ASR engines. Although this disclosure describes receiving audio input in a particular manner, this disclosure contemplates receiving audio input in any suitable manner.

[0131] In a particular embodiment, assistant system 140 may generate multiple transcriptions corresponding to audio input. The transcription may be a text-to-text conversion of the audio input. In a particular embodiment, assistant system 140 may have a process for selecting which ASR engine to send the audio input to for transcription. In a particular embodiment, assistant system 140 may use multiple ASR engines (e.g., sending audio input to multiple ASR engines or all available ASR engines) to generate multiple transcriptions. ASR engines may be stored locally, on a remote computing system, or both. Assistant system 140 may use a combination of locally stored and / or remotely stored ASR engines. In a particular embodiment, an ASR engine may be a third-party ASR engine associated with third-party system 170. The third-party ASR engine may be separate from and outside assistant system 140. In a particular embodiment, one or more ASR engines may be added for use by assistant system 140. The additional ASR engines may be stored locally and / or on a remote computing system. In a particular embodiment, each ASR engine may be associated with a specific domain among multiple domains. As an example, and not a limitation, an ASR engine for shopping-related functions may be associated with a shopping domain, or an ASR engine for music-related functions may be associated with a music domain or a more general media domain. Each domain may include one or more agents specific to that domain. As an example, and not a limitation, a music domain may be associated with an online music application agent. Each domain may be associated with a set of tasks specific to that domain. As an example, and not a limitation, tasks associated with a shopping domain may include purchase tasks or add-to-cart tasks. In a particular embodiment, two or more separate ASR engines may be combined to produce a combined ASR engine. Each of these previous separate ASR engines may be associated with a single domain. As an example, and not a limitation, an ASR engine for music-related functions and an ASR engine for video-related functions may be combined to generate an ASR engine for media-related functions. Each ASR engine may be associated with a set of agents from a plurality of agents. In a particular embodiment, agents may include one or more first-party agents or third-party agents. In a particular embodiment, a domain may be associated with a plurality of agents, each operable to perform domain-specific tasks. As an example, and not a limitation, a music domain may be associated with an online music application agent capable of playing songs. In a particular embodiment, generating multiple transcriptions may include sending audio input to each ASR engine and receiving multiple transcriptions from each ASR engine. In a particular embodiment, the assistant system 140 may send audio input to a third-party ASR engine to generate one or more transcriptions. The third-party ASR engine may return the generated transcriptions. The assistant system 140 may select from the received transcriptions to send to an NLU module to determine the intent and slot associated with the transcription.Although this disclosure describes generating multiple transcripts corresponding to an audio input in a particular manner, this disclosure contemplates generating multiple transcripts corresponding to an audio input in any suitable manner.

[0132] In a particular embodiment, assistant system 140 can determine a combination of intents and slots associated with transcription. Assistant system 140 can collect each generated transcription and determine one or more intents and one or more slots associated with that transcription. By way of example and not limitation, for the transcription "Where is the nearest gas station?", assistant system 140 can determine that the intent is [IN: find location] and the slot is [SL: gas station]. In a particular embodiment, assistant system 140 can send transcriptions received from the ASR engine to the NLU module to determine one or more intents and one or more slots associated with the transcription. Although this disclosure describes determining the combination of intents and slots associated with transcription in a particular manner, this disclosure contemplates determining the combination of intents and slots associated with transcription in any suitable manner.

[0133] In a particular embodiment, assistant system 140 may select one or more combinations of intents and slots associated with audio input from a plurality of combinations. After generating multiple combinations of intents and slots from multiple transcriptions, assistant system 140 may select a combination associated with the audio input to determine what the user has requested. By way of example, and not limitation, assistant system 140 may determine whether the audio input corresponds to a request to perform a search, play a song, etc., by selecting a combination of intents and slots representing the audio input. In a particular embodiment, assistant system 140 may use a meta-speech engine to select one or more combinations of intents and slots. For example, a meta-speech engine may perform the selection process described herein. In a particular embodiment, assistant system 140 may identify a domain for each combination of intents and slots. By way of example, and not limitation, an intent [IN: find location] and a slot [SL: grocery store] may be associated with a navigation domain. Assistant system 140 may perform a mapping of the domains of the combinations of intents and slots to domains associated with multiple ASR engines. By way of example, and not limitation, for two ASR engines, an ASR engine for shopping-related functions and an ASR engine for music-related functions, assistant system 140 may map the domains of each combination of intents and slots to each ASR engine. Therefore, if a combination of intent and slot is associated with a shopping domain, then that particular combination can be mapped to an ASR engine for shopping-related functions. In a particular embodiment, if a combination of intent and slot is mapped to an ASR engine, the assistant system 140 can select to use the combination mapped to the ASR engine. If multiple mappings exist (e.g., multiple combinations based on the domain to different ASR engines), the assistant system 140 can select the combination with the maximum number of mappings. As an example, and not a limitation, if three combinations are mapped to ASR engines for shopping-related functions based on the domain, but one combination is mapped to an ASR engine for music-related functions based on the domain, then the assistant system 140 can select to use the combination mapped to the ASR engine for shopping-related functions.

[0134] In a particular embodiment, an ASR engine for default system operation may exist. If combinations of intents and slots determined from transcriptions generated by the ASR engine for default system operation are mapped to the domain of the ASR engine for default system operation, then combinations of intents and slots can be selected to be associated with the audio input. In a particular embodiment, combinations associated with the ASR engine for default system operation can be used if a conflict or error exists. As an example, and not a limitation, combinations of intents and slots associated with transcriptions generated from the ASR engine for default system operation can be used to represent audio input. In a particular embodiment, the assistant system 140 can use a machine learning model to rank combinations of intents and slots. The assistant system 140 can identify features of combinations of intents and slots, each of which can indicate whether the combination has a particular attribute. As an example, and not a limitation, the features of a combination can relate to whether the combination references a location. If a user is determined to be traveling (e.g., in a vehicle), the user may be interested in the direction and combination of reference locations, which may be ranked higher than combinations that do not reference locations. The assistant system 140 can select combinations based on their ranking, such that higher-ranked combinations can be selected to represent the audio input. In a particular embodiment, the assistant system 140 can identify identical combinations of intents and slots and rank these identical combinations based on the number of such combinations. As an example, and not a limitation, if an ASR engine for general functions generates transcriptions associated with shopping domain combinations of intents and slots, and an ASR engine for shopping-related functions generates transcriptions associated with shopping domain combinations of intents and slots, then shopping domain combinations of intents and slots can be ranked higher than other combinations. While this disclosure describes selecting one or more combinations of intents and slots associated with audio input from a plurality of combinations in a particular manner, this disclosure contemplates selecting one or more combinations of intents and slots associated with audio input from a plurality of combinations in any suitable manner.

[0135] In a particular embodiment, assistant system 140 may generate a response to audio input based on a selected combination. After selecting a combination of intent and slot associated with the audio input, assistant system 140 may send the combination of intent and slot to inference module 222 to parse the intent and slot and generate a response. As an example, and not a limitation, for the intent [IN: find location] and slot [SL: gas station], assistant system 140 may identify the nearest gas station or multiple nearby gas stations. Assistant system 140 may generate a response including a list of gas stations for the user to view. In a particular embodiment, assistant system 140 may send the selected combination to multiple agents for parsing. Assistant system 140 may receive multiple responses from each agent and sort the responses from the agents. Assistant system 140 may select one response from the multiple responses based on the sorting. In a particular embodiment, the response may be an action performed by assistant system 140 or a result generated from a query. As an example, and not a limitation, the user may have requested to turn on the lights or requested information about the weather. Although this disclosure describes generating a response to audio input in a particular manner, this disclosure contemplates generating a response to audio input in any suitable manner.

[0136] In a particular embodiment, assistant system 140 may send instructions to present a response to audio input. Assistant system 140 may send instructions to client system 130 to present a response to audio input to a user. In a particular embodiment, the instructions for presenting the response include notification of an action to be performed or a list of one or more results. By way of example, and not limitation, assistant system 140 may send instructions to client system 130 to perform an action and / or notify the user that the action has been completed. For example, if a user requests to turn on a light, the system may notify the user that the light is on. By way of another example, and not limitation, if a user requests weather information, assistant system 140 may send weather information to client system 130 to present to the user. Although this disclosure describes sending instructions to present a response to audio input in a particular manner, this disclosure contemplates sending instructions to present a response to audio input in any suitable manner.

[0137] Figure 5An exemplary flowchart 500 is shown illustrating the use of multiple Automatic Speech Recognition (ASR) engines 208 to generate transcriptions of audio input. In a particular embodiment, a client system 130 may receive audio input from a user. The client system 130 may send the audio input to an assistant orchestrator 206. The assistant orchestrator 206 may send the audio input to a meta-engine 502. The meta-engine 502 may reside on the client system 130, be locally stored on the assistant system 140, or be stored on a remote computing system. The meta-engine 502 may process the received audio input and send it to multiple ASR engines 208. In a particular embodiment, the ASR engines 208 may be ASR engine 216 and / or a combination of both. Each ASR engine 208 may generate one or more transcriptions of the audio input to be sent to an NLU module 210. Although only three ASR engines 208 are shown, any number of ASR engines 208 may be present, such as two, four, five, etc. The NLU module 210 may determine one or more intentions and one or more slot combinations associated with each transcription. NLU module 210 can send combinations of intents and slots to meta-engine 502. Meta-engine 502 can select one or more combinations of intents and slots associated with audio input. In a particular embodiment, meta-engine 502 can send the selected combinations to inference module 214. Inference module 214 can generate a response to the audio input based on the selected combinations. In a particular embodiment, if there are multiple combinations of selections, inference module 214 can parse all combinations of intents and slots. Inference module 214 can sort each result and select a result based on that sorting. As an example and not a limitation, one combination of selections may include the intent [IN: Play music] and [SL: Song_1], while another combination of selections may include the intent [IN: Query] and slot [SL: Song_1]. Inference module 214 can determine that client system 130 was previously playing music and therefore sort the selected combination with intent [IN: Play music] above the other combination. Inference module 214 can determine that the result is to perform an action, instruct client system 130 to perform an action, or send the result to client system 130 to present to the user. In a particular embodiment, the inference module 214 may determine that a notification instructing the user to perform an action needs to be sent, and send an instruction to the client system 130 to present the notification to the user.

[0138] Figure 6An exemplary flowchart 600 is shown, illustrating the use of multiple selection strategies to choose combinations of intents and slots to generate a response. In a particular embodiment, meta-engine 502 may include a meta-turn 602 interfacing with ASR engines 208a-208b, a meta-runtime client 604, multiple ASR strategies 606, and a meta-turn 608 performing the selection of combinations of intents and slots. In a particular embodiment, meta-turn 602 may receive audio input from client system 130. By way of example and not limitation, meta-turn 602 may receive audio input from assistant orchestrator 206. In a particular embodiment, flowchart 600 may occur on client system 130 or a remote computing system. After receiving audio input, meta-turn 602 may send the audio input to all available ASR engines 208 to process the audio input and generate multiple transcriptions. ASR engines 208 may return the transcriptions to meta-runtime client 604. Although only two ASR engines are shown, any number of ASR engines 208 may be present. In a particular embodiment, meta-runtime client 604 may determine the intent and slot associated with each transcription. The meta-runtime client 604 can receive transcriptions from the ASR engine 208 and send the transcriptions to the NLU module 210 to determine the intent and slot associated with each transcription. The NLU module 210 can return to the meta-runtime client 604 a determined combination of intents and slots associated with each transcription. In a particular embodiment, the meta-runtime client 604 can determine an ASR strategy 606 to select the combination of intents and slots. In a particular embodiment, the meta-runtime client 604 may have a sequence of which strategy is initially tried or initiated. In a particular embodiment, the meta-runtime client 604 may initially perform a mapping of the domains associated with the combination to the ASR engine 208 via ASR strategy 1 606a. For this purpose, ASR strategy 1 606a can use ontology 440 to map the domains of the combination of intents and slots to the domains of the ASR engine. In a particular embodiment, the meta-runtime client 604 can use the combination associated with the domains mapped to the ASR engine 208. As an example, and not a limitation, if ASR engine 208 is associated with a music domain, and a combination of intents and slots is identified as associated with a music domain, then the meta-runtime client 604 can select that combination of intents and slots. In a particular embodiment, if the domain of the combination of intents and slots is not mapped to ASR engine 208, then the meta-runtime client 604 can use another ASR strategy 606. ASR strategy 606 may include multiple selection strategies as described herein. In a particular embodiment, an ASR strategy 606 may include conflict or error resolution. As an example, and not a limitation, in the presence of conflict or error, a combination of intents and slots generated from transcription from the default ASR engine 208 may be selected as associated with the audio input. Although only three ASR strategies 606 are shown, any number of ASR strategies 606 may be available.After selecting a combination of intent and slot via one of the ASR strategies 606, the combination can then be sent to the metawheel 608. The metawheel 608 can then send the selected combination to the inference module 214 or another component of the assistant system 140. The output of the metawheel 608 can be processed to determine the response to the audio input.

[0139] Figure 7 An example mapping 700 is shown of the Automatic Speech Recognition (ASR) engine 208 to domain 702 and its corresponding agent 704 and available tasks 706. While mapping 700 shows a certain number of ASR engines 208, domains 702, agents 704, and tasks 706, any number of each can be combined with the others. In a particular embodiment, the mapping includes ASR engines 1 208a and 2 208b mapped to domains 1 702a and 2 702b, respectively. In a particular embodiment, each domain 702 may have its own domain-specific corresponding agent 704. In a particular embodiment, each agent 704 may perform a task 706 specific to that domain 702. In a particular embodiment, one or more agents 704 may share the same task 706. By way of example and not limitation, agent 2 704b may perform the same task 1 706d as task 1 706f of agent 3 704c. In a particular embodiment, there may be an ASR engine 208 that may have a domain 702 overlapping with other ASR engines 208. By way of example and not limitation, a generic ASR engine 208 may share domain 702 with other ASR engines 208. In a particular embodiment, two separate ASR engines 208 may be combined to generate a combined ASR engine 208. In a particular embodiment, each domain 702 may have its own set of agents 704 and tasks 706.

[0140] Figure 8An example process 800 for generating transcriptions for audio input using multiple automatic speech recognition engines is illustrated. In a particular embodiment, user 802 may say “find the best pitcher son” or something similar to audio input 804 (because assistant system 140 may not initially determine what the user wants to say). Audio input 804 may be captured by client system 130 and sent to all available ASR engines 208. By way of example, and not limitation, client system 130 may send audio input 804 to all ASR engines 208 stored on local and / or remote computing systems. Each ASR engine 208 may generate multiple transcriptions 806. By way of example, and not limitation, ASR engine 208a for general functions may generate transcriptions 806 including “find the best pitcher on…” 806a, “find the best picture son” 806b, and “find the best pitcher son” 806c, and ASR engine 208b for music-related functions may generate transcriptions 806 including “find the best picture song” 806d and “find the best pitcher song” 806e. In a particular embodiment, transcription 806 may be sent to NLU module 210 to generate multiple combinations 808 of intents and slots. Each transcription 806 may have multiple combinations 808. In a particular embodiment, assistant system 140 may use an ASR strategy to select combinations of intents and slots associated with audio input 804. As an example, and not a limitation, assistant system 140 may sort each of the combinations of intents and slots and select combinations based on that sorting. For example, assistant system 140 may determine that the user is not interested in motion and may remove transcription 1 806c from the list of combinations to be selected. As another example, and not a limitation, assistant system 140 may identify the amount of each slot and intent to select combination 808. For example, “best pitcher” is mentioned more often than “best picture” is mentioned more often and is therefore more likely to be an entity in the slot. Additionally, the number of transcriptions 806 that include “best pitcher” is greater than the number of transcriptions 806 that include “best picture”. Assistant system 140 may also determine that the number of query intents is greater than the number of play music intents and therefore determines that the intent is more likely to be a query. In a particular embodiment, assistant system 140 may determine a combination 808 comprising the two most frequently mentioned intents and slots. As an example, and not a limitation, assistant system 140 may select combination 808c. In a particular embodiment, combination 808b may refer to another team the user has not contacted, but combination 808c may refer to a team the user has contacted (e.g., having attended a game, performed a previous query, etc.). In a particular embodiment, assistant system 140 may generate a response based on the selected combination 808 and send an instruction to client system 130 to present that response to the user.

[0141] Figure 9An example method 900 for generating transcriptions for audio input using multiple automatic speech recognition engines is illustrated. The method may begin at step 910, where an assistant system 140 receives a first audio input from a client system associated with a first user. At step 920, the assistant system 140 may generate multiple transcriptions corresponding to the first audio input based on the multiple automatic speech recognition (ASR) engines. In a particular embodiment, each ASR engine may be associated with a corresponding domain among multiple domains. At step 930, the assistant system 140 may determine a combination of one or more intentions and one or more slots associated with each transcription. At step 940, the assistant system 140 may select one or more combinations of intentions and slots associated with the first user input from the multiple combinations via a meta-speech engine. At step 950, the assistant system 140 may generate a response to the first audio input based on the selected combination. At step 960, the assistant system 140 may send instructions to the client system for presenting the response to the first audio input. Where appropriate, the particular embodiment may be repeated. Figure 9 One or more steps of the method. Although this disclosure will Figure 9 The specific steps of the method are described and shown as occurring in a specific order, but this disclosure contemplates... Figure 9 Any suitable steps of the method occur in any suitable order. Furthermore, although this disclosure describes and illustrates including... Figure 9 The present disclosure describes an example method that uses multiple automatic speech recognition engines to generate a transcript of audio input using specific steps. However, this disclosure contemplates any suitable method that includes any appropriate steps for generating a transcript of audio input using multiple automatic speech recognition engines, and where appropriate, the method may include... Figure 9 All, some, or none of the methods Figure 9 The steps are described and illustrated in this disclosure. Furthermore, although this disclosure describes and illustrates the implementation... Figure 9 The method may refer to a specific component, device, or system of a particular step, but this disclosure contemplates the execution of... Figure 9 Any suitable step of the method, any suitable component, device or system, or any suitable combination thereof.

[0142] Social media

[0143] Figure 10An example social graph 1000 is illustrated. In a particular embodiment, the social network system 160 may store one or more social graphs 1000 in one or more data stores. In a particular embodiment, the social graph 1000 may include multiple nodes—which may include multiple user nodes 1002 or multiple concept nodes 1004—and multiple edges 1006 connecting these nodes. Each node may be associated with a unique entity (i.e., a user or a concept), and each entity may have a unique identifier (ID), such as a unique number or username. For educational purposes, a two-dimensional visual map representation is shown. Figure 10 The example social graph 1000 is shown. In a particular embodiment, a social network system 160, a client system 130, an assistant system 140, or a third-party system 170 can access the social graph 1000 and related social graph information for appropriate applications. The nodes and edges of the social graph 1000 can be stored as data objects in, for example, a data storage device (e.g., a social graph database). Such a data storage device may include one or more searchable or queryable indexes of the nodes or edges of the social graph 1000.

[0144] In a particular embodiment, user node 1002 may correspond to a user of social networking system 160 or assistant system 140. By way of example and not limitation, a user may be an individual (human user), entity (e.g., a business, company, or third-party application), or group (e.g., an individual or entity) that interacts or communicates with or through social networking system 160 or assistant system 140. In a particular embodiment, when a user registers an account with social networking system 160, social networking system 160 may create a user node 1002 corresponding to the user and store user node 1002 in one or more data storage devices. The user and user node 1002 described herein may, where appropriate, refer to a registered user and the user node 1002 associated with the registered user. Alternatively, where appropriate, the user and user node 1002 described herein may refer to a user who has not registered with social networking system 160. In a particular embodiment, user node 1002 may be associated with information provided by the user or information collected by various systems, including social networking system 160. As an example, and not a limitation, a user may provide his or her name, profile picture, contact information, date of birth, gender, marital status, family status, occupation, educational background, preferences, interests, or other demographic information. In a particular embodiment, user node 1002 may be associated with one or more data objects corresponding to information associated with the user. In a particular embodiment, user node 1002 may correspond to one or more web interfaces.

[0145] In a particular embodiment, concept node 1004 may correspond to a concept. By way of example and not limitation, a concept may correspond to a location (such as, for example, a movie theater, restaurant, landmark, or city); a website (such as, for example, a website associated with social networking system 160 or a third-party website associated with a web application server); an entity (such as, for example, an individual, business, group, sports team, or celebrity); a resource (such as, for example, an audio file, video file, digital photograph, text file, structured document, or application), which may reside on a server (e.g., a web application server) within or outside social networking system 160; real estate or intellectual property (such as, for example, sculpture, painting, film, game, song, idea, photograph, or written work); a game; an activity; an idea or theory; another suitable concept; or two or more such concepts. Concept node 1004 may be associated with information about a concept provided by a user or information collected by various systems (including social networking system 160 and assistant system 140). By way of example and not limitation, concept information may include a name or title; one or more images (e.g., an image of a book cover); location (e.g., an address or geographic location); a website (which may be associated with a URL); contact information (e.g., a phone number or email address); other suitable concept information; or any suitable combination of such information. In a particular embodiment, concept node 1004 may be associated with one or more data objects corresponding to information associated with concept node 1004. In a particular embodiment, concept node 1004 may correspond to one or more web interfaces.

[0146] In a particular embodiment, nodes in social graph 1000 may represent or be represented by a web interface (which may be referred to as a “profile interface”). The profile interface may be hosted by or accessible to social network system 160 or assistant system 140. The profile interface may also be hosted on a third-party website associated with third-party system 170. By way of example, and not limitation, a profile interface corresponding to a particular external web interface may be that particular external web interface, and the profile interface may correspond to a particular concept node 1004. The profile interface may be viewable by all other users or a selected subset of other users. By way of example, and not limitation, user node 1002 may have a corresponding user profile interface, where the corresponding user can add content, make statements, or otherwise express himself or her. By way of another example, and not limitation, concept node 1004 may have a corresponding concept profile interface, where one or more users can add content, make statements, or express themselves, particularly regarding the concept corresponding to concept node 1004.

[0147] In a particular embodiment, concept node 1004 may represent a third-party web interface or resource hosted by third-party system 170. Among other elements, the third-party web interface or resource may include content, selectable or other icons, or other interactive objects representing actions or activities. By way of example and not limitation, the third-party web interface may include selectable icons such as “like,” “check-in,” “eat,” “recommend,” or other suitable actions or activities. A user viewing the third-party web interface can perform an action by selecting one of the icons (e.g., “check-in”), causing client system 130 to send a message instructing the user’s action to social network system 160. In response to this message, social network system 160 may create an edge (e.g., a check-in type edge) between user node 1002 corresponding to the user and concept node 1004 corresponding to the third-party web interface or resource, and store edge 1006 in one or more data stores.

[0148] In a particular embodiment, a pair of nodes in social graph 1000 can be connected to each other via one or more edges 1006. The edge 1006 connecting a pair of nodes can represent the relationship between the pair of nodes. In a particular embodiment, edge 1006 may include or represent one or more data objects or attributes corresponding to the relationship between the pair of nodes. As an example, and not a limitation, a first user can indicate that a second user is a "friend" of the first user. In response to this indication, social network system 160 can send a "friend request" to the second user. If the second user confirms the "friend request," social network system 160 can create an edge 1006 in social graph 1000 connecting the first user's user node 1002 to the second user's user node 1002, and store edge 1006 as social graph information in one or more data stores 164. Figure 10In the example, social graph 1000 includes edges 1006 indicating a friendship relationship between user nodes 1002 of user "A" and user "B", and edges indicating a friendship relationship between user nodes 1002 of user "C" and user "B". Although this disclosure describes or illustrates a specific edge 1006 with a specific attribute relating to a particular user node 1002, this disclosure contemplates any suitable edge 1006 relating to user nodes 1002 with any suitable attribute. By way of example and not limitation, edge 1006 may represent friendship, family relationship, business or employment relationship, fan relationship (including, for example, likes, etc.), follower relationship, visitor relationship (including, for example, visit, view, check-in, share, etc.), subscriber relationship, superior / subordinate relationship, reciprocal relationship, non-reciprocal relationship, another suitable type of relationship, or two or more such relationships. Furthermore, although this disclosure generally describes nodes as being associated, this disclosure also describes users or concepts as being associated. In this document, references to connected users or concepts may, where appropriate, refer to the nodes in the social graph 1000 that are connected by one or more edges 1006. The separation degree between two objects represented by two nodes is the count of edges in the shortest path connecting the two nodes in the social graph 1000. As an example and not a limitation, in the social graph 1000, user node 1002 of user "C" is connected to user node 1002 of user "A" via multiple paths, such as a first path directly passing through user node 1002 of user "B", a second path passing through concept node 1004 of company "A1me" and user node 1002 of user "D", and a third path passing through user node 1002 and concept node 1004 representing school "Stateford", user "G", company "A1me", and user "D". User "C" and user "A" have a separation degree of two because the shortest path connecting their corresponding nodes (i.e., the first path) includes two edges 1006.

[0149] In a particular embodiment, edge 1006 between user node 1002 and concept node 1004 may represent a specific action or activity performed by a user associated with user node 1002 on a concept associated with concept node 1004. This is intended as an example, not a limitation. Figure 10As shown, users can "like," "attend," "play," "listen," "cook," "work at," or "read," each of which can correspond to an edge type or subtype. The concept profile interface corresponding to concept node 1004 can include, for example, an optional "check-in" icon (such as, for example, a clickable "check-in" icon) or an optional "add to favorites" icon. Similarly, after a user clicks these icons, the social network system 160 can create a "favorites" edge or a "check-in" edge in response to the user's action corresponding to the respective action. As another example, and not as a limitation, a user (user "C") can use a specific application (a third-party online music application) to listen to a specific song ("Imagine"). In this case, the social network system 160 can create a "listen" edge 1006 and a "use" edge (such as, for example, a clickable "check-in" icon) between the user node 1002 corresponding to the user and the concept node 1004 corresponding to the song and application. Figure 10 As shown), to indicate that the user has listened to the song and used the application. Furthermore, the social network system 160 can create a "play" edge 1006 (as shown) between the concept nodes 1004 corresponding to the song and the application. Figure 10 As shown), to indicate that a specific song is played by a specific application. In this case, the "play" edge 1006 corresponds to the action performed by an external application (a third-party online music application) on an external audio file (the song "Imagine"). Although this disclosure describes a specific edge 1006 with a specific attribute connecting user node 1002 and concept node 1004, this disclosure contemplates any appropriate edge 1006 with any appropriate attribute connecting user node 1002 and concept node 1004. Furthermore, although this disclosure describes an edge between user node 1002 and concept node 1004 representing a single relationship, this disclosure contemplates edges between user node 1002 and concept node 1004 representing one or more relationships. By way of example and not limitation, edge 1006 could represent that a user likes and uses a specific concept. Alternatively, another edge 1006 could represent a relationship between user node 1002 and concept node 1004 (such as...). Figure 10 As shown, there are various types of relationships (or multiple single relationships) between user node 1002 of user "E" and concept node 1004 of "Online Music Application".

[0150] In a particular embodiment, the social networking system 160 may create an edge 1006 between user node 1002 and concept node 1004 in the social graph 1000. As an example, and not a limitation, (such as, for example, using a web browser or dedicated application hosted by the user's client system 130) a user viewing a concept profile interface may indicate that he or she likes the concept represented by concept node 1004 by clicking or selecting a "like" icon. This may cause the user's client system 130 to send a message to the social networking system 160 indicating that the user likes the concept associated with the concept profile interface. In response to this message, the social networking system 160 may create an edge 1006 between the user node 1002 and concept node 1004 associated with the user, as shown by the "like" edge 1006 between the user node and concept node 1004. In a particular embodiment, the social networking system 160 may store the edge 1006 in one or more data stores. In a particular embodiment, the edge 1006 may be automatically formed by the social networking system 160 in response to a specific user action. As an example, and not a limitation, if a first user uploads a picture, watches a movie, or listens to a song, an edge 1006 can be formed between the user node 1002 corresponding to the first user and the concept node 1004 corresponding to those concepts. Although this disclosure describes the formation of a particular edge 1006 in a specific manner, this disclosure contemplates the formation of any suitable edge 1006 in any suitable manner.

[0151] Vector space and embedding

[0152] Figure 11 An example view of vector space 1100 is shown. In a particular embodiment, objects or n-grams can be represented in a d-dimensional vector space, where d represents any suitable dimension. Although vector space 1100 is shown as a three-dimensional space, this is merely for illustrative purposes, as vector space 1100 can have any suitable dimension. In a particular embodiment, n-grams can be represented as vectors in vector space 1100, referred to as term embeddings. Each vector can include coordinates corresponding to a specific point in vector space 1100 (i.e., the endpoint of the vector). This is intended as an example and not as a limitation. Figure 11 As shown, vectors 1110, 1120, and 1130 can be represented as points in vector space 1100. n-grams can be mapped to their corresponding vector representations. This is shown as an example, not a restriction, by applying functions defined by a dictionary. n-gramt1 and n-gramt2 can be mapped to vectors in vector space 1100, respectively. and Make and As another example, and not a limitation, a dictionary trained to map text to vector representations can be utilized, or such a dictionary can be generated through training. As another example, and not a limitation, a word embedding model can be used to map n-grams to vector representations in vector space 1100. In a particular embodiment, n-grams can be mapped to vector representations in vector space 1100 using a machine learning model (e.g., a neural network). The machine learning model may have been trained using sequences of training data (e.g., corpora of multiple objects, each comprising an n-gram).

[0153] In a particular embodiment, an object may be represented as a vector in vector space 1100, which is referred to as a feature vector or object embedding. This is done as an example, not as a limitation, by applying a function. Objects e1 and e2 can be mapped to vectors in vector space 1100, respectively. and Make and In certain embodiments, an object can be mapped to a vector based on one or more characteristics, attributes, or features of the object, the object's relationship to other objects, or any other suitable information associated with the object. This is an example, not a limitation, of the function. Objects can be mapped to vectors through feature extraction, which can start from an initial measurement dataset and construct derived values ​​(e.g., features). As an example, and not a limitation, objects, including those in videos or images, can be mapped to vectors by using algorithms to detect or isolate various desired parts or shapes of the object. Features used to compute the vectors can be based on information obtained from edge detection, corner detection, blob detection, ridge detection, scale-invariant feature transforms, edge orientation, change intensity, autocorrelation, motion detection, optical flow, thresholding, blob extraction, template matching, Hough transforms (e.g., lines, circles, ellipses, arbitrary shapes), or any other suitable information. As another example, and not a limitation, objects including audio data can be mapped to vectors based on features (e.g., spectral slope, pitch coefficient, audio spectral centroid, audio spectral envelope, Mel-frequency cepstrum, or any other suitable information). In certain embodiments, when the object has data that is too large to be efficiently processed or includes redundant data, the function... An object can be mapped to a vector using a transformed, streamlined feature set (e.g., feature selection). In a particular embodiment, the function... An object e can be mapped to a vector based on one or more n-grams associated with it. Although this disclosure describes the representation of n-grams or objects in a vector space in a particular manner, this disclosure contemplates the representation of n-grams or objects in a vector space in any suitable manner.

[0154] In a particular embodiment, the social network system 160 can compute a similarity measure of vectors in the vector space 1100. The similarity measure can be cosine similarity, Minkowski distance, Mahalanobis distance, Jaccard similarity coefficient, or any suitable similarity measure. This is provided as an example and not as a limitation. and The similarity measure can be cosine similarity. As another example, rather than as a limitation and Similarity can be measured by Euclidean distance. A similarity measure between two vectors can represent the degree of similarity between two objects or n-grams corresponding to the two vectors, as measured by the distance between the two vectors in vector space 1100. As an example, and not a limitation, vectors 1110 and 1120 can correspond to objects that are more similar to each other than the objects corresponding to vectors 1110 and 1130, based on the distance between the corresponding vectors. Although this disclosure describes the computation of a similarity measure between vectors in a particular manner, this disclosure contemplates the computation of a similarity measure between vectors in any suitable manner.

[0155] More information on vector spaces, embeddings, feature vectors, and similarity measures can be found in U.S. Patent Application No. 14 / 949436, filed November 23, 2015; U.S. Patent Application No. 15 / 286315, filed October 5, 2016; and U.S. Patent Application No. 15 / 365789, filed November 30, 2016.

[0156] Artificial Neural Networks

[0157] Figure 12An example artificial neural network (“ANN”) 1200 is illustrated. In a particular embodiment, an ANN may refer to a computational model comprising one or more nodes. The example ANN 1200 may include an input layer 1210, hidden layers 1220, 1230, 1240, and an output layer 1250. Each layer of the ANN 1200 may include one or more nodes, such as node 1205 or node 1215. In a particular embodiment, each node of the ANN may be connected to another node of the ANN. By way of example and not limitation, each node of the input layer 1210 may be connected to one or more nodes of the hidden layer 1220. In a particular embodiment, one or more nodes may be bias nodes (e.g., nodes in a layer that are not connected to any node in the previous layer and do not receive input from them). In a particular embodiment, each node in each layer may be connected to one or more nodes in the previous or next layer. Although Figure 12 This disclosure describes a specific ANN with a specific number of layers, a specific number of nodes, and specific relationships between nodes; however, this disclosure contemplates any suitable ANN with any suitable number of layers, any suitable number of nodes, and any suitable relationships between nodes. This is provided as an example and not as a limitation, although... Figure 12 The relationships between each node in the input layer 1210 and each node in the hidden layer 1220 are depicted, but one or more nodes in the input layer 1210 may not be related to one or more nodes in the hidden layer 1220.

[0158] In a particular embodiment, the ANN may be a feedforward ANN (e.g., an ANN without loops or cycles, where communication between nodes flows in one direction starting from the input layer and progressing to successive layers). By way of example, and not limitation, the input to each node of hidden layer 1220 may include the outputs of one or more nodes of input layer 1210. By way of another example, and not limitation, the input to each node of output layer 1250 may include the outputs of one or more nodes of hidden layer 1240. In a particular embodiment, the ANN may be a deep neural network (e.g., a neural network including at least two hidden layers). In a particular embodiment, the ANN may be a deep residual network. A deep residual network may be a feedforward ANN, which includes hidden layers organized into residual blocks. The input to each residual block after the first residual block may be a function of the output and input of the previous residual block. By way of example, and not limitation, the input to residual block N may be F(x) + x, where F(x) may be the output of residual block N-1, and x may be the input to residual block N-1. Although this disclosure describes a specific ANN, it envisions any suitable ANN.

[0159] In a particular embodiment, the activation function may correspond to each node of the ANN. The activation function of a node may define the node's output for a given input. In a particular embodiment, the node's input may include a set of inputs. By way of example and not limitation, the activation function may be an identity function, a binary step function, a logic function, or any other suitable function. By way of another example and not limitation, the activation function of node k may be a sigmoid function. hyperbolic tangent function Rectifier F k (s k ) = max(0, s k or any other suitable function F k (s k ), where s k This can be a valid input to node k. In a particular embodiment, the input to the activation function corresponding to the node can be weighted. Each node can generate an output using the corresponding activation function based on the weighted input. In a particular embodiment, each relationship between nodes can be associated with a weight. As an example, and not a limitation, the relationship 1225 between nodes 1205 and 1215 can have a weighting coefficient of 0.4, which can indicate that the output of node 1205 multiplied by 0.4 is used as the input to node 1215. As another example, and not a limitation, the output y of node k... k It can be y k =F k (s k ), where F k It can be the activation function corresponding to node k, s k =∑ j (w jk x j x can be a valid input to node k. j It can be the output of node j connected to node k, and w jk This can be a weighting coefficient between node j and node k. In a particular embodiment, the input to a node in the input layer can be based on a vector representing the object. Although this disclosure describes specific inputs and outputs of nodes, it considers any suitable inputs and outputs of nodes. Furthermore, although this disclosure may describe specific relationships and weights between nodes, it considers any suitable relationships and weights between nodes.

[0160] In certain embodiments, training data can be used to train an ANN. By way of example, and not limitation, the training data may include the inputs and expected outputs of an ANN 1200. By way of another example, and not limitation, the training data may include vectors, each representing a training object and the expected label for each training object. In certain embodiments, training an ANN may include modifying the weights associated with the connections between nodes of the ANN by optimizing an objective function. By way of example, and not limitation, training methods (e.g., conjugate gradient, gradient descent, stochastic gradient descent) may be used to backpropagate the sum of squared errors as a distance measurement between each vector representing a training object (e.g., using a cost function that minimizes the sum of squared errors). In certain embodiments, dropout techniques may be used to train the ANN. By way of example, and not limitation, one or more nodes may be temporarily ignored during training (e.g., not receiving input and not generating output). For each training object, one or more nodes of the ANN may have a certain probability of being ignored. The nodes ignored for a particular training object may differ from the nodes ignored for other training objects (e.g., nodes may be temporarily ignored object-by-object). Although this disclosure describes training ANNs in a particular manner, this disclosure envisions training ANNs in any suitable manner.

[0161] privacy

[0162] In certain embodiments, one or more objects of a computing system (e.g., content or other types of objects) may be associated with one or more privacy settings. One or more objects may be stored on or otherwise associated with any suitable computing system or application, such as, for example, a social networking system 160, a client system 130, an assistant system 140, a third-party system 170, a social networking application, an assistant application, a messaging application, a photo-sharing application, or any other suitable computing system or application. Although the examples discussed herein are in the context of an online social network, these privacy settings can be applied to any other suitable computing system. The object's privacy settings (or "access settings") may be stored in any suitable manner (e.g., in association with the object, at an index on an authorization server, in another suitable manner, or any suitable combination thereof). The object's privacy settings may specify how the object (or specific information associated with the object) can be accessed, stored, or otherwise used (e.g., viewed, shared, modified, copied, performed, surfaced, or identified) within the online social network. An object can be described as "visible" relative to a specific user or other entity when its privacy settings allow access to that object by that user or other entity. As an example, and not a limitation, users of an online social network can specify privacy settings for their profile pages that identify a set of users who can access their work experience information on the profile page, thus excluding other users from accessing that information.

[0163] In certain embodiments, the privacy settings of an object may specify a “blocked list” of users or other entities that should not be allowed to access certain information associated with the object. In certain embodiments, the blacklist may include third-party entities. The blocked list may specify one or more users or entities to whom the object is not visible. As an example, and not a limitation, a user may specify a group of users who cannot access an album associated with that user, thereby excluding those users from accessing the album (while potentially allowing access to the album to some users not in the specified user group). In certain embodiments, privacy settings may be associated with a specific social graph element. The privacy settings of a social graph element (e.g., a node or edge) may specify how the online social network can be used to access the social graph element, information associated with the social graph element, or objects associated with the social graph element. As an example, and not a limitation, a specific conceptual node 1004 corresponding to a specific photo may have privacy settings specifying that the photo can only be accessed by the user tagged in the photo and the friends of the user tagged in the photo. In certain embodiments, privacy settings may allow users to opt in or out so that their content, information, or actions are stored / recorded by the social network system 160 or assistant system 140 or shared with other systems (e.g., third-party system 170). Although this disclosure describes the use of a particular privacy setting in a particular manner, this disclosure contemplates the use of any suitable privacy setting in any suitable manner.

[0164] In certain embodiments, privacy settings can be based on one or more nodes or edges of the social graph 1000. Privacy settings can be specified for one or more edges 1006 or edge types of the social graph 1000, or for one or more nodes 1002, 1004 or node types of the social graph 1000. Privacy settings applied to a specific edge 1006 connecting two nodes can control whether the relationship between two entities corresponding to those two nodes is visible to other users of the online social network. Similarly, privacy settings applied to a specific node can control whether a user or concept corresponding to that node is visible to other users of the online social network. As an example, and not a limitation, a first user can share an object with the social network system 160. The object can be associated with a concept node 1004 connected to the first user's user node 1002 via edge 1006. The first user can specify privacy settings applied to a specific edge 1006 connected to the concept node 1004 of the object, or can specify privacy settings applied to all edges 1006 connected to the concept node 1004. As another example, and not a limitation, a first user can share a collection of objects of a specific object type (e.g., a collection of images). The first user can specify specific privacy settings for all objects of that particular object type associated with the first user (e.g., specifying that all images posted by the first user are only visible to the first user's friends and / or users tagged in the images).

[0165] In a particular embodiment, the social networking system 160 may present a "privacy wizard" (e.g., within a webpage, module, one or more dialog boxes, or any other suitable interface) to a first user to help the first user specify one or more privacy settings. The privacy wizard may display instructions, appropriate privacy-related information, current privacy settings, one or more input fields for accepting changes or confirmations of the specified privacy settings from the first user, or any suitable combination thereof. In a particular embodiment, the social networking system 160 may provide a "dashboard" function to the first user, which displays the first user's current privacy settings. The dashboard function may be displayed to the first user at any appropriate time (e.g., after input from the first user who invoked the dashboard function, or after a specific event or triggering action occurs). The dashboard function may allow the first user to modify one or more of their current privacy settings at any time and in any suitable manner (e.g., redirecting the first user to the privacy wizard).

[0166] Privacy settings associated with an object can specify any suitable granularity for allowing or denying access. As an example, and not a limitation, access can be specified for specific users (e.g., only me, my roommate, my boss), users within a specific separation (e.g., friends, friends of friends), user groups (e.g., gaming clubs, my family), user networks (e.g., employees of a specific employer, students or alumni of a specific university), all users (“public”), no users (“private”), users of third-party systems, specific applications (e.g., third-party applications, external websites), other suitable entities, or any suitable combination thereof. While this disclosure describes specific granularities for allowing or denying access, this disclosure contemplates any suitable granularity for allowing or denying access.

[0167] In a particular embodiment, one or more servers 162 may be authorization / privacy servers for implementing privacy settings. In response to a request from a user (or other entity) for a specific object stored in data storage 164, the social networking system 160 may send a request for that object to data storage 164. The request may identify the user associated with the request, and the object may only be sent to the user (or the user's client system 130) if the authorization server determines, based on the privacy settings associated with the object, that the user is authorized to access the object. If the requesting user is not authorized to access the object, the authorization server may prevent the requested object from being retrieved from data storage 164 or from being sent to the user. In a search-query context, an object may be offered as a search result only if the querying user is authorized to access the object, for example, if the object's privacy settings allow it to be displayed to the querying user, discovered by the querying user, or otherwise visible to the querying user. In a particular embodiment, the object may represent content visible to the user through the user's feed. By way of example and not limitation, one or more objects may be visible to a user's "Trending" page. In certain embodiments, an object may correspond to a specific user. The object may be content associated with a specific user, or it may be a specific user's account or information stored on a social networking system 160 or other computing system. As an example, and not a limitation, a first user may view one or more second users on an online social network through the "People You May Know" feature or by viewing the first user's friend list. As an example, and not a limitation, a first user may specify that they do not wish to see objects associated with a particular second user in their feed or friend list. If an object's privacy settings do not allow it to be exposed to, discovered by, or visible to a user, that object may be excluded from search results. Although this disclosure describes implementing privacy settings in a particular manner, this disclosure contemplates implementing privacy settings in any suitable manner.

[0168] In certain embodiments, different objects of the same type associated with a user may have different privacy settings. Different types of objects associated with a user may have different types of privacy settings. As an example, and not a limitation, a first user may specify that the first user's status updates are public, but any images shared by the first user are only visible to the first user's friends on an online social network. As another example, and not a limitation, a user may specify different privacy settings for different types of entities (e.g., individual users, friends of friends, followers, user groups, or corporate entities). As another example, and not a limitation, a first user may specify a group of users who can view videos posted by the first user, while preventing the videos from being visible to the first user's employer. In certain embodiments, different privacy settings may be provided for different user groups or user demographics. As an example, and not a limitation, a first user may specify that other users attending the same university as the first user can view the first user's photos, but other users who are family members of the first user cannot view those same photos.

[0169] In a particular embodiment, the social networking system 160 may provide one or more default privacy settings for each object of a specific object type. The privacy settings of an object set as the default can be changed by the user associated with that object. By way of example and not limitation, all images posted by a first user may have a default privacy setting that is visible only to the first user's friends, and for a particular image, the first user may change the privacy settings of that image to be visible to friends and friends of friends.

[0170] In certain embodiments, privacy settings may allow a first user to specify (e.g., by opting out or not opting in) whether the social networking system 160 or assistant system 140 may receive, collect, record, or store specific objects or information associated with the user for any purpose. In certain embodiments, privacy settings may allow a first user to specify whether a particular application or process may access, store, or use specific objects or information associated with the user. Privacy settings may allow the first user to opt in or out, allowing objects or information to be accessed, stored, or used by a particular application or process. The social networking system 160 or assistant system 140 may access such information to provide specific functionality or services to the first user, but the social networking system 160 or assistant system 140 may not access the information for any other purpose. Before accessing, storing, or using such objects or information, the social networking system 160 or assistant system 140 may prompt the user to provide privacy settings that specify which applications or processes (if any) may access, store, or use the objects or information before allowing any such action. As an example, and not as a limitation, a first user may transmit messages to a second user via an application associated with an online social network (e.g., a messaging app), and may specify privacy settings that the social network system 160 or the assistant system 140 should not store such messages.

[0171] In certain embodiments, a user may specify whether the social networking system 160 or the assistant system 140 can access, store, or use a specific type of object or information associated with a first user. As an example, and not a limitation, the first user may specify that images sent by the first user through the social networking system 160 or the assistant system 140 cannot be stored by the social networking system 160 or the assistant system 140. As another example, and not a limitation, the first user may specify that messages sent from the first user to a specific second user cannot be stored by the social networking system 160 or the assistant system 140. As yet another example, and not a limitation, the first user may specify that all objects sent via a specific application can be saved by the social networking system 160 or the assistant system 140.

[0172] In certain embodiments, privacy settings may allow a first user to specify whether specific objects or information associated with the first user can be accessed from a specific client system 130 or a third-party system 170. Privacy settings may allow the first user to opt in or out of accessing objects or information from a specific device (e.g., the user's phonebook on their smartphone), a specific application (e.g., a messaging app), or a specific system (e.g., an email server). The social networking system 160 or assistant system 140 may provide default privacy settings for each device, system, or application, and / or may prompt the first user to specify specific privacy settings for each context. As an example, and not a limitation, the first user may utilize the location service features of the social networking system 160 or assistant system 140 to provide recommendations for restaurants or other places near the user. The first user's default privacy settings may specify that the social networking system 160 or assistant system 140 may use location information provided from the first user's client device 130 to provide location-based services, but the social networking system 160 or assistant system 140 may not store the first user's location information or provide it to any third-party system 170. The first user can then update their privacy settings to allow third-party image-sharing apps to use location information to geotag photos.

[0173] In certain embodiments, privacy settings may allow a user to specify one or more geographic locations from which they can access an object. Access to or denial of access to an object may depend on the geographic location of the user attempting to access the object. As an example, and not a limitation, a user may share an object and specify that only users in the same city can access or view the object. As another example, and not a limitation, a first user may share an object and specify that the object is only visible to a second user when the first user is in a specific location. If the first user leaves that specific location, the object may no longer be visible to the second user. As another example, and not a limitation, a first user may specify that the object is only visible to second users within a threshold distance of the first user. If the first user subsequently changes location, the second user who originally had access to the object may lose access, and a new group of second users may gain access when they reach within the threshold distance of the first user.

[0174] In a particular embodiment, the social networking system 160 or assistant system 140 may have the capability to use a user's personal or biometric information as input for user authentication or experience personalization purposes. Users may choose to utilize these capabilities to enhance their experience on the online social network. By way of example, and not limitation, a user may provide personal or biometric information to the social networking system 160 or assistant system 140. A user's privacy settings may specify that such information is only available for specific processes (such as authentication) and that such information cannot be shared with any third-party system 170 or used for other processes or applications associated with the social networking system 160 or assistant system 140. By way of another example, and not limitation, the social networking system 160 may provide a user with the capability to provide a voiceprint recording to the online social network. By way of example, and not limitation, if a user wishes to utilize this capability of the online social network, the user may provide a voice recording of their own voice to provide status updates on the online social network. The voice input recording can be compared with the user's voiceprint to determine what words the user spoke. A user's privacy settings can specify that such voice recordings can only be used for voice input purposes (e.g., authenticating users, sending voice messages, improving voice recognition for using voice operation features on online social networks), and also specify that such voice recordings cannot be shared with any third-party system 170, or used by other processes or applications associated with the social network system 160. As another example, and not as a limitation, the social network system 160 can provide users with the ability to provide reference images (e.g., facial contours, retinal scans) to the online social network. The online social network can compare the reference image with later received image input (e.g., for user authentication, tagging users in photos). The user's privacy settings can specify that such images can only be used for limited purposes (e.g., authentication, tagging users in photos), and also specify that such images cannot be shared with any third-party system 170, or used by other processes or applications associated with the social network system 160.

[0175] Systems and Methods

[0176] Figure 13An example computer system 1300 is illustrated. In a particular embodiment, one or more computer systems 1300 perform one or more steps of one or more methods described or illustrated herein. In a particular embodiment, one or more computer systems 1300 provide the functionality described or illustrated herein. In a particular embodiment, software running on one or more computer systems 1300 performs one or more steps of one or more methods described or illustrated herein, or provides the functionality described or illustrated herein. Specific embodiments include one or more portions of one or more computer systems 1300. Herein, references to computer systems may include computing devices and vice versa, where appropriate. Furthermore, references to computer systems may include one or more computer systems, where appropriate.

[0177] This disclosure contemplates any suitable number of computer systems 1300. The computer systems 1300 are contemplated to take any suitable physical form. By way of example and not limitation, the computer system 1300 may be an embedded computer system, a system-on-a-chip (SOC), a single-board computer system (SBC) (such as, for example, a computer-on-module (COM) or system-on-module (SOM)), a desktop computer system, a laptop or notebook computer system, an interactive kiosk, a mainframe, a computer system mesh, a mobile phone, a personal digital assistant (PDA), a server, a tablet computer system, or a combination of two or more of these. Where appropriate, the computer system 1300 may include one or more computer systems 1300; may be monolithic or distributed; may span multiple locations; may span multiple machines; may span multiple data centers; or may reside in a cloud, which may include one or more cloud components in one or more networks. Where appropriate, one or more computer systems 1300 may perform one or more steps of the methods described or shown herein without substantial spatial or temporal limitations. By way of example and not limitation, one or more computer systems 1300 may execute one or more steps of the methods described or illustrated herein in real time or in batch mode. Where appropriate, one or more computer systems 1300 may execute one or more steps of the methods described or illustrated herein at different times or at different locations.

[0178] In a particular embodiment, computer system 1300 includes a processor 1302, a memory 1304, a storage device 1306, an input / output (I / O) interface 1308, a communication interface 1310, and a bus 1312. Although this disclosure describes and illustrates a particular computer system having a particular number of particular components in a particular arrangement, this disclosure contemplates any suitable computer system having any suitable number of any suitable components in any suitable arrangement.

[0179] In a particular embodiment, processor 1302 includes hardware for executing instructions (e.g., those that constitute a computer program). By way of example, and not limitation, to execute instructions, processor 1302 may retrieve (or fetch) instructions from internal registers, internal caches, memory 1304, or storage device 1306; decode and execute them; and then write one or more results to internal registers, internal caches, memory 1304, or storage device 1306. In a particular embodiment, processor 1302 may include one or more internal caches for data, instructions, or addresses. Where appropriate, this disclosure contemplates that processor 1302 may include any suitable number of suitable internal caches. By way of example, and not limitation, processor 1302 may include one or more instruction caches, one or more data caches, and one or more translation lookup buffers (TLBs). Instructions in the instruction cache may be copies of instructions in memory 1304 or storage device 1306, and the instruction cache may accelerate the retrieval of those instructions by processor 1302. The data in the data cache may be: a copy of data in memory 1304 or storage device 1306 for operating instructions executed at processor 1302; the result of a previous instruction executed at processor 1302 for access by a subsequent instruction executed at processor 1302 or for writing to memory 1304 or storage device 1306; or other suitable data. The data cache can accelerate read or write operations performed by processor 1302. The TLB can accelerate virtual address translation with respect to processor 1302. In a particular embodiment, processor 1302 may include one or more internal registers for data, instructions, or addresses. Where appropriate, this disclosure contemplates that processor 1302 may include any suitable number of suitable internal registers. Where appropriate, processor 1302 may include one or more arithmetic logic units (ALUs); be a multi-core processor; or include one or more processors 1302. Although this disclosure describes and illustrates specific processors, this disclosure contemplates any suitable processor.

[0180] In a particular embodiment, memory 1304 includes main memory for storing instructions for execution by processor 1302 or data for operation of processor 1302. By way of example and not limitation, computer system 1300 may load instructions from storage device 1306 or another source (e.g., another computer system 1300) into memory 1304. Processor 1302 may then load instructions from memory 1304 into internal registers or internal caches. To execute instructions, processor 1302 may retrieve instructions from internal registers or internal caches and decode them. During or after instruction execution, processor 1302 may write one or more results (which may be intermediate or final results) to internal registers or internal caches. Processor 1302 may then write one or more of these results to memory 1304. In a particular embodiment, processor 1302 executes only instructions in one or more internal registers or internal caches or in memory 1304 (rather than memory device 1306 or elsewhere), and operates only on data in one or more internal registers or internal caches or in memory 1304 (rather than memory device 1306 or elsewhere). One or more memory buses (each of which may include an address bus and a data bus) couple processor 1302 to memory 1304. As described below, bus 1312 may include one or more memory buses. In a particular embodiment, one or more memory management units (MMUs) reside between processor 1302 and memory 1304 and facilitate access to memory 1304 requested by processor 1302. In a particular embodiment, memory 1304 includes random access memory (RAM). Where appropriate, the RAM may be volatile memory. Where appropriate, the RAM may be dynamic RAM (DRAM) or static RAM (SRAM). Furthermore, where appropriate, the RAM may be single-port RAM or multi-port RAM. This disclosure contemplates any suitable RAM. Where appropriate, memory 1304 may include one or more memories 1304. Although this disclosure describes and illustrates specific memories, this disclosure contemplates any suitable memory.

[0181] In a particular embodiment, storage device 1306 includes a mass storage device for data or instructions. By way of example and not limitation, storage device 1306 may include a hard disk drive (HDD), a floppy disk drive, flash memory, an optical disk drive, a magneto-optical disk drive, magnetic tape, or a Universal Serial Bus (USB) drive, or a combination of two or more of these. Where appropriate, storage device 1306 may include removable or non-removable (or fixed) media. Where appropriate, storage device 1306 may be internal or external to computer system 1300. In a particular embodiment, storage device 1306 is a non-volatile solid-state memory. In a particular embodiment, storage device 1306 includes read-only memory (ROM). Where appropriate, the ROM may be a mask-programmable ROM, a programmable ROM (PROM), an erasable PROM (EPROM), an electrically erasable PROM (EEPROM), an electrically variable ROM (EAROM), or flash memory, or a combination of two or more of these. This disclosure contemplates mass storage device 1306 employing any suitable physical form. Where appropriate, storage device 1306 may include one or more storage device control units that facilitate communication between processor 1302 and storage device 1306. Where appropriate, storage device 1306 may include one or more storage devices. Although this disclosure describes and illustrates specific storage devices, any suitable storage device is contemplated herein.

[0182] In a particular embodiment, I / O interface 1308 includes hardware, software, or both that provide one or more interfaces for communication between computer system 1300 and one or more I / O devices. Where appropriate, computer system 1300 may include one or more of these I / O devices. One or more of these I / O devices can enable communication between a person and computer system 1300. By way of example and not limitation, I / O devices may include a keyboard, keypad, microphone, monitor, mouse, printer, scanner, speaker, still camera, stylus, tablet computer, touchscreen, trackball, video camera, another suitable I / O device, or a combination of two or more of these. I / O devices may include one or more sensors. This disclosure contemplates any suitable I / O devices and any suitable I / O interface 1308 for them. Where appropriate, I / O interface 1308 may include one or more device or software drivers that enable processor 1302 to drive one or more of these I / O devices. Where appropriate, I / O interface 1308 may include one or more I / O interfaces 1308. Although this disclosure describes and illustrates specific I / O interfaces, this disclosure contemplates any suitable I / O interface.

[0183] In a particular embodiment, communication interface 1310 includes hardware, software, or both providing one or more interfaces for communication (e.g., packet-based communication) between computer system 1300 and one or more other computer systems 1300 or one or more networks. By way of example and not limitation, communication interface 1310 may include a network interface controller (NIC) or network adapter for communicating with Ethernet or other wired-based networks, or a wireless NIC (WNIC) or wireless adapter for communicating with wireless networks (e.g., Wi-Fi networks). This disclosure contemplates any suitable network and any suitable communication interface 1310 for it. By way of example and not limitation, computer system 1300 may communicate with one or more portions of an ad hoc network, a personal area network (PAN), a local area network (LAN), a wide area network (WAN), a metropolitan area network (MAN), or the Internet, or a combination of two or more of these. One or more of these networks may be wired or wireless. As an example, computer system 1300 may communicate with a wireless PAN (WPAN) (e.g., Bluetooth WPAN), a Wi-Fi network, a Wi-Fi Max network, a cellular telephone network (e.g., a Global System for Mobile Communications (GSM) network), or other suitable wireless networks, or a combination of two or more of these. Where appropriate, computer system 1300 may include any suitable communication interface 1310 for any of these networks. Where appropriate, communication interface 1310 may include one or more communication interfaces 1310. Although specific communication interfaces are described and shown in this disclosure, any suitable communication interface is contemplated in this disclosure.

[0184] In a particular embodiment, bus 1312 includes hardware, software, or both, that couple components of computer system 1300 to each other. By way of example and not limitation, bus 1312 may include an Accelerated Graphics Port (AGP) or other graphics bus, an Extended Industry Standard Architecture (EISA) bus, a Front Side Bus (FSB), a HyperTransport (HT) interconnect, an Industry Standard Architecture (ISA) bus, an Infiniband interconnect, a Low Pin Count (LPC) bus, a memory bus, a Micro Channel Architecture (MCA) bus, a Peripheral Component Interconnect (PCI) bus, a PCI-Express (PCIe) bus, a Serial Advanced Technology Attachment (SATA) bus, a Video Electronics Standards Association Local (VLB) bus, or any other suitable bus, or a combination of two or more of these. Where appropriate, bus 1312 may include one or more buses 1312. Although this disclosure describes and illustrates specific buses, this disclosure contemplates any suitable bus or interconnect.

[0185] In this document, where appropriate, one or more computer-readable non-transitory storage media may include one or more semiconductor-based or other integrated circuits (ICs) (e.g., field-programmable gate arrays (FPGAs) or application-specific integrated circuits (ASICs)), hard disk drives (HDDs), hybrid hard disk drives (HHDs), optical disks, optical disk drives (ODDs), magneto-optical disks, magneto-optical disk drives, floppy disks, floppy disk drives (FDDs), magnetic tape, solid-state drives (SSDs), RAM drives, secure digital cards or drives, any other suitable computer-readable non-transitory storage media, or any suitable combination of two or more of these. Where appropriate, computer-readable non-transitory storage media may be volatile, non-volatile, or a combination of volatile and non-volatile.

[0186] Other miscellaneous items

[0187] In this document, unless otherwise expressly indicated or indicated by the context, "or" is inclusive rather than exclusive. Therefore, in this document, unless otherwise expressly indicated or indicated by the context, "A or B" means "A, B, or both." Furthermore, unless otherwise expressly indicated or indicated by the context, "and" is both joint and separate. Therefore, in this document, unless otherwise expressly indicated or indicated by the context, "A and B" means "A and B, jointly or separately."

[0188] This invention includes all changes, substitutions, variations, alterations, and modifications to the exemplary embodiments described or shown herein that will be understood by those skilled in the art within the scope of the appended claims. The invention is defined in the appended claims and is not limited to the exemplary embodiments described or shown herein. Furthermore, although this disclosure describes and illustrates corresponding embodiments herein as including specific components, elements, features, functions, operations, or steps, any of these embodiments may include any combination or substitution of any component, element, feature, function, operation, or step described or shown anywhere herein that will be understood by those skilled in the art. Furthermore, references in the appended claims to a means or system or component of a means or system suitable for, arranged to, capable of, configured to, implemented to, operable to, or operated to perform a particular function include that means, system, or component, whether or not it or that particular function is activated, turned on, or unlocked, provided that the means, system, or component is so adapted, arranged, enabled, configured, implemented, operable, or operated. Furthermore, although this disclosure describes or illustrates specific embodiments as providing particular advantages, specific embodiments may provide some, all, or none of these advantages.

Claims

1. A method comprising, by one or more computing systems: receiving, from a client system associated with a first user, a first audio input; generating, based on a plurality of automated speech recognition (ASR) engines, a plurality of transcriptions corresponding to the first audio input, wherein each ASR engine is associated with a respective domain of a plurality of domains; determining, for each transcription, a combination of one or more intents and one or more slots associated with the transcription; selecting, by a meta-voice engine, one or more combinations of intents and slots associated with the first audio input from the plurality of combinations; for each combination of intents and slots, identifying a domain of the plurality of domains, wherein selecting one or more combinations of intents and slots comprises mapping the domain of each combination of intents and slots to a domain associated with one of the plurality of ASR engines, wherein one or more combinations of intents and slots are selected when the domain of the respective combination of intents and slots matches the domain of one of the plurality of ASR engines; generating, based on the selected combinations, a response to the first audio input; and sending, to the client system, instructions for presenting the response to the first audio input.

2. The method of claim 1, wherein each ASR engine is associated with one or more agents of a plurality of agents specific to the respective ASR engine.

3. The method of claim 1 or claim 2, wherein each domain of the plurality of domains comprises one or more agents specific to the respective domain; wherein the agents comprise one or more of first-party agents or third-party agents.

4. The method of claim 1 or claim 2, wherein each domain of the plurality of domains comprises a set of tasks specific to the respective domain.

5. The method of claim 1 or claim 2, wherein the plurality of domains are associated with a plurality of agents, and wherein each agent is operable to perform one or more tasks specific to one or more of the domains.

6. The method of claim 1 or claim 2, wherein generating the plurality of transcriptions comprises: sending, to each of the plurality of ASR engines, the first audio input; and receiving, from the plurality of ASR engines, the plurality of transcriptions.

7. The method of claim 1 or claim 2, wherein one or more of the ASR engines of the plurality of ASR engines are third-party ASR engines associated with a third-party system separate from and external to the one or more computing systems, the method further comprising: sending, to one of the third-party ASR engines, the first audio input to generate one or more transcriptions; and receiving, from one of the third-party ASR engines, one or more transcriptions generated by the third-party ASR engine, wherein generating the plurality of transcriptions comprises selecting the one or more transcriptions generated by the third-party ASR engine to determine the combination of intents and slots associated with each respective transcription.

8. The method of claim 1 or claim 2, further comprising: ​ ​ ​ one or more features of each combination of intent and slot, wherein the one or more features indicate whether the combination of intent and slot has a property; and ranking the plurality of combinations based on their respective identified features, wherein selecting the one or more combinations of intent and slot includes selecting the one or more combinations of intent and slot based on the ranking of the plurality of combinations.

9. The method of claim 1 or claim 2, further comprising: identifying one or more identical combinations of intent and slot from the plurality of combinations; and ranking the one or more identical combinations of intent and slot based on the number of identical combinations of intent and slot, wherein selecting the one or more combinations of intent and slot includes using the ranking of the one or more identical combinations of intent and slot.

10. The method of claim 1 or claim 2, wherein generating a response to the first audio input includes: sending the selected combination to a plurality of agents; receiving a plurality of responses from the plurality of agents corresponding to the selected combination; ranking the plurality of responses received from the plurality of agents; and selecting a response from the plurality of responses based on the ranking of the plurality of responses.

11. The method of claim 1 or claim 2, wherein one of the plurality of ASR engines is a combined ASR engine based on two or more separate ASR engines, and wherein each of the two or more separate ASR engines is associated with a separate domain of the plurality of domains.

12. The method of claim 1 or claim 2, and one or more of the following: wherein the response includes one or more of the following: an action to perform, or one or more results generated from the query; wherein the instructions for presenting the response include a notification of the action to perform or a list of the one or more results.

13. One or more computer-readable non-transitory storage media embodying software that is operable when executed to: receive a first audio input from a client system associated with a first user; generate a plurality of transcriptions corresponding to the first audio input based on a plurality of automatic speech recognition (ASR) engines, wherein each ASR engine is associated with a respective domain of a plurality of domains; determine, for each transcription, a combination of one or more intents and one or more slots associated with the transcription; select, by a meta-voice engine, one or more combinations of intent and slot from the plurality of combinations that are associated with the first user input; for each combination of intent and slot, identify a domain of the plurality of domains, wherein selecting the one or more combinations of intent and slot includes mapping the domain of each combination of intent and slot to a domain associated with one of the plurality of ASR engines, wherein the one or more combinations of intent and slot are selected when the domain of the respective combination of intent and slot matches the domain of one of the plurality of ASR engines; generate a response to the first audio input based on the selected combinations; and send, to the client system, instructions for presenting the response to the first audio input.

14. A system comprising: one or more processors; and a non-transitory memory coupled to the processor, the memory comprising instructions executable by the processor, when executed, the processor operable to: receive a first audio input from a client system associated with a first user; generate, based on a plurality of automated speech recognition (ASR) engines, a plurality of transcriptions corresponding to the first audio input, wherein each ASR engine is associated with a respective domain of a plurality of domains; determine, for each transcription, a combination of one or more intents and one or more slots associated with the transcription; select, by a meta speech engine, one or more combinations of intents and slots associated with the first user input from the plurality of combinations; for each combination of intents and slots, identify a domain of the plurality of domains, wherein selecting one or more combinations of intents and slots comprises mapping the domain of each combination of intents and slots to a domain associated with one of the plurality of ASR engines, wherein one or more combinations of intents and slots are selected when the domain of the respective combination of intents and slots matches the domain of one of the plurality of ASR engines; generate, based on the selected combination, a response to the first audio input; and send, to the client system, instructions for presenting the response to the first audio input.

Citation Information

Patent Citations

  • Predicting labels using a deep-learning model

    US10387464B2

  • Automated cinematic decisions based on descriptive models

    US10511808B2

  • Search ranking and recommendations for online social networks based on reconstructed embeddings

    US10579688B2

  • Assisting users with personalized and contextual communication content

    US10782986B2

  • Resolving entities from multiple data sources for assistant systems

    US10803050B1