Restricting third-party application access to audio data content

A machine learning model in automated assistants determines appropriate access to audio data for third-party applications, addressing resource and security issues by restricting access based on interaction context, thus conserving resources and enhancing privacy.

JP7789992B2Active Publication Date: 2025-12-22GOOGLE LLC
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
JP2025514814
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2022-09-12
Filing Date
2022-12-08
Publication Date
2025-12-22
Estimated Expiration
2042-12-08

AI Technical Summary

Technical Problem

Existing automated assistants face challenges in managing third-party application access to audio data content, leading to unnecessary resource consumption and security risks due to unfettered access to captured audio data.

Method used

Implement a machine learning model to determine the appropriateness of providing audio data content to applications based on the context and intent of user interactions, restricting access when inappropriate to conserve resources and enhance security.

Benefits of technology

This approach reduces unnecessary processing and transmission of audio data, conserves computing and network resources, and enhances user privacy by preventing unauthorized access to sensitive information.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007789992000001
    Figure 0007789992000001
  • Figure 0007789992000002
    Figure 0007789992000002
  • Figure 0007789992000003
    Figure 0007789992000003
Patent Text Reader

Abstract

Embodiments relate to restricting an application's access to audio data content captured after rendering content to a user at the application's request. The application can generate the content rendered to the user using an additional request to receive audio data content from audio data captured immediately after rendering the content. The content can be processed using a trained machine learning model that generates as an output an indication that providing the audio data content after rendering the content from the application may have been inappropriate. If the application improperly requests the audio data content, the application may be restricted from providing the audio data content and / or subsequent audio data content.
Need to check novelty before this filing date? Find Prior Art

Description

[Background technology]

[0001] Humans can engage in human-to-computer interactions using interactive software applications referred to herein as “automated assistants” (also referred to as “digital agents,” “interactive personal assistants,” “intelligent personal assistants,” “assistant applications,” “conversational agents,” etc.). For example, a human (sometimes referred to as a “user” when interacting with an automated assistant) can provide commands and / or requests to the automated assistant using verbal natural language input (i.e., utterances), which can be converted to text and then processed, possibly by providing textual (e.g., typed) natural language input and / or by touch and / or non-speech physical movement(s) (e.g., hand gesture(s), gaze, facial movements, etc.). The automated assistant responds to the request by providing responsive user interface output (e.g., audible and / or visual user interface output), controlling one or more smart devices, and / or controlling one or more function(s) of the device implementing the automated assistant (e.g., controlling other application(s) on the device).

[0002] As described above, many automated assistants are configured to interact via verbal utterances. To protect user privacy and / or conserve resources, the automated assistant refrains from performing one or more automated assistant functions based on all verbal utterances present in audio data detected via a microphone(s) of a client device that (at least partially) implements the automated assistant. Rather, certain processing based on verbal utterances occurs only in response to determining that certain condition(s) exist.

[0003] For example, many client devices that include and / or interface with automated assistants include a hotword detection model. When the microphone(s) of such a client device are not inactive, the client device can continuously process audio data detected via the microphone(s) using the hotword detection model and generate a predictive output indicating the presence of one or more hotwords (including multi-word expressions), such as “Hey, Assistant,” “Okay, Assistant,” and / or “Assistant.” When the predictive output indicates the presence of a hotword, any audio data that follows within a threshold time (and optionally is determined to include voice activity) can be processed by one or more on-device and / or remote automated assistant components, such as speech recognition component(s) and / or voice activity detection component(s). Audio data predicted to include a hotword can also be processed by other on-device and / or remote automated assistant component(s). Additionally, recognized text (from the speech recognition component(s)) can be processed using natural language understanding engine(s) and / or action(s) can be performed based on the output of the natural language understanding engine(s). The action(s) may include, for example, generating and providing a response and / or controlling one or more application(s) and / or smart device(s). Other hot words (e.g., "no," "stop," "cancel," "volume up," "volume down," "next track," "previous track," etc.) may be mapped to various commands, and when the predicted output indicates the presence of one of these hot words, the mapped command may be processed by the client device. However, when the predicted output indicates the absence of the hot word, the corresponding audio data is discarded without further processing, thereby conserving resources and preserving user privacy.

[0004] An automated assistant executing at least in part on a client device can communicate with application(s) installed on the computing device and / or applications running on other devices (e.g., cloud-based applications or skills for the automated assistant). As part of their functionality, the application(s) can selectively request access to audio data content of the audio data captured by the automated assistant.

[0005] For example, an application can generate prompts that are rendered to a user as part of an interaction between the user and the application. For example, the interaction can be mediated by an automated assistant, and prompts can be provided to the automated assistant, which can render the prompts as part of the interaction. The application can further request access to audio data content of audio data captured following a prompt during an interaction with a user. For example, if a prompt from an application is, "What kind of music would you like to listen to?", the application can request access to the audio data content that follows the prompt so that the application can ascertain the user's response (e.g., a response specifying a genre of music). The audio data content provided to the application can include the audio data itself, an automatic speech recognition (ASR) transcription of the audio data (e.g., generated by an ASR engine of the automated assistant), and / or a structured representation generated based on the transcription (e.g., NLU data generated by a natural language understanding (NLU) engine of the automated assistant).

[0006] However, providing unfettered access to audio data content when requested by an application can result in various drawbacks. For example, communicating audio data content to an application when the audio data content is not truly required for interaction can utilize the bandwidth of constrained communication resources. For example, transmitting the audio data content can utilize network resources when the application is running on a remote server(s). As another example, communicating audio data content to an application when the audio data content is not truly required for interaction can result in security concerns for content inadvertently captured within the audio data content. Summary of the Invention

[0007] Some embodiments disclosed herein relate to determining whether to restrict an application's access to audio data content of audio data captured by a microphone during an interaction with a user based on whether the application inappropriately requested access to the audio data content. Some embodiments include distinguishing between content generated by the application that is rendered to the user during an interaction turn and a request from the application to be provided with audio data content corresponding to audio data captured after at least the interaction turn. The content of the interaction turn is processed using one or more machine learning models trained to generate an output indicative of the likelihood that the application properly received the audio data content (e.g., the user can expect, following an output instance, that the microphone was active, the client device captured audio data, and that content be provided to the application). If the output does not exceed a threshold of appropriateness, the instance of the application may be restricted from accessing the requested audio data content and / or content corresponding to subsequently captured audio data.

[0008] As an example, an application may generate "What is the capital of France?" content to be rendered during a dialog turn in a dialogue with a user. The application may request and / or otherwise cause the content of the captured audio data to be provided after the dialog turn (e.g., activate a microphone, refrain from deactivating the microphone if it is already active, request access to the captured audio data, request the content of the captured audio data). For a dialog turn that includes the content "What is the capital of France?", the user likely expects that the audio immediately following it will be captured by the device and that the content of the audio data (e.g., the audio data, a textual representation of the audio data) will be provided to the application because the content of the dialog turn is a question directed to the user. Thus, in this example, providing the audio data content to the application is likely appropriate given the purpose of the dialog turn (i.e., to elicit a response).

[0009] As another example, an application may generate and provide "Thanks for playing. Bye" content during a dialogue turn of a dialogue, and further provide the application with audio data content captured at least after rendering the dialogue turn. In this case, the dialogue turn includes a sign-off message, and the user is unlikely to expect audio data content to be provided to the application. Thus, in this example, providing audio data content to the application after rendering the dialogue turn is unlikely to be appropriate.

[0010] In some implementations, a machine learning model can be utilized to determine whether it is / was appropriate to provide the application with audio data content captured after content rendered to a user during a dialogue turn based on the content generated by the application. For example, the machine learning model can be trained using training instances, each including content from a dialogue turn rendered to a user and subsequently provided to the application with captured audio data content. Further, each training instance can include an indication of whether providing the audio data content was appropriate given the intent of the content of the corresponding dialogue turn. The machine learning model can provide an output indicating the likelihood that providing the audio data content was appropriate. For example, the trained machine learning model can provide as an output a numerical value between 0.0 (e.g., full confidence that providing the audio data content after the dialogue turn was rendered was appropriate) and 1.0 (e.g., full confidence that providing the audio data content was inappropriate).

[0011] In some implementations, a text representation of the content of a dialogue turn that was initially provided to the user as audio may be utilized as input to the machine learning model. For example, during a dialogue turn, content including "What is the capital of France?" may be rendered by an automated assistant generated by an application and communicating with the application (e.g., an automated assistant that facilitates communication between the application and the user). The text representation may be, for example, a transcription of what was provided to the user as audio during the dialogue.

[0012] In some implementations, the vector representation of the content of a dialogue turn can be used as input to a machine learning model. For example, in some implementations, processing can be performed that utilizes one or more other machine learning models, resulting in vectors in an embedding space that represent the content of the dialogue. The vectors can then be provided as input to the machine learning model for further processing.

[0013] In some implementations, in addition to the content of the dialogue turn, additional information may be provided as input to the machine learning model. In some implementations, the additional information may include an application type indicating the type of application that provided the dialogue turn being processed. In some cases, a user may expect some applications to receive audio data content, while other applications are less likely to access the audio data content. The application type may indicate, for example, that the application is a quiz game, a mapping application, a ride-sharing application, a restaurant reservation application, and / or one or more other applications that may provide interaction to the user and may cause activation of the client device's microphone.

[0014] In some implementations, the additional information may include an indication of the context of the corresponding dialogue turn. For example, the training instance may include an indication of whether the corresponding dialogue turn was the first turn of the dialogue (i.e., the first response from the application following an invocation of the application), whether it was the final turn of the dialogue (i.e., the application did not provide a follow-up dialogue turn), or whether the dialogue turn was in the middle of the dialogue. Also, for example, the context may include information related to other dialogue turns of the dialogue (e.g., content immediately before, after, and / or content from one or more other dialogue turns of the dialogue).

[0015] In some implementations, the additional information may include the source of the application that generated the content rendered during the interaction turn. For example, one or more sources of the application (e.g., application developer, application seller) may be provided as input to the machine learning model along with the content. In some cases, the source of the application may be an indication of whether the application is trusted (e.g., known developer vs. unknown or suspicious developer).

[0016] In some implementations, the additional information may include one or more user settings related to confidentiality of sharing of audio data content. For example, a user may set one or more settings as part of their user account that indicate the degree of confidentiality of applications the user expects when receiving audio data content.

[0017] In addition to providing content as input to the machine learning model, the training instances can include any or all of the additional information described above as additional input to the machine learning model. For example, the training instances can include the type of application that generated the corresponding content, the context of the content, and / or other additional information related to the content of one or more rendered dialogue turns. Furthermore, in some implementations, the training instances can be generated based on other sources of dialogue. For example, content from one or more discussion websites, such as message boards, can be used to generate training instances that each include user-generated comments and an indication of whether the comments appropriately elicit a response. The indication of appropriateness / inappropriateness of providing a response to the application that generated the content of the training instances can be determined, for example, by one or more human curators reviewing the content and determining whether the user would expect a response to be captured and provided to the application based on an assumption that the content originated from the application.

[0018] Once trained, the machine learning model can be utilized to process content provided during one or more dialogue turns of a target application, such as a third-party application, and can determine whether the target application needs to access audio data content from audio data captured by a client device(s) running an instance of the target application. For example, content from multiple dialogue turns provided in response to a request from the target application to provide content, in which the application subsequently receives the audio data content, can be identified and processed using the machine learning model. The output of the machine learning model can be utilized to determine the likelihood that the audio data was appropriately provided to the target application.

[0019] As an example, 1000 requests generated by a target application to provide content and receive subsequent audio data content may be identified from a log of requests to render content by multiple instances of the target application. For each rendered dialogue turn of content, the output of the machine learning model may range from 0.0 to 1.0, as described above. For example, if the output from processing the content of any of the 1000 requests indicates that the corresponding request for subsequent audio data content may have been inappropriate (e.g., output exceeding a threshold), the target application may be flagged for manual review and / or the instance of the target application may be restricted from accessing subsequent audio data content (e.g., the audio data content may not be provided, or the audio data content may be provided but the user may be more prominently alerted about the transmission), either permanently or until further manual review is performed. Also, for example, if the output generated from a threshold number of requests (e.g., content from more than 10 out of 1000 processed requests) exceeds a threshold, the target application may be flagged for further manual review and / or the instance of the application may be restricted from accessing subsequent audio data content, including uninstalling and / or disabling the application.

[0020] In some implementations, the requests may come from multiple client devices, each client device running an instance of the target application. For example, logs from multiple devices running instances of the application may be identified, and requests to render content that further include requests for subsequent audio data content to be provided may be used to determine whether any of the instances of the application are inappropriately accessing subsequent audio data. In some implementations, content from one or more dialogue turns may be processed as it is rendered (or before it is rendered), so that inappropriate requests for audio data content may be identified as (or before it occurs) when the audio data content is provided to the application. For example, as each dialogue turn of an application is rendered and subsequently provided with audio data content, a trained machine learning model may be used to process the dialogue turn content to determine whether the provision of the audio data was appropriate. Also, for example, for each request to render content as a dialogue turn with a request for subsequent audio data content, the content may be processed before being rendered to determine whether the request for the subsequent audio data content is appropriate. In either case, if the output from the machine learning model indicates that it is / was inappropriate to provide the audio data content to the application, the application may be restricted from further access to the audio data and / or restricted from providing further content to the user.

[0021] Thus, the methods described herein can mitigate the need to unnecessarily process audio data content by preventing it from being provided for processing in instances where providing it is inappropriate. Furthermore, not transmitting audio data and / or audio data content when inappropriate and / or when the transmitted audio data content is not subsequently utilized reduces computing resource and / or network usage. Still further, security can be improved by mitigating transmission of audio data content when a user does not expect the content to be provided to one or more applications. By determining, prior to providing the audio data content, that providing it is inappropriate given the content of a dialogue turn rendered to a user, personal information that may be included in the audio data content can remain secure and inaccessible to third parties when unnecessary.

[0022] The above description is provided as a summary of only some of the embodiments disclosed herein. These and other embodiments are described in further detail herein.

[0023] It will be understood that all combinations of the above concepts, and additional concepts described in more detail herein, are contemplated as being part of the subject matter disclosed herein, for example, all combinations of claimed subject matter listed at the end of this disclosure are contemplated as being part of the subject matter disclosed herein. [Brief explanation of the drawings]

[0024] [Figure 1] FIG. 1 illustrates an exemplary environment in which embodiments disclosed herein may be practiced. [Figure 2] Illustrates the interaction between the application and the user. [Figure 3]1 is a flowchart illustrating training of a machine learning model and utilization of the trained machine learning model to process one or more dialogue turns provided by an application. [Figure 4] 1 is a flowchart illustrating an exemplary method according to various embodiments disclosed herein. [Figure 5] 10 is a flowchart illustrating another exemplary method according to various embodiments disclosed herein. [Figure 6] 10 is a flowchart illustrating another exemplary method according to various embodiments disclosed herein. [Figure 7] 1 illustrates an exemplary architecture of a computing device. DETAILED DESCRIPTION OF THE INVENTION

[0025] Referring initially to Figure 1, an exemplary environment in which various embodiments may be implemented is shown. Figure 1 includes a client device 100 running an automated assistant that runs an instance of an automated assistant client 120. One or more cloud-based automated assistant components may be implemented on one or more computing systems (collectively referred to as "cloud" computing systems) communicatively connected to client device 100 via one or more local and / or wide area networks (e.g., the Internet). An instance of automated assistant client 120, optionally through interaction(s) with one or more of the cloud-based automated assistant components, can form what, from a user's perspective, appears to be an instance of a logical automated assistant with which the user can engage in human-computer interaction.

[0026] Client device 100 may be, for example, a desktop computing device, a laptop computing device, a tablet computing device, a mobile phone computing device, a vehicle computing device (e.g., an in-vehicle communication system, an in-vehicle entertainment system, an in-vehicle navigation system), a standalone interactive speaker, a smart appliance such as a smart television, and / or a wearable device that includes a computing device (e.g., a watch with a computing device, glasses with a computing device, a virtual reality or augmented reality computing device).

[0027] Automated assistant 120 engages in a human-computer interactive session with a user via user interface input and output devices of client device 100. To protect user privacy and / or conserve resources, in many situations, a user must explicitly invoke automated assistant 120 before the automated assistant can fully process the verbal utterance. Explicit invocation of automated assistant 120 can occur in response to specific user interface input received at assistant device 100. For example, user interface inputs that can invoke automated assistant 120 via client device 100 can optionally include the operation of hardware and / or virtual buttons on client device 100. Additionally, automated assistant client can include one or more local engines, such as invocation engine 130 operable to detect the presence of one or more spoken common invocation wake words. Invocation engine 130 can invoke automated assistant 120 in response to detecting one of the spoken invocation wake words. For example, invocation engine 130 can invoke automated assistant 100 in response to detecting a verbal invocation wake word such as "Hey, Assistant," "Okay, Assistant," and / or "Assistant." Invocation engine 130 can continuously process a stream of audio data frames based on output from one or more microphones 140 of client device 100 to monitor for the occurrence of a verbal invocation phrase. While monitoring for the occurrence of a verbal invocation phrase, invocation engine 130 discards (e.g., after temporarily storing in a buffer) all audio data frames that do not contain the verbal invocation phrase. However, if invocation engine 130 detects the occurrence of the verbal invocation phrase in a processed audio data frame, the invocation engine can invoke automated assistant 120.As used herein, "invoking" automated assistant 120 may include activating one or more previously inactive features of automated assistant 120. For example, invoking automated assistant 120 may include operating one or more local engines and / or cloud-based automated assistant components to further process the audio data frame from which the invocation phrase was detected and / or one or more subsequent audio data frames (where prior to the invocation, no further processing of the audio data frames had occurred). For example, the local and / or cloud-based components may process the captured audio data using an ASR model in response to the invocation of automated assistant 120.

[0028] 1 is illustrated as including an automatic speech recognition (ASR) engine 122, a natural language understanding (NLU) engine 124, a text-to-speech (TTS) engine 126, and a fulfillment engine 128. In some implementations, one or more of the illustrated engines may be omitted (e.g., instead implemented solely by the cloud-based automated assistant component(s) 140) and / or additional engines may be provided (e.g., the invocation engine described above).

[0029] The ASR engine 122 can process audio data capturing verbal utterances to generate recognitions of the verbal utterances. For example, the ASR engine 122 can process the audio data utilizing one or more ASR machine learning models to generate predictions of recognized text corresponding to the utterances. In some of these implementations, the ASR engine 122 can generate, for each of one or more recognized terms, a corresponding confidence measure indicating the confidence that the predicted term corresponds to the verbal utterance.

[0030] The TTS engine 126 can convert text into synthetic speech and, in doing so, can rely on one or more speech synthesis neural network models. The TTS engine 126 can be utilized, for example, to convert text responses into audio data that includes a synthesized version of the text, which is audibly rendered via the hardware speaker(s) 150 of the client device 100.

[0031] The NLU engine 124 determines the semantic meaning(s) of the audio and / or text converted from the audio by the ASR engine and determines assistant action(s) corresponding to the semantic meaning(s). In some implementations, the NLU engine 124 determines the assistant action(s) as intent(s) and / or parameter(s) determined based on the recognition(s) of the ASR engine 122. In some situations, the NLU engine 124 can resolve the intent(s) and / or parameter(s) based on a single utterance of the user; in other situations, prompts can be generated based on the unresolved intent(s) and / or parameter(s), those prompts are rendered to the user, and the user response(s) to those prompt(s) are utilized by the NLU engine 124 in resolving the intent(s) and / or parameter(s). In these situations, the NLU engine 124 can optionally work in coordination with the dialogue manager engine 170 to determine the outstanding intent(s) and / or parameter(s) and / or generate the corresponding prompt(s). The NLU engine 124 can utilize one or more NLU machine learning models in determining the intent(s) and / or parameter(s).

[0032] The fulfillment engine 128 can cause the execution of assistant action(s) determined by the NLU engine 124. For example, if the NLU engine 124 determines the assistant action to "turn on the kitchen light," the fulfillment engine 128 can cause corresponding data to be sent (either directly to the light or to a remote server associated with the light's manufacturer) to "turn on" the "kitchen light." As another example, if the NLU engine 124 determines the assistant action to "provide a summary of the user's meetings for today," the fulfillment engine 128 can access the user's calendar, summarize the user's meetings for the day, and cause the summary to be visually and / or audibly rendered at the client device 100.

[0033] Automated assistant 120 further includes application interface 160 that can communicate with one or more third-party applications running on client device 100. Through application interface 160, the application (not shown) can provide information to automated assistant 120 to cause automated assistant 120 to perform one or more actions and / or the application can be authorized to directly access one or more components of client device 100. For example, in some implementations, the application can provide text to automated assistant 120, which can then utilize one or more components to process the text (e.g., perform ASR, NLU, TTS). Also, for example, the application can be authorized to directly receive audio data content corresponding to audio data captured by microphone 140 and / or provide output via speaker 150.

[0034] In some implementations, automated assistant 120 can continue to interact with the user via client device 110 through microphone 140 and / or speaker 150. For example, referring to FIG. 2, an interaction between a user and automated assistant 120 is shown. Before the interaction begins, invocation engine 130 can monitor audio data captured by microphone 140 to detect the presence of one or more spoken general invocation wake words in the audio data. For example, in interaction turn 205, the user utters the phrase "Okay, assistant." This phrase can be a wake word indicating the user's intent to initiate an interaction with automated assistant 120.

[0035] Dialogue turn 205 further includes the user uttering, "Open the quiz application." Dialogue manager 170 can determine the intent based on one or more techniques described herein and perform one or more actions based on the intent. For example, dialogue manager 170 can determine that the user's intent is to interact with an application called "quiz application" running on client device 100 and / or running on one or more other devices. In response, fulfillment engine 128 can send some or all of the audio data content from the audio data captured by microphone 140 to the application. In response, the application can provide a response (e.g., text) to automated assistant 120 via application interface 160 to be provided to the user as audio output via speaker 150. For example, the application can provide the text "Welcome to the quiz application. What would you like to play?" to automated assistant 120, which can utilize TTS engine 126 to process the text and then render dialogue turn 210 via speaker 150.

[0036] In some implementations, the application can further send an indication to automated assistant 120 via application interface 160 that a response is expected from the provided content and that subsequent audio is to be captured. In response, automated assistant 120 can activate microphone 140, thereby capturing subsequent audio, and generate audio data that can be further processed. For example, after providing dialogue turn 210, automated assistant 120 can activate microphone 140 and / or refrain from deactivating microphone 140. The audio data captured after dialogue turn 210 can be further processed by automated assistant 120 (e.g., ASR, NLU), and one or more actions can be performed. For example, the content of the captured audio data, e.g., a text representation of the audio data (after ASR processing) and / or the user's intent (after NLU), can be provided to the application as audio data content.

[0037] For example, in dialogue turn 215, the user utters, "Let's try country capitals." In some implementations, audio data including the utterance may be provided to an application for further processing. In some implementations, the automated assistant 120 may first perform ASR to generate text for "Let's try country capitals," and the text may be provided to the application as audio data content for further processing. In some implementations, the automated assistant 120 may first determine the intent by performing NLU and provide the intent (e.g., "quiz = country capitals") to the application as audio data content.

[0038] The dialogue may continue until the user indicates that they want to stop the dialogue, as shown in FIG. 2 . For example, at dialogue turn 235, the user indicates “play is over” as a sign-off to the dialogue. In response, dialogue manager 170 can determine that the user intends to stop the dialogue, and / or the dialogue turn can be provided to the application as audio data content for further processing. In some implementations, the application can determine that the dialogue is complete and that no subsequent audio data content should be provided to the application. This can include, for example, providing an indication to automated assistant 120 to stop sending audio data content.

[0039] Thus, in some dialog turns provided to a user while interacting with application 160, the user may expect microphone 140 to be active and capturing audio data. For example, if a dialog turn solicits a response (e.g., a question), the user may expect any subsequent audio data to be captured and provided to the application to determine whether the audio data includes a response to the question. However, in some dialog turns provided to a user, the user will not expect any subsequent audio data content to be provided to the application. Because the application may be a third-party application, the user may expect that third parties will not have access to audio data content that may include, for example, personal information unrelated to the interaction with the application.

[0040] 1 , the exemplary environment further includes an analysis device 110 that can make decisions based on analyzing a dialogue turn that includes content generated by an application, where audio data content is provided to the application following its generation by the application. For example, if an application generates content that is provided during a dialogue turn, where subsequent audio data content is provided to the application following the dialogue turn, the content may be stored in one or more databases, such as dialogue turn database 112, for subsequent processing.

[0041] In some implementations, client device 100 can provide content from a dialogue turn either as it occurs or periodically in batches, following which the application that generated the content is provided with the subsequent audio data content. For example, the content of dialogue turn 210 in FIG. 2 can be provided to analysis device 110 in real time (e.g., starting as soon as it is rendered), or it can be stored on client device 100 for a period of time and then sent in batches, for example, with other dialogue turns of the dialogue shown in FIG. 2 (and / or with turns from other dialogues). In some implementations, content from dialogue turns provided by other client devices in addition to client device 100 can be stored in dialogue turn database 112 along with the content of dialogue turns rendered by client device 100.

[0042] The dialogue turn processor 114 can receive dialogue turn content and process the content to determine whether it was appropriate to provide audio data content following rendering of the dialogue turn. For example, with reference to FIG. 3 , a flowchart illustrating processing content of one or more dialogue turns is provided. As shown, the dialogue turn processor 114 can receive dialogue turn content from one or more client devices 100 and / or identify dialogue turn content from the dialogue turn database 112. In some implementations, the dialogue turn content can be stored with and / or provided with additional information. For example, in some implementations, the dialogue turn content can be associated with an application type indicating the category of application that provided the dialogue turn content (e.g., a quiz application, a ride-sharing application, a restaurant reservation application). The application type can be utilized in further processing of the dialogue turn, as described in more detail herein. Also, for example, in some implementations, the dialogue turn content can be provided with contextual information regarding the dialogue of which the dialogue turn was a part (e.g., whether the dialogue turn was rendered at the beginning, end, or middle of the dialogue, information regarding the content of other dialogue turns in the corresponding dialogue).

[0043] In some implementations, the dialogue turn content may be provided to the dialogue turn processor 114 and / or stored in the dialogue turn database 112 as a text representation of the synthetic speech provided to the user during the dialogue. For example, in some implementations, an application may generate text to provide to the user as dialogue turn content during the dialogue, and the automated assistant 120 may utilize the TTS engine 126 to generate the synthetic speech to provide to the user. In some implementations, the dialogue turn content may include the intent of the synthetic speech provided to the user as part of the dialogue. For example, natural language understanding may be performed on the text representation of the dialogue turn content using one or more machine learning models that generate the intent of the content as output. In some implementations, the dialogue turn content may include a vector representation of the content, such as using Word2vec and / or one or more other models to generate vectors to be embedded in a semantic embedding space.

[0044] In some implementations, a machine learning model (MLM) 121 can be utilized to determine whether it was appropriate to provide the captured audio data content after rendering a given dialogue turn. For example, the machine learning model 121 can receive as input the content of the dialogue turn and, optionally, additional information, such as the type of application that generated the content and / or the context for the dialogue, as described above. The machine learning model 121 can generate as output an indication of whether it was appropriate for the application that generated the content to be provided with the subsequent audio data content. For example, the machine learning model 121 can generate a number between 0.0 and 1.0 that indicates whether it was inappropriate to provide the audio data content, where 0.0 indicates complete certainty that it was appropriate to provide the audio data content and 1.0 indicates complete certainty that it was inappropriate to provide the audio data content (or vice versa, depending on how the model was trained).

[0045] In some implementations, the machine learning model 121 can receive as input dialogue turn content, including a textual representation of what was presented to the user in the dialogue. For example, referring again to FIG. 3 , the dialogue turn processor 114 can receive text of the dialogue turn content from the client device 100 and / or identify the text in the dialogue turn database 112 and process the text by providing it as input to the machine learning model 121. The resulting generated output 320 can be provided to the output analysis module 118 for analysis to determine whether the application that generated the content should restrict the application from providing audio data content generated from audio captured in subsequent interactions between the user and the application. In some implementations, the machine learning model 121 can provide as the dialogue turn content the intent of the dialogue turn presented to the user by the automated assistant 120. In some implementations, as described above, the client device 100 can provide as the dialogue turn content a vector embedded in a semantic embedding space representing what was rendered to the user. The machine learning model 121 can generate as output the generated output 320, which can be further analyzed by the output analysis module 118. In some implementations, the dialogue turn processor 114 may receive and / or identify dialogue turn content that is in text form, first perform natural language processing on the text, and provide the output of the natural language processing as input to the machine learning model 121. Thus, in implementations in which the machine learning model 121 receives vectors as input, the text may be processed using natural language processing to generate the vectors by one or more other components, and / or the machine learning model 121 may be utilized to process the text by the dialogue turn processor 114 prior to further processing.

[0046] The machine learning model 121 may be trained with one or more training instances that, when processed, refine the generated output of the machine learning model 121. As an example, referring to FIG. 3 , the training instances 300 each include dialogue turn content 305, which may include content generated by an application and provided to a user via a dialogue turn. Each training instance 300 further includes a supervised indication 310 determined by curation of the dialogue turn content 305. For example, a human may review the content and, based on the content, determine whether it is appropriate to provide subsequent audio data content. The curator may assign a value (e.g., yes if appropriate, 0 if inappropriate, or vice versa) to indicate whether it would be appropriate to provide the subsequent audio data content if the dialogue turn content 305 were provided within the dialogue. In some implementations, the training instance 300 may be generated from content previously provided as part of the dialogue (e.g., content from the dialogue turn database 112 or other sources of generated content). In some implementations, one or more other content sources may be utilized to determine the dialogue turn content 305 of the training instance 300. For example, a messaging application and / or discussion board website can be used to identify content that may be provided to users in an interaction, and one or more curators can determine, based on the content, a supervised indication 310 to assign to the training instance 300. Optionally, the training instance can include an application type 315 that indicates the type of application that generated the interaction turn content. Optionally, the training instance 300 can also include an interaction context 320 that indicates additional information about the interaction that includes the interaction turn content 305.For example, the dialogue context 320 may include content from one or more other dialogue turns of the dialogue, the position of the dialogue turn in the dialogue, and / or other information of the dialogue that is not in the content rendered in the dialogue turn.

[0047] The output analysis module 118 can determine whether to restrict an application's access to audio data content generated from the captured audio data based on the generated output 320. As described above, the generated output 320 can be, for example, a numeric value between 0.0 (full confidence that the previous provision of audio data to the application 160 was appropriate) and 1.0 (full confidence that the provision of audio data was inappropriate). The output analysis module 118 can determine, based on the processing output generated by the machine learning model 121, whether the application inappropriately requested access to audio data content generated from the captured audio data at least after the rendering of the content when requesting rendering of the content as a dialogue turn. If the output analysis module 118 determines that the application requesting the audio data content was inappropriate, the application can immediately restrict access to subsequent audio data, the application can be flagged for manual review, and / or one or more other actions can be taken to otherwise prevent the application from continuing to receive audio data inappropriately.

[0048] In some implementations, the generated output 320 can be utilized to determine whether to increase scrutiny of subsequent requests by the application to access audio data content of audio data captured in response to rendering the content. For example, if the generated output exceeds a threshold of inappropriateness, subsequent interaction with the application can be accompanied by an indication rendered to the user that the application is receiving audio data content. The indication can be, for example, a visual indication that the microphone is active and that audio data is being processed, a message provided to the user while the user is interacting with the application, an audio cue alerting the user that the application may be processing audio data, and / or one or more other indications that are more prominent than indications that may be present in other applications that the application may be processing audio data content.

[0049] In some implementations, if the generated output 320 meets a threshold of inappropriateness, the application may be restricted from accessing the audio data content and / or from accessing subsequent audio data content. For example, the automated assistant 120 may first provide the rendered content (and optionally additional information as described herein) to the dialogue turn processor 114 for processing before providing the audio data content to the application. If the generated output 320 meets the threshold of appropriateness (or does not meet the threshold of inappropriateness), the audio data content may be provided to the application. However, if the generated output 320 indicates that it would be inappropriate to provide the audio data content after rendering the content, the audio data content may not be sent to the application.

[0050] In some implementations, content from multiple dialogue turns, following which the application that generated the content is provided with audio data content, may be identified, and the content from each of the identified dialogue turns may be processed using the dialogue turn processor 114. For example, the dialogue turn database 112 may contain content from multiple dialogue turns (either from a single dialogue or multiple dialogues generated from a single instance of an application or multiple instances of an application). In some implementations, if the output 320 generated from one or more of the processed content exceeds a threshold, the application may be restricted from subsequent audio data content. For example, if the generated output 320 exceeds a 0.7 threshold for inappropriateness, the request for audio data content may be deemed inappropriate, and the application may be restricted. Also, for example, an application may be restricted from subsequent audio data content in multiple ways, checking the generated output 320 to determine whether it meets multiple thresholds, each with a different restriction method. For example, an application that generates inappropriate content at a generated output 320 of 0.7 may be blocked entirely from providing subsequent audio data content, while an application that generates inappropriate content at a generated output 320 of 0.4 may be restricted by having a higher alert level to the user associated with subsequent instances of providing audio data content. Other examples of restricting subsequent audio data content from being provided to an application may include uninstalling the application and / or instances of the application, alerting the user to instances in which audio data content was provided inappropriately, and / or flagging the application for further manual review.

[0051] 4 shows a flowchart illustrating an example method for processing dialogue turn content using a trained machine learning model. For convenience, the operations of the method are described with reference to a system that performs the operations, such as the system shown in FIG. 1. The system of the method includes one or more processors and / or other component(s) of a client device. Furthermore, although the operations of the method are shown in a particular order, this is not meant to be limiting. One or more operations may be reordered, omitted, or added.

[0052] In step 405, a request for the content to be rendered and subsequent audio data content is received from an application. The content and request may be received by a component that shares one or more characteristics with automated assistant 120 of FIG. 1. For example, automated assistant 120 may receive a request via application interface 160 that the content received with the request be rendered to an interacting user. Additionally, the application may request via application interface 160 that the audio data content be provided from the captured audio data at least after the content is rendered.

[0053] In step 410, the content is processed using the trained machine learning model to generate an output. The machine learning model may share one or more characteristics with machine learning model 121 of FIG. 1. In some implementations, a component that shares one or more characteristics with dialogue turn processor 114 may provide the content received in step 405 as input to machine learning model 121. In some implementations, the content may be provided as input to machine learning model 121 with additional information. For example, the content may be provided to the machine learning model along with a type indicating the category of the application that provided the content, contextual information related to the user interaction (e.g., where in the interaction the dialogue turn occurred), and / or other additional information. The content may be provided, for example, as text data, vectors generated via natural language processing, and / or the intent of the content. In some implementations, automated assistant 120 may process the content before rendering it in a dialogue turn. In some implementations, automated assistant 120 may provide the content first and then process the content thereafter, but before providing subsequent audio data content to the application.

[0054] At step 415, providing the output is determined to be inappropriate based on the output generated from the machine learning model. In some implementations, the machine learning model may provide as the generated output a value indicating the likelihood that providing the subsequent audio data content was appropriate given the content generated by the application. For example, the generated output may indicate that providing the audio data content after the given content is rendered is 0.7, indicating a 70% likelihood that providing is inappropriate. If the value exceeds a threshold, at step 420 the application is restricted from accessing the audio data content.

[0055] 5 shows a flowchart illustrating an example method for processing dialogue turn content using a trained machine learning model. For convenience, the operations of the method are described with reference to a system that performs the operations, such as the system shown in FIG. 1. The system of the method includes one or more processors and / or other component(s) of a client device. Furthermore, although the operations of the method are shown in a particular order, this is not meant to be limiting. One or more operations may be reordered, omitted, or added.

[0056] At step 505, a diagram is identified that includes the content to be rendered to the user, followed by the application that generated the content to provide the audio data content. The dialogue turn content may be identified by a component that shares one or more characteristics with the dialogue turn processor 114. In some implementations, the dialogue turn content may be identified in a database that stores content generated by one or more applications, such as the dialogue turn database 112. In some implementations, the client device 100 may provide the content from the rendered dialogue turn. The dialogue turn content may be provided with additional information as described herein, including with reference to step 405 of FIG. 4.

[0057] At step 510, the content is processed using the machine learning model and an output is generated that indicates the likelihood that providing the audio data content was appropriate. The processing of the dialogue turn content may share one or more characteristics with step 410. For example, content from a dialogue turn may be processed and the generated output may indicate the likelihood that providing the audio data content following the application was inappropriate.

[0058] At step 515, it is determined that providing audio data content to the application is inappropriate. Step 515 may share one or more characteristics with step 415 of FIG. 4. For example, the output generated from processing a given dialogue turn can be compared to a threshold, and if the generated output exceeds the threshold, it can be determined that providing audio data content is inappropriate. Thus, for 10 processed dialogue turns, the 10 resulting outputs can be utilized to determine, for example, the number of times that, when rendered, the audio data content was inappropriately provided. For example, if a threshold number of processed dialogue turns are determined to be inappropriate, then at step 520 the application is restricted from accessing subsequent audio data content.

[0059] 6 shows a flowchart illustrating an example method for processing dialogue turn content using a trained machine learning model. For convenience, the operations of the method are described with reference to a system that performs the operations, such as the system shown in FIG. 1. This system of the method includes one or more processors and / or other component(s) of a client device. Furthermore, although the operations of the method are shown in a particular order, this is not meant to be limiting. One or more operations may be reordered, omitted, or added.

[0060] At step 605, content of multiple dialog turns is identified, followed by the application that generated the content to be rendered to the user and for which audio data content is provided. The dialog turn content may be identified by a component that shares one or more characteristics with the dialog turn processor 114. In some implementations, the dialog turn content may be identified in a database that stores content generated by one or more applications, such as the dialog turn database 112. In some implementations, the client device 100 may provide content from multiple rendered dialog turns. The dialog turn content may be provided with additional information as described herein, including with reference to steps 405 of FIG. 4 and 505 of FIG. 5.

[0061] At step 610, each of the content is processed using the machine learning model to generate an output for each dialogue turn content. The processing of the dialogue turn content may share one or more characteristics with step 410 and / or step 510. For example, content from a dialogue turn may be processed, and the generated output may indicate that providing audio data content following the application may have been inappropriate. Each dialogue turn may be processed in this manner, whereby for each dialogue turn content processed, a corresponding output is generated.

[0062] At step 615, the generated output is utilized to determine whether a threshold value of the output indicates that it would be inappropriate to provide audio data output to the application. Step 615 may share one or more characteristics with step 415 of FIG. 4 and / or step 515 of FIG. 5. For example, the generated output from processing a given dialogue turn can be compared to a threshold value, and if the generated output exceeds the threshold value, it can be determined that providing audio data content is inappropriate. Thus, for 10 processed dialogue turns, the 10 resulting generated outputs can be utilized to determine, for example, the number of audio data content that, when rendered, would be inappropriately provided. For example, if a threshold number of processed dialogue turns are determined to be inappropriate, then at step 520 the application is restricted from accessing subsequent audio data content.

[0063] 7 is a block diagram of an exemplary computing device 710 that may optionally be utilized to perform one or more aspects of the techniques described herein. The computing device 710 typically includes at least one processor 714 that communicates with multiple peripherals via a bus subsystem 712. These peripherals may include, for example, a storage subsystem 724 including a memory subsystem 725 and a file storage subsystem 726, a user interface output device 720, a user interface input device 722, and a network interface subsystem 716. The input and output devices enable user interaction with the computing device 710. The network interface subsystem 716 provides an interface to external networks and is connected to corresponding interface devices in other computing devices.

[0064] The user interface input devices 722 may include pointing devices such as keyboards, mice, trackballs, touchpads, graphics tablets, scanners, touchscreens integrated into displays, audio input devices such as voice recognition systems, microphones, and / or other types of input devices. In general, use of the term "input device" is intended to include all possible types of devices and methods for inputting information into the computing device 710 or communications network.

[0065] The user interface output devices 720 may include a display subsystem, a printer, a fax machine, or a non-visual display such as an audio output device. The display subsystem may include a flat panel device such as a cathode ray tube (CRT), a liquid crystal display (LCD), a projection device, or other mechanism for producing a visible image. The display subsystem may also provide a non-visual display, such as via an audio output device. In general, use of the term "output device" is intended to include any type of device or method for outputting information from the computing device 710 to a user or other machine or computing device.

[0066] Storage subsystem 724 stores programming and data structures that provide the functionality of some or all of the modules described herein. For example, storage subsystem 724 may include logic to perform selected aspects of the methods of Figures 4-6 and / or to implement various components shown in Figures 1 and 3.

[0067] These software modules typically execute on the processor 714 alone or in combination with other processors. The memory 725 used by the storage subsystem 724 may include multiple memories, such as a main random access memory (RAM) 730 for storing instructions and data during program execution and a read-only memory (ROM) 732 in which fixed instructions are stored. The file storage subsystem 726 may provide persistent storage of program files and data files and may include a hard disk drive, a floppy disk drive with associated removable media, a CD-ROM drive, an optical drive, or a removable media cartridge. Modules that implement the functionality of an embodiment may be stored by the file storage subsystem 726 in the storage subsystem 724 or on another machine accessible by the processor(s) 714.

[0068] Bus subsystem 712 provides a mechanism that allows the various components and subsystems of computing device 710 to communicate with each other in an intended manner. Although bus subsystem 712 is shown illustratively as a single bus, other implementations of a bus subsystem may use multiple buses.

[0069] Computing device 710 can be a variety of types of devices, such as a workstation, a server, a computing cluster, a blade server, a server farm, or any other data processing system or computing device. Due to the ever-changing nature of computers and networks, the description of computing device 710 shown in Figure 7 is intended only as a specific example to illustrate some implementations. Many other configurations of computing device 710 may have more or fewer components than the computing device shown in Figure 7.

[0070]

[0006] Embodiments disclosed herein include a method that includes receiving, by an automated assistant and from the application during an interaction session occurring via a client device between a user and an application and mediated by the automated assistant, content to be rendered by the automated assistant on behalf of the application in an interaction turn of the interaction session, and a request that the application be provided with audio data content corresponding to audio data captured by a microphone of the client device during and / or after rendering of the content in the interaction turn. The method further includes processing the content using a trained machine learning model to generate an output based on the content and indicating whether it is appropriate to provide the audio data content to the application. In response to determining that providing the audio data content is inappropriate, the method includes restricting access to the audio data content by the application.

[0071] These and other implementations of the technology disclosed herein may include one or more of the following features.

[0072] In some implementations, processing the content is performed by one or more other computing devices in addition to the client device.

[0073] In some embodiments, determining that providing the audio data is inappropriate includes determining that the output does not meet an appropriateness threshold. In some of these embodiments, the method further includes selecting an appropriateness threshold from a plurality of candidate appropriateness thresholds based on one or more attributes of the application. In some of these embodiments, the one or more attributes of the application based on which the appropriateness threshold is selected include the type of application. In other of these embodiments, the method further includes selecting an appropriateness threshold from a plurality of candidate appropriateness thresholds based on a data security level previously specified for the client device and / or the user.

[0074] In some implementations, restricting the audio data includes disabling the application.

[0075] In some implementations, prior to the interaction session, the method includes training the machine learning model using supervised training instances, each including corresponding content rendered during a corresponding previous interaction turn with a corresponding user and a corresponding supervised indication of whether it is appropriate to provide audio data content in response to the rendering of the corresponding content.

[0076] Other embodiments disclosed herein include a method including: identifying content of a dialog turn based on content rendered during the dialog turn and based on an application that generated the content receiving audio data content corresponding to audio data captured by a microphone of a rendering client device after at least the content was rendered in the dialog turn; processing the content of the dialog turn using a machine learning model to generate an output based on the content and indicating whether it was appropriate to provide the audio data content to the application; and in response to determining that the generated output indicates it was inappropriate to provide the audio data content to the application, restricting access by one or more application instances of the target application to subsequent audio data content corresponding to the subsequent audio data.

[0077] These and other implementations of the technology disclosed herein may include one or more of the following features.

[0078] In some implementations, processing the content of the dialogue turn includes determining an application type of the target application and providing the application type as an input to a machine learning model.

[0079] In some implementations, processing the content of the dialog turn includes identifying a text representation of the content of the dialog turn and providing the text representation of the content of the dialog turn as input to a machine learning model.

[0080] In some embodiments, processing the content of the dialog turn includes processing the content of the dialog turn to determine an intent of the content of the dialog turn and providing the intent of the content of the dialog turn as an input to a machine learning model. In some of these embodiments, processing the content of the dialog turn includes processing the content of the dialog turn using an additional machine learning model to generate a vector and providing the vector as an input to the machine learning model.

[0081] In some embodiments, the method further includes identifying a context for the dialog turn within the dialog and providing the context for the dialog turn as an additional input to the machine learning model. In some of these embodiments, the context indicates that the dialog turn is the final dialog turn in the dialog. In other of these embodiments, the context indicates the content of one or more previous and / or subsequent dialog turns in the dialog.

[0082] Still other embodiments disclosed herein include a method including: identifying content of multiple dialogue turns based on the content of each dialogue turn rendered during the corresponding dialogue turn and subsequent instances of an application that generated the corresponding content receiving audio data content captured as audio data by a microphone of a corresponding rendering client device at least after the corresponding content has been rendered; processing the content of the dialogue turns using a trained machine learning model to generate an output that is based on the corresponding content and indicates whether it was appropriate to provide the audio data content to the application; and causing one or more current application instances of the application to limit subsequent audio data content in response to determining that a threshold value of the generated output indicates that it was inappropriate to provide at least a portion of the audio data content to the application.

[0083] These and other implementations of the technology disclosed herein may include one or more of the following features.

[0084] In some implementations, a given generated output indicates that it was inappropriate to provide the corresponding audio data content to an application if the given generated output exceeds a threshold value.

[0085] In some implementations, multiple dialogue turns were part of a single dialogue.

[0086] In some implementations, multiple dialogue turns were provided by a single instance of the target application among one or more instances of the target application.

[0087] In some implementations, the method further includes identifying an application type of the target application and determining a threshold for the generated output based on the application type.

[0088] In some implementations, the method further includes identifying a third party associated with the application and determining a threshold for the generated output based on the third party.

[0089] In situations where certain implementations described herein may collect or use personal information about users (e.g., data about users extracted from other electronic communications, information about the user's social network, the user's location, the user's time, the user's biometric information, the user's activity and demographic information, relationships between users, etc.), users are provided with one or more opportunities to control whether information is collected, whether personal information is stored, whether personal information is used, and how information about them is collected, stored, and used. In other words, the systems and methods described herein collect, store, and / or use a user's personal information only with explicit permission from the relevant user.

[0090] For example, a user may be given control over whether a program or feature collects user information about that particular user or about other users associated with the program or feature. Each user about whom personal information is collected is presented with one or more options to control the collection of information related to that user and to grant permission or approval for whether and what information is collected. For example, a user may be provided with one or more such control options over a communications network. Additionally, certain data may be processed in one or more ways to remove personally identifiable information before it is stored or used. As one example, a user's identity may be treated so that personally identifiable information cannot be determined. As another example, a user's geographic location may be generalized to a broader area, making it impossible to determine the user's specific location.

[0091] While several embodiments have been described and illustrated herein, various other means and / or structures can be utilized to perform the functions and / or obtain the results and / or obtain one or more advantages described herein, and each such variation and / or modification is deemed to be within the scope of the embodiments described herein. More generally, all parameters, dimensions, materials, and configurations described herein are intended to be illustrative, and the actual parameters, dimensions, materials, and / or configurations will depend on the specific application(s) for which the teachings are used. Those skilled in the art will recognize, or be able to ascertain using no more than routine experimentation, many equivalents to the specific embodiments described herein. Accordingly, it should be understood that the foregoing embodiments are presented by way of example only, and that, within the scope of the appended claims and their equivalents, embodiments may be practiced other than as specifically described and claimed. Embodiments of the present disclosure are directed to each individual feature, system, article, material, kit, and / or method described herein. Furthermore, any combination of two or more such features, systems, articles, materials, kits, and / or methods is within the scope of the present disclosure, unless such features, systems, articles, materials, kits, and / or methods are mutually inconsistent.

Claims

1. A method executed by one or more processors, comprising: During an interaction session occurring through a client device, between a user and an application, and mediated by an automated assistant, by the automated assistant and from the application, content rendered by the automated assistant on behalf of the application in an interaction turn of the interaction session; a request that the application be provided with audio data content corresponding to audio data captured by a microphone of the client device during and / or after rendering of the content in the dialogue turn; receiving the processing the content using a trained machine learning model to generate an output that is based on the content and indicates whether it is appropriate for the audio data content to be provided to the application; and In response to determining that it is inappropriate to provide the audio data content, restricting access to the audio data content by the application; A method comprising:

2. The method of claim 1 , wherein processing the content is performed by one or more other computing devices in addition to the client device.

3. The method of claim 1 , wherein determining that it is inappropriate to provide audio data comprises determining that the output does not meet a threshold of appropriateness.

4. selecting the suitability threshold from a plurality of candidate suitability thresholds based on one or more attributes of the application; The method of claim 3 further comprising:

5. The method of claim 4 , wherein the one or more attributes of the application from which the suitability threshold is selected include the type of the application.

6. selecting the appropriateness threshold from a plurality of candidate appropriateness thresholds based on a previously specified data security level for the client device and / or the user; The method of claim 3 further comprising:

7. The method of claim 1 , wherein restricting the audio data includes disabling the application.

8. Prior to said interactive session, Each of them, corresponding content that was rendered during a corresponding previous interaction turn with the corresponding user; and a corresponding supervised indication of whether it is appropriate to provide audio data content in response to rendering the corresponding content; and training the machine learning model using supervised training instances, The method of claim 1 further comprising:

9. A method executed by one or more processors, comprising: identifying content of an interaction turn based on content rendered during an interaction turn of an interaction session and based on a target application that generated the content receiving audio data content corresponding to audio data captured by a microphone of a rendering client device at least after the content has been rendered in the interaction turn; processing the content of the dialogue turn using a machine learning model to generate an output based on the content and indicating whether it was appropriate to provide the audio data content to the target application; and In response to determining that the generated output indicates that it was inappropriate to provide the audio data content to the target application, restricting access by the target application to subsequent audio data content corresponding to the subsequent audio data; A method comprising:

10. Processing the content of the dialogue turn comprises: determining an application type of the target application; providing the application type as an input to the machine learning model; 10. The method of claim 9, comprising:

11. Processing the content of the dialogue turn comprises: identifying a textual representation of the content of the dialogue turn; providing the textual representation of the content of the dialogue turn as an input to the machine learning model; 10. The method of claim 9, comprising:

12. Processing the content of the dialogue turn comprises: processing the content of the dialogue turn to determine an intent of the content of the dialogue turn; providing the intent of the content of the dialogue turn as an input to the machine learning model; 10. The method of claim 9, comprising:

13. Processing the content of the dialogue turn comprises: processing the content of the dialogue turn using an additional machine learning model to generate a vector; providing the vector as an input to the machine learning model; 10. The method of claim 9, comprising:

14. identifying a context of the dialog turn within the dialog session; providing the context of the dialogue turn as an additional input to the machine learning model; 10. The method of claim 9, further comprising:

15. The method of claim 14 , wherein the context indicates that the dialog turn is a final dialog turn in the dialog session.

16. The method of claim 14 , wherein the context indicates the content of one or more previous and / or subsequent dialogue turns of the dialogue session.

17. A method executed by one or more processors, comprising: identifying content of a plurality of dialogue turns based on respective content of the dialogue turns rendered during the corresponding dialogue turns and subsequently receiving audio data content captured as audio data by a microphone of a corresponding rendering client device at least after the corresponding content is rendered; processing the content of the dialogue turn using a trained machine learning model to generate an output that is based on the corresponding content and indicates whether it was appropriate to provide the audio data content to the application; and In response to determining that the generated output threshold indicates that it was inappropriate to provide at least a portion of the audio data content to the application, restricting subsequent access to audio data content by said application; A method comprising:

18. 20. The method of claim 17, wherein a given generated output indicates that it was inappropriate to provide the corresponding audio data content to the application if the given generated output exceeds a threshold value.

19. The method of claim 17 , wherein the multiple dialogue turns were part of a single dialogue.

20. The method of claim 17 , wherein the plurality of dialogue turns are provided by the application.

21. identifying an application type of the application; determining the threshold value of the generated output based on the application type; 21. The method of claim 20, further comprising:

22. Identifying a third party associated with the application; determining the threshold value of the generated output based on the third party; 20. The method of claim 17, further comprising:

23. A computer program comprising instructions which, when executed by one or more processors of a computing system, cause the computing system to perform the method of any one of claims 1 to 22.

24. One or more computing devices configured to perform the method of any one of claims 1 to 22.

Citation Information

Patent Citations

  • Management layer for multiple intelligent personal assistant services

    JP2018181330A

  • Digital device and speech to text conversion processing method thereof

    US20160357987A1