Automatic round turn description in multi-round dialog

By dynamically updating context-dependent durations and combining acoustic and natural language features, the session computing interface can more accurately detect the end of a user turn, solving the problems of user interruption or response delay in existing technologies and improving the user experience.

CN114762038BActive Publication Date: 2025-10-28MICROSOFT TECHNOLOGY LICENSING LLC
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202080081979.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2019-11-26
Filing Date
2020-10-28
Publication Date
2025-10-28
Estimated Expiration
2040-10-28

AI Technical Summary

Technical Problem

Existing session computing interfaces struggle to accurately detect the end of a user turn, potentially leading to premature interruptions or response delays, thus impacting user experience.

Method used

By dynamically updating the context-dependent duration, the system automatically detects the end of a round based on silence and conversation history in the user's speech, and uses a previously trained model combined with acoustic and natural language features to assess whether a round has ended.

Benefits of technology

It improves the accuracy and response speed of the session computing interface in multi-turn dialogues, avoids interrupting users, ensures timely responses, and enhances the user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114762038B_ABST
    Figure CN114762038B_ABST
Patent Text Reader

Abstract

A method for automatically describing turns in a multi-turn dialogue between a user and a session computing interface. The method involves receiving audio data encoded from the user's speech in the multi-turn dialogue. The audio data is analyzed to identify utterances in the user's speech followed by silence. In response to the silence exceeding a context-dependent duration dynamically updated based on the session history of the multi-turn dialogue and features of the received audio, the utterance is identified as the final utterance in the turn of the multi-turn dialogue, wherein the session history includes one or more previous turns of the multi-turn dialogue performed by the user and one or more previous turns of the multi-turn dialogue performed by the session computing interface.
Need to check novelty before this filing date? Find Prior Art

Description

Background Technology

[0001] A session computing interface can engage in dialogue with one or more users, such as by assisting users by answering queries. The session computing interface can conduct dialogues across multiple rounds, including user rounds in which the user utters one or more utterances, and computer rounds in which the session computing interface responds to one or more previous user rounds. Summary of the Invention

[0002] This summary is provided to introduce the selection of concepts in a simplified form, which will be further described in the detailed description below. This summary is not intended to identify key or essential features of the claimed subject matter, nor is it intended to limit the scope of the claimed subject matter. Furthermore, the claimed subject matter is not limited to implementations that address any or all of the shortcomings pointed out in any part of this disclosure.

[0003] In a method for automatically describing turns in a multi-turn dialogue between a user and a session computing interface, audio data encoded from the user's speech in the multi-turn dialogue is received. The audio data is analyzed to identify utterances in the user's speech that are followed by silence. In response to the silence exceeding a context-dependent duration dynamically updated based on the session history of the multi-turn dialogue and features of the received audio, the utterance is identified as the final utterance in the turn of the multi-turn dialogue. The session history includes one or more previous turns of the multi-turn dialogue conducted by the user and one or more previous turns of the multi-turn dialogue conducted by the session computing interface. Attached Figure Description

[0004] Figure 1 An exemplary data flow for a session computing interface is shown.

[0005] Figure 2A A method is shown to evaluate whether there is silence in user speech that exceeds a dynamically updated context-dependent duration following user speech.

[0006] Figure 2B A method for automatically describing turn boundaries in multi-turn dialogues is shown.

[0007] Figures 3A-3B An example of a dialogue is shown, which includes user utterances, silence, and responses from the session computing interface.

[0008] Figure 4 An exemplary computing system is shown. Detailed Implementation

[0009] Conversational computing interfaces enable human users to interact with computers in a natural way. A properly trained conversational computing interface can handle natural user interactions, such as spoken user utterances or written user commands, without requiring the user to use a specific syntax defined by the computer. This allows human users to use natural language when speaking to the computer. For example, a user can interact using natural language by asking the computer to answer a question or by giving the computer commands. In response, the conversational computing interface is trained to automatically perform actions, such as answering questions or otherwise assisting the user (e.g., booking an airline ticket in response to a user utterance such as "Book a flight at 8 am tomorrow").

[0010] The session computing interface can be configured to respond to a variety of user interaction events. Non-limiting examples of events include user utterances in the form of voice and / or text, button presses, network communication events (e.g., receiving the result of an application programming interface (API) call), and / or gesture input. More generally, events include any event that may be associated with user interaction and can be detected by the session computing interface, such as via input / output hardware (e.g., microphone, camera, keyboard, and / or touchscreen), communication hardware, etc. Such events can be represented in a computer-readable format, such as a text string generated by a speech-to-text interpreter or screen coordinate values ​​associated with touch input.

[0011] A conversational computing interface can respond to user speech or any other user interaction event by performing appropriate actions. These actions can be implemented by executing any suitable computer operation, such as outputting synthesized speech via a speaker and / or interacting with other computer technologies via an API. Non-limiting examples of actions include searching for search results, making purchases, and / or scheduling appointments. The conversational computing interface can be configured to perform any suitable actions to assist the user. Non-limiting examples of actions include performing calculations, controlling other computer and / or hardware devices (e.g., by calling an API), communicating over a network (e.g., by calling an API), receiving user input (e.g., in the form of any detectable event), and / or providing output (e.g., in the form of displayed text and / or synthesized speech). More generally, actions can include any behavior that a computer system is configured to perform. Other non-limiting examples of actions include: controlling electronic devices (e.g., turning lights on / off in a user's home, adjusting a thermostat, and / or playing multimedia content via a monitor / speaker), interacting with commercial and / or other services (e.g., calling an API to schedule a ride via a ride-hailing service and / or order food / parcels via a delivery service), and / or interacting with other computer systems (e.g., accessing information from a website or database, sending an email, and / or accessing a user's schedule in a calendar program).

[0012] In some examples, the conversational computation interface can engage in multi-turn dialogues with one or more users, for example, to assist users by answering queries. The conversational computation interface can conduct dialogues across multiple turns, including user turns where the user utters one or more utterances, and computer turns where the conversational computation interface responds to one or more previous user turns. "Turn" herein refers to any set of user interface events input by the user in a multi-turn dialogue. For example, a turn may include a set of one or more utterances. "Eutterance" herein may be used to refer to any span of audio and / or text representing the user's speech, such as the audio span for which an automatic speech recognition (ASR) system produces a single result. Non-limiting examples of utterances include words, sentences, paragraphs, and / or any suitable representation of the user's speech. In some examples, utterances can be obtained by processing audio data using any suitable ASR system. The ASR system can be configured to segment the user's speech into one or more individual utterances (e.g., separated by silence).

[0013] To facilitate better multi-turn dialogue, the session computation interface should be able to accurately detect the end of a user turn. However, a single user turn may consist of more than one individual utterance separated by periods of user silence. As used herein, “silence” can refer to a time span in which no particular user’s speech of interest is detected in the audio data, even if other voices or even the speech of other human speakers can be detected in the audio data during that time span. In some examples, silence may optionally include a time span in which the ASR system fails to recognize a particular utterance of the user, even if the user may be making nonverbal sounds (e.g., murmurs, sighs, or filler words such as “hm” or “umm”). However, in some examples, such user sounds and / or filler words may be treated in the same way as recognized words.

[0014] While ASR systems may be suitable for segmenting user speech into utterances, their segmentation may not correspond to appropriate turn boundaries. For example, when a user's turn comprises more than one utterance, although an ASR system can be configured to segment the turn into multiple utterances, it will not be configured to evaluate which of the multiple utterances actually corresponds to the end of the user's turn. Thus, if the end of an utterance detected by the ASR system is considered the end of a user's turn, the end of the turn may be prematurely determined before the user has finished speaking, potentially interrupting the user. For example, it is believed that 25-30% of human conversation "turns" (e.g., where a user attempts to complete a specific plan or goal) in task-oriented sessions include silences longer than 0.5 seconds. Thus, while ASR systems may be suitable for segmenting human speech into utterances, silences between utterances may need to be considered to accurately characterize user turns.

[0015] The session computation interface can be configured to avoid interrupting the user during a user turn and wait for a triggered response / action after the user finishes speaking (e.g., to allow the user to potentially continue speaking in subsequent utterances). In some examples, the session computation interface can detect the end of a user turn based on silences in the user's speech that exceed a predefined duration (e.g., a user turn is considered to have ended after 2 seconds of silence). However, a user turn may include silences in the middle of the user's speech. For example, the user may stop to think, or the user may be interrupted by a different task or by interaction with another speaker. More generally, a user turn may include one or more utterances separated by silences. Although the user may be silent between each utterance, the user turn may not end until after the final utterance. Thus, the end of a turn may be detected prematurely before the user has finished speaking, resulting in the user being interrupted. Furthermore, if the predefined silence duration is made longer (e.g., 10 seconds), the response of the session computation interface may be delayed due to the longer silence duration, even if the user is not interrupted frequently. When interacting with such a conversational computing interface, users may experience frustration due to interruptions, waiting, and / or pre-planning of speech to avoid waiting or interruptions (e.g., trying to speak without pausing).

[0016] Therefore, a method for automatically describing turns in a multi-turn dialogue between a user and a session computing interface dynamically updates a context-specific duration, which is used to evaluate the end of a user turn based on other contexts established based on the user's utterances and / or the session history of the multi-turn dialogue.

[0017] As disclosed herein, automatic puff description advantageously enables the session computation interface to wait for an appropriate duration for the user's speech to complete, thereby improving the user experience. For example, accurate puff description helps avoid interrupting the user by waiting an appropriate amount of time when the user may not have finished speaking, while also responding quickly to the user by waiting a shorter duration when the user may have finished speaking.

[0018] As a non-limiting example, the context-dependent duration can be dynamically updated based on the specific content of the user's utterance. For example, when a user is responding to a specific question (e.g., providing specific details related to performing a previously requested session computation interface), the user can be given more time to consider their answer. As another example, if the user uses filler words (e.g., "hmm" or "um"), a longer context-dependent duration can be dynamically selected.

[0019] Figure 1 An exemplary dataflow architecture 100 is illustrated, including a session computation interface 102 configured to automatically characterize turn boundaries by automatically detecting the end of a user turn in a multi-turn dialogue. The session computation interface 102 is configured to receive audio data 104, such as audio data of user speech output from a microphone listening to one or more speakers. The session computation interface 102 is configured to maintain a session history 106 of multi-turn dialogues between the session computation interface 102 and one or more users. The session computation interface 102 automatically detects the end of a user turn based on silences in the user's speech exceeding a context-dependent duration 120, which is dynamically updated based on the audio data 104 and further based on the session history 106. By using the dynamically updated context-dependent duration 120, the session computation interface 102 can avoid interrupting the user when silences exist in the user's speech, while also quickly processing user requests as the user finishes speaking.

[0020] The session computing interface 102 may include one or more microphones configured to output audio data. For example, the session computing interface 102 may include a directional microphone array configured to listen to speech, output audio data, and evaluate the spatial location corresponding to one or more utterances in the audio data. In some examples, the evaluated spatial location may be used to identify the speaker in the audio data (e.g., to distinguish nearby speakers based on the evaluated spatial location). In some examples, the session computing interface 102 may include one or more audio speakers configured to output audio responses (e.g., utterances by the session computing interface in response to user utterances, such as a description of an action performed by the session computing interface).

[0021] The session computation interface 102 is configured to dynamically update the context-dependent duration 120 based on features of the audio data 104 and / or based on the session history 106. Typically, the features can be provided to a previously trained model 130 configured to dynamically select the context-dependent duration 120 based on any suitable set of features. For example, the features may include acoustic features of the audio data 104 and / or natural language and / or text features derived from the audio data 104. Furthermore, the session history 106 can track any suitable acoustic, natural language, and / or text features throughout the multi-turn dialogue. However, although the session history 106 can track various features throughout the multi-turn dialogue, in general, the features in the session history 106 (such as natural language and / or text features) may not fully reflect the features of the user's voice and / or speech delivery. Therefore, the previously trained model 130 is configured to analyze the acoustic features of the audio data 104 to take into account the features of the user's voice and / or speech delivery. By taking into account both audio data 104 and conversation history 106, the previously trained model 130 can achieve relatively higher accuracy in assessing whether a round has ended, compared to assessments based solely on audio data 104 or solely on conversation history 106.

[0022] In some examples, the audio data 104 is characterized by one or more acoustic features. For example, the one or more acoustic features may include the user's instantaneous speech pitch, the user's baseline speech pitch, and / or the user's intonation relative to the user's baseline speech pitch. The user's baseline speech pitch can be determined in any suitable manner, for example, based on a baseline speech pitch tracked about the user's most recent utterances and / or tracked in the user's current session and / or (one or more) previous sessions against the user's session history 106. In examples, the instantaneous speech pitch can be determined based on any suitable acoustic analysis, for example, based on the Mel-frequency cepstral coefficients of the audio data 104 data. In examples, the baseline speech pitch can be evaluated based on a rolling window average and / or the maximum and / or minimum pitch range of the instantaneous speech pitch. The user's intonation relative to the user's baseline speech pitch can be evaluated in any suitable manner, for example, by the frequency difference between the user's instantaneous speech pitch and the user's baseline speech pitch during the audio data 104.

[0023] In some examples, the features of audio data 104 (e.g., acoustic features of audio data 104) include the user's speech rate. For example, speech rate can be determined by counting the number of words detected by the ASR system and dividing by the duration of the utterance. In some examples, the one or more acoustic features include one or more of the following: 1) a baseline speech rate for the user, 2) the user's speech rate in the utterance, and 3) the difference between the user's speech rate in the utterance and the baseline speech rate for the user. Alternatively or additionally, the features of audio data 104 may include the speech rate or speech rate difference at the end of the utterance. The speech rate at the end of the utterance can be selectively measured in any suitable manner, for example, by measuring the speech rate of a predefined portion of the utterance (e.g., the last quarter or the last half), by measuring the speech rate of a predefined duration of the utterance (e.g., the last 3 seconds), and / or based on measuring the speech rate after an inflection point and / or a change in speech rate (e.g., after the speech rate begins to rise or fall).

[0024] In addition to speech tone and / or speech rate, the context-dependent duration 120 can be set based on any other suitable acoustic features of the audio data 104. Non-limiting examples of acoustic features include the duration of utterances, speech amplitude, and / or the duration of words within an utterance.

[0025] In some examples, in addition to the acoustic features of the audio data 104, the previously trained model 130 is configured to evaluate textual and / or natural language features (e.g., lexical, syntactic, and / or semantic features derived from one or more user utterances in the audio data 104) identified for the audio data 104. Therefore, the context duration can be set based on features of the audio data 104, including one or more text-based features automatically derived from the audio data 104. For example, the textual and / or natural language features can be identified by any suitable natural language model. As a non-limiting example, natural language features can include the sentence-end probability of the final n-gram in the utterance. In some examples, the sentence-end probability of the final n-gram can be derived from a natural language model operating on the audio data 104. In some examples, the sentence-end probability of the final n-gram can be derived from a natural language model operating on the output of an ASR system operating on the audio data 104 (e.g., the natural language model can process the utterance based on the identified text rather than and / or other than the audio data). For example, words like “and” and “the” are less likely to appear at the end of an English sentence, having a lower associated sentence-end probability. More generally, natural language features can include the syntactic role of the last word in a sentence, such as an adjective, noun, or verb. For example, the end of a turn may be more closely associated with some syntactic roles (e.g., nouns may be more common at the end of a user's turn than adjectives). Therefore, the features of audio data 104 can include the syntactic properties of the final n-gram in the discourse.

[0026] In some examples, the features of audio data 104 include automatically identified filler words from a predefined list of filler words. For example, the predefined list of filler words might include non-verbal utterances such as “um,” “uh,” or “er.” As another example, the predefined list of filler words might include words and / or phrases associated with the user completing a thought, such as “wait a moment” or “let me think.” In some examples, the textual and / or natural language features of audio data 104 might include automatically identified filler words within audio data 104. (As in...) Figure 1 As shown, the context-dependent duration 120 can allow for a relatively long silence 108B2 based on the utterance 108B1, which includes the filler word “hmmm”.

[0027] In addition to the acoustic and / or natural language / text features of the audio data 104, the context-related duration 120 can be dynamically set based on any suitable context established in the conversation history 106 of a multi-turn dialogue, for example, based on any suitable features of previous user turns and / or conversation computation interface responses. In some examples, the context-related duration is dynamically set based on the conversation history 106 based at least on a semantic context derived from one or both of the following: 1) the user's previous utterances in the conversation history 106, and 2) the conversation computation interface's previous utterances in the conversation history 106. The semantic context can be evaluated by any suitable model (e.g., artificial intelligence (AI), machine learning (ML), natural language processing (NLP), and / or statistical models). As a non-limiting example of using semantic context to dynamically set a suitable context-related duration, if the conversation computation interface 102 is waiting for specific information (e.g., an answer to a clarifying question, a request for confirmation, and / or any other details from the user), a relatively long silence may be allowed before the conversation computation interface 102 obtains a turn. For example, as in Figure 1 As shown, the session computation interface 102 is waiting for specific information about preferred flight schedules and allows a relatively long silence 108B2 between user utterance 108B1 (where the user has not yet provided all the required information) and user utterance 108B3, while also allowing a shorter silence 108B4 after user utterance 108B3 (at which point the user has provided the required information).

[0028] More generally, for any acoustic and / or textual features described herein (e.g., speech tone, speech rate, word selection, etc.), the features can be extended to include derived features based on comparisons relative to a baseline for the feature, comparisons of feature changes over time, selective measurement of features toward the end of the user's utterance, such as those described above regarding baseline speech tone and speech rate at the end of the user's utterance. In some examples, the features of audio data 104 include one or more features of the utterance, selectively measured for the end portion of the utterance. For example, speech tone (similar to speech rate) can be selectively measured at the end of the utterance. In some examples, the context-dependent duration is dynamically set based on the conversation history 106, at least based on measurements of changes in features occurring throughout the utterance. As a non-limiting example, speech tone can be measured after an inflection point where the speech tone stops rising and begins to fall, or selectively measured after a point where the speech tone falls by more than a threshold amount. In some examples, the features of audio data 104 can include preprocessing results, including one or more utterances obtained from an ASR system. For example, the one or more utterances may include portions of audio segmented by the ASR system and / or corresponding natural language information associated with the audio data 104. In some examples, the ASR system is a pre-configured ASR system (e.g., the ASR system does not need to be specifically trained and / or configured to work with the session computing interface and / or the previously trained model described herein). In some examples, utterances from the ASR system may be preprocessed using speaker identification and / or listener identification to determine the speaker who uttered the user utterance for each user utterance, and the listener the speaker intends to address in the utterance. For example, the ASR system may be configured to evaluate the speaker and / or listener identity associated with each utterance. In some examples, the preprocessed utterances may be selectively fed to a previously trained model for analyzing the audio data 104 to dynamically evaluate context-dependent duration. For example, the preprocessed utterances may be filtered to provide only utterances spoken to the session computing interface by a specific user, while excluding utterances spoken by other users and / or to entities other than the session computing interface. In other examples, both pre-processed and / or unprocessed utterances can be fed into a previously trained model for analysis, for example, where the input features represent the speaker and / or listener for each utterance.

[0029] In some examples, the time span during which the ASR system fails to recognize a specific user utterance can be considered silence in the user's speech, even if the user may be uttering nonverbal sounds (e.g., murmurs, sighs, or filler words such as "hm" or "umm"). In other examples, nonverbal sounds and / or other user interface events may not be considered silence. For example, nonverbal sounds can be considered as acoustic features of audio data 104 (e.g., features used to measure pitch, volume, and / or speech rate of nonverbal sounds). For example, if the user begins to murmur or speak to a different entity (e.g., to another person), the previously trained model is configured to assess that the user has ended a turn based on features of audio data 104 and / or conversation history 106. For example, if the user's utterance is fully actionable through the conversation computation interface and the user immediately begins to murmur, the previously trained model can be configured to assess that the turn has ended (e.g., because the user is no longer speaking to the conversation computation interface). Alternatively, in other examples (e.g., based on different features of the conversation history 106), the previously trained model can be configured to infer that the user did not complete a turn but instead murmured during a period of silence (e.g., while thinking). Similarly, in some examples, if the user begins to interact with a different entity, the previously trained model is configured to assess the end of the user's turn based on features of the audio data 104 and / or the conversation history 106 (e.g., when the user begins to talk to someone else after completing a fully actionable utterance). In other examples, if the user begins to speak to a different entity, the previously trained model is configured to assess that the speech delivered to other entities at the user's turn was silent based on features of the audio data 104 and / or the conversation history 106 (e.g., the user can continue speaking to the conversational computing interface after paying attention to other entities).

[0030] In some examples, audio can be received via a directional microphone array, and the session computing interface can be configured to evaluate directional / location information associated with utterances in the audio based on the audio. For example, the directional / location information associated with utterances can be characteristics used to determine the speaker and / or receiver associated with the utterance. For example, the location associated with utterances from a user can be tracked throughout the session to evaluate whether the utterance originated from the user. Alternatively or additionally, the user's location can be evaluated based on any other suitable information, such as the location of one or more sensors associated with the user's mobile device. Similarly, the location of other speakers can be tracked based on audio data 104, and the location information can be used to evaluate the speaker identity for utterances in audio data 104. As another example, the session computing interface can be configured to evaluate the direction associated with utterances (e.g., based on sounds associated with utterances received at different microphones with different volumes and / or reception times). Thus, the direction associated with the utterance can be used as a characteristic for evaluating the speaker and / or receiver of the utterance. For example, if a user is speaking to a session computing interface, they may primarily emit speech in a first direction (e.g., towards a microphone facing the session computing interface), while when the user is speaking to a different entity, they may primarily emit speech in a second, different direction. Similarly, the frequency, amplitude / volume, and / or other characteristics of sound captured by one or more microphones can be evaluated to determine the speaker identity and / or recipient associated with the utterance in any suitable manner. For example, a user may speak to the session computing interface at a first volume level while speaking to other entities at other volume levels (e.g., speaking directly to the microphone when speaking to the session computing interface, and speaking at a different angle relative to the microphone when speaking to other entities). Furthermore, recipient detection can be based on the timing of the utterance; for example, if a user's utterance briefly follows the utterances of a different entity, the user's utterance may be a response to the utterances of that different entity.

[0031] The previously trained model 130 may include any suitable set or other combination of one or more previously trained AI, ML, NLP, and / or statistical models. The following will discuss... Figure 4 A non-restrictive example describing the model.

[0032] The previously trained model 130 can be trained on any suitable training data. Typically, the training data includes exemplary audio data and corresponding session history data, as well as any suitable features derived from the audio and / or session history. In other words, the training data includes features similar to those in the audio data 104 and / or session history 106. In an example, the training data includes one or more exemplary utterances and corresponding audio data and session history data, wherein each utterance is labeled to indicate whether it is the final utterance of a user turn or a non-final utterance of a user turn. Therefore, the model can be trained to reproduce the labels based on the features (e.g., predict whether a complete turn has occurred). For example, the previously trained model 130 is configured to predict whether a turn has ended based on features of the audio data 104 and / or based on the session history 106 regarding the duration of silence following the utterance in the audio data 104.

[0033] As an example, in response to a user uttering the specific utterance "Schedule a flight to Boston tomorrow," the previously trained model 130 may initially assess that a turn has not yet ended when the user has just finished speaking. However, if there is a brief silence following the user's utterance (e.g., after 0.5 seconds of silence), the session computation interface 102 is configured to reoperate the previously trained model 130 to reassess whether a turn has occurred regarding the duration of the silence following the utterance. Thus, the previously trained model 130 can determine that a turn has ended based on the silence following the utterance. Even if the previously trained model 130 determines that a turn has not yet ended, the session computation interface 102 is configured to continue using the previously trained model 130 to reassess whether a turn has ended. For example, if a longer silence occurs after the utterance (e.g., 2 seconds of silence), the previously trained model 130 may determine that the turn has ended.

[0034] In some examples, the training data may be domain-specific, such as travel plans, calendars, shopping, and / or appointment schedules. In some examples, the training data may include multiple domain-specific training datasets for different domains. In some examples, the training data may be domain-agnostic, for example, by including one or more different sets of domain-specific training data. In some examples, the model(s) may include one or more models trained on different sets of domain-specific training data. In some examples, the previously trained natural language model is a domain-specific model of multiple domain-specific models, wherein each domain-specific module is trained on multiple exemplary task-oriented dialogues corresponding to the domain. In some examples, the model(s) may include combinations and / or sets of different domain-specific models.

[0035] For example, the training data may include multiple exemplary task-oriented dialogues, each including one or more turns of an exemplary user. For example, the previously trained natural language model may be configured to classify utterances followed by silence as final utterances (e.g., a final utterance describing the end of a user's turn and any subsequent silence) or non-final utterances (e.g., user turns including additional utterances separated by silence). Exemplary task-oriented dialogues can be derived from any suitable session. As a non-limiting example, an exemplary dialogue can be derived from a task-oriented dialogue between two people. Alternatively or additionally, models(one) may be trained on task-oriented dialogues between three or more individuals, task-oriented dialogues between one or more individuals and one or more session computing interfaces, and / or non-task-oriented dialogues between a person and / or a session computing interface.

[0036] The session history 106 can be represented in any suitable form, such as a list, array, tree, and / or graphical data structure. Generally, the session history 106 represents one or more user rounds 108 and one or more computer rounds 110. In some examples, the user rounds 108 and / or computer rounds 110 can be recorded sequentially and / or organized according to the chronological order in which the rounds occurred, for example, using timestamps and / or indexes indicating the chronological order.

[0037] Conversation history 106 Figure 1 The right side is expanded to show a non-limiting example for a multi-turn dialogue. The session history 106 includes: a first user turn 108A; a first computer turn 110A, wherein the session computing interface 102 responds to the user based on the first user turn 108A; a second user turn 108B; and a second computer turn 110B, wherein the session computing interface 102 responds to the user based on the previous turn.

[0038] although Figure 1A session history 106 is shown that includes an explicit representation of silence (e.g., a silence duration indicating the amount of time during which no user speech was detected), but the session history 106 may represent silence and / or omit silence entirely in any suitable manner. In some examples, silence may be implicitly recorded in the session history 106 based on information associated with utterances in the session history 106. For example, each utterance may be associated with a start time, stop time, and / or duration, such that the silence duration is implicitly represented by the difference between the time / duration associated with the utterance. Alternatively or additionally, each word within each utterance may be associated with time / duration information, and thus, silence between utterances may be implicitly represented by the difference between the time associated with the word at the end of the utterance and the time associated with the word at the beginning of the subsequent utterance. For example, the session history may use a sequence of utterances to represent user turns, rather than using a sequence of utterances separated by silences. Furthermore, user turns may include one or more utterances in any suitable representation. For example, although the conversation history 106 presents utterances as text, utterances can be represented within a user turn using any suitable combination of text, speech audio, speech data, natural language embeddings, etc. More generally, a turn can include any number of utterances, optionally separated by silences, including utterance 108B1, silence 108B2, utterance 108B3, and silence 108B4, just like user turns 108BB.

[0039] Figure 1 The session history 106 is shown, depicting the state after several rounds between the user and the session computing interface 102. Although not in... Figure 1 It is explicitly shown that, however, a conversation history 106 can be established upon receiving audio data 104, for example, by automatically recognizing speech from audio data 104 to store user utterance text corresponding to user turns. For example, conversation history 106 may initially be empty and then progressively expanded to include user turns 108A, computer turns 110A, user turns 108B, and computer turns 110B as they occur. Furthermore, the conversation history can be maintained in any suitable manner for multi-turn dialogues, including one or more conversation computing interfaces and one or more users, across any suitable number of turns. For example, each utterance in the conversation history can be tagged to indicate the specific user who uttered the utterance (e.g., determined by speech recognition, facial recognition, and / or a directional microphone). For example, conversation history 106 can be expanded using subsequent turns via conversation computing interface 102 and one or more users. Furthermore, although… Figure 1The session history 106 is shown, in which user turns have been described, but user turns can be described dynamically based on the identification of the final utterance followed by silence (e.g., with the receipt of audio data 104), wherein the duration of the silence exceeds the context-dependent duration 120, which is dynamically updated as described herein.

[0040] In some examples, a round may include a session computing interface response, such as responses 110A1 and 110B1 for computer rounds 110A and 110B, respectively. Alternatively or additionally, a computer round may include any other suitable information relating to the session computing interface action, such as executable code for responding to a user, and / or data values ​​generated by performing actions, running code, and / or calling APIs.

[0041] The session computation interface 102 is configured to automatically describe turns in a multi-turn dialogue. To describe turns, the session computation interface evaluates whether a user's utterance is the final utterance in a turn. If the utterance is the final utterance, its end marks the end of a turn and the beginning of a subsequent turn. The session computation interface does not use static schemes to describe turns, such as silence following a speech exceeding a fixed duration. Instead, the session computation interface evaluates whether a utterance is the final utterance in a particular turn based on a context-specific duration 120 that is dynamically updated / changed according to features of the session history 106 and audio data 104.

[0042] Based on the user's utterance in the audio data 104 and based on the conversation history 106, the context-dependent duration 120 can be dynamically updated to a longer or shorter duration. For example, the conversation computation interface 102 can be configured to dynamically update the context-dependent duration 120 based at least on more recent acoustic features of the audio data 104 and / or the conversation history 106. For example, a previously trained model 130 can be configured to selectively evaluate different portions of the audio data 104. In an example, the previously trained model 130 can be configured to compute an overall evaluation based on a weighted combination of individual evaluations of different portions of the audio data. Alternatively or additionally, the previously trained model 130 can be configured to incorporate other features to evaluate timestamp features associated with the audio data 104 (e.g., such that the previously trained model 130's evaluation of other features is contextualized by timestamp feature data). Thus, the previously trained model 130 can be configured to base the overall evaluation on an appropriate weighted combination that selectively weights one or more specific portions of the audio data. For example, weight parameters for weighted combinations can be learned during training, such as by adjusting the weight parameters to selectively evaluate one or more portions of the audio data according to a suitable weighted combination. In the example, the previously trained model 130 is configured to selectively evaluate the end portion of the audio data 104 and / or the conversation history 106 (e.g., the last portion of the audio data defined in any suitable manner, such as duration, such as 1 second, 2 seconds, or 10 seconds; or the last portion of a utterance in the audio data, such as the last 30% of the utterance or the time corresponding to the last three words in the utterance). In some examples, the previously trained model 130 is configured to selectively evaluate multiple distinct portions of the audio data (e.g., features corresponding to the first 20%, last 20%, last 10%, and the last word of the audio data, as a non-limiting example). Thus, the previously trained model 130 can be configured (e.g., via training) to evaluate context-dependent duration 120 based on features for different portions of the audio data.

[0043] If a silence of at least the context-relevant duration is detected, the session computation interface 102 is configured to determine the end of a user turn. By selecting a relatively long context-relevant duration 120 for a specific context, the session computation interface 102 allows the user to remain silent for a relatively long time and then resume speaking without being interrupted by the session computation interface 102. For example, the relatively long context-relevant duration 120 may be selected based on the specific context associated with the longer silence (e.g., because the user may spend more time considering an answer about that context). For example, when the user is in the middle of expressing details about a decision such as an airline booking in utterance 108B1, the session computation interface 102 is configured to recognize that the user turn ends after a silence with a relatively long context-relevant duration 120 (e.g., a 17-second silence 108B2, or even longer). For example, as described herein, the session computation interface 102 is configured to operate a previously trained model 130 to evaluate whether a utterance followed by silence is the final utterance, describing the end of the user turn. In some examples, the previously trained model 130 may evaluate whether a utterance followed by silence is a non-final utterance (e.g., the round has not yet ended). If the silence lasts for a longer duration, the session computation interface 102 may be configured to continuously operate the previously trained model 130 to re-evaluate whether the utterance following the longer duration of silence is the final utterance of the round. Thus, the context-dependent duration 120 can be effectively determined by the previously trained model 130 based on the continuous re-evaluation of whether the utterance following the silence is the final utterance as the silence continues and the duration of the silence increases.

[0044] Similarly, by selecting a relatively short context-dependent duration 120 for a specific context, the session computation interface 102 can quickly process user requests without causing the user to wait. For example, when the user has expressed relevant details that can be manipulated by the session computation interface 102, such as in the utterance 108B3 specifying an 8:00 AM flight, the session computation interface 102 is configured to select a relatively short context-dependent duration 120 so that it can proceed to the computer round 110B after a brief silence (e.g., a 1-second silence 108B4, or even a shorter silence).

[0045] The session computation interface 102 can evaluate whether user speech includes silence greater than the context-specific duration 120 in any suitable manner. As a non-limiting example, the session computation interface 102 can maintain a round-end timer 122. The round-end timer 120 can take any duration value between zero (e.g., zero seconds) and the context-specific duration 120. Typically, when user speech is detected in audio data 104, the context-specific duration 120 can be continuously updated based on the user speech and the session history 106, for example, such that the context-specific duration 120 is a dynamically updated value dependent on the context established by the user speech and / or the session history 106. Furthermore, each time user speech is detected in audio data 104, the round-end timer 122 can be reset to zero, for example, because by definition, the user speech detected in audio data 104 is not silence in the user speech. If no user voice is detected in the audio data 104 (e.g., if silence occurs in the user voice), the round-end timer 122 can be increased and / or advanced in any suitable manner, for example, to reflect the observed duration of silence in which no user voice was detected. If no user voice is detected in the audio data 104 for a sufficiently long duration, the round-end timer 122 may exceed the context-dependent duration 120. In response to the round-end timer 122 exceeding the context-dependent duration 120, the session computation interface 102 is configured to evaluate whether silence of the context duration 120 has occurred in the audio data 104. Thus, the session computation interface 102 can determine that the user round has ended. This disclosure includes non-limiting examples of evaluating whether silence of at least dynamically updated context-dependent duration 120 has occurred by maintaining the round-end timer 122. Alternatively or additionally, evaluating whether silence of at least dynamically updated context-dependent duration 120 has occurred can be implemented in any suitable manner. For example, the session computing interface 102 can evaluate the duration of silence based on the timestamp value associated with the user's voice and / or audio data 104, rather than maintaining the round end timer 122. For example, the session computing interface 102 can compare the current timestamp value with a previous timestamp value associated with the user's most recently detected voice, so as to evaluate the duration of silence as the difference between the current timestamp value and the previous timestamp value.

[0046] Figure 2AAn exemplary method 20 for assessing whether there is silence following a user's speech that exceeds a dynamically updated context-related duration is illustrated. Method 20 includes maintaining the dynamically updated context-related duration and maintaining a round-end timer indicating silence observed in the audio data, and comparing the dynamically updated context-related duration with the round-end timer. For example, method 20 can be used to determine whether a user's speech is the final utterance of a user round based on the identification of silence following at least the context-related duration. Figure 2A Method 20 can be derived from Figure 1 The session computation interface 102 is used to implement this. For example, the session computation interface 102 can execute method 20 to maintain dynamically updated context-specific duration 120 and round end timer 122.

[0047] At 22, method 20 includes analyzing audio data including user speech. For example, the analysis may include evaluating one or more acoustic features of the audio data, performing ASR on the audio data, and / or evaluating natural language and / or text features of the audio data (e.g., based on the performed ASR). As an example, the audio data may be some of all audio data received in the session. Alternatively or additionally, the audio data may include an incoming audio data stream, such as audio data captured and output by a microphone. For example, the incoming audio data stream may be sampled to obtain a first portion of the incoming audio, and then subsequently resampled to obtain a second subsequent portion of the incoming audio following the first received portion. Thus, the analysis may be based on any suitable portion of the incoming audio data stream. Analyzing the audio data may include detecting silence in the user speech, using any suitable criteria for identifying silence (e.g., as described above). At 24, method 20 includes updating the session history of a multi-turn dialogue based on the analysis of the audio data. For example, updating the session history may include recording one or more words in the speech, recording acoustic features of the speech, and / or recording semantic context based on the user's expressed intent in the speech.

[0048] At point 26, method 20 includes dynamically updating the context-dependent duration based on analysis of both the audio data and the conversation history. For example, the context-dependent duration can be determined based on an assessment of whether the user is likely to have completed the conversation based on the audio data; for instance, a longer context-dependent duration is assigned if the user is assessed as likely not to have completed the conversation. As another example, the context-dependent duration can be dynamically updated to a larger value if the user's recent speech includes filler words such as "umm". The context-dependent duration can be dynamically updated based on any suitable assessment of the features of the audio data and also on the conversation history, for example, by operating a previously trained model on the audio data and the conversation history. Figure 1 The session computation interface 102 is configured to provide audio data 104 and session history 106 to the previously trained model 130 so as to dynamically update the context-specific duration 120 based on the results computed by the previously trained model 130.

[0049] At 28, method 20 includes resetting the round end timer based on the context-dependent duration, for example, by instantiating the timer by setting the current time of the timer to zero and configuring the timer to expire after the context-dependent duration.

[0050] At point 30, method 20 includes analyzing additional audio data to assess whether additional user speech occurred before the round-end timer expired. For example, analyzing additional audio data may include, for instance, continuously receiving additional audio data output from a microphone (e.g., an incoming audio stream) to analyze user speech captured by the microphone in real time as the user speaks.

[0051] In response to detecting more user speech in the audio data before the round-end timer expires, method 20 includes determining at 32 that the user's speech is a non-final utterance of the user's round. Therefore, method 20 also includes returning to 22 to analyze additional audio data (e.g., additional audio data including one or more subsequent utterances) and dynamically re-evaluating the context-dependent duration.

[0052] In response to the absence of detected user speech in the audio data before the round end timer expires, method 20 includes: at 34, determining that the user speech is the final utterance of the user round based on a silence of at least a context-related duration following the recognition of the user speech. Therefore, the user round can be automatically described as ending with the determined final utterance.

[0053] Method 20 can be operated continuously to analyze audio data, such as analyzing audio data while receiving user speech in real time via a microphone. For example, Method 20 can be operated on the initial portion of the audio data to dynamically evaluate the context-dependent duration. A round-end timer runs, and additional audio is received / analyzed while the timer is running. As additional audio is received / analyzed, Method 20 can be operated to determine whether additional user speech has been detected before the round-end timer expires (e.g., causing a re-evaluation of the context-dependent duration and resetting the timer). If the round-end timer has elapsed based on the context-dependent duration without detecting user speech, the user's speech is determined to be the final utterance of the user's round.

[0054] As an example, return to Figure 1Method 20 can operate on the initial portion of the audio data 104 corresponding to the user's speech in user turn 108A, for example, the portion of the user's utterance 108A1 of "arrange a flight". Based on the audio data 104 and the session history 106, the session computation interface 102 is configured to provide the audio data 104 and the session history 106 as input to a previously trained model 130. Therefore, based on the audio data 104 and the session history 106, the previously trained model 130 can dynamically determine an initial context-relevant duration of 3 seconds, during which the user can specify additional information related to the flight. Thus, the turn end timer can be initially reset to 3 seconds. The turn end timer can expire partially, but before the full 3 seconds expire, the audio data 104 can include additional speech from the user, saying "...to Boston tomorrow". Thus, based on the additional speech in the audio data 104, the session computation interface 102 can dynamically determine an updated context-relevant duration of 2 seconds and reset the turn end timer to 2 seconds. Therefore, if the turn-end timer expires, such as when no further user speech occurs within a 2-second duration as indicated by silence 108A2, the session calculation interface 102 can determine that the user's speech is the final utterance of the user's turn, based on the recognition that there is at least a context-related duration of silence following the user's speech. Thus, the user turn can be automatically described as ending with the determined final utterance.

[0055] As shown in session history 106, the multi-turn dialogue includes user utterance 108A1, in which the user requests session computation interface 102 to "schedule a flight to Boston tomorrow." After a 2-second silence 108A2, session computation interface 102 responds in computation turn 110A, including a response 110A1 informing the user that "there are two options: 8 AM tomorrow and 12 PM tomorrow," and asking the user which option they prefer. Although in Figure 1 Not shown, but the session calculation interface 102 can perform any appropriate action to ensure that the response 110A1 is relevant, such as using one or more APIs associated with airline bookings to determine that the 8 a.m. and 12 p.m. options are available, or using one or more APIs associated with user configuration settings to ensure that the options are consistent with previously specified user preferences (e.g., preferred airline, cost limits).

[0056] As described above, the session computation interface 102 is configured to provide a previously trained model 130 with session history 106 and audio data 104, the previously trained model 130 being configured to dynamically update a context-specific duration 120 when the user speaks and / or when silence 108B2 occurs. Figure 1As shown, as a non-limiting example, the context-dependent duration 120 can be dynamically updated to 2 seconds. As a non-limiting example of possible factors in a silence duration of 2 seconds evaluated by one or more previously trained models, one or more previously trained models could determine that the turn ends after a relatively short 2-second silence based on the user utterance 108A1 being a complete sentence containing an actionable request to schedule a flight to Boston.

[0057] Therefore, in computer turn 110A, session computing interface 102 informs the user of two flight scheduling options and asks the user to select a preferred option. In user turn 108B, the user initially begins speaking with utterance 108B1 and then stops speaking during a 17-second silence 108B2. However, despite the relatively long silence of 17 seconds, session computing interface 102 is configured to recognize that the user has not yet completed user turn 108B. (See above reference...) Figure 2A As described, the session computation interface 102 can be configured to dynamically update the context-dependent duration to a relatively large value (e.g., 20 seconds, 30 seconds, or any other suitable time) based on the evaluation of features of audio data 104 and session history 106 by a previously trained model 130.

[0058] As a non-limiting example of possible factors in such an evaluation, user utterance 108B1 ends with a filler word (“hmmm”) and utterance 108B1 is not a complete sentence. Furthermore, utterance 108B1 does not specify a preference for flight scheduling, but based on session history 106, session computation interface 102 can anticipate a user-specified preference based on the response 110A1 in a previous computer round 110A. Therefore, session computation interface 102 is configured to recognize that user round 108B includes additional utterance 108B3. Moreover, since utterance 108B3 does specify a preference for flight scheduling, namely “8 AM flight,” session computation interface 102 is configured to recognize that the user may have finished speaking (e.g., because session computation interface 102 has been given all the information necessary to complete the action of scheduling a flight). Therefore, based on the evaluation of the features of the audio data 104 and the session history 106 by the previously trained model 130, the session computing interface 102 is configured to respond after a brief silence 108B4 of only 1 second, wherein the computer round 110B includes a response 110B1 indicating that the flight has been successfully scheduled.

[0059] As in Figure 1As shown, the session calculation interface 102 is configured to dynamically update the context-dependent duration 120 to various durations based on audio data 104 and / or session history 106. For example, after utterances 108A2 and 108B3, the session calculation interface 102 is configured to dynamically evaluate relatively short (e.g., less than 3 seconds) silence durations, for example, based on features of the audio data 104 associated with the user who is finishing speaking and / or the session history 106. However, after utterance 108B1, even if the user stops speaking for a relatively long silence 108B2 of 17 seconds, the session calculation interface 102 is also configured to dynamically allow even longer silences (e.g., by dynamically updating the context-specific duration to larger values, such as 20 seconds, and / or by continuously re-evaluating whether the turn has ended based on audio data 104 and / or session history 106 when silence 108B2 occurs). Therefore, the session computing interface 102 avoids interrupting the user between utterances 108B1 and utterances 108B2, thereby allowing the user to complete the speech.

[0060] although Figure 1 The illustration shows silences of exemplary lengths (such as 1 second, 2 seconds, or 17 seconds), but the dynamically updated context-dependent duration 120 can be any suitable duration, such as 0.25 seconds, 0.5 seconds, or any other suitable duration. In some examples, based on the characteristics of the audio data 104 and / or the session history 106, the session computation interface 102 can be configured to evaluate very short silences (e.g., 0.25 seconds, 0.1 seconds, or less) based on characteristics associated with the user who is finishing speaking. For example, after the user selects “8 a.m. flight” in utterance 108B3, the session computation interface 102 can alternatively be configured to begin a computer round 110B after a very short silence of only 0.1 seconds, for example, based on the user fully specifying the “8 a.m.” option even before the user finishes saying the word “flight,” thereby allowing the session computation interface 102 to output a response 110B1 quickly after the user finishes speaking. For example, in some examples, the context-dependent duration 120 could be zero seconds (e.g., no detectable silence duration) or a very small duration (e.g., 1 millisecond or 1 microsecond). For example, a previously trained model 130 could assess whether a user turn has ended before and / or immediately after the end of the user's utterance and before any silence appears in the user's speech. As an example, if the user finishes speaking to the session computing interface 102 and subsequent user speech is not directed to the session computing interface 102 (e.g., because the user is addressing a different entity, or is "thinking aloud" by talking to themselves).

[0061] although Figure 1An example of a session history 106 with interactions via user utterances and computer-responded utterances is shown, but the session computation interface 102 can also be configured to receive other user interaction events. Therefore, although the examples in this disclosure concern features derived from audio data and / or spoken text, the methods of this disclosure can be similarly applied to any suitable user interaction event and its features. For example, detecting the end of a round for a session computation interface 102 configured to receive a user interaction event in the form of a button press can be based on any suitable feature, such as a button press, the frequency of button presses toward the end of a time span in which a button press was detected. Although the methods described herein concern processing speech audio, text features, and / or natural language features, it is believed that the methods described herein (e.g., using one or more previously trained models trained with labeled data, such as...) Figure 1 The previously trained model 130 can similarly be applied to other user interaction events. For example, the natural language model can be configured to process sequential data representing any appropriate set of user interaction events occurring over a time span, including user utterances, gestures, button presses, gaze direction, etc. For example, user gaze direction can be considered when detecting the recipient of a speaker's and / or user's and / or other speaker's utterances. As another example, the previously trained model can be configured to assess whether a user turn has not yet ended (because the user is using the computer device to look up information related to the user turn) based on features indicating that the user's gaze is associated with the computer device (e.g., a mobile phone) associated with the session computing interface 102.

[0062] The methods disclosed herein (e.g., method 20) can be applied to describe round boundaries for any suitable session computation interface based on audio data and any suitable representation of session history. Figure 1 A non-limiting example of the dataflow architecture for the session computation interface 102 and a non-limiting example of the session history 106 are shown. However, the description of the round boundaries according to this disclosure can be performed with respect to any suitable session computation interface having any suitable dataflow architecture (e.g., using a previously trained model configured to use any suitable set of input features from the session computation interface and / or the session history). Figure 2B An exemplary method 200 is shown for describing turns in a multi-turn dialogue between a user and a session computing interface.

[0063] At 202, method 200 includes receiving audio data encoded from a user's speech in a multi-turn dialogue. For example, the audio data may include one or more utterances of the user separated by silence. The one or more utterances in the audio data may correspond to one or more turns. Therefore, method 200 can detect the end of a turn based on audio data from a session computing interface and / or session history.

[0064] At 204, method 200 includes analyzing the audio data to identify utterances followed by silence in the user's speech. Analyzing the audio data may include operating an ASR system. For example, the utterance may be one of a plurality of utterances separated by silence in a round, and the plurality of utterances may be provided automatically by the ASR system. In some examples, the ASR system may perform some initial segmentation of the utterances based on silence in the user's speech. For example, the ASR system may be configured to automatically segment the audio into individual utterances based on identifying silences of a predefined threshold duration. However, as described above, a single user round may include a plurality of utterances separated by silence. Therefore, the methods of this disclosure may be operated to appropriately depict rounds comprising a plurality of utterances provided by an ASR system.

[0065] At 206, method 200 includes dynamically updating silences exceeding the context-dependent duration in response to features based on the conversation history of the multi-turn dialogue and received audio, identifying the utterance as the final utterance in a turn of the multi-turn dialogue. As a non-limiting example, assessing whether a silence exceeds the context-dependent duration can be based on, as in... Figure 2A The method 20 shown is used to formulate the session history. The session history can be maintained in any suitable format, for example, represented as in... Figure 1 The non-limiting example of session history 106 shown includes one or more previous rounds of a multi-turn dialogue performed by the user and one or more previous rounds of a multi-turn dialogue performed by the session calculation interface. Return to Figure 2B At 208, the context-related duration is dynamically updated based on the conversation history of the multi-turn dialogue and also based on the characteristics of the received audio. At 210, in response to silence exceeding the context-related duration evaluated at 206, method 200 further includes identifying the utterance as the final utterance in a user turn of the multi-turn dialogue. At 212, in response to silence not exceeding the context-related duration, method 200 further includes determining that the turn of the multi-turn dialogue includes additional utterances that begin after the silence. Therefore, after 212, method 200 may include returning to 204 to further analyze the audio data to identify additional utterances followed by further silences.

[0066] Context-dependent durations can be dynamically updated based on any suitable analysis of any feature set of the received audio data and session history. For example, this can be based on operations on one or more previously trained models (e.g., AI, ML, NLP, and / or statistical models), such as... Figure 1 The previously trained model 130 can be used to dynamically update the context-dependent duration. For example, the context-dependent duration can be dynamically updated based on features of the user's voice / speech delivery, the words spoken by the user, and / or any other suitable user interface events that occur while receiving audio data and / or are recorded in the session history.

[0067] In some examples, such as in Figure 3A As shown, users may remain silent for a period of time, even if they haven't finished speaking. Thus, the conversation computation interface may detect the end of a turn, even though the turn may not be fully actionable based on information provided by the user. Therefore, the conversation computation interface can be configured to determine whether a turn is fully actionable independently of whether it has ended, and to generate a response utterance based on the final utterance in response to a turn not being fully actionable. For example, a previously trained model can be trained based on training data including both fully actionable and incompletely actionable turns. In some examples, identifying whether a turn is fully actionable may include operating a previously trained model to generate actions for the response utterance and determining a confidence value for prediction. For example, if the confidence value for prediction exceeds a predefined threshold, the turn can be considered fully actionable; while if the confidence value for prediction does not exceed a predefined threshold, the turn can be considered incompletely actionable.

[0068] Session computing interface 102 is configured as an operation model 130 to describe, for example, in Figure 3A The example scenario shown illustrates round boundaries. For example, based on operation method 20 and / or method 200 as described above, round boundaries can be detected between silence 108A2 and response 110A1, between silence 108B2 and response 110B1, and between silence 108C2 and response 110C1.

[0069] Following user utterance 108B1, since user utterance 108B1 has not yet provided the information needed to fully respond to the user, the session computation interface allows for long silences. However, as in Figure 3AAs shown, even after a 45-second silence, the user did not resume speaking. For example, the user might have been engaged in a different task. Therefore, the session computation interface can be configured to detect the end of the user's turn and output response 110B1 instead of waiting indefinitely. As shown, response 110B1 informs the user that the session computation interface still needs to know which flight time the user prefers and restates the option. Subsequently, the user re-engages in the conversation and selects the "8 AM" option in utterance 1080. Therefore, the session computation interface is configured to immediately resume processing the user's request using the provided information, outputting response 110C1 after a brief 1-second silence 108C2.

[0070] As in Figure 3B As shown, in some examples, in addition to the voice of the primary user described in the examples above, the received audio also includes the voices of one or more different users. The voices of one or more different users may be unrelated to the session between the primary user and the session computing interface, thus creating potential difficulties in describing rounds. For example, as in... Figure 3B As shown, user turn 108B may include interruptions from other speakers, such as a barista in a coffee shop responding to a user's other speaker utterances 108B2 regarding a coffee order. Furthermore, in addition to interruptions from other speakers, user turn 108C may include utterances from the primary speaker who is not speaking at the session computing interface, for example, responding to a question posed by the barista in other speaker utterances 108C2.

[0071] However, the session computation interface according to this disclosure is configured to describe rounds in a multi-speaker scenario, wherein, Figure 3B The sessions depicted are non-limiting examples. For instance, a previously trained model of a session computation interface could be trained on a session history where a user's turn is interrupted, such as in... Figure 3B The conversation history described in document 106 is used to configure the previously trained model to identify similar situations.

[0072] For example, as described above, the session computation interface can be configured to preprocess audio data and / or session history data to indicate the speaker and / or the listener associated with each utterance in the audio data and / or session history. Therefore, in some examples, evaluating turn boundaries can be based on features derived from a filtered subset of utterances provided to a previously trained model, e.g., only based on the utterances of the main user in which the main user speaks to the session computation interface.

[0073] Alternatively or additionally, a previously trained model describing turns for the primary user can be trained based on evaluation features derived from utterances from which the primary user is not a speaker and / or from which the conversational computing interface is not a listener. In some examples, the previously trained model can be trained based on exemplary audio data and conversation history associated with multiple different speakers and / or in which speakers address multiple different listeners. For example, the conversation history and / or audio data can be associated with labels indicating turns for the primary user (e.g., indicating a time span and / or a set of utterances associated with the turn, which may include interruptions such as utterances from other users). For example, exemplary conversation history may include conversations where the primary user is interrupted in the middle of a turn (e.g., as shown in user turns 108B and 108C), and / or in which the primary user speaks to different entities in the middle of a turn (e.g., as shown in user turn 108C).

[0074] The methods and processes described herein can be attached to a computing system of one or more computing devices. Specifically, such methods and processes can be implemented as an executable computer application, a network-accessible computing service, an API, a library, or a combination of the above and / or other computing resources.

[0075] Figure 4 A simplified representation of a computing system 400 is illustrated schematically. The computing system 400 is configured to provide any computing functionality described herein. The computing system 400 may take the form of one or more personal computers, network-accessible server computers, tablet computers, home entertainment computers, gaming devices, mobile computing devices, mobile communication devices (e.g., smartphones), virtual / augmented / mixed reality computing devices, wearable computing devices, Internet of Things (IoT) devices, embedded computing devices, and / or other computing devices. The computing system 400 is designed for use as described herein. Figure 1 The following is a non-limiting example of an embodiment of the data flow architecture 100 of the session computing interface 102 shown.

[0076] The computing system 400 includes a logic subsystem 402 and a storage subsystem 404. The computing system 400 may optionally include a display subsystem 406, an input subsystem 408, a communication subsystem 410, and / or... Figure 4 Other subsystems not shown.

[0077] Logical subsystem 402 includes one or more physical devices configured to execute instructions. For example, the logical subsystem may be configured to execute instructions as part of one or more applications, services, or other logical structures. The logical subsystem may include one or more hardware processors configured to execute software instructions. Alternatively or additionally, the logical subsystem may include one or more hardware or firmware devices configured to execute hardware or firmware instructions. The processor of the logical subsystem may be single-core or multi-core, and the instructions running thereon may be configured for sequential, parallel, and / or distributed processing. Individual components of the logical subsystem may optionally be distributed across two or more separate devices that may be remotely located and / or configured for coordinated processing. Aspects of the logical subsystem may be virtualized and run by remotely accessible networked computing devices configured in a cloud computing configuration.

[0078] Storage subsystem 404 includes one or more physical devices configured to temporarily and / or permanently store computer information, such as data and instructions executable by the logical subsystem. When the storage subsystem includes two or more devices, the devices may be co-located and / or remotely located. Storage subsystem 404 may include volatile, non-volatile, dynamic, static, read / write, read-only, random access, sequential access, location-addressable, file-addressable, and / or content-addressable devices. Storage subsystem 404 may include removable and / or built-in devices. The state of storage subsystem 404 may be changed—for example, to store different data—when the logical subsystem executes instructions.

[0079] Various aspects of the logic subsystem 402 and the storage subsystem 404 can be integrated together into one or more hardware logic components. Such hardware logic components may include, for example, application-specific integrated circuits (PASIC / ASIC), application-specific standard products (PSSP / ASSP), system-on-a-chip (SOC), and complex programmable logic devices (CPLD).

[0080] The logical subsystem and storage subsystem can collaborate to instantiate one or more logical machines. As used herein, the term "machine" is used collectively to refer to a combination of hardware, firmware, software, instructions, and / or any other components that collaborate to provide computer functionality. In other words, "machine" is never an abstract concept and always has a tangible form. A machine can be instantiated by a single computing device, or a machine can include two or more sub-components instantiated by two or more different computing devices. In some implementations, a machine includes local components (e.g., software applications running on a computer processor) that collaborate with remote components (e.g., cloud computing services provided by a network of server computers). The software and / or other instructions that give a particular machine its functionality may optionally be stored as one or more non-executing modules on one or more suitable storage devices.

[0081] The machine can be implemented using any suitable combination of existing and / or future ML, AI, statistical and / or NLP techniques, such as via one or more previously trained ML, AI, NLP and / or statistical models. Non-limiting examples of techniques that can be incorporated into the implementation of one or more machines include support vector machines, multilayer neural networks, convolutional neural networks (e.g., including spatial convolutional networks for processing images and / or videos, temporal convolutional neural networks for processing audio signals and / or natural language sentences, and / or any other suitable convolutional neural network configured to convolve and pool features across one or more temporal and / or spatial dimensions), recurrent neural networks (e.g., long short-term memory networks), associative memories (e.g., lookup tables, hash tables, Bloom filters, neural graphs). These include neural random access memory (NRAM), word embedding models (e.g., GloVe or Word2Vec), unsupervised spatial and / or clustering methods (e.g., nearest neighbor algorithms, topological data analysis and / or k-means clustering), graph models (e.g., (hidden) Markov models, Markov random fields, (hidden) conditional random fields and / or AI knowledge bases), and / or natural language processing techniques (e.g., tokenization, stemming, region selection and / or dependency parsing, and / or intent recognition, segmentation models and / or super-segmentation models (e.g., hidden dynamic models)).

[0082] In some examples, the methods and processes described herein can be implemented using one or more differentiable functions, wherein the gradient of the differentiable function can be computed and / or estimated with respect to the input and / or output of the differentiable function (e.g., with respect to training data, and / or with respect to a target function). Such methods and processes can be determined at least in part by a set of trainable parameters. Therefore, the trainable parameters for a particular method or process can be tuned by any suitable training procedure to continuously improve the functionality of the method or process. As described herein, the model can be “configured” based on training, for example, by training the model using multiple instances of training data suitable for causing adjustments to the trainable parameters, thereby producing the described configuration.

[0083] Non-limiting examples of training processes used to tune trainable parameters include: supervised training (e.g., using gradient descent or any other suitable optimization method), zero-shot, few-shot, unsupervised learning methods (e.g., classification based on categories derived from unsupervised clustering methods), reinforcement learning (e.g., feedback-based deep Q-learning) and / or generative adversarial neural network training methods, belief propagation, RANSAC (random sample consensus), context robbery methods, maximum likelihood methods, and / or expectation maximization. In some examples, multiple methods, processes, and / or components of the system described herein can be trained simultaneously with respect to an objective function that measures the performance of the collective function of multiple components (e.g., with respect to reinforcement feedback and / or with respect to labeled training data). Simultaneous training of multiple methods, processes, and / or components can improve this collective function. In some examples, one or more methods, processes, and / or components can be trained independently of other components (e.g., offline training on historical data).

[0084] Language models can leverage lexical features to guide word sampling / search for speech recognition. For example, a language model can be defined at least in part by the statistical distribution of words or other lexical features. For instance, a language model can be defined by the statistical distribution of an n-gram, defining the transition probabilities between candidate words based on lexical statistics. Language models can also be based on any other suitable statistical features, and / or utilize the results of processing those statistical features using one or more machine learning and / or statistical algorithms (e.g., confidence values ​​generated by such processing). In some examples, the statistical model can constrain which words can be recognized from an audio signal, for example, based on the assumption that the words in the audio signal come from a specific vocabulary.

[0085] Alternatively or additionally, the language model may be based on one or more previously trained neural networks to represent audio input and words in a shared latent space, for example, a vector space learned by one or more audio and / or word models (e.g., wav2letter and / or word2vec). Therefore, finding candidate words may include searching the shared latent space based on vectors encoded by the audio model to find audio input, so as to find candidate word vectors for decoding using the word model. For one or more candidate words, the shared latent space may be used to evaluate the confidence level of the candidate words in the speech audio.

[0086] Language models can be used in conjunction with acoustic models configured to evaluate the confidence level of candidate words in speech audio contained within the audio signal based on word-based acoustic features (e.g., Mel-frequency cepstral coefficients, formants, etc.). Optionally, in some examples, language models can be combined with acoustic models (e.g., the evaluation and / or training of the language model can be based on the acoustic model). An acoustic model defines a mapping between acoustic signals and basic sound units such as phonemes, e.g., tagged speech audio. Acoustic models can be based on any suitable combination of existing or future ML and / or AI models, such as: deep neural networks (e.g., Long Short-Term Memory, Temporal Convolutional Neural Networks, Restricted Boltzmann Machines, Deep Belief Networks), Hidden Markov Models (HMMs), Conditional Random Fields (CRFs) and / or Markov Random Fields, Gaussian Mixture Models, and / or other graphical models (e.g., Deep Bayesian Networks). Audio signals processed using an acoustic model can be preprocessed in any suitable manner, such as encoding with any appropriate sampling rate, Fourier transform, bandpass filter, etc. The acoustic model can be trained to recognize the mapping between acoustic signals and sound units based on training using labeled audio data. For example, the acoustic model can be trained based on labeled audio data including speech audio and corrected text to learn the mapping between speech audio signals and sound units represented by corrected text. Therefore, the acoustic model can be continuously improved to enhance its effectiveness in correctly recognizing speech audio.

[0087] In some examples, in addition to statistical models, neural networks, and / or acoustic models, language models can incorporate any suitable graphical model, such as a Hidden Markov Model (HMM) or a Conditional Random Field (CRF). Given speech audio and / or other words identified so far, the graphical model can utilize statistical features (e.g., transition probabilities) and / or confidence values ​​to determine the probability of recognizing a word. Therefore, the graphical model can leverage statistical features, previously trained machine learning models, and / or acoustic models to define the transition probabilities between states represented in the graphical model.

[0088] The models(s) described herein can combine any suitable combination of AI, ML, NLP, and / or statistical models. For example, the models can include one or more models configured as a whole, and / or one or more models configured in any other suitable manner. The models(s) can be trained on any suitable data. For example, models according to this disclosure can include one or more AI, ML, NLP, and / or statistical models trained on labeled data from task-oriented dialogues between people. In some examples, the models can be trained in an "end-to-end" manner, for example, regarding the accuracy of predicting whether a utterance is the final utterance in a user turn.

[0089] When included, display subsystem 406 can be used to present a visual representation of data stored by storage subsystem 404. This visual representation may take the form of a graphical user interface (GUI). Display subsystem 406 may include one or more display devices utilizing virtually any type of technology. In some implementations, display subsystem may include one or more virtual reality, augmented reality, or mixed reality displays.

[0090] When included, the input subsystem 408 may include or interface with one or more input devices. Input devices may include sensor devices or user input devices. Examples of user input devices include keyboards, mice, touchscreens, or game controllers. In some embodiments, the input subsystem may include or interface with selected Natural User Input (NUI) components. Such components may be integrated or peripheral, and the translation and / or processing of input actions may be handled on-board or off-board. Exemplary NUI components may include one or more microphones (e.g., a directional microphone array) for speech and / or speech recognition; infrared, color, stereo, and / or depth cameras for machine vision and / or gesture recognition; and head trackers, eye trackers, accelerometers, and / or gyroscopes for motion detection and / or intent recognition.

[0091] When included, the communication subsystem 410 can be configured to communicatively couple the computing system 400 to one or more other computing devices. The communication subsystem 410 may include wired and / or wireless communication devices compatible with one or more different communication protocols. The communication subsystem can be configured to communicate via a personal area network, a local area network, and / or a wide area network.

[0092] In one example, a method for automatically describing turns in a multi-turn dialogue between a user and a session computing interface includes: receiving audio data encoded from the user's speech in the multi-turn dialogue. In this example or any other example, the method further includes: analyzing the received audio to identify a utterance followed by silence in the user's speech. In this example or any other example, the method further includes: identifying the utterance as the final utterance in a turn of the multi-turn dialogue in response to the silence exceeding a context-dependent duration dynamically updated based on the session history of the multi-turn dialogue and features of the received audio, wherein the session history includes one or more previous turns of the multi-turn dialogue performed by the user and one or more previous turns of the multi-turn dialogue performed by the session computing interface. In this example or any other example, the features of the received audio include one or more acoustic features. In this example or any other example, the one or more acoustic features include the user's intonation relative to the user's baseline speech tone. In this example or any other example, the one or more acoustic features include one or more of the following: 1) a baseline speech rate for the user, 2) the user's speech rate in the utterance, and 3) the difference between the user's speech rate in the utterance and the baseline speech rate for the user. In this example or any other example, the features of the received audio include the sentence-end probability of the final n-gram in the utterance, the sentence-end probability of the final n-gram being derived from a language model. In this example or any other example, the features of the received audio include syntactic properties of the final n-gram in the utterance. In this example or any other example, the features of the received audio include automatically identified filler words from a predefined list of filler words. In this example or any other example, the context-related duration is dynamically updated based on the conversation history of the multi-turn dialogue, at least based on semantic context derived from one or both of the following: 1) previous utterances by the user in the conversation history of the multi-turn dialogue, and 2) previous utterances by the conversation computation interface in the conversation history of the multi-turn dialogue. In this example or any other example, the features of the received audio include one or more features of the utterance, which are selectively measured for the end portion of the utterance. In this example or any other example, the context-dependent duration is dynamically updated based on the session history of the multi-turn dialogue, at least based on measurements of feature changes occurring throughout the utterance. In this example or any other example, the method further includes: determining whether a turn in the multi-turn dialogue is fully operable based on the final utterance of the turn, and generating a response utterance based on the final utterance in response to the turn not being fully operable.In this example or any other example, in addition to the user's speech, the received audio also includes the speech of one or more different users, and the method further includes: distinguishing the user's utterances from the speech of the one or more different users, and parsing the utterances from the user from the received audio, wherein the turns of the multi-turn dialogue only include the user's utterances. In this example or any other example, the utterance is one of a plurality of utterances separated by silence in the turns, the plurality of utterances being automatically provided by an automatic speech recognition system configured to automatically segment audio into individual utterances. In this example or any other example, analyzing the received audio includes operating a previously trained natural language model trained on a plurality of exemplary task-oriented dialogues, each exemplary task-oriented dialogue including one or more turns of an exemplary user. In this example or any other example, the previously trained natural language model is a domain-specific model of a plurality of domain-specific models, wherein each domain-specific module is trained on a corresponding plurality of domain-specific exemplary task-oriented dialogues.

[0093] In this example, a computing system includes a logic subsystem and a storage subsystem. In this example, or any other example, the storage system stores instructions that can be executed by the logic subsystem. In this example, or any other example, the instructions are executable to receive audio-encoded speech of a user in a multi-turn dialogue between a user and a session computing interface. In this example, or any other example, the instructions are executable to analyze the received audio to identify a utterance followed by silence in the user's speech. In this example, or any other example, the instructions are executable to identify the utterance as the final utterance in a turn of the multi-turn dialogue in response to a silence exceeding a context-dependent duration dynamically updated based on the session history of the multi-turn dialogue and characteristics of the received audio, wherein the session history includes one or more previous turns of the multi-turn dialogue performed by the user and one or more previous turns of the multi-turn dialogue performed by the session computing interface. In this example or any other example, the computing system further includes a directional microphone array, via which received audio is received, the received audio including, in addition to the user's speech, the speech of one or more different users, and the instructions are also executable to evaluate the spatial location associated with a utterance in the received audio, to distinguish the user's utterance from the speech of one or more different users based on the evaluated spatial location, and to parse the utterance from the user from the received audio, wherein the turns of the multi-turn dialogue include only the user's utterance. In this example or any other example, the computing system further includes an audio speaker, wherein the instructions are also executable to determine whether the turn of the multi-turn dialogue is fully operable based on the final utterance in the turn, and in response to the turn being not fully operable, to output a response utterance based on the final utterance via the audio speaker. In this example or any other example, the computing system further includes a microphone, wherein the utterance is one of a plurality of utterances separated by silence in the turn, the plurality of utterances being automatically provided by an automatic speech recognition system configured to receive audio via the microphone and to automatically segment the received audio into individual utterances.

[0094] In the example, automatically describing turns in a multi-turn dialogue between a user and a session computing interface includes receiving audio data encoded from the user's speech in the multi-turn dialogue. In this example or any other example, the method further includes analyzing the received audio to identify utterances followed by silence in the user's speech. In this example or any other example, the method further includes dynamically updating the context-dependent silence duration based on the session history of the multi-turn dialogue, features of the received audio, and semantic context derived from one or both of the following: 1) previous utterances of the user in the session history of the multi-turn dialogue, and 2) previous utterances of the user by the session computing interface in the session history of the multi-turn dialogue. In this example or any other example, the method further includes identifying the utterance as the final utterance in the turn of the multi-turn dialogue in response to the silence exceeding the context-dependent duration. In this example or any other example, the method further includes identifying the utterance as a non-final utterance of multiple utterances in the turn of the multi-turn dialogue in response to the silence not exceeding the context-dependent duration, and further analyzing the received audio to identify one or more additional utterances followed by one or more additional silences.

[0095] It should be understood that the configurations and / or methods described herein are exemplary in nature, and these specific embodiments or examples should not be considered limiting, as many variations are possible. The specific routines or methods described herein may represent one or more of any number of processing strategies. Thus, the various actions illustrated and / or described may be performed in the illustrated and / or described sequence, in other sequences, in parallel, or omitted. Similarly, the order of the above processes may be changed.

[0096] The subject matter of this disclosure includes all novel and non-obvious combinations and sub-combinations of various processes, systems and configurations, as well as other features, functions, actions and / or characteristics disclosed herein, and any and all their equivalents.

Claims

1. A method for automatically describing turns in a multi-turn dialogue between a user and a session computing interface, comprising: Receive audio data encoded from the user's speech during the multi-turn dialogue; Analyze the received audio to identify utterances followed by silence in the user's speech; as well as In response to the silence exceeding a context-dependent duration dynamically updated based on the conversation history of the multi-turn dialogue and features of the received audio, the utterance is identified as the final utterance in a turn of the multi-turn dialogue, wherein the conversation history includes one or more previous turns of the multi-turn dialogue performed by the user and one or more previous turns of the multi-turn dialogue performed by the conversation calculation interface.

2. The method according to claim 1, wherein, The characteristics of the received audio include one or more acoustic features.

3. The method according to claim 2, wherein, The one or more acoustic features include the user's intonation relative to the user's baseline speech pitch.

4. The method according to claim 2, wherein, The one or more acoustic features include one or more of the following: 1) a baseline speech rate for the user, 2) the user's speech rate in the utterance, and 3) the difference between the user's speech rate in the utterance and the baseline speech rate for the user.

5. The method according to claim 1, wherein, The features of the received audio include the sentence-end probability of the final n-gram in the utterance, which is derived from a language model.

6. The method according to claim 1, wherein, The features of the received audio include the syntactic properties of the final n-gram in the utterance.

7. The method according to claim 1, wherein, The characteristics of the received audio include automatically identified filler words from a predefined list of filler words.

8. The method according to claim 1, wherein, The context-related duration is dynamically updated based on the session history of the multi-turn dialogue, based at least on the semantic context derived from one or both of the following: 1) the user's previous utterances in the session history of the multi-turn dialogue, and 2) the session calculation interface's previous utterances in the session history of the multi-turn dialogue.

9. The method according to claim 1, wherein, The features of the received audio include one or more features of the utterance, selectively measured for the end portion of the utterance.

10. The method according to claim 1, wherein, The context-related duration is dynamically updated based on the conversation history of the multi-turn dialogue, at least based on measurements of feature changes occurring throughout the utterance.

11. The method according to claim 1, further comprising: Based on the final utterance in the round of the multi-turn dialogue, it is determined whether the round of the multi-turn dialogue is fully operable, and in response to the round not being fully operable, a response utterance is generated based on the final utterance.

12. The method according to claim 1, wherein, In addition to the user's speech, the received audio also includes the speech of one or more different users, and the method further includes: distinguishing the user's speech from the speech of the one or more different users, and parsing the speech from the user from the received audio, wherein the rounds of the multi-turn dialogue only include the user's speech.

13. The method according to claim 1, wherein, The utterance is one of a plurality of utterances separated by silence in the round, the plurality of utterances being automatically provided by an automatic speech recognition system configured to automatically segment audio into individual utterances.

14. The method according to claim 1, wherein, The analysis of the received audio includes operations on a previously trained natural language model trained on a number of exemplary task-oriented dialogues, each of which includes one or more turns of an exemplary user.

15. The method according to claim 14, wherein, The previously trained natural language model is a domain-specific model among multiple domain-specific models, wherein each domain-specific module is trained on multiple exemplary task-oriented dialogues corresponding to the domain.

Citation Information

Patent Citations

  • Adaptive speech endpoint detector

    US20180090127A1

  • Context-based detection of end-point of utterance

    US20190318759A1

  • Automated speech recognition using a dynamically adjustable listening timeout

    US20190348065A1