Natural speaker understanding

EP4690183A1Pending Publication Date: 2026-02-11CERENCE OPERATING CO
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
EP2024718689
Authority / Receiving Office
EP · EP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-03-30
Filing Date
2024-03-25
Publication Date
2026-02-11

AI Technical Summary

Technical Problem

Current natural language understanding systems struggle to replicate human-like communication by failing to account for contextual information beyond spoken words and have difficulty managing interactions with multiple human clients, leading to stilted and remote interactions.

Method used

The system employs real-time extraction and analysis of audio streams to identify speakers, their attributes, and environmental context, using speaker segmentation, attribute detection, and auditory scene analysis to create interaction packages that integrate timing information, enabling secure, personalized, and human-like interactions.

Benefits of technology

This approach allows for full duplex communication, precise intent inference, and effective management of multiple speakers, enhancing the naturalness and security of human-machine interactions by providing real-time information on speaker identities, languages, and environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure US2024021266_03102024_PF_FP_ABST
    Figure US2024021266_03102024_PF_FP_ABST
Patent Text Reader

Abstract

A method includes providing interaction packages for consumption by an application that engages in speech interaction with a human client in an environment. The interaction packages include a speaker event and a scene event, both of which have been tagged with timing information. The method includes continuously listening to the environment to obtain a stream of audio data, partitioning it into audio segments and using those audio segments to obtain the scene events and the speaker events for the interaction packages.
Need to check novelty before this filing date? Find Prior Art

Description

NATURAL SPEAKER UNDERSTANDINGCross-Reference to Related Applications

[0001] This application claims priority to U.S. Provisional Application No. 63 / 455,601, filed March 30, 2023, the content of which is hereby incorporated by reference in its entirety.Background

[0002] In the process of natural language understanding, an application receives an utterance from a human client. The application processes the utterance and formulates a suitable reply, which it then utters. The human client may then follow up with another utterance. The nature of the interaction is thus one in which the application and the human client take turns speaking.

[0003] As a result, the interaction between application and human client develops the flavor of communication via walkie-talkie. The interaction seems stilted and remote from the natural way that humans constantly interrupt each other.

[0004] An additional shortcoming of interactions that include an application arises from the application’s difficulty in accommodating multiple human clients. A typical application faces difficulty in keeping individual human clients separated and associating the correct intent with each human client.Summary

[0005] The invention relies in part on the observation that communication between humans relies on more than words. It is, in fact, a multi-sensory experience that relies only in part on the spoken word. The information beyond the spoken word, i.e., context information from the human client’s environment, is often used to resolve ambiguities that would otherwise arise if only the spoken word were available.

[0006] Unlike a conventional application, human beings are aware of the existence of multiple speakers. At a given instant, human being is aware of who is speaking. This results in a mental picture of the conversational environment. This awareness also provides the information necessary for taking advantage of ebbs and flows in a conversation to engage more naturally with other individuals.

[0007] The apparatus and methods disclosed herein attempt to mimic this multi- sensory communication method by supplementing the spoken word with information concerning the environment in which the words were spoken. In addition topromoting the resolution of ambiguities that are so prevalent in speech, this also tends to promote more human-like communication between humans and machines.

[0008] The invention features circuitry that carries out real-time extraction of information concerning speakers participating in a conversation. It does so based on real-time audio streams that include both speech utterances and other sounds. The circuitry combines a variety of functions in pursuit of this goal. These functions include speaker segmentation, which identifies how many participants are in a conversation, and speaker- attribute detection, which provides information about those participants. Examples of speaker- attribute detection include voice biometric methods that identify individuals based on the sounds of their respective voices, age-detection methods that estimate the age of a speaker based on qualities of that speaker’s voice, and language detection methods, to assist in the natural language understanding process.

[0009] Also included is circuitry that carries out auditory scene analysis. Such auditory scene analysis provides contextual information about the participants in the conversation.

[0010] The circuitry described and claimed herein thus differs from natural language understanding by using information above and beyond mere speech to more precisely infer the meaning intended by a speaker.

[0011] An objective of the invention is to enable digital systems and human machine interfaces to provide secure, contextual, personalized, and human-like interactions to users and to do so by providing real-time information about the identities of the speakers, what language they are speaking, when they are speaking, what age group they are in, where they are speaking, and who the primary speaker is. In doing so, the circuitry disclosed herein analyzes audio streams in real-time to provide such functionality.

[0012] In one aspect, the invention features a method that includes providing a service to an application that engages in speech interaction with a human client within an environment. The providing of this service includes causing an interaction packager to execute certain steps, among which is the step of providing a stream of interaction packages for consumption by the application, each of the interaction packages comprising a scene event and a speaker event. This step includes continuously listening to the environment to obtain a stream of audio data and partitioning the audio data into audio segments. Then, for each audio segment that includes voice activity, the method continues with extracting speaker- specific segments, each of which corresponds to speech by a speaker in the audioenvironment, generating the speaker event based at least in part on one of the speakerspecific segments, generating an audio scene prints, and after having generated the audio scene print event and the speaker event, tagging the speaker event and the audio scene print event with timing information, the audio scene print thus defining a scene event. The method then concludes with integrating the audio scene events and the speaker events according to the timing information to form the interaction packages.

[0013] In some practices, extracting speaker- specific segments comprises extracting a speaker- specific segment corresponding to synthesized speech from the application and ignoring the speaker- specific segment. As a result, the interaction packages omit speaker events arising from the synthesized speech.

[0014] In other practices, providing the stream of interaction packages occurs concurrently with continuously listening to the audio environment. This enables full duplex speech communication between the applicant and one or more human clients thereof.

[0015] In still other practices, for each of the audio segments, generating the scene print comprises using background audio present in the audio segment.

[0016] Also among the practices are those in which generating the scene event comprises adding location information and environmental information to the scene event, the location being indicative of a location at which the speaker event took place and the environmental information being indicative of an audio environment in which the speaker event took place.

[0017] The application that receives the interaction package can execute at a variety of different locations. Among the practices of the invention are those in which the application comprises an automotive assistant that is executing on an infotainment system of an automobile, those in which the application comprises a digital assistant is executing in a processing system of a kiosk, and those in which the application executes within a processing system of a robot.

[0018] In still other practices, the application comprises a natural language understanding unit.

[0019] Some practices of the method further include storing the speaker- specific segments in an episodic memory. In such practices, storing the speaker- specific segments comprises storing only speaker- specific segments arising from natural speech. This would exclude segments arising from speech by the application itself.

[0020] Also among the practices of the method are those in which generating the scene event comprises continuously monitoring the environment to receive non-audio information therefrom and using the non-audio information in the scene event.

[0021] Still other practices include identifying an audio segment that lacks voice activity and using the audio segment only for generating the scene event.

[0022] Among the other practices are those that include collection of speaker- specific attributes. In those cases in which each of the speaker-specific segments corresponds to a speaker, some practices of the method include, for each of the speaker- specific segments, extracting attribute information about the speaker. In such cases, generating the speaker event based at least in part on one of the speaker- specific segments comprises including the attribute information in the speaker event.

[0023] Still other practices in which each of the speaker- specific segments corresponds to a speaker include those in which generating the speaker event based at least in part on one of the speaker- specific segments comprises tagging the speaker event with information identifying the speaker.

[0024] Further practices also include causing the application to use the interaction packages to engage in full duplex communication with the client.

[0025] Other practices include the use of non-audio information to provide further context information. Among these are practices in which generating the scene event comprises continuously monitoring the environment to receive non-audio information therefrom and using the non-audio information to identify a location of a client during the speaker event.

[0026] In another aspect, the invention features an interaction packager that provides a stream of interaction packages for consumption by an application that is engaging in speech interaction with a human client in an environment. Each of the interaction packages include a scene event and a speaker event. The interaction packager comprises a speaker- segmentation module that separates incoming audio into speakerspecific segments, a speaker analyzer that tags the speaker- specific segments with speaker attributes, a scene analyzer that provides an audio snapshot of the environment, an episodic memory that adds timing information to the audio snapshot and the speaker- specific segments, and an integrator that receives the speaker- specific segments and the audio snapshots and constructs, therefrom, the interaction packages.

[0027] While it may appear to the uninitiated that the steps disclosed and claimed herein can be entirely performed in the human mind, this has turned out to be an illusion. Through numerous controlled and double-blind experiments attempting to doso, it was discovered that each attempt to perform the steps in the human mind was a failure. Attempts were also made to execute the steps on a generic computer. These also ended in failure. In the end, it was found to be necessary to build a specialpurpose computer that had been configured specifically to perform the steps described and claimed herein.

[0028] These and other features of the invention will be apparent from the following detailed description and its accompanying figures, in which:Description of Drawings

[0029] FIG. 1 shows the interaction packager interacting with an application that is engaging in speech interaction with a client;

[0030] FIG. 2 shows details of the interaction packager shown in FIG. 1;

[0031] FIG. 3 shows a scene event and a speaker event integrated into an interaction package; and

[0032] FIG. 4 shows a later scene event and speaker event integrated into a subsequent interaction package.Detailed Description

[0033] FIG. 1 shows an interaction packager 10 that receives a stream of audio data 12 from an audio input 14 and a stream of non-audio data 16 from a sensor input 18. The interaction packager 10 uses information from the streams of audio data 12 and non-audio data 16 to produce interaction packages 20 for consumption. In FIG. 1, the interaction packages 20 are consumed by an ASR / NLU (“automatic speech recognition and natural language understanding”) unit 22 that provides speechprocessing services to a digital assistant 24. However, in some embodiments, an ASR / NLU unit 22 is incorporated into the interaction packager 10, as a result of which the interaction packages 20 are consumed directly by the digital assistant 24.

[0034] The digital assistant 24 provides an audio signal 26 to a loudspeaker 28 to produce synthesized speech 30 for consumption by one or more human clients 32 (hereafter “clients”). As a result, the audio input 14 receives both synthesized speech 30 from the digital assistant 24 and natural speech 34 from the human clients 32. The audio input 14 thus receives a superposition of natural speech 34 and synthesized speech 30.

[0035] In addition, the audio input 14 receives background audio 36 from various noise sources 38 in the environment. However, instead of attempting to filter thisbackground audio 36 or otherwise suppress it, the interaction packager 10 uses it as a basis for carrying out audio scene analysis. The result of this audio scene analysis is incorporated into the interaction packages 20 to provide context for use by the ASR / NLU unit 22 when attempting to ascertain the intent of the human speaker 32.

[0036] Instead of attempting to suppress background audio 36, the interaction packager 10 suppresses the synthesized speech 30. This simple expedient has the surprising effect of promoting the digital assistant’s ability to process speech by multiple human users 32 at the same time even when the utterances of the human users 32 are temporally interleaved with each other and with the synthesized speech 30.

[0037] Referring now to FIG. 2, the interaction packager 10 features a stream processor 40 that carves data bytes from the streaming audio 12 and assembles those bytes into audio segments 42. The stream processor 40 then carries out certain preprocessing steps. These steps include identifying voice activity in an audio segment 42. Those audio segments 42 that contain voice activity are provided to a segmentation module 44, to a multi-speaker detector 46, and to a scene analyzer 48. Those audio segments 42 that lack voice activity are provided only to the scene analyzer 48.

[0038] For each such audio segment 42, the multi- speaker detector 46 outputs a speaker-count 50. The speaker-count 50 estimates how many speakers are speaking in that audio segment 42. The multi-speaker detector 46 then provides the speaker-count 50 to the segmentation module 44 and to the scene analyzer 48.

[0039] The segmentation module 44, which has received the same audio segment 42, uses this speaker-count 50 as a basis for dividing the content of the audio segment 42 into one or more speaker- specific segments 52. A suitable implementation of a segmentation module 44 relies on diarization. However, other methods are known for this purpose.

[0040] The segmentation module 44 sends the speaker-specific segments 52 downstream to corresponding instances of a speaker analyzer 54.

[0041] Unlike the embodiment shown in FIG. 1, the embodiment shown in FIG. 2 has its own built-in ASR / NLU unit 22. As a result, in this embodiment, the segmentation module 44 also sends speaker- specific segments 52 to the ASR / NLU 22.

[0042] Each speaker analyzer 54 analyzes its corresponding speaker- specific segment 52 to determine speaker attributes. In doing so, the speaker analyzer 54 communicates with a multi-attribute detector 56.

[0043] The attribute detector 56 features a variety of components for detecting particular attributes.

[0044] Among the attribute detector’s components is a speaker recognizer 58 that performs speaker recognition to verify an individual’s identity. A useful way to do so is to compare a speaker’s voice with a voiceprint known to belong to that speaker.

[0045] Also among the attribute detector’s components is a language detector 60, an age detector 62, and an emotion detector 63 that carries out sentiment analysis.

[0046] In an attempt to fool the interaction packager 10 into believing that a particular speaker is present, it is possible to play a mere recording of a speaker’s voice to the interaction packager 10. To provide some protection against this possibility, the attribute detector 46 includes a spoof detector 64.

[0047] An important function of the speaker analyzer 54 is that of identifying the synthesized speech 30 from the digital assistant 24. By tagging a speaker- specific segment 52 for a digital assistant’s voice, the speaker analyzer 54 enables the remainder of the interaction packager to ignore the digital assistant’s voice. Doing so promotes the ability to listen continuously to the environment. This, in turn, makes it possible to carry out full duplex communication with the digital assistant. As a result, it becomes unnecessary to wait for a digital assistant 24 to finish speaking.

[0048] In the meantime, the scene analyzer 48, which has also received the audio segment 42, extracts the background audio 36 from the audio segment 42. It uses this background audio 36 while carrying out auditory scene analysis. The scene analyzer 48 uses background audio 36 for, among other things, inferring locations of clients 32 and identifying other audio elements present in the background. Examples of inferring locations include inferring whether the clients 32 are in a vehicle and if, so, which one. Examples of other audio elements that may be in the background include music, construction noises, the hum of an air-conditioner, furnace, or other appliance, the purring of a cat, or the din of a cocktail party or other social gathering.

[0049] The scene analyzer 48 uses the foregoing information to create an audio- scene print 66 that identifies a location based on the acoustic properties and surrounding audio environment. The audio- scene print 66 is useful for determining a client’s environment, such as whether the client 32 is in a vehicle or at home or whether the client 32 is in a known or unknown environment.

[0050] The process by which the scene analyzer 48 identifies a location includes the collection of numerous sound samples, each of which is then labelled as belonging to one of a set of sonic environments.

[0051] Given a set of sound samples that all belong to a sonic environment, it becomes possible to define a “feature vector” of representative sonic features for that sonic environment. These form a set of representative feature vectors.

[0052] By defining a suitable metric, it also becomes possible to define a metric that expresses a similarity between two feature vectors. A particularly suitable metric is the cosine distance, which can be regarded as an angle between two feature vectors.

[0053] Upon receiving an audio segment 42, the scene analyzer 48 forms a corresponding feature vector for that audio segment 42. It then determines the value of the metric between that feature vector and each of the representative feature vectors. The representative feature vector for which that value is minimized is then inferred as representing the sonic environment from which the audio segment 42 arose. This provides a basis for determining the audio-scene print 66.

[0054] The foregoing process requires collecting the relevant sound samples and carrying out the classification of those sound samples. This is typically carried out by training a deep neural network.

[0055] Similar techniques are used for other components of FIG. 2 that classify an input as belonging to one of two or more classes.

[0056] The audio-scene print 66 thus serves as a snapshot of the background in which communication takes place. Knowledge of the background conditions is useful, for example, for determining the manner of communication. As an example, a digital assistant 34 that receives an audio- scene print 66 reporting a background dominated by a low frequency hum could modulate its voice to a higher range of frequencies to avoid interference. Alternatively, an audio-scene print 66 that indicates excessive broadband noise would prompt a digital assistant 24 to communicate visually, for example by displaying text.

[0057] Both the scene analyzer 48 and the speaker analyzer 54 provide information to an episodic memory 68.

[0058] The episodic memory 50 is what allows the interaction packager 10 to mimic the human ability to remember who said what and when during the course of a conversation. The episodic memory 68 tags each speaker- specific segment 52 with a time of occurrence and with speaker attributes collected by the speaker analyzer 54. The episodic memory 68 outputs a script 70. The script 70 provides information on what each client 32 said and when the client 32 said it. In addition, the script 70 includes information about each client 32.

[0059] This script 70, however, must still be brought to life with information about the environment in which it arose. This information is provided by the scene analyzer 48 and a sensor integrator 72 that receives the non-audio information 16 from the sensor input 18.

[0060] Examples of non-audio information 16 include information from a camera, information from weight sensors in seats, and information concerning locations of microphones. Such non-audio information 16 is useful for determining where a particular client 32 is located. For example, in a vehicle with weight sensors at each seat, it is possible to infer whether a client 32 is a passenger or driver. The sensor integrator 72 packages this non-audio information 16 into a non-audio scene print 74 that complements the audio- scene print 66 provided by the scene analyzer 48.

[0061] The script 70, the audio-scene print 66, and the non-audio scene print 52 are provided to an integrator 76.

[0062] In a broad sense, the integrator 76 functions as an electronic “dramaturge.” In the theater, a dramaturge’s role involves adding context to a script to bring it to life. Such context includes subject matter analogous to what the scene analyzer 48 and the sensor integrator 72 provide. It is the action of the integrator 76 that lifts the output of the interaction packager 10 into something more than merely the dry semantic meaning of the words in the audio segments 42.

[0063] The integrator 76 consolidates its inputs in a machine-readable way and provides interaction packages 20 for consumption by an external application, such as a digital assistant 24 in FIG. 2 or an ASR / NEU unit 22 as shown in FIG. 1. Embodiments include those in which the application includes a digital assistant 24 embedded in an infotainment system, in a processing system within a kiosk, or in a robot.

[0064] In a typical embodiment, the intelligence provided by the integrator 76 comprises consumable events that have been further tagged with semantic labels and attributes based on the inputs to the integrator 76. The digital assistant 24 uses this intelligence to provide personalized, contextual, and secure voice interactions.

[0065] A particular advantage of the interaction packager 10 is that it listens to its environment continuously. In contrast, a typical digital assistant 24 cannot listen while it is speaking. As a result, a client’s experience interacting with a digital assistant 24 is much like using a walkie-talkie. The communication is only half duplex.

[0066] The output of the interaction packager 10 provides a way for the digital assistant 24 to mimic the act of listening continuously. As a result, a client who interacts with a digital assistant 24 that uses the output of the interaction packager 10 would enjoy the benefit of full duplex communication. This more closely approximates natural speech between people, in which different people interrupt each other and attempt to talk over each other with ever increasing vigor in an exercise that often verges on the edge of cacophony.

[0067] The interaction packager 10 achieves this by listening continuously, even while the digital assistant 24 is itself speaking. As a result of its speaker-recognition ability, the interaction packager 10 is able to recognize and ignore the digital assistant’s synthesized voice and to continuously update its intelligence on the natural speech 34 in real time. This permits the interaction packager 10 to alert the digital assistant 24 to a change in speaker or to a speaker who turns out to be an impostor. It also permits the interaction packager 10 to provide real-time information on a speaker’s identity, preferred language, age, as well as any changes in attributes of the auditory scene or the non- auditory scene.

[0068] The interaction packager 10 described herein provides a way for a digital assistant 24 to determine that more than one client 32 is attempting to interact with the digital assistant 24 and to also receive information about the identity of each client 32 and when these client 32 speak, even when they speak at the same time. Using this information, the digital assistant 24 receives expressions of intent from multiple clients 32 and matches each expression of intent with the correct client 32. This in turn can be used to personalize the interaction between the digital assistant 24 and each of the clients 32 and to thereby more closely mimic human interaction.

[0069] As an example, consider the case of a first client who asks the digital assistant 24 to arrange for food delivery and a second client who then instructs the digital assistant 24 to bill the food to a particular account, all in the same voice interaction.

[0070] A digital assistant 24 that continuously receives packaged events from the interaction packager 10 is able to handle this complexity in a single voice transaction without the need for the two clients 32 to formally take turns speaking to the digital assistant 24 and without the need to provide different user profiles to the digital assistant 24. This results in a fluid multi-speaker interaction that is barely distinguishable from normal human interaction but with the added bonus of built-in security through speaker recognition.

[0071] The ability to provide speaker segmentation, speaker recognition and identification, and to store chronologically arranged events in the episodic memory 68result in a synergy that allows the interaction packager 10 to support continuous interaction with active listening. This, in turn, results in the ability to provide spoken feedback in a natural way without the need to formally take turns speaking to different people.

[0072] This combination of features, when combined with the ability to identify and ignore synthesized speech 30, also has the unexpected result of making it possible to interleave natural speech 34 with synthetic speech 30. As a result, a digital assistant 24 that relies on the interaction packager 10 is able to, in effect, listen constantly. A digital assistant 24 so enabled is therefore able to interject remarks and to transition into and out of a dialog smoothly without delays. This, in turn, results in human-like natural conversational flow, thereby rendering obsolete the current requirement of rigid tum-by-turn interactions with digital assistants 24.

[0073] The invention’s episodic memory 50 enables the event packager 10 to remember what was said by different clients 32 and also to associate this information with the attributes of the clients 32. This is significantly different from known systems, in which having more than one client 32 speaking at a time tends to confuse the digital assistant 24.

[0074] Known digital assistants can listen to one client 12, process that client’s utterances through an ASR / NLU module 22 to extract intent, process that client’s utterances to identify the client, and then associate the client’s intent with the client’s identity. However, if conversation proceeds in a natural way with multiple clients speaking in such a way that their utterances are interleaved in time, the digital assistant 24 becomes disoriented and confused.

[0075] An interaction packager 10 as described herein provides services needed by a digital assistant 22 for preserving context of multiple clients and the auditory scene during a voice interaction session without requiring individual sessions with each client. This ability arises as a result of the interaction of the components shown in FIG. 2 and the ability to associate multiple intents with corresponding clients and doing so while preserving a chronological sequence of interactions. The resulting event packager 10 thus results in a natural speaker understanding tool rather than merely a natural language understanding tool.

[0076] In another embodiment, the digital assistant 24 provides a voice interface for a kiosk. Such kiosks are often found in public places to provide traveler assistance. In such cases, a user utters a question in a preferred language. The digital assistant 24 recognizes the language and responds in the appropriate language. This is carried out without requiring the user to make an on-screen selection.

[0077] A difficulty that arises with some voice interfaces is that of continuing to use an audio interface even if the audio environment makes it pointless to do so. An interaction packager 10 as described herein, with its sensitivity to the audio environment enables the digital assistant 24 to respond in a different mode. For example, if the scene analyzer 48 detects a particularly high level of background noise 38, the digital assistant 24 uses this information to transition into using a graphical user interface instead of a loudspeaker 28 for interacting with the client 32.

[0078] The interaction packager 10 also provides enough intelligence for the digital assistant 24 to explain to a client 32 why voice interaction cannot be carried out, thereby preventing an often-frustrating user experience. The output of the interaction packager 10 provides a digital assistant 24 with the basis for explaining an interruption in service. Examples of such reasons that can be provided include an excessive number of speakers, excessively loud background audio 36, or an excessively weak audio signal, resulting, for example, from a client 32 being too far from a microphone when speaking. Based on information provided by the scene analyzer 48 or the sensor integrator 52, the interaction packager 10 is able to provide hints based on context, i.e., “contextual hinting.” In the foregoing examples, this might include asking the client 32 to come closer to the microphone or asking the client 32 to take steps to reduce background audio 36, e.g., by turning the music down or requesting that people speak more quietly.

[0079] FIG. 3 shows examples of interaction events provided to the integrator 76. The events include a scene event 78 and a speaker event 80.

[0080] The scene event 78 includes location information 82 that indicates where an interaction took place, environmental information 84 that indicates the state of the acoustic environment, and timing information 86 for identifying the scene event 78 within a sequence of events and by a time interval during which it took place within the sequence of events. The speaker event 80 includes speaker information 88 and similar timing information 86. The integrator 76 combines the scene event 78 and the speaker event 80 to produce an event package 20 that is provided to the digital assistant 24.

[0081] The location information 82 and the environmental information 84 in the scene event 78 are the result of the scene analyzer’s having processed audio segments 42 provided by the stream processor 40. The speaker information 88 in the speaker event 80 results from information provided by the speaker analyzer 54. The sequencing information 86 again results from a tag applied at the episodic memory 68.

[0082] Interaction packages similar to that depicted in FIG. 3 are sent continuously. In particular, FIG. 4 shows the continuation of an apparent conversation that is taking place in the same vehicle. This can be seen in the timing information 86, which indicates that these events took place immediately after the events shown in FIG. 3.

[0083] As is apparent from the environmental information 84, Alex’s jeep has since come to a stop (i.e., “stationary”: true) and the music has been turned off. In addition, a new speaker has emerged. Apparently, not much is known about who “Alex” is speaking with. The speaker information 88 in FIG. 4 only indicates that the speaker is also speaking English, and that the speaker is a real person. Nevertheless, this information may develop in subsequent speaker events 80 as more samples of that speaker’s speech arrive at the speaker analyzer 54.

[0084] Having described the invention and a preferred embodiment thereof, what is claimed as new and secured by letters patent is:

Claims

What is claimed is:

1. A method comprising providing a service to an application that engages in speech interaction with a human client within an environment, wherein providing said service to said application comprises causing an interaction packager to execute steps of: providing a stream of interaction packages for consumption by said application, each of said interaction packages comprising a scene event and a speaker event, wherein providing said stream of interaction packages comprises continuously listening to said environment to obtain a stream of audio data, partitioning said audio data into audio segments, for each audio segment that includes voice activity, extracting speaker-specific segments, each of which corresponds to speech by a speaker in said audio environment, generating said speaker event based at least in part on one of said speaker- specific segments, generating an audio scene print, and after having generated said audio scene print and said speaker event, tagging said speaker event and said audio scene print with timing information, wherein said audio scene print, when tagged, becomes a scene event, and integrating said scene events and said speaker events according to said timing information to form said interaction packages.

2. The method of claim 1, wherein extracting speaker-specific segments comprises extracting a speaker-specific segment corresponding to synthesized speech from said application and ignoring said speaker- specific segment, whereby said interaction packages omit speaker events arising from said synthesized speech.

3. The method of claim 1, wherein providing said stream of interaction packages occurs concurrently with continuously listening to said audio environment, thereby enabling full duplex speech communication between said applicant and one or more human clients thereof.

4. The method of claim 1, wherein, for each of said audio segments, generating said audio scene print comprises using background audio present in said audio segment.

5. The method of claim 1, wherein generating said scene event comprises adding location information and environmental information to said scene event, said location being indicative of a location at which said speaker event took place and said environmental information being indicative of an audio environment in which said speaker event took place.

6. The method of claim 1, wherein said application comprises an automotive assistant that is executing on an infotainment system of an automobile.

7. The method of claim 1, wherein said application comprises a digital assistant that is executing in a processing system of a kiosk.

8. The method of claim 1, wherein said application executes within a processing system of a robot.

9. The method of claim 1, wherein said application comprises a natural language understanding unit.

10. The method of claim 1, further comprising storing said speaker- specific segments in an episodic memory, wherein storing said speaker- specific segments comprises storing only speaker- specific segments arising from natural speech.

11. The method of claim 1, wherein generating said scene event comprises continuously monitoring said environment to receive non-audio information therefrom and using said non-audio information in said scene event.

12. The method of claim 1, further comprising identifying an audio segment that lacks voice activity and using said audio segment only for generating said scene event.

13. The method of claim 1, wherein each of said speaker- specific segments corresponds to a speaker, said method further comprising, for each of said speaker- specific segments, extracting attribute information about said speaker and wherein generating said speaker event based at least in part on one of said speaker- specific segments comprises including said attribute information in said speaker event.

14. The method of claim 1, wherein each of said speaker- specific segments corresponds to a speaker and wherein generating said speaker event based at least in part on one of said speaker- specific segments comprises tagging said speaker event with information identifying said speaker.

15. The method of claim 1, further comprising causing said application to use said interaction packages to engage in full duplex communication with said client.

16. The method of claim 1, wherein generating said scene event comprises continuously monitoring said environment to receive non-audio information therefrom and using said non-audio information to identify a location of a client during said speaker event.

17. An apparatus comprising an interaction packager that provides a stream of interaction packages for consumption by an application that is engaging in speech interaction with a human client in an environment, each of said interaction packages comprising a scene event and a speaker event, wherein said interaction packager comprises a speaker-segmentation module that separates incoming audio into speaker- specific segments, a speaker analyzer that tags said speaker- specific segments with speaker attributes, a scene analyzer that provides an audio snapshot of said environment, an episodic memory that adds timing information to said audio snapshot and said speakerspecific segments, and an integrator that receives said speaker- specific segments and said audio snapshots and constructs, therefrom, said interaction packages.