Natural speaker understanding

By extracting voice and environmental information in real time, interactive packages are generated to resolve ambiguities in the interaction between the application and multiple human clients, achieving full-duplex voice communication and natural interaction, thus enhancing the interactive capabilities of the digital assistant.

CN121039733APending Publication Date: 2025-11-28CERENCE OPERATING CO
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202480023403.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2023-03-30
Filing Date
2024-03-25
Publication Date
2025-11-28

AI Technical Summary

Technical Problem

Existing applications struggle to adapt to the intentions of different clients when interacting with multiple human clients, leading to interaction difficulties and ambiguity, and making it difficult to simulate the multi-sensory experience of natural human communication.

Method used

By extracting speech and environmental information in real time, and utilizing speaker segmentation, attribute detection, and auditory scene analysis, interactive packages are generated to provide contextual and personalized voice interaction. This includes speech segmentation, speaker recognition, environmental analysis, and integration of non-audio information, simulating multi-sensory communication.

Benefits of technology

It enables full-duplex voice communication with multiple human clients, improves the natural interaction capabilities of digital assistants, reduces ambiguity, provides a personalized and context-sensitive interactive experience, and enhances security and naturalness.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121039733A_ABST
    Figure CN121039733A_ABST
Patent Text Reader

Abstract

A method includes providing an interaction package for use by an application participating in voice interaction with a human client in an environment. The interaction packet includes speaker events and scene events, both of which have been marked with timing information. The method includes continuously listening to an environment to obtain a stream of audio data, segmenting it into audio segments, and using these audio segments to obtain scene events and speaker events of an interaction packet.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Cross-references to related applications

[0002] This application claims priority to U.S. Provisional Application No. 63 / 455,601, filed on March 30, 2023, the contents of which are incorporated herein by reference in their entirety. Background Technology

[0003] In natural language understanding, the application receives utterances from a human client. The application processes the utterance and formulates an appropriate response, which it then sends out. The human client can then follow up with another utterance. Therefore, the nature of the interaction is that the application and the human client take turns speaking.

[0004] As a result, the interaction between the application and the human client produced a flavor of walkie-talkie communication. This interaction appears static and far removed from the natural way in which humans constantly interfere with each other.

[0005] An additional drawback of application interaction arises from the difficulty of adapting applications to multiple human clients. Typical applications face the challenge of keeping individual human clients separate and associating the correct intent with each human client. Summary of the Invention

[0006] This invention is partly based on the discovery that communication between humans does not solely rely on language. In fact, it is a multisensory experience that depends only partially on spoken language. Information beyond spoken language, namely contextual information from the human client's environment, is often used to eliminate ambiguity that may arise when relying solely on spoken language.

[0007] Unlike traditional applications, people are aware of the presence of multiple speakers. At any given moment, humans know who is speaking. This creates a mental image of the conversational environment. This perception also provides the necessary information for the system to take advantage of the rhythmic fluctuations in the dialogue, thus enabling more natural communication with others.

[0008] The apparatus and method disclosed in this paper attempt to simulate this multi-sensory communication method by supplementing speech content with relevant information from the speech environment. Besides helping to resolve the ambiguity prevalent in speech, this also contributes to promoting communication between humans and machines that is closer to natural human communication.

[0009] The invention features circuitry that performs real-time extraction of information about speakers participating in a conversation. It does so based on real-time audio streams that include speech utterances and other sounds. The circuitry combines various functions to achieve this goal. These functions include speaker segmentation, which identifies how many participants are in a conversation, and speaker attribute detection, which provides information about those participants. Examples of speaker attribute detection include voice biometrics methods that identify individuals based on their respective voices, age detection methods that estimate the age of a speaker based on the quality of their voice, and language detection methods to assist natural language understanding processes.

[0010] Also included is circuitry that performs auditory scene analysis. Such auditory scene analysis provides contextual information about participants in a conversation.

[0011] Thus, the circuitry described and claimed herein differs from natural language understanding in that it utilizes information beyond mere speech to more accurately infer the intent of a speaker.

[0012] It is an object of the invention to enable digital systems and human-machine interfaces to provide safe, contextual, personalized, and human-like interaction to users and to do so by providing real-time information about who the speaker is, what language they are speaking, when they are speaking, what age group they are in, where they are speaking, and who the primary speaker is. In doing so, the circuitry disclosed herein analyzes audio streams in real-time to provide such functionality.

[0013] In one aspect, the invention features a method that includes providing a service to an application that is participating in a voice interaction with a human client in an environment. Providing the service includes causing an interaction packager to perform certain steps, including inter alia, steps to provide a stream of interaction packages for use by the application, each interaction package including a scene event and a speaker event. The steps include continuously listening to the environment to obtain a stream of audio data and segmenting the audio data into audio segments. Then, for each audio segment that includes speech activity, the method proceeds to extract speaker-specific segments, each speaker-specific segment corresponding to a speech of a speaker in the audio environment, generate a speaker event based at least in part on one of the speaker-specific segments, generate an audio scene print, and after the audio scene print event and the speaker event have been generated, tag the speaker event and the audio scene print event with timing information, the audio scene print thereby defining the scene event. The method then ends with integrating the audio scene event and the speaker event to form an interaction package according to the timing information.

[0014] In some practices, extracting the speaker-specific segments includes extracting the speaker-specific segments corresponding to the synthesized speech from the application and ignoring the speaker-specific segments. As a result, the interaction package ignores the speaker events caused by the synthesized speech.

[0015] In other practices, the stream of interaction packages occurs concurrently with the continuous listening audio environment. This enables full duplex speech communication between the application and its human client(s).

[0016] In other practices, for each audio segment, generating the scene imprint includes using background audio present in the audio segment.

[0017] There are also practices of the method in which generating the scene event includes adding location information and environment information to the scene event, the location information indicating a location where the speaker event occurred, and the environment information indicating an audio environment where the speaker event occurred.

[0018] The application that receives the interaction package can be executing in a variety of different locations. In practices of the invention, there are those practices in which the application comprises a car assistant executing on an infotainment system of a car, those practices in which the application comprises a digital assistant executing in a processing system of a kiosk, and those practices in which the application executes within a processing system of a robot.

[0019] In other practices, the application includes a natural language understanding unit.

[0020] Some practices of the method also include storing the speaker-specific segments in a scenario memory. In such practices, storing the speaker-specific segments includes storing only speaker-specific segments that originate from natural speech. This would exclude segments generated by the application itself from speech.

[0021] There are also practices of the method in which generating the scene event includes continuously monitoring the environment to receive non-audio information from the environment, and using the non-audio information in the scene event.

[0022] There are also other practices that include identifying audio segments that lack speech activity, and using only that audio segment to generate the scene event.

[0023] In other practices, a set of speaker-specific attributes is included. In cases where each speaker-specific segment corresponds to a speaker, some practices of the method include, for each speaker-specific segment, extracting attribute information about the speaker. In such cases, generating the speaker event based at least in part on one of the speaker-specific segments includes including the attribute information in the speaker event.

[0024] Further practices include those in which generating a speaker event based at least in part on one of the speaker-specific segments includes tagging the speaker event with information that identifies the speaker.

[0025] Further practices also include causing the application to use the interaction package to engage in full-duplex communication with the client.

[0026] Other practices include using non-audio information to provide further contextual information. In these practices, in which generating a scene event includes continuously monitoring the environment to receive non-audio information therefrom, and using the non-audio information to identify the location of the client during the speaker event.

[0027] In another aspect, the invention features an interaction packager that provides a stream of interaction packages for use by an application that is engaging in voice interaction with a human client in an environment. Each interaction package includes a scene event and a speaker event. The interaction packager includes a speaker segmentation module that divides incoming audio into speaker-specific segments, a speaker analyzer that tags the speaker-specific segments with speaker attributes, a scene analyzer that provides audio snapshots of the environment, a context store that adds timing information to the audio snapshots and the speaker-specific segments, and an integrator that receives the speaker-specific segments and the audio snapshots and constructs interaction packages therefrom.

[0028] Although the steps disclosed and claimed herein might appear to an outsider to be capable of being performed entirely in a human brain, this has proven to be an illusion. Controlled double-blind experiments in which attempts were made to perform the steps in a human brain have failed each time. Attempts to perform the steps on a general-purpose computer have also failed. It has finally been discovered that a special-purpose computer must be constructed that is specially configured to perform the steps described and claimed herein. BRIEF DESCRIPTION OF DRAWINGS

[0029] These and other features of the invention will be apparent from the following detailed description, and from the drawings, in which:

[0030] Figure 1 An interaction packager is shown that is interacting with an application that is engaging in voice interaction with a client;

[0031] Figure 2 An interaction packager is shown that is interacting with an application that is engaging in voice interaction with a client; Figure 1 Details of the interaction packager are shown;

[0032] Figure 3 Scene events and speaker events that are integrated into an interaction package are shown; and

[0033] Figure 4Later scene events and speaker events that are integrated into the subsequent interaction package are shown. DETAILED DESCRIPTION

[0034] Figure 1 An interaction packager 10 is shown that receives a stream 12 of audio data from an audio input 14 and a stream 16 of non-audio data from a sensor input 18. The interaction packager 10 uses information from the stream 12 of audio data and the stream 16 of non-audio data to generate an interaction package 20 for use. In Figure 1 The interaction package 20 is used by an ASR / NLU ("automatic speech recognition and natural language understanding") unit 22 that provides speech processing services to a digital assistant 24. However, in some embodiments, the ASR / NLU unit 22 is incorporated into the interaction packager 10, with the result that the interaction package 20 is used directly by the digital assistant 24.

[0035] The digital assistant 24 provides an audio signal 26 to a speaker 28 to generate synthesized speech 30 for use by one or more human clients 32 (hereinafter "clients"). As a result, the audio input 14 receives the synthesized speech 30 from the digital assistant 24 and natural speech 34 from the human clients 32. The audio input 14 thus receives an overlay of natural speech 34 and synthesized speech 30.

[0036] In addition, the audio input 14 receives background audio 36 from various noise sources 38 in the environment. However, rather than attempting to filter or otherwise suppress this background audio 36, the interaction packager 10 uses it as a basis for performing an audio scene analysis. The results of this audio scene analysis are incorporated into the interaction package 20 to provide context for use by the ASR / NLU unit 22 in attempting to discern the intent of the speaker 32.

[0037] The interaction packager 10 suppresses the synthesized speech 30 rather than attempting to suppress the background audio 36. This simple expedient has the surprising effect of enhancing the ability of the digital assistant to simultaneously process the speech of multiple human users 32 even when the utterances of the human users 32 are interwoven in time with one another and with the synthesized speech 30.

[0038] Referring now to Figure 2 The interaction packager 10 is characterized by a stream processor 40 that partitions data bytes from the stream audio 12 and combines these bytes into audio segments 42. The stream processor 40 then performs certain pre-processing steps. These steps include identifying speech activity in the audio segments 42. Those audio segments 42 that contain speech activity are provided to a segmentation module 44, a multiple speaker detector 46, and a scene analyzer 48. Those audio segments 42 that do not contain speech activity are provided only to the scene analyzer 48.

[0039] For each such audio segment 42, the multi-speaker detector 46 outputs a speaker count 50. The speaker count 50 estimates how many speakers are speaking in that audio segment 42. The multi-speaker detector 46 then provides the speaker count 50 to the segmentation module 44 and the scene analyzer 48.

[0040] The segmentation module 44, having received the same audio segment 42, uses the speaker count 50 as a basis to divide the content of the audio segment 42 into one or more speaker-specific segments 52. A suitable implementation of the segmentation module 44 relies on voiceprint clustering techniques (diarization). However, other methods known in the art can be used to achieve this purpose.

[0041] Segmentation module 44 sends speaker-specific segments 52 downstream to the corresponding instance of speaker analyzer 54.

[0042] and Figure 1 The embodiments shown are different. Figure 2 The illustrated embodiment has its own built-in ASR / NLU unit 22. As a result, in this embodiment, the segmentation module 44 also sends speaker-specific segments 52 to the ASR / NLU 22.

[0043] Each speaker analyzer 54 analyzes its corresponding speaker-specific segment 52 to determine speaker attributes. In doing so, the speaker analyzer 54 communicates with the multi-attribute detector 56.

[0044] The attribute detector 56 is characterized by various components used to detect specific attributes.

[0045] The attribute detector components include, in particular, a speaker recognizer 58, which performs speaker identification to verify an individual's identity. One useful way to do this is to compare the speaker's voice with a voiceprint known to belong to that speaker.

[0046] The components of the attribute detector also include, in particular, a language detector 60, an age detector 62, and a sentiment detector 63 that performs sentiment analysis.

[0047] In an attempt to deceive the interactive packager 10 into believing that a specific speaker exists, the interactive packager 10 may simply play a recording of the speaker's voice. To provide some protection against this possibility, the attribute detector 46 includes a deception detector 64.

[0048] A key function of the speaker analyzer 54 is to identify the synthesized speech 30 from the digital assistant 24. By tagging speaker-specific segments 52 of the digital assistant's speech, the speaker analyzer 54 enables the remainder of the interactive packager to ignore the digital assistant's speech. This improves the ability to continuously monitor the environment. As a result, full-duplex communication with the digital assistant is possible. Consequently, it is not necessary to wait for the digital assistant 24 to finish speaking.

[0049] Simultaneously, the scene analyzer 48, which has also received audio segment 42, extracts background audio 36 from audio segment 42. It uses the background audio 36 while performing auditory scene analysis. The scene analyzer 48 uses the background audio 36 to infer, in particular, the location of the client 32 and to identify other audio elements present in the background. Examples of inferring location include inferring whether the client 32 is in a vehicle, and if so, inferring that it is. Examples of other audio elements that may be in the background include music, construction noise, the sound of air conditioning / heating equipment / other appliances running, a cat purring, or the noise of a cocktail party or other social gathering.

[0050] Scene analyzer 48 uses the aforementioned information to create an audio scene imprint 66, which identifies location based on acoustic properties and the surrounding audio environment. The audio scene imprint 66 is useful for determining the customer's environment, such as whether the customer 32 is in a vehicle or at home, or whether the customer 32 is in a known or unknown environment.

[0051] The process of scene analyzer 48 identifying location involves collecting multiple sound samples and then labeling each sample as belonging to a set of sound environments.

[0052] Given a set of sound samples that all belong to a sound environment, it becomes possible to define "feature vectors" that represent the sound features of that sound environment. These form a set of representative feature vectors.

[0053] By defining an appropriate metric, a measure can also be defined to represent the similarity between two feature vectors. A particularly suitable metric is the cosine distance, which can be considered as the angle between two feature vectors.

[0054] Upon receiving audio segment 42, the scene analyzer 48 forms a corresponding feature vector for that audio segment 42, and then determines the value of a metric between that feature vector and each representative feature vector. The representative feature vector whose metric is minimized is then inferred to represent the acoustic environment from which the audio segment 42 originates. This provides the basis for determining the audio scene imprint 66.

[0055] The process described above requires collecting relevant sound samples and classifying these samples. This is typically done by training a deep neural network.

[0056] Similar technologies are used Figure 2 The input is categorized into other components that belong to one of two or more categories.

[0057] The audio scene imprint 66 is thus used as a snapshot of the background in which communication occurs. Knowledge of the background conditions is useful, for example, for determining the mode of communication. As an example, a digital assistant 34 receiving an audio scene imprint 66 reporting a background dominated by low-frequency hum can modulate its speech to a higher frequency range to avoid interference. Alternatively, an audio scene imprint 66 indicating excessive broadband noise will prompt the digital assistant 24 to engage in visual communication, for example, by displaying text.

[0058] Both the scene analyzer 48 and the speaker analyzer 54 provide information to the scene memory 68.

[0059] Context memory 50 allows the interactive packager 10 to mimic the human ability to remember who said what and when they said it during a conversation. Context memory 68 tags each speaker-specific segment 52 with the occurrence time and speaker attributes collected by the speaker analyzer 54. Context memory 68 outputs a script 70. Script 70 provides information about what each client 32 said and when they said it. Additionally, script 70 includes information about each client 32.

[0060] However, the script 70 must still be brought into the context of its occurrence. This information is provided by the scene analyzer 48 and the sensor integrator 72, which receives non-audio information 16 from the sensor input 18.

[0061] Examples of non-audio information 16 include information from the camera, information from the weight sensor in the seat, and information about the microphone's location. This non-audio information 16 is useful for determining where a particular client 32 is located. For example, in a vehicle with a weight sensor at each seat, it can be inferred whether the client 32 is a passenger or a driver. The sensor integrator 72 packages this non-audio information 16 into a non-audio scene imprint 74 that complements the audio scene imprint 66 provided by the scene analyzer 48.

[0062] Script 70, audio scene imprint 66, and non-audio scene imprint 52 are provided to integrator 76.

[0063] In a broader sense, integrator 76 acts as an electronic "drama." In a theater, the role of drama involves adding context to the script to bring it to life. Such context includes topics similar to those provided by scene analyzer 48 and sensor integrator 72. It is the role of integrator 76 that makes the output of interactive packager 10 no longer merely the dry semantics of words in audio clip 42.

[0064] Integrator 76 merges its inputs in a machine-readable manner and provides interaction package 20 for use by external applications, such as... Figure 2 Digital Assistant 24 or Figure 1 The ASR / NLU unit 22 shown is illustrated.

[0065] In a typical embodiment, the intelligence provided by integrator 76 includes consumable events, which have been further tagged with semantic labels and attributes based on input to integrator 76. Digital assistant 24 uses this intelligence to provide personalized, context-sensitive, and secure voice interaction. Embodiments include digital assistant 24 whose application is embedded in an infotainment system, a digital assistant built into a self-service kiosk processing system, or a digital assistant integrated into a robot.

[0066] A key advantage of the interactive packager 10 is its continuous listening to its environment. In contrast, a typical digital assistant 24 cannot listen while speaking. As a result, the customer experience of interacting with the digital assistant 24 is more like using a walkie-talkie. Communication is only half-duplex.

[0067] The output of the interactive packer 10 provides the digital assistant 24 with a way to simulate continuous listening actions. As a result, clients interacting with the digital assistant 24 using the output of the interactive packer 10 will enjoy the benefits of full-duplex communication. This is closer to natural speech between people, where different people interrupt each other and try to talk to each other with increasing effort in an exercise that often borders on dissonance.

[0068] The interactive packager 10 achieves this by continuously listening, even while the digital assistant 24 is speaking. As a result of its speaker recognition capabilities, the interactive packager 10 is able to identify and ignore the digital assistant's synthesized speech and is able to continuously update its intelligence regarding the natural speech 34 in real time. This allows the interactive packager 10 to warn the digital assistant 24 of speaker changes or to warn of speakers that have been proven to be imposters. It also allows the interactive packager 10 to provide real-time information about any changes in speaker identity, preferred language, age, and attributes of the auditory or non-auditory context.

[0069] The interaction packager 10 described herein provides a way for the digital assistant 24 to determine that more than one client 32 is attempting to interact with the digital assistant 24, and also receives information about the identity of each client 32 and when these clients 32 are speaking, even if they are speaking simultaneously. Using this information, the digital assistant 24 receives expressions of intent from the multiple clients 32 and matches each expression of intent with the correct client 32. This can then be used to personalize the interaction between the digital assistant 24 and each client 32, thereby more closely mimicking human interaction.

[0070] As an example, consider the scenario of a first client and a second client. The first client requests the digital assistant 24 to arrange food delivery, and the second client then instructs the digital assistant 24 to bill the food to a specific account, all within the same voice interaction.

[0071] The digital assistant 24, which continuously receives packaged events from the interaction packager 10, can handle this complexity in a single voice transaction without requiring two clients 32 to take turns speaking to the digital assistant 24, and without providing the digital assistant 24 with different user profiles. The resulting smooth multi-speaker interaction is almost indistinguishable from normal human interaction, and also has the built-in security advantages achieved through speaker recognition technology.

[0072] The ability to provide speaker segmentation, speaker identification and identification, and to store chronologically ordered events in the context memory 68 works synergistically, allowing the interactive packager 10 to support continuous interaction with active listening. This, in turn, results in the ability to provide verbal feedback in a natural manner without requiring a formal turn-taking of different people.

[0073] When combined with the ability to recognize and ignore synthesized speech 30, this feature combination also has the unexpected result of interleaving natural speech 34 with synthesized speech 30. As a result, the digital assistant 24, which relies on the interaction packager 10, is able to listen effectively and continuously. Therefore, the digital assistant 24 thus enabled is able to insert comments and smoothly transition into and out of conversations without delay. This, in turn, results in a human-like natural conversational flow, thereby eliminating the need for the current rigid turn-based interaction with the digital assistant 24.

[0074] The context memory 50 of this invention enables the event packager 10 to remember what different clients 32 say and associate that information with the attributes of the clients 32. This is significantly different from known systems, where multiple clients 32 speaking simultaneously often leads to confusion for the digital assistant 24.

[0075] Known digital assistants can listen to a client 12, process the client's utterances through an ASR / NLU module 22 to extract intent, process the client's utterances to identify the client, and then associate the client's intent with the client's identity. However, if the conversation proceeds naturally, with multiple client statements intertwined in time, the digital assistant 24 becomes disoriented and confused.

[0076] As described herein, the interaction packager 10 provides the services required by the digital assistant 22 to maintain the context and auditory scene for multiple clients during a voice interaction session without requiring a separate session with each client. This capability is due toFigure 2 The interaction of the components shown, and the ability to associate multiple intents with corresponding clients while maintaining the temporal order of the interactions, arise from this. Therefore, the resulting event packager 10 produces a natural speaker understanding tool, not just a natural language understanding tool.

[0077] In another embodiment, digital assistant 24 provides a voice interface for the self-service kiosk. Such kiosks are typically found in public places to provide assistance to travelers. In this case, the user speaks their question in their preferred language. Digital assistant 24 recognizes the language and responds in the appropriate language. This is done without requiring the user to make selections on an on-screen display.

[0078] One difficulty with some voice interfaces is that they continue to use the audio interface even when the audio environment makes it pointless. The interaction packager 10 described herein leverages its sensitivity to the audio environment to enable the digital assistant 24 to respond in different modes. For example, if the scene analyzer 48 detects a particularly high level of background noise 38, the digital assistant 24 uses this information to switch to using a graphical user interface instead of the speaker 28 to interact with the client 32.

[0079] The interaction packager 10 also provides the digital assistant 24 with sufficient intelligence to explain to the client 32 why voice interaction cannot be performed, thus preventing often frustrating user experiences. The output of the interaction packager 10 provides the digital assistant 24 with a basis for explaining service interruptions. Examples of these reasons that can be provided include too many speakers, excessive background audio 36, or a weak audio signal, for example, causing the audio signal from the client 32 to be too far from the microphone when speaking. Based on information provided by the scene analyzer 48 or the sensor integrator 52, the interaction packager 10 is able to provide context-based prompts, i.e., “contextual prompts.” In the example above, this could include asking the client 32 to move closer to the microphone or asking the client 32 to take steps to reduce the background audio 36, such as by turning off music or asking people to speak more quietly.

[0080] Figure 3 Examples of interaction events provided to integrator 76 are shown. These events include scene event 78 and speaker event 80.

[0081] Scene event 78 includes location information 82 indicating the occurrence of the interaction, environmental information 84 indicating the state of the acoustic environment, and timing information 86 for identifying scene event 78 within an event sequence and by the time interval between scene events occurring within the event sequence. Speaker event 80 includes speaker information 88 and similar timing information 86. Integrator 76 combines scene event 78 and speaker event 80 to generate event package 20 provided to digital assistant 24.

[0082] The location information 82 and environment information 84 in scene event 78 are the results of the scene analyzer processing the audio segment 42 provided by the stream processor 40. The speaker information 88 in speaker event 80 is the result of information provided by the speaker analyzer 54. The sorting information 86 is again the result of the tags applied at scene memory 68.

[0083] and Figure 3 Similar interactive packets to those described in the text are sent continuously. Specifically, Figure 4 This demonstrates the continuation of a clear session occurring within the same vehicle. This can be seen in timing information 86, which indicates that these events occurred within... Figure 3 The event shown occurs immediately afterward.

[0084] From environmental information 84, it's clear that Alex's Jeep has stopped (i.e., "stationary": true) and the music is off. Furthermore, a new speaker has appeared. Clearly, the system has very little information about the identity of the person speaking to Alex. Figure 4 The speaker information 88 only indicates that the speaker is also speaking English and that the speaker is a real person. However, as more speech samples from the speaker are sent to the speaker analyzer 54, subsequent speaker events 80 may gradually refine this information.

[0085] The present invention and its preferred embodiments have been described, and the technical solutions defined in the claims are protected.

Claims

1. A method of providing services to an application that engages in voice interaction with a human client within an environment, wherein providing the services to the application includes causing an interaction packager to perform the steps of: providing a stream of interaction packages for use by the application, each of the interaction packages including a scene event and a speaker event, The stream that provides the interaction package includes: Continuously listen to the environment to obtain a stream of audio data. The audio data is segmented into audio segments. For each audio segment that includes speech activity, Extract speaker-specific segments, each extracted speaker-specific segment corresponding to the speaker's speech in the audio environment. The speaker event is generated at least in part based on one of the speaker's specific segments. Generate audio scene imprints, and After generating the audio scene imprint and the speaker event, the speaker event and the audio scene imprint are marked with timing information, wherein the audio scene imprint becomes a scene event when it is marked. Based on the timing information, the scene events and the speaker events are integrated to form the interaction package.

2. The method of claim 1, wherein extracting speaker-specific segments comprises extracting speaker-specific segments corresponding to synthesized speech from the application and ignoring the speaker-specific segments, thereby the interaction package ignores speaker events caused by the synthesized speech.

3. The method of claim 1, wherein the stream providing the interaction package occurs simultaneously with continuous listening to the audio environment, thereby enabling full-duplex voice communication between the application and one or more human clients.

4. The method according to claim 1, wherein, For each audio segment in the audio clips, generating the audio scene imprint includes using background audio present in the audio clips.

5. The method of claim 1, wherein generating the scene event includes adding location information and environmental information to the scene event, the location information indicating the location where the speaker event occurs, and the environmental information indicating the audio environment in which the speaker event occurs.

6. The method of claim 1, wherein the application includes a car assistant executed on the car's infotainment system.

7. The method of claim 1, wherein the application includes a digital assistant executed in the processing system of the self-service machine.

8. The method of claim 1, wherein the application is executed within the robot's processing system.

9. The method of claim 1, wherein the application includes a natural language understanding unit.

10. The method of claim 1, further comprising storing the speaker-specific segment in a contextual memory, wherein storing the speaker-specific segment includes storing only speaker-specific segments derived from natural speech.

11. The method of claim 1, wherein generating the scene event comprises continuously monitoring the environment to receive non-audio information therefrom, and using the non-audio information in the scene event.

12. The method of claim 1, further comprising identifying audio segments lacking speech activity, and using only the audio segments to generate the scene event.

13. The method according to claim 1, wherein, Each of the speaker-specific segments corresponds to a speaker, and the method further includes, for each of the speaker-specific segments, extracting attribute information about the speaker, and wherein generating the speaker event based at least in part on one of the speaker-specific segments includes including the attribute information in the speaker event.

14. The method according to claim 1, wherein, Each of the speaker-specific segments corresponds to a speaker, and generating the speaker event based at least in part on one of the speaker-specific segments includes tagging the speaker event with information that identifies the speaker.

15. The method of claim 1, further comprising enabling the application to use the interaction package to participate in full-duplex communication with the client.

16. The method of claim 1, wherein generating the scene event comprises continuously monitoring the environment to receive non-audio information therefrom, and using the non-audio information to identify the location of the client during the speaker event.

17. An apparatus including an interaction packager that provides a stream of interaction packages for use by an application engaging in voice interaction with a human client in an environment, each of the interaction packages including a scene event and a speaker event, wherein the interaction packager includes a speaker segmentation module for separating incoming audio into speaker-specific segments, a speaker analyzer for tagging the speaker-specific segments with speaker attributes, a scene analyzer for providing an audio snapshot of the environment, a scene memory for adding timing information to the audio snapshot and the speaker-specific segments, and an integrator for receiving the speaker-specific segments and the audio snapshot and constructing the interaction packages therefrom.