System for providing interruption response policy, method and medium using same

Detect dialogue interrupts through hardware processors and machine learning models, combined with natural language generators, provide appropriate response strategies, solve the unnatural problems of AI roles during interruption, and improve the naturalness and fluency of interactions.

CN120257036APending Publication Date: 2025-07-04DISNEY ENTERPRISES INC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202411530299.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2024-01-04
Filing Date
2024-10-30
Publication Date
2025-07-04

AI Technical Summary

Technical Problem

The prior art is unable to effectively detect and respond to interruptions in conversations by AI characters, resulting in awkward and unnatural experience when interacting with humans.

Method used

Hardware processors, machine learning models, and natural language generators are employed to detect interrupts in conversations and identify appropriate response strategies based on context and personality of AI roles, including retaining, giving up or negotiating conversation rounds, and dynamically generate conversation content.

Benefits of technology

It realizes the natural response of AI characters when interrupted, improves the fluency and consistency of interactions, and reduces the sense of embarrassment.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120257036A_ABST
    Figure CN120257036A_ABST
Patent Text Reader

Abstract

A system includes a hardware processor; a memory for storing software codes; and a machine learning (ML) model trained to detect an interruption of the conversation. The system detects sound emitted by a human and / or another AI character during a session round when the human and / or another AI character interacts with the AI character, and classifies the sound as an interruption or interaction-independent using an ML model. When the sound is independent of the interaction, the dialogue round of the AI role is continued. When the sound is an interruption, the system identifies a response policy to continue the interaction and executes the response policy, the response policy including at least one of (i) preserving a session round, (ii) discarding a session round, or (iii) negotiating with a human and / or another AI role to determine a preserved or discarded session round.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to a system for providing an interruption response strategy, in particular to a system for providing an interruption response strategy for an AI character, a method used by the system for providing an interruption response strategy for an AI character, and a computer-readable non-transitory medium. Background Art

[0002] A typical feature of most interpersonal interactions is spontaneity, including how individuals in a conversation respond to interruptions in their speech. In order for non-human social media embodied in the form of, for example, an artificial intelligence (AI) interaction character to be able to converse in a natural human way, it is very important to provide the ability for the non-human social media to respond appropriately when interrupted.

[0003] Conventional solutions for providing speech to non-human social media may fail to detect interruptions to the speech or lack sufficient response capabilities when an interruption is detected. For example, an interruption is simply ignored when not detected, and the social media continues its speech, that is, "continue speaking despite the interruption". As an option, conventional responses to detected speech interruptions in non-human social media include completely stopping the speech when the interruption occurs. These responses may be appropriate in some cases but not in others. Unfortunately, both of these conversation scenarios may lead to embarrassment and unnaturalness when humans interact with the social media, especially when the interruption goes unnoticed and the social media continues to speak despite the interruption, which may make humans feel ignored. Therefore, there is a need in the art for a solution that enables non-human social media to respond to interruptions in its speech in a manner consistent with the spontaneity in human conversation behavior. Summary of the Invention

[0004] A system for providing an interruption response strategy for an AI character, comprising:

[0005] A hardware processor;

[0006] A system memory that stores software code; and

[0007] A machine learning (ML) model that is trained to detect conversation interruptions of conversation participants;

[0008] The hardware processor is configured to execute the software code so as to:

[0009] When interacting with at least one of a human or a second AI character by a first artificial intelligence (AI) character, detect a sound emitted by at least one of the human or the second AI character during a conversation turn by the first artificial intelligence (AI) character;

[0010] Use the ML model to classify the voice as at least one of an interruption of the interaction by a human or a second AI character, or as unrelated to the interaction;

[0011] When the voice is classified as unrelated to the interaction:

[0012] Continue using the first AI character for the dialogue turn of the first AI character;

[0013] When the voice is classified as an interruption of the interaction:

[0014] Identify a response strategy for continuing the interaction; and

[0015] Execute the identified response strategy, which includes at least one of the following: (i) retain the dialogue turn of the first AI character, (ii) abandon the dialogue turn of the first AI character, or (iii) negotiate with at least one of a human or a second AI character to determine whether to retain or abandon the dialogue turn of the first AI character.

[0016] According to one or more embodiments, when the voice is classified as an interruption, the interruption causes a change in the prosody of the speech of the first AI character during the dialogue turn.

[0017] According to one or more embodiments, when the voice is classified as an interruption, the interruption causes at least one of a speech fluency disorder or speech hesitation in the first AI character during the dialogue turn.

[0018] According to one or more embodiments, the response strategy for continuing the interaction is identified at least in part based on the role personality of the first AI character.

[0019] According to one or more embodiments, when the voice is classified as an interruption, at least one of the words, gestures, facial expressions, or gaze avoidance of the first AI character is used to confirm the interruption.

[0020] According to one or more embodiments, when the identified response strategy includes retaining the dialogue turn of the first AI character, the hardware processor is further configured to execute the software code so as to:

[0021] Modify one or more planned speech lines of the dialogue turn in response to the interruption.

[0022] According to one or more embodiments, further includes:

[0023] A natural language generator (NLG);

[0024] Wherein, when the recognized response strategy includes retaining the conversation turn of the first AI role, the hardware processor is further configured to execute the software code to:

[0025] In response to the interruption, use NLG to dynamically generate one or more dialogue lines for the first AI role to complete the conversation turn.

[0026] A method used by a system including a hardware processor and a system memory for providing an interruption response strategy for an AI role, the system memory storing software code and a machine learning (ML) model, the machine learning (ML) model being trained to detect interruptions to a conversation by participants in the conversation, the method including:

[0027] During a conversation turn in an interaction between a first artificial intelligence (AI) role and at least one of a human or a second AI role, detect a sound emitted by at least one of the human or the second AI role through the software code executed by the hardware processor;

[0028] Use the ML model through the software code executed by the hardware processor to classify the sound as an interruption to the interaction by at least one of the human or the second AI role, or as unrelated to the interaction;

[0029] When the sound is classified as unrelated to the interaction:

[0030] Continue the conversation turn of the first AI role through the software code executed by the hardware processor and using the first AI role;

[0031] When the sound is classified as an interruption:

[0032] Identify a response strategy for continuing the interaction through the software code executed by the hardware processor; and

[0033] Execute the recognized response strategy through the software code executed by the hardware processor, the response strategy including at least one of the following: (i) retaining the conversation turn of the first AI role, (ii) abandoning the conversation turn of the first AI role, or (iii) negotiating with at least one of the human or the second AI role to determine whether to retain or abandon the conversation turn of the first AI role.

[0034] According to one or more embodiments, when the sound is classified as an interruption, the interruption causes a change in the prosody of the speech of the first AI role in the conversation turn.

[0035] According to one or more embodiments, when the sound is classified as an interruption, the interruption causes at least one of speech fluency disorder or speech hesitation in the first AI character during the dialogue turn.

[0036] According to one or more embodiments, the response strategy for continuing the interaction is identified at least in part based on the character personality of the first AI character.

[0037] According to one or more embodiments, when the sound is classified as an interruption, at least one of the words, gestures, facial expressions or gaze avoidance of the first AI character is used to confirm the interruption.

[0038] According to one or more embodiments, when the identified response strategy includes retaining the dialogue turn of the first AI character, the method further includes:

[0039] Modifying one or more planned speech lines of the dialogue turn by the software code executed by the hardware processor in response to the interruption.

[0040] According to one or more embodiments, the system further includes a natural language generator (NLG), and when the identified response strategy includes retaining the dialogue turn of the first AI character, the method further includes:

[0041] Dynamically generating one or more dialogue lines by the software code executed by the hardware processor in response to the interruption to complete the dialogue turn of the first AI character.

[0042] A computer-readable non-transitory medium storing instructions that, when executed by a hardware processor, instantiate a method, the method including:

[0043] During a dialogue turn in an interaction between a first artificial intelligence (AI) character and at least one of a human or a second AI character, detecting a sound emitted by at least one of the human or the second AI character;

[0044] Using a trained machine learning (ML) model to classify the sound to detect a dialogue interruption of a dialogue participant, classifying the sound as an interruption of at least one of the human or the second AI character, or classifying it as irrelevant to the interaction;

[0045] When the sound is classified as irrelevant to the interaction:

[0046] Using the first AI character to continue the dialogue turn of the first AI character;

[0047] When the sound is classified as an interruption:

[0048] Identify a response strategy for continuing the interaction; and

[0049] Execute the identified response strategy, the response strategy including at least one of the following: (i) retain the conversation turn of the first AI role, (ii) abandon the conversation turn of the first AI role, or (iii) negotiate with at least one of the human or the second AI role to determine whether to retain or abandon the conversation turn of the first AI role.

[0050] According to one or more embodiments, when the voice is classified as an interruption, the interruption causes at least one of a change in the prosody of speech, a speech fluency disorder, or speech hesitation in the conversation turn of the first AI role.

[0051] According to one or more embodiments, the response strategy for continuing the interaction is identified at least partially based on the role personality of the first AI role.

[0052] According to one or more embodiments, when the voice is classified as an interruption, at least one of the words, gestures, facial expressions, or gaze avoidance of the first AI role is used to confirm the interruption.

[0053] According to one or more embodiments, when the identified response strategy includes retaining the conversation turn of the first AI role, the method further includes:

[0054] Modify one or more planned speech lines of the conversation turn in response to the interruption.

[0055] According to one or more embodiments, when the identified response strategy includes retaining the conversation turn of the first AI role, the method further includes:

[0056] Dynamically generate one or more conversation lines using a natural language generator in response to the interruption to complete the conversation turn of the first AI role. Description of the Drawings

[0057] Figure 1 Shows an exemplary system for providing an interruption response strategy for an artificial intelligence (AI) role according to one embodiment;

[0058] Figure 2A Shows a more detailed diagram of an input unit according to one embodiment, the input unit being suitable for use as Figure 1 a component of the system shown;

[0059] Figure 2Bshows a more detailed diagram of an output unit according to an embodiment, which is suitable for use as Figure 1 a component of the system shown;

[0060] Figure 3 shows a schematic diagram of an interrupt response policy recognition engine for software code suitable for use in Figure 1 the system shown; and

[0061] Figure 4 shows a flowchart presenting an exemplary method for use in a system according to an embodiment, the method providing an interrupt response policy for an AI persona. DETAILED DESCRIPTION

[0062] The following description includes specific information related to embodiments in the present disclosure. Those skilled in the art will recognize that the present disclosure may be implemented in ways different from those specifically discussed herein. The accompanying drawings and the detailed description thereof in this application are directed only to exemplary embodiments. Unless otherwise specified, similar or corresponding elements in the drawings may be represented by similar or corresponding reference numerals. Additionally, the drawings and illustrations in this application are generally not drawn to scale and are not intended to correspond to actual relative sizes.

[0063] This application discloses a system and method for providing an interrupt response policy for an artificial intelligence (AI) persona, aiming to address and overcome deficiencies in conventional techniques, enabling the AI persona to respond to interruptions in its speech in a manner consistent with human conversation behavior, while also being consistent with the communication goals of the AI persona and its personality or "role persona". Additionally, existing solutions for providing an interrupt response policy for an AI persona can be advantageously implemented as automated systems and methods.

[0064] As used in this application, the terms "automation", "automated", and "automation process" refer to systems and processes that do not require the participation of a human system administrator. Although in some implementations, the interrupt response policies identified by the systems and methods of the present disclosure may be reviewed and even modified by a human editor or system administrator, such human participation is optional. Accordingly, the methods described in this application may be executed under the control of the hardware processing components of the disclosed systems.

[0065] In addition, the AI character defined in this application refers to a non-human social medium that exhibits behaviors and intelligence that can be perceived by a person interacting with the AI character, and acts as a unique individual with its own personality. The AI character can be implemented as a machine or other physical device (such as a robot or a toy), or can be a virtual entity such as a digital character presented through animation on a screen. The AI character can communicate with its unique voice (for example, vocalization, pitch, loudness, speech rate, dialect, accent, rhythm, intonation, etc.) so that a human observer can recognize the AI character as a unique individual. The AI character can exhibit the characteristics of a real or historical person, a fictional character in literature, film, etc., or simply a unique individual that exhibits a personality pattern recognizable to humans.

[0066] It should be noted that, according to the definition of this application, the expression "real-time" refers to a time interval that enables an interaction (such as a conversation) to be responded to and confirmed by the AI character without unnatural delay after the AI character's speech is interrupted by a human speaker. For example, "real-time" can refer to the response time of the AI character being within a range of one hundred milliseconds or less. It should also be noted that the term "non-verbal vocalization" refers to vocalizations not based on language, such as grunts, sighs or laughter, while "non-vocal sounds" refer to sounds made by hand clapping or other manual means. It should also be noted that the term "prosody" used herein has its conventional meaning, referring to the stress, rhythm and intonation of spoken language.

[0067] Generally speaking, this application discloses a novel and creative solution for providing an interruption response strategy for an AI character, which solution goes beyond conventional methods and advances the state of the art in the following ways: classifying speech into interrupted input and non-interrupted input, using context dialogue information to determine an appropriate response strategy for interruption or other response strategies applicable to a specific situation, and parsing the response strategy to generate an appropriate inserted dialogue, repeating part of the speech that the AI character has already spoken, and abandoning part of the planned speech.

[0068] This solution provides a system and method that can react to interruption signals from a human speaker or other AI characters, and then determine whether the interrupted AI character intends to abandon its turn in the conversation. Based on these interruption signals, the interruption response strategy recognition engine contains logic to modify the existing speech playback of the interrupted AI character, selectively re-render parts of that speech, or both. According to the definition of this application, the term "turn in the conversation" refers to a segment of speech by a participant in a conversation that is intended not to be interrupted, and at the same time the participant is considered to "have the floor".

[0069] For example, when an interruption signal is detected and classified during the dialogue turn of an AI character, the interruption response strategy recognition engine may take one of several actions. If the dialogue turn is to be abandoned (i.e., the AI character will stop speaking and allow the human speaker or another AI character to speak), the voice stream of the AI character may be paused and a brief fade-out applied, for example by reducing volume and pitch, to prevent audible clicks in the voice. To produce natural behavior, natural language processing (NLP) may be applied to the output to identify appropriate stopping points, such as at the end of a clause or before a stop word. Optionally, disfluencies may be inserted to simulate the cognitive load of a human trying to speak and listen simultaneously. Alternatively, as a supplement, the content of the AI character's voice stream may be modified by inserting transitional phrases to identify the interruption and provide a polite transition, such as "If you have something to say, please go ahead", to improve dialogue fluency.

[0070] Alternatively, if the dialogue turn of the AI character is to be retained, the remaining speech of the AI character may be re-rendered from where the current interruption occurred, with additional pauses or amplitude modulation added to the first few words to reflect the category of disfluency exhibited by humans when interrupted. Alternatively, or in addition, the content of the AI character's voice stream may be modified by inserting transitional phrases to identify the interruption and provide a polite transition, such as "Please let me continue and I'll make sure you have a chance to speak afterwards". The decision-making process for the modification may utilize a heuristic rule set based on observed human speech patterns, where humans typically speak louder and at a higher pitch when vying for a dialogue turn, or a trained machine learning (ML) model may be used for the decision. It should be noted that when transitional elements are needed, such as confirmation of the interruption, this transition can benefit from a specially trained voice style that reflects the intonation characteristics used by human speakers in these situations.

[0071] The current interruption solution provides an interruption response strategy for the AI character, utilizing word-level timestamps and streaming playback to execute the recognized interruption response strategy, but typically uses a shorter time offset independent of words to perform the recognition of the interruption response strategy. When an interruption is thus detected and classified, the system can precisely track the position of the playback head in the speech of the AI character and then re-render appropriately, possibly with or without deliberate repetition applied. The voice streaming method is also able to maintain low latency during re-rendering, thus maintaining credibility.

[0072] The ML model can be trained to classify detected sounds as interruptions to a participant in a conversation. For example, the ML model can be trained to distinguish interruptive speech from normal conversation speech or background noise. Additionally, in some embodiments, the ML model trained to detect interruptions in a conversation can be applied to more than just speech and can also be used to interpret facial expressions, gaze cues, or other non-verbal expressions to reflect the intent of a human speaker to interrupt the dialogue turn of an AI character.

[0073] Figure 1 A schematic diagram of a system 100 according to one embodiment is shown, which is used to provide an interruption response strategy for an AI character. As Figure 1 shown, the system 100 includes a computing platform 102 having a hardware processor 104, an input unit 130, an output unit 140 including a display 108, a transceiver 148, and a system memory 106 implemented as a non-transitory storage medium. According to this exemplary embodiment, the system memory 106 stores software code 110 including an interruption response strategy recognition engine 160, an AI character personality database 120 including AI character personalities 122a and 122b, an interaction history database 124 including interaction histories 126a, 126b, and 126c, and an ML model 128 trained to detect dialogue interruptions of participants in a conversation. Additionally, Figure 1 AI characters 154a and 154b are shown, and the software code 110, when executed by the hardware processor 104, can provide an interruption response strategy for AI characters 154a and 154b.

[0074] As Figure 1 shown, the system 100 is implemented in a usage environment that includes a communication network 150 providing a network communication link 152 and a natural language generator (NLG) 129, which can be or include a large language ML model and is communicatively connected to the system 100 via the communication network 150 and the network communication link 152. Figure 1 Also shown are a human speaker 112, AI characters 154a and 154b, a dialogue turn 114 by the AI character 154a or 154b, a sound 116 emitted by the human speaker 112 or the AI character 154a or 154b during the dialogue turn 114, and an optional confirmation 118 of the sound 116 when the sound 116 is classified by the ML model 128 of the system 100 as an interruption to a conversation between the human speaker 112 and one or both of the AI characters 154a and 154b, or a conversation between the AI characters 154a and 154b (excluding the human speaker 112).

[0075] Note that, according to the definition of the present application, the term "ML model" refers to a computational model that makes predictions based on patterns learned from data samples or "training data". Various learning algorithms can be used to map the correlation between input data and output data. These correlations form a computational model that can be used for future predictions on new input data. Such a prediction model can include, for example, one or more logistic regression models, Bayesian models, or artificial neural networks (NNs). Additionally, in the context of deep learning, a "deep neural network" can refer to a neural network that utilizes multiple hidden layers between the input layer and the output layer, which may allow learning based on features not explicitly defined in the original data. In the present application, any feature identified as a neural network refers to a deep neural network.

[0076] It should also be noted that although Figure 1 AI character 154a is depicted as a digital character presented on display 108 and AI character 154b is depicted as a robot, these representations are provided only as examples. In other embodiments, AI characters 154a and 154b can be instantiated by devices such as audio speakers, displays, or figurines, or only by audio speakers or displays, to name just a few. It should also be noted that AI character 154b generally corresponds to AI character 154a and can include any features attributed to AI character 154a. Additionally, although Figure 1 not shown, similar to computing platform 102, in some embodiments, AI character 154b can include a hardware processor 104, an input unit 130, an output unit 140, and a system memory 106, where the system memory 106 stores software code 110 including an interrupt response policy recognition engine 160, an AI character personality database 120, an interaction history database 124, and an ML model 128.

[0077] Furthermore, although Figure 1 a human speaker 112 and two AI characters 154a and 154b are depicted, this representation is only an example. In other embodiments, one AI character, two AI characters, or multiple AI characters can engage in conversations with one or more humans corresponding to human speaker 112. Alternatively, in various embodiments, two or more AI characters, such as AI characters 154a and 154b, can participate in a conversation without human participation, or be excluded, or humans can participate only as non-speaking observers. It should also be noted that although Figure 1 two character personalities 122a and 122b and three interaction histories 126a, 126b, and 126c are depicted, the AI character personality database 120 typically stores dozens or hundreds of character personalities, and the interaction history database 124 typically stores hundreds or thousands of interaction histories.

[0078] In addition, it should be noted that each of the interaction histories 126a, 126b, and 126c can be an interaction history specifically for the cumulative interactions between an AI character and the same person (such as the human speaker 112), or for one or more different time sessions in which the interactions between one or more AI characters and the human speaker 112 continue. Further, although in some embodiments, the interaction histories stored in the interaction history database 124 can be a comprehensive record of the interactions between a human speaker and the AI characters 154a, 154b, or both the AI characters 154a and 154b, in other embodiments, the interaction histories stored in the interaction history database 124 can retain only a predetermined number of the most recent interactions between a human speaker and the AI characters 154a, 154b, or both the AI characters 154a and 154b.

[0079] It should be emphasized that the data describing the previous interactions and stored in the interaction history database 124 preferably does not contain personally identifiable information (PII) of the human speaker interacting with the AI characters 154a and 154b. Thus, although the AI characters 154a and 154b are generally able to distinguish between an anonymous human speaker who has had a previous interaction with the AI characters 154a and 154b and an anonymous human speaker who has no previous interaction experience with the AI characters 154a or 154b, the interaction history database 124 is not required to retain information describing the age, gender, race, ethnicity, or any other PII of any human speaker who has conversed or otherwise interacted with the AI characters 154a or 154b.

[0080] Although this application mentions that the software code 110, the AI character personality database 120, the interaction history database 124, and the ML model 128 are stored in the system memory 106 for the sake of clarity of concept, more generally, the system memory 106 can take the form of any computer-readable non-transitory storage medium. The "computer-readable non-transitory storage medium" as defined in this application refers to any medium that does not include a carrier wave or other transient signal that provides instructions to the hardware processor 104 of the computing platform 102. Thus, the computer-readable non-transitory medium can correspond to various types of media, such as volatile media and non-volatile media. Volatile media can include dynamic memory, such as dynamic random access memory (dynamic RAM), and non-volatile memory can include optical, magnetic, or electrostatic storage devices. Common forms of computer-readable non-transitory storage media include, for example, optical discs, RAM, programmable read-only memory (PROM), erasable PROM (EPROM), and flash memory.

[0081] In addition, in some embodiments, system 100 may utilize a decentralized secure digital ledger outside of system memory 106. Examples of such decentralized secure digital ledgers can include blockchains, hash graphs, directed acyclic graphs (DAGs), and ledgers, to name just a few. In the case of using a blockchain ledger as the decentralized secure digital ledger, it may be advantageous or desirable for the decentralized secure digital ledger to utilize a consensus mechanism with a proof-of-stake (PoS) protocol rather than the more energy-consuming proof-of-work (PoW) protocol.

[0082] It should also be noted that although Figure 1 software code 110, AI character personality database 120, interaction history database 124, and ML model 128 are depicted as coexisting in system memory 106, this representation is only for the sake of conceptual clarity. More generally, system 100 may include one or more computing platforms 102, such as computer servers, which may be co-located or form an interconnected but distributed system, such as a cloud-based system. Thus, hardware processor 104 and system memory 106 may correspond to distributed processor and memory resources within system 100. Accordingly, in some embodiments, software code 110, AI character personality database 120, interaction history database 124, and ML model 128 may be stored remotely from each other in the distributed memory resources of system 100. In addition, although Figure 1 NLG 129 is depicted as a remote resource accessible by system 100 via communication network 150, in some embodiments, NLG 129 may be a component of system 100 and may be stored in the memory resources of system 100.

[0083] Computing platform 102 may be a desktop computer or any other suitable mobile or fixed computing system with sufficient data processing capabilities to provide a user interface and implement the functions of computing platform 102 described herein. For example, in other embodiments, computing platform 102 may be a laptop computer, tablet computer, smartphone, or an augmented reality (AR) or virtual reality (VR) device, e.g., providing a display 108. Display 108 may be a liquid crystal display (LCD), light-emitting diode (LED) display, organic light-emitting diode (OLED) display, quantum dot (QD) display, or any other suitable display screen that performs the physical conversion of signals to light.

[0084] It should also be noted that although Figure 1The display input unit 130 and the output unit 140 are both located on the computing platform 102. However, this representation is only an example. In other embodiments including a full audio interface, for example, the input unit 130 can be implemented as a microphone, and the output unit 140 can be an audio speaker. Additionally, in the case where the AI character 154b is implemented in the form of a robot or other type of machine, the input unit 130 and / or the output unit 140 can be integrated with the AI character 154b instead of the computing platform 102. In other words, in some embodiments, the AI character 154b can include one or both of the input unit 130 and the output unit 140.

[0085] The hardware processor 104 can include multiple hardware processing units, such as one or more central processing units, one or more graphics processing units, one or more tensor processing units, one or more field programmable gate arrays (FPGAs), custom hardware for machine learning training or inference, and an application programming interface (API) server, for example. By definition, the terms "central processing unit" (CPU), "graphics processing unit" (GPU), and "tensor processing unit" (TPU) used in this application have their conventional meanings in the art. That is, the CPU includes an arithmetic logic unit (ALU) for performing arithmetic and logical operations of the computing platform 102 and a control unit (CU) for retrieving programs such as software code 110 from the system memory 106, while the GPU can be implemented to reduce the processing overhead of the CPU by performing compute-intensive graphics or other processing tasks. The TPU is an application-specific integrated circuit (ASIC) specifically configured for artificial intelligence applications such as machine learning modeling.

[0086] The transceiver 148 can be implemented as a wireless communication unit, which is configured to be used with one or more wireless communication protocols. For example, the transceiver 148 can include a fourth-generation (4G) wireless transceiver and / or a 5G wireless transceiver. Additionally, or alternatively, the transceiver 148 can be configured to communicate using one or more wireless fidelity global microwave interoperability low energy (BLE), radio frequency identification (RFID), near field communication (NFC), and 60 GHz wireless communication methods.

[0087] Figure 2A A more detailed illustration of the input unit 230 according to one embodiment is shown. The input unit 230 is suitable for use as Figure 1 a component of the illustrated system 100. As Figure 2AAs shown, the input unit 230 may include a speech-to-text (STT) module 232, a variety of sensors 234, one or more microphones 236 (hereinafter referred to as "microphone 236"), and an analog-to-digital converter (ADC) 238. As Figure 2A shown, the sensors 234 of the input unit 230 may include one or more cameras 234a (hereinafter referred to as "camera 234a"), an automatic speech recognition (ASR) sensor 234b, a radio frequency identification (RFID) sensor 234c, a face recognition (FR) sensor 234d, and an object recognition (OR) sensor 234e. Thus, the input unit 230 generally corresponds to Figure 1 the input unit 130 in

[0088]

[0089] Figure 2B Figure 1 Note that the specific sensors shown in the sensors 234 of the input unit 130 / 230 are only examples. In other embodiments, the sensors 234 of the input unit 130 / 230 may include more or fewer sensors than the camera 234a, ASR sensor 234b, RFID sensor 234c, FR sensor 234d, and OR sensor 234e. In addition, in some embodiments, the sensors 234 may include one or more sensors other than the camera 234a, ASR sensor 234b, RFID sensor 234c, FR sensor 234d, and OR sensor 234e. It should also be noted that when the camera 234a is included in the sensors 234 of the input unit 130 / 230, various types of cameras may be included, such as red, green, and blue (RGB) still image and video cameras, RGB-D cameras including depth sensors, and infrared (IR) cameras, etc. A more detailed schematic diagram of the output unit 240 according to an embodiment is shown. The output unit 240 is suitable for use as Figure 1 a component of the system 100 shown. As Figure 2B shown, the output unit 240 may include a combination of one or more TTS modules 242, one or more audio speakers 244 (hereinafter referred to as "audio speaker 244"), and the display 208. As Figure 2B shown, in some embodiments, the output unit 240 may include one or more mechanical actuators 246 (hereinafter referred to as "mechanical actuator 246"). When serving as a component of the output unit 240, the mechanical actuator 246 may be used to generate gesture, facial expression, or gaze cue prompts of the AI character 154b and / or to control the activities of one or more joints or limbs of the AI character 154b. The output unit 240 and the display 208 generally correspond to Figure 1The output unit 140 and the display 108 in []. Therefore, the output unit 140 and the display 108 may share any features attributed to the output unit 240 and the display 208 in the present disclosure, and vice versa.

[0090] It should be noted that the specific features shown in the output unit 140 / 240 are only examples. In other embodiments, the output unit 140 / 240 may include more or fewer features than the TTS module 242, the audio speaker 244, the display 208, and the mechanical actuator 246. Additionally, in other embodiments, the output unit 140 / 240 may include one or more features other than the TTS module 242, the audio speaker 244, the display 208, and the mechanical actuator 246. As described above, the displays 108 / 208 of the output unit 140 / 240 may be implemented as LCDs, LED displays, OLED displays, QD displays, or any other suitable display screen that performs a physical conversion of signals to light.

[0091] Figure 3 shows a schematic diagram of an interrupt response policy recognition engine 360 included in the software code 110 suitable for Figure 1 the system 100 shown. As Figure 3 shown, the interrupt response policy recognition engine 360 is configured to receive a dialogue turn 314 of an AI role and an interruption 372 from a human speaker, and use a context interpretation code group 362, a dialogue turn importance prediction code group 364, an interruption importance prediction code group 366, and a confirmation and response policy generation code group 368 to generate one or both of an optional confirmation 318 of the interruption 372 and a response policy 378 to the interruption 372. Figure 3 Dialogue context data 373, a dialogue turn importance score 374, and an interruption importance score 376 are also shown in [].

[0092] It should be noted that the context interpretation code group 362 of the interrupt response policy recognition engine 360 may be or include an ML model of a trained context classifier. Further note that the dialogue turn importance prediction code group 364 and the interruption importance prediction code group 366 may be or include respective predictive ML models.

[0093] The interrupt response policy recognition engine 360, the dialogue turn 314, and the optional confirmation 318 generally correspond to Figure 1 the interrupt response policy recognition engine 160, the dialogue turn 114, and the optional confirmation 118 in []. Therefore, the interrupt response policy recognition engine 160, the dialogue turn 114, and the optional confirmation 118 may share any features attributed to their respective interrupt response policy recognition engine 360, dialogue turn 314, and optional confirmation 318 in this application, and vice versa. That is, althoughFigure 1 Although not shown in [the figure], similar to the interruption response policy recognition engine 360, the interruption response policy recognition engine 160 may include features corresponding to the context interpretation code group 362, the dialogue turn importance prediction code group 364, the interruption importance prediction code group 366, and the confirmation and response policy generation code group 368, respectively.

[0094] The functions of the software code 110 including the interruption response policy recognition engines 160 / 360 will be described further with reference to Figure 4 below. Figure 4 FIG. 480 is a flowchart showing an example method of presenting an interruption response policy for providing an AI role to a system according to one embodiment. For the method outlined in Figure 4 some details and features have been omitted in FIG. 480 to avoid obscuring the discussion of the inventive features in this application.

[0095] With reference to Figure 4 and further reference to Figure 1 and Figure 2A FIG. 480 includes detecting a sound 116 (action 481) emitted by the human speaker 112 and / or the second AI role (i.e., one of the AI roles 154a or 154b) during a dialogue turn 114 when the first AI role (i.e., one of the AI roles 154a or 154b) interacts with the human speaker 112 and / or the second AI role (i.e., the other of the AI roles 154a or 154b). The sound 116 may include the speech of the human speaker 112 and / or the second AI role, non-verbal vocalizations (such as grunts, sighs, or laughter of the human speaker 112 and / or the second AI role), and non-verbal sounds emitted by the human speaker 112 and / or the second AI role. As described above, non-verbal vocalizations refer to vocalizations not based on language, such as grunts, sighs, or laughter, while non-verbal sounds refer to sounds such as palm slaps or other manually emitted sounds. The sound 116 can be detected in action 481 by the software code 110 executed by the hardware processor 104 of the system 100 and using the microphone 236 of the input unit 130 / 230.

[0096] With reference to Figure 1 、 Figure 2A 、 Figure 3 and Figure 4In combination, flowchart 480 further includes classifying the sound 116 as an interruption 372 of the human speaker 112 and / or the second AI character by using the ML model 128, or as a sound unrelated to the interaction of the first AI character with the human speaker 112 and / or the second AI character (action 482). As described above, the ML model 128 is trained to detect interruptions in the conversations of the participants in the dialogue. For example, the ML model 128 can be trained to distinguish between interruptive speech and normal conversational speech. Additionally, in some embodiments, the ML model 128 can be applied not only to speech but also to interpret gestures, facial expressions, gaze cues, or other non-verbal expressions to reflect the intention of the human speaker 112 and / or the second AI character to interrupt the dialogue turn 114 of the first AI character. The sound 116 can be classified as an interruption 372 of the interaction of the first AI character with the human speaker 112 and / or the second AI character or as a classification unrelated to the interaction in action 482 by software code 110 executed by the hardware processor 104 of the system 100 and using the STT module 232 and the ASR sensor 234b of the input unit 130 / 230, the ML model 128 (and in some embodiments using one or more cameras 234a, FR sensors 234d, and OR sensors 234e of the input unit 130 / 230).

[0097] Combined Figure 1 and Figure 4 , flowchart 480 further includes: when the sound 116 is classified as unrelated to the interaction of the first AI character with the human speaker 112 and / or the second AI character, continuing the dialogue turn 114 of the first AI character using the first AI character (action 483). That is, in some use cases, when the sound 116 is classified by the ML model 128 in action 482 as unrelated to the interaction of the first AI character with the human speaker 112 and / or the second AI character, the dialogue turn 114 is continued as if the sound 116 had not been detected in action 481. However, in other use cases, although the dialogue turn 114 of the first AI character continues in action 483, its continuation can include speech hesitation or speech fluency disorders. For example, when the sound 116 is a cough or other reflexive or involuntary sound of the human speaker 112 and not an interruption of the dialogue turn 114 and thus unrelated to the content of the dialogue turn 114, the first AI character may pause briefly, hesitate, or exhibit a fluency disorder when continuing the dialogue turn 114. Action 483 can be performed by software code 110 executed by the hardware processor 104 of the system 100.

[0098] Refer to Figure 1 , Figure 3 and Figure 4In combination, flowchart 480 further includes: when voice 116 is classified as interruption 372, identifying a response strategy 378 (action 484) for continuing the interaction. When voice 116 is classified as interruption 372, the identification of the response strategy 378 for the first AI character to continue interacting with the human speaker 112 and / or the second AI character can be based on one or more features of the interaction in action 484, and the one or more features include the context of the conversation between the first AI character and the human speaker 112 and / or the second AI character, the predicted importance of the conversation turn 114 / 314, and the predicted importance of the interruption 372.

[0099] The identification of the response strategy 378 for the first AI character to continue interacting can be performed in action 484 by software code 110 executed by the hardware processor 104 of system 100 and using the interruption response strategy identification engine 160 / 360. As described above, in some embodiments, the context interpretation code group 362 of the interruption response strategy identification engine 160 / 360 can be or include an ML model trained as a context classifier. Further as described above, the conversation turn importance prediction code group 364 and the interruption importance prediction code group 366 of the interruption response strategy identification engine 160 / 360 can be or include their respective predictive ML models.

[0100] As Figure 3 shown, the context interpretation code group 362 of the interruption response strategy identification engine 160 / 360 can be configured to receive the conversation turn 114 / 314 and the interruption 372 as inputs, interpret the conversation context including the conversation turn 114 / 314, and provide the conversation turn 114 / 314, the interruption 372, and the conversation context data 373 as outputs. The conversation turn importance prediction code group 364 of the interruption response strategy identification engine 160 / 360 can be configured to receive the conversation turn 114 / 314 from the context interpretation code group 362, predict the importance of the conversation turn 114 / 314, and output the conversation turn importance score 374 to the confirmation and response strategy generation code group 368 of the interruption response strategy identification engine 160 / 360.

[0101] The interruption importance prediction code group 366 of the interruption response policy recognition engine 160 / 360 can be configured to receive an interruption 372 from the context interpretation code group 362, predict the importance of the interruption 372, and output an interruption importance score 376 to the confirmation and response policy generation code group 368. In various embodiments, the software code 110, when executed by the hardware processor 104 of the system 100, can utilize the confirmation and response policy generation code group 368 of the interruption response policy recognition engine 160 / 360 and use one, some, or all of the conversation context data 373, the conversation turn importance score 374, and the interruption importance score 376 to identify a response policy 378 for continuing the interaction between the first AI role and the human speaker 112 and / or the second AI role.

[0102] For example, in a use case where the conversation turn includes the first AI role providing a safety instruction to the human speaker 112, the conversation turn importance score 374 may be predicted to be high. As another example, in a use case where the interruption 372 contains an urgent or time-sensitive statement such as: "I have to leave now in order to make my next appointment on time.", the interruption importance score 376 may be predicted to be high.

[0103] In some embodiments, in addition to one or more of the conversation context data 373, the conversation turn importance score 374, and the interruption importance score 376, the response policy 378 can also be identified in the action 384 based on the interaction history between the first AI role and the human speaker 112 and / or the second AI role stored in the interaction history database 124, the role personality of the first AI role or the second AI role stored in the AI role personality database 120, or both the interaction history between the first AI role and the human speaker 112 and / or the second AI role and the role personalities of the first AI role and the second AI role.

[0104] For example, in a case where the interaction history between the first AI role and the human speaker 112 and / or the second AI role indicates that the human speaker 112 and / or the second AI role is a persistent interrupter, the hardware processor 104 of the system 100 can execute the software code 110, obtain this information from the interaction history database 124, and further determine the response policy 378 based on this data for continuing the interaction between the first AI role and the human speaker 112 and / or the second AI role. Alternatively, or in addition, when the personality of the first AI role is extroverted and confident or, conversely, introverted and compliant, the hardware processor 104 of the system 100 can execute the software code 110, obtain these personality traits from the AI role personality database 120, and further determine the response policy 378 based on this information for continuing the interaction between the first AI role and the human speaker 112 and / or the second AI role.

[0105] Reference Figure 1 、 Figure 3 and Figure 4 In combination, in some embodiments, flowchart 480 may further include: when sound 116 is classified as interruption 372, using one or more of speech, gestures, facial expressions, or gaze cues (such as gaze aversion) to confirm interruption 372 through the first AI character (action 485). In some embodiments, the optional confirmation 118 / 318 of interruption 372 may be or include speech by the first AI character explicitly confirming interruption 372, such as the statement: "You seem to want to say something now." Alternatively, or in addition, the optional confirmation 118 / 318 may include a change in the prosody of the speech of the first AI character in dialogue turn 114 / 314, such as a change in speech rate, a change in volume or timbre, or any combination thereof.

[0106] As another alternative or addition, in some embodiments, the optional confirmation 118 / 318 of interruption 372 may include speech fluency disorders of the first AI character in dialogue turn 114 / 314, such as stuttering, slurring, or making sounds such as "um" or "uh". As yet another alternative or addition, in some embodiments, the optional confirmation 118 / 318 of interruption 372 may be manifested as speech hesitation of the first AI character in dialogue turn 114 / 314. In addition, in some embodiments, as described above, the optional confirmation 118 / 318 of interruption 372 may further include one or more gestures, facial expressions, or gaze cues of the first AI character in dialogue turn 114 / 314, such as gaze aversion.

[0107] Note that action 485 is optional and may be omitted from the method summarized by flowchart 480 in some embodiments. In embodiments where optional action 485 is omitted, action 486 described below may directly follow action 484. However, reference Figure 1 、 Figure 3 and Figure 4 , and further reference Figure 2B , in embodiments where interruption 372 is confirmed, action 485 may be performed by software code 110 executed by the hardware processor 104 of system 100 and using the confirmation and response policy generation code group 368 of the confirmation and response policy recognition engine 160 / 360 and the TTS module 242 and audio speaker 244 of the output unit 140 / 240. In addition, in some embodiments, the mechanical actuator 246 of the output unit 140 / 240 may also be used to generate one or more of the gestures, facial expressions, or gaze cues of the first AI character to perform action 485.

[0108] Continue to refer to Figure 1 、 Figure 3 and Figure 4 In combination, flowchart 480 also includes executing response strategy 378 identified in action 485, where response strategy 378 includes at least one of the following: (i) retaining the dialogue turn 114 / 314 of the first AI role, (ii) relinquishing the dialogue turn 114 / 314 of the first AI role, or (iii) the first AI role negotiating with the human speaker 112 and / or the second AI role to determine whether to retain or relinquish the dialogue turn 114 / 314 of the first AI role (action 486). The execution of response strategy 378 can be accomplished in action 486 through software code 110 executed by the hardware processor 104 of system 100 and using the confirmation and response strategy generation code group 368 of the interrupt response strategy recognition engine 160 / 360.

[0109] In the case where the response strategy 378 identified in action 485 includes retaining the use of the dialogue turn 114 / 314 by the first AI role, the remaining part of the dialogue turn 114 / 314 of the first AI role can be re - presented from where the interruption 372 occurred, with additional pauses or amplitude modulations added to the first few words to reflect the category of disfluency exhibited by the human when interrupted. For the decision - making process regarding the modification of the remaining part of the dialogue turn 114 / 314 after the interruption 372, a heuristic rule set observed based on human speech patterns can be used. These humans usually increase their volume and pitch or change their speech content when competing for the dialogue turn. The decision - making process can be dynamically performed using NLG 129. As described above, NLG 129 can be or include a large - language ML model, or a combination of heuristic rules and dynamic NLG can be used to contribute to the decision - making.

[0110] Thus, in some embodiments, when response strategy 378 includes retaining the dialogue turn 114 / 314 of the first AI role, the hardware processor 104 of system 100 can further execute software code 110 to modify one or more predetermined script lines of the dialogue turn 114 / 314 in response to the interruption 372 using the confirmation and response strategy generation code group 368 of the interrupt response strategy recognition engine 160 / 360. Alternatively, in some embodiments where response strategy 378 includes retaining the dialogue turn 114 / 314 of the first AI role, the hardware processor 104 of system 100 can further execute software code 110 to dynamically generate one or more dialogue lines using the confirmation and response strategy generation code group 368 of the interrupt response strategy recognition engine 160 / 360 and using NLG 129 in response to the interruption 372 to complete the dialogue turn 114 / 314 of the first AI role.

[0111] If the dialogue turn 114 / 314 is to be relinquished (i.e., the first AI role is to stop speaking and allow the human speaker 112 and / or the second AI role to speak), the speech stream of the first AI role may be paused and a brief fade-out applied, e.g., by reducing the volume and pitch of the first AI role, to prevent any audible clicks in the speech. To produce natural behavior, natural language processing (NLP) can be used to identify appropriate pause points, such as at the end of a clause or before a stop word. Optionally, disfluencies can be inserted to mimic the cognitive load of a human speaking and listening simultaneously.

[0112] In an implementation where the response strategy 378 includes negotiation by the first AI role for retaining or relinquishing the dialogue turn 114 / 314, the negotiation can be conducted using a predefined scripted negotiation statement used by the first AI role, or can be conducted using negotiation terms dynamically generated by the NLG 129. In the case where negotiation occurs, the first AI role's retaining or relinquishing of the dialogue turn 114 / 314 as described above can follow the negotiation.

[0113] It should be noted that the response strategy 378 can be executed in real time with respect to the detected sound 116. As defined above, in the current context, real time refers to a time interval that enables an interaction such as a dialogue to occur without an unnatural delay apparent between an interruption 372 by the human speaker 112 and / or the second AI role and the first AI role's response to the interruption 372. For example, real time can refer to a response time in the range of one hundred milliseconds or less.

[0114] In combination Figure 1 and Figure 4 it should be noted that for the method outlined by the flowchart 480, the actions 481, 482, 483, 484 (hereinafter referred to as "actions 481 - 484") and action 486, or actions 481 - 484, optional action 485 and action 486 can be performed as an automated process from which human participation other than the interaction between the human speaker 112 and the AI role 154a or 154b can be omitted.

[0115] Accordingly, the present application discloses a system and method for providing an interruption response strategy for an AI role, aiming to solve and overcome the deficiencies in the conventional technology. As described above, the novel and creative solution disclosed in the present application for providing an interruption response strategy for an AI role classifies speech into interrupted and non - interrupted inputs, uses context dialogue information to determine an appropriate response strategy for the interruption, and parses the response strategy to generate appropriate inserted dialogue, repeat part of the speech that the AI role has already said, or relinquish part of the planned speech.

[0116] It is apparent from the above description that various techniques can be used to implement the concepts described in this application without departing from the scope of these concepts. Additionally, although these concepts have been specifically described with reference to some embodiments, those of ordinary skill in the art will recognize that changes can be made in form and detail without departing from the scope of these concepts. Therefore, the described embodiments should be considered exemplary in all respects and not restrictive. It should also be understood that this application is not limited to the specific embodiments described herein, but rather many rearrangements, modifications, and substitutions can be made without departing from the scope of the disclosure.

Claims

1. A system for providing an interruption response strategy for an AI character, comprising: A hardware processor; A system memory that stores software code; And A machine learning (ML) model that is trained to detect interruptions in a conversation by a conversation participant; The hardware processor is configured to execute the software code so as to: During an interaction between a first artificial intelligence (AI) character and at least one of a human or a second AI character, detect a sound emitted by at least one of the human or the second AI character during a conversation turn of the first artificial intelligence (AI) character; Use the ML model to classify the sound as an interruption of the interaction by at least one of the human or the second AI character, or classify it as irrelevant to the interaction; When the sound is classified as irrelevant to the interaction: Continue the conversation turn of the first AI character using the first AI character; When the sound is classified as an interruption of the interaction: Identify a response strategy for continuing the interaction; And Execute the identified response strategy, the response strategy including at least one of the following: (i) retain the conversation turn of the first AI character, (ii) abandon the conversation turn of the first AI character, or (iii) negotiate with at least one of the human or the second AI character to determine whether to retain or abandon the conversation turn of the first AI character.

2. The system for providing an interruption response strategy for an AI character according to claim 1, wherein when the sound is classified as an interruption, the interruption causes a change in the prosody of the speech of the first AI character during the conversation turn.

3. The system for providing an interruption response strategy for an AI character according to claim 1, wherein when the sound is classified as an interruption, the interruption causes at least one of a speech fluency disorder or speech hesitation in the first AI character during the conversation turn.

4. The system for providing an interruption response strategy for an AI character according to claim 1, wherein the response strategy for continuing the interaction is identified at least in part based on the character personality of the first AI character.

5. The system for providing an interruption response strategy for an AI character according to claim 1, wherein when the sound is classified as an interruption, at least one of the words, gestures, facial expressions or gaze avoidance of the first AI character is used to confirm the interruption.

6. The system for providing an interruption response strategy for an AI character according to claim 1, wherein when the identified response strategy includes retaining the conversation turn of the first AI character, the hardware processor is further configured to execute the software code so as to: Modify one or more planned speech lines of the conversation turn in response to the interruption.

7. The system for providing an interruption response strategy for an AI character according to claim 1, further comprising: A natural language generator (NLG); Wherein, when the identified response strategy includes retaining the conversation turn of the first AI character, the hardware processor is further configured to execute the software code so as to: In response to the interruption, use NLG to dynamically generate one or more dialogue lines for the first AI character to complete the dialogue turn.

8. A method used by a system for providing an interruption response strategy for an AI character, the system including a hardware processor and a system memory, the system memory storing software code and a machine learning (ML) model, the machine learning (ML) model being trained to detect interruptions to a dialogue by participants in the dialogue, the method including: During a dialogue turn in an interaction between a first artificial intelligence (AI) character and at least one of a human or a second AI character, detect a sound emitted by at least one of the human or the second AI character by the software code executed by the hardware processor; Using the ML model through the software code executed by the hardware processor, classify the sound as an interruption to the interaction by at least one of the human or the second AI character, or as unrelated to the interaction; When the sound is classified as unrelated to the interaction: Continue the dialogue turn of the first AI character through the software code executed by the hardware processor and using the first AI character; When the sound is classified as an interruption: Identify a response strategy for continuing the interaction through the software code executed by the hardware processor; and Execute the identified response strategy through the software code executed by the hardware processor, the response strategy including at least one of the following: (i) retain the dialogue turn of the first AI character, (ii) abandon the dialogue turn of the first AI character, or (iii) negotiate with at least one of the human or the second AI character to determine whether to retain or abandon the dialogue turn of the first AI character.

9. The method according to claim 8, wherein when the sound is classified as an interruption, the interruption causes a change in the prosody of the speech of the first AI character in the dialogue turn.

10. The method according to claim 8, wherein when the sound is classified as an interruption, the interruption causes at least one of a speech fluency disorder or speech hesitation in the first AI character in the dialogue turn.

11. The method according to claim 8, wherein the response strategy for continuing the interaction is identified at least in part based on the character personality of the first AI character.

12. The method according to claim 8, wherein when the sound is classified as an interruption, at least one of the words, gestures, facial expressions, or gaze avoidance of the first AI character is used to confirm the interruption.

13. The method according to claim 8, wherein when the identified response strategy includes retaining the dialogue turn of the first AI character, the method further includes: Modify one or more planned speech lines of the dialogue turn in response to the interruption through the software code executed by the hardware processor.

14. The method according to claim 8, wherein the system further comprises a natural language generator (NLG), and wherein when the identified response strategy includes retaining the dialogue turn of the first AI role, the method further comprises: Dynamically generating, in response to the interruption, one or more dialogue lines by the software code executed by the hardware processor to complete the dialogue turn of the first AI role.

15. A computer-readable non-transitory medium having instructions stored thereon that, when executed by a hardware processor, instantiate a method, the method comprising: During a dialogue turn of an interaction between a first artificial intelligence (AI) role and at least one of a human or a second AI role, detecting a sound emitted by at least one of the human or the second AI role; Classifying the sound using a trained machine learning (ML) model to detect a dialogue interruption of a dialogue participant, classifying the sound as an interruption of at least one of the human or the second AI role, or classifying the sound as irrelevant to the interaction; When the sound is classified as irrelevant to the interaction: Continuing the dialogue turn of the first AI role using the first AI role; When the sound is classified as an interruption: Identifying a response strategy for continuing the interaction; and Executing the identified response strategy, the response strategy including at least one of the following: (i) retaining the dialogue turn of the first AI role, (ii) abandoning the dialogue turn of the first AI role, or (iii) negotiating with at least one of the human or the second AI role to determine whether to retain or abandon the dialogue turn of the first AI role.

16. The computer-readable non-transitory medium according to claim 15, wherein when the sound is classified as an interruption, the interruption causes at least one of a change in the prosody of speech, a speech fluency disorder, or speech hesitation of the first AI role during the dialogue turn.

17. The computer-readable non-transitory medium according to claim 15, wherein the response strategy for continuing the interaction is identified at least in part based on the role personality of the first AI role.

18. The computer-readable non-transitory medium according to claim 15, wherein when the sound is classified as an interruption, at least one of the words, gestures, facial expressions, or gaze avoidance of the first AI role is used to confirm the interruption.

19. The computer-readable non-transitory medium according to claim 15, wherein when the identified response strategy includes retaining the dialogue turn of the first AI role, the method further comprises: Modifying one or more planned speech lines of the dialogue turn in response to the interruption.

20. The computer-readable non-transitory medium according to claim 15, wherein when the identified response strategy includes retaining the dialogue turn of the first AI role, the method further comprises: Dynamically generating, in response to the interruption, one or more dialogue lines using a natural language generator to complete the dialogue turn of the first AI role.