Training and development of multi-sourced machine learning model-based artificial intelligence character
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- DISNEY ENTERPRISES INC
- Filing Date
- 2025-01-10
- Publication Date
- 2026-06-04
AI Technical Summary
Existing AI character training systems face challenges in estimating target audience interactions, providing realistic testing conditions, and incorporating user feedback to adapt to evolving content and interactions, lacking effective mechanisms for human-friendly insights and updates.
A multi-source machine learning model-based system for AI character training, utilizing interaction data to generate and compare interaction graphs, modify behavioral models, and integrate qualitative and quantitative feedback for continuous improvement.
Ensures realistic and adaptable AI character interactions by simulating participant behaviors, refining models based on feedback, and enhancing user experience through iterative learning.
Smart Images

Figure 00000000_0000_ABST
Abstract
Description
[Technical Field]
[0001] The present invention relates to the training and development of multi-source machine learning model-based artificial intelligence characters. [Background technology]
[0002] Advances in artificial intelligence (AI) have led to the development of diverse systems that provide AI characters that simulate social agents. However, constructing dialogue for an AI character requires not only understanding what the AI character should say or do, but also anticipating how one or more human users will react during a particular interaction. Given linguistic variability, personality profiles, demographics, and the context in which the interaction takes place, generating realistic interaction behavior by an AI character during a single interaction can require processing hundreds or thousands of data points. As a result, it is impossible for humans to validate all possible interaction scenarios without relying on one or more sophisticated computational models. Consequently, there is a need in the art for a multi-source machine learning model-based solution for training and developing AI characters. Summary of the Invention
[0003] The following description includes specific information regarding the implementation of the present invention. Those skilled in the art will recognize that the present invention may be implemented in ways other than those specifically discussed herein. The drawings and their accompanying detailed description herein are directed to exemplary implementations only. Unless otherwise specified, the same or corresponding elements among the figures are designated by the same or corresponding reference numerals. Furthermore, the drawings and figures herein are generally not to scale and are not intended to correspond to actual relative dimensions.
[0004] This application discloses a system and method for implementing multi-source machine learning (ML) model-based artificial intelligence (AI) character training and development. Furthermore, in some implementations, the present solution for implementing multi-source ML model-based AI character training and development can be advantageously implemented as an automated system and method.
[0005] As used herein, the terms "automated," "automated," and "automating" refer to systems and processes that do not require the participation of a human system administrator. In some implementations, the AI character training and development solutions disclosed herein may be monitored or managed by a human AI character system designer, although human involvement is optional. Thus, the methods described herein may be performed under the control of hardware processing components of the disclosed systems.
[0006] As defined herein, an AI character refers to a non-human social agent that exhibits behaviors and intelligence that can be perceived as unique by a human user interacting with the AI character. An AI character may be implemented as a machine or other physical device, such as a robot or toy, or it may be a virtual entity, such as a visually projected digital character, depicted by an animation on a screen. An AI character may speak with its own distinctive voice (e.g., voicing, pitch, loudness, rate, directness, accent, rhythm, intonation, etc.) such that a human observer recognizes the AI character as a unique individual. An AI character may represent traits of living or historical characters, fictional characters from literature, film, etc., or simply a unique individual who exhibits patterns that humans can recognize as personality traits.
[0007] It should be noted that, as defined herein, the term "non-vocal" refers to non-linguistically based vocalizations such as grunts, sighs, or laughs, to name a few, and the term "non-vocal" refers to handclaps or other manually generated sounds. It should also be noted that the term "prosody" as used herein has its conventional meaning and refers to the stress, rhythm, and intonation of spoken language.
[0008] Furthermore, as defined herein, it should be noted that the term "participant cohort" refers to a particular human individual, an individual AI character, or a group of human and / or AI characters that share communication habits or other characteristic traits specific to that participant cohort. Thus, a "participant cohort personality profile" may include a personality type or character (e.g., introversion vs. extroversion), as well as, in some implementations, an individual's real or simulated prior interactions with one or more AI characters. Furthermore, the term "context for action" may refer to speech directed at a participant cohort or an event in which the participant cohort engages, the participant cohort's goals or motivations, environmental factors such as weather and location, and activity by the participant cohort prior to or concurrent with the subject matter of communication with the participant cohort.
[0009] As an overview, note that the process of continuing to train AI characters for social interaction faces three major challenges. First, before an AI character is deployed for use, it is typically difficult for an AI character designer to estimate and prepare what cohort of participants in the target audience for the AI character are likely to interact with the AI character and how. Second, even after the AI character is deployed, there are currently few mechanisms for gaining insights that human users typically find comfortable, generate misunderstandings, or are simply cumbersome for human users to navigate. Third, a particular AI character may have an association with an actively expanding or otherwise evolving content franchise or may have a prior interaction history with a human user or cohort of users, which may give the human user expectations about what the AI character "perceives" and the ability to respond to those changes over time. The above challenges require (i) the development of realistic testing conditions to ensure successful deployment of AI characters and updates to their interactions and content, and (ii) effective feedback loops that allow the incorporation of experience from live interactions with human users into the AI characters' behavior, as well as the content contained in the AI characters' knowledge bases. [Brief explanation of the drawings]
[0010] [Figure 1] An exemplary system for multi-source machine learning (ML) model-based artificial intelligence (AI) character training and development is presented, according to one implementation. [Figure 2] 1 shows an exemplary diagram depicting multi-source ML model-based training of an AI character prior to deployment of the trained AI character, according to one implementation. [Figure 3A]1 shows a flowchart presenting an exemplary method for performing multi-source ML model-based training of an AI character prior to deployment of the trained AI character, according to one embodiment. [Figure 3B] 3B illustrates additional operations to extend the method outlined in FIG. 3A, according to one embodiment. [Figure 4] 1 shows an exemplary diagram illustrating multi-source ML model-based AI character training and development after deployment of the trained AI character, according to one implementation. [Figure 5] 1 shows a flowchart presenting an exemplary method for performing multi-source ML model-based AI character training and development after deployment of a trained AI character, according to one implementation. DETAILED DESCRIPTION OF THE INVENTION
[0011] FIG. 1 illustrates an exemplary system 100 for implementing multi-source ML model-based AI character training and development, according to one implementation. As shown in FIG. 1, the system 100 includes a computing platform 102 having a hardware processor 104 and a system memory 106 implemented as a non-transitory storage medium. According to this exemplary implementation, the system memory 106 stores software code 110, a participant cohort personality database 120 including participant cohort personality profiles 122a, 122b, 122c, and 122d, an interaction history database 124 (including interaction histories 126a, 126b, and 126c), and one or more trained ML models 128 (hereinafter, “trained ML models” 128). This database may include, for example, one or more multimodal foundation models and / or one or more large language models.
[0012] It should be noted that, as defined herein, the expression "ML model" refers to a computational model for making predictions based on patterns learned from samples of data or "training data." Various learning algorithms can be used to map correlations between interaction data and output data. These correlations form a computational model that can be used to make future predictions regarding new interaction data. Such predictive models may include one or more logistic regression models, Bayesian models, or artificial neural networks (NNs), large language models, multimodal foundation models, as well as various classical AI models, to name a few.
[0013] 1, interaction bitopic system 100 is implemented within a use environment that includes a communications network 140 providing network communications links 142 and one or more AI behavioral models 130 (hereinafter, “AI behavioral models” 130). The AI behavioral models may be or may include, for example, one or more multimodal foundation models and / or one or more large language models communicatively coupled to system 100 via communications network 140 and network communications links 142. Also shown in FIG. 1 are AI characters 154a and 154b, display device 108, human users 112 interacting with one or both of AI characters 154a and 154b, interaction data 114 received by software code 110 from an interaction history database, qualitative feedback 116 received by system 100 from at least some of human users 112, interaction quality assessment data 117, and any quantitative feedback data 118 received from at least some of human users 112. With respect to the qualitative feedback 116, it should be noted that the qualitative feedback 116 may include audio or written responses to questions or other prompts provided by at least some of the human users 112, or text comments submitted by at least some of the human users 112.
[0014] While FIG. 1 illustrates AI character 154a as a digital character rendered on display device 108 and AI character 154b as a robot, these representations are provided merely as examples. In other implementations, one or both of AI characters 154a and 154b are represented by various devices, such as audio speakers, displays, graphics, or augmented reality (AR) or virtual reality (VR) devices, to name a few. Note that AI character 154b generally corresponds to AI character 154a and may include any of the characteristics attributed to AI character 154a. Additionally, although not shown in FIG. 1, similar to computing platform 102, in some implementations, AI character 154b may include a hardware processor 104 and a system memory 106 that stores software code 110, a participant cohort personality database 120, an interaction history database 124, and a trained ML model 128.
[0015] 1 depicts three human users 112 and two AI characters 154a and 154b, this representation is merely exemplary. In other implementations, one AI character, two AI characters, or more than two AI characters may be involved in interactions with one or more human users corresponding to human user 112. Also, while FIG. 1 depicts four participant cohort personality profiles 122a, 122b, 122c, and 122d and three interaction histories 126a, 126b, and 126c, it should be noted that participant cohort personality database 120 would typically store tens, hundreds, or thousands of participant cohort personality profiles, while interaction history database 124 would typically store hundreds or thousands of interaction histories.
[0016] Further, it should be noted that each of interaction histories 126a, 126b, and 126c may be an interaction history dedicated to the cumulative interactions of an AI character with the same participant cohort, or an interaction history dedicated to one or more different time sessions spanning the interactions of one or more AI characters and participant cohorts. Furthermore, in some implementations, the interaction history stored in interaction history database 124 may be comprehensive with respect to interactions by a participant cohort with a particular AI character, while in other implementations, the interaction history stored in interaction history database 124 may retain only a predetermined number of the most recent interactions by a participant cohort with an AI character.
[0017] Note that the data describing previous interactions and maintained in the interaction history database 124 preferably excludes personally identifiable information (PII) of individual members of the participant cohort. Thus, the interaction history database 124 is not required to maintain information describing the age, gender, race, ethnicity, or other PII of any members of the participant cohort with whom the AI character has conversed or otherwise interacted.
[0018] This application refers to the software code 110, the participant cohort personality database 120, the interaction history database 124, and the trained ML model 128 as being stored in the system memory 106 for conceptual clarity, but more generally, the system memory 106 can take the form of any computer-readable, non-transitory storage medium. As defined herein, the term "computer-readable, non-transitory storage medium" refers to any medium other than a carrier wave or other transitory signal that provides instructions to the hardware processor 104 of the computing platform 102. Accordingly, computer-readable, non-transitory media may correspond to various types of media, such as volatile and non-volatile media. Volatile media can include dynamic memory, such as dynamic random access memory (dynamic RAM), while non-volatile memory can include optical, magnetic, or electrostatic storage devices. Common forms of computer-readable, non-transitory storage media include, for example, optical disks, RAM, programmable read-only memory (PROM), erasable PROM (EPROM), and flash memory.
[0019] Additionally, in some implementations, the system 100 may utilize a distributed secured digital ledger in addition to the system memory 106. Examples of such distributed secured digital ledgers may include blockchains, hashgraphs, directed acyclic graphs (DAGs), and Holochain® ledgers, to name a few. In use cases where the distributed secured digital ledger is a blockchain ledger, it may be advantageous or desirable for the distributed secured digital ledger to utilize a consensus mechanism with a proof-of-stake (PoS) protocol rather than the more energy-intensive proof-of-work (PoW) protocol.
[0020] Note that while FIG. 1 depicts the software code 110, the participant cohort personality database 120, the interaction history database 124, and the trained ML model 128 as being co-located within the system memory 106, that representation is also provided merely as an aid to conceptual clarity. More generally, the system 100 can include one or more computing platforms 102, such as, for example, co-located computer servers, or can form an interactively linked but distributed system, such as, for example, a cloud-based system. As a result, the hardware processor 104 and the system memory 106 can correspond to distributed processor and memory resources within the system 100. As a result, in some implementations, the software code 110, the participant cohort personality database 120, the interaction history database 124, and the trained ML model 128 can be stored remotely from one another on the distributed memory resources of the system 100. Additionally, while FIG. 1 illustrates the AI behavioral models 130 as one or more remote resources accessible by the system 100 communication network 140 and network communication links 142, in some implementations, one or more of the AI behavioral models 130 may be components or components of the system 100 and may be stored within the system memory 106.
[0021] The hardware processor 104 may include multiple hardware processing units, such as, for example, one or more central processing units, one or more graphics processing units, and one or more tensor processing units, one or more field-programmable gate arrays (FPGAs), custom hardware for machine learning training or inference, and an application programming interface (API) server. By way of definition, as used herein, the terms “central processing unit” (CPU), “graphics processing unit” (GPU), and “tensor processing unit” (TPU) have their customary meanings in the art. That is, a CPU includes an arithmetic logic unit for performing arithmetic and logical operations on the computing platform 102 and a control unit for retrieving programs, such as software code 110, from the system memory 106, while a GPU may be implemented to reduce the processing overhead of the CPU by performing computationally intensive graphics or other processing tasks. A TPU is an application-specific integrated circuit (ASIC) specifically configured for AI applications such as machine learning modeling.
[0022] In some implementations, computing platform 102 may correspond to one or more web servers accessible via a packet-switched network, such as the Internet. Alternatively, computing platform 102 may correspond to one or more computer servers included in a private wide area network (WAN), local area network (LAN), or other type of limited distribution or private network. Additionally or alternatively, in some implementations, system 100 may utilize a local area broadcast method, such as User Datagram Protocol (UDP) or Bluetooth. Furthermore, in some implementations, system 100 may be implemented virtually, such as in a data center. For example, in some implementations, system 100 may be implemented in software or as a virtual machine. Furthermore, in some implementations, communication network 140 may be a high-speed network suitable for high-performance computing (HPC), such as a 10 GigE network or an Infiniband network.
[0023] 2 shows an example diagram 200 depicting multi-source ML model-based training of an AI character prior to deployment of the trained AI character, according to one implementation. According to the example data-driven test scenario development implementation shown in FIG. 2, interaction data 214 is extracted from interaction history database 224 to store interaction histories 226a, 226b, and 226c.
[0024] The interaction data 214 may be used to generate an interaction graph 232 of the participant cohort's behavior during interactions with the previously placed AI character. Further, the interaction data 214 may be provided as a training input to one of the AI behavior models 230. The AI behavior model trained with the interaction data 214 may then be used to simulate the behavior of the participant cohort represented by the interaction data 214 to provide a predicted interaction graph 234 of the participant cohort's behavior. The predicted interaction graph 234 may then be compared with the generated interaction graph 232 to determine a similarity score of the predicted interaction graph 234 relative to the generated interaction graph 232.
[0025] In use cases where the similarity score identified by comparing the predicted interaction graph 234 with the generated interaction graph 232 meets the similarity criteria, for example, the trained AI behavior model may be used to deploy as an interactive AI character to train a new AI character. However, in use cases where the similarity score does not satisfy the similarity criteria, the AI behavior model may be modified based on one or more differences between the predicted interaction graph 234 and the generated interaction graph 232, the modified AI behavior model may be used to run a simulation that provides a new predicted interaction graph, and the new predicted interaction graph may be compared with the generated interaction graph 232. The process is repeated until the comparison of the predicted interaction graph with the generated interaction graph 232 meets the similarity criteria, at which point the modified AI behavior model may be used to train a new AI character for interaction for one or more times.
[0026] 2 , interaction history database 224 storing interaction data 214, interaction histories 226a, 226b, and 226c, and AI behavior model 230 generally corresponds to interaction history database 124 storing interaction data 114, interaction histories 226a, 126b, and 126c, and AI behavior model 130, respectively, in FIG. 1 . Therefore, interaction data 214, interaction history database 224, interaction histories 226a, 226b, and 226c, and AI behavior model 230 can share any of the characteristics attributed to the respective interaction data 114, interaction history database 124, interaction histories 126a, 126b, and 126c, and AI behavior model 130, according to the present invention, and vice versa. Thus, for example, AI behavior model 130, AI behavior model 230, etc., may be or include one or more multimodal foundation models and / or one or more large language models.
[0027] The process depicted by diagram 200 of Figure 2, and the functionality of software code 110, when executed by hardware processor 104 of system 100 of Figure 1, are further described with reference to Figures 3A and 3B. Figure 3A shows a flowchart 360 presenting an exemplary method for performing multi-source ML model-based AI character training prior to placement of the trained AI character, according to one embodiment, while Figure 3B shows additional operations for extending the method outlined in Figure 3A, according to one embodiment. Note that with respect to the method outlined in Figures 3A and 3B, certain details and features have been left out of flowchart 360 so as not to obscure the discussion of inventive features herein.
[0028] 3A and with further reference to FIGS. 1 and 2, a flowchart 360 includes receiving interaction data 114 / 214 and identifying behaviors and personality profiles where the interaction data 114 / 214 respectively correspond to multiple participant cohorts within the behavior (operation 361). In various use cases, the behaviors identified by the interaction data 114 / 214 may include one or more of the same voices directed to each of the participant cohorts and the same events undertaken by each of the participant cohorts.
[0029] In use cases where the actions identified by the interaction data 114 / 214 include the same event being undertaken by each of the participant cohort, such event may take the form of entering a room, cruise ship cabin, or other venue and / or forming a physical interaction with one or more objects within such venue by each of the participant cohort. For example, the actions may include entering a hotel room or cruise ship cabin, occupying a piece of furniture located therein, and turning electronic devices such as a television (TV) or stereo on or off.
[0030] In some implementations, personality profiles corresponding to each of the multiple participant cohorts in the operation identified in operation 361 may be included in participant cohort personality profiles 122a, 122b, 122c, and 122d stored in participant cohort personality profile database 120. As described above, those participant cohort personality profiles may include multiple personality types or personalities (e.g., introversion vs. extroversion) and, in some implementations, may include an individual's prior real or simulated interactions with one or more AI characters.
[0031] 1 , the interaction data 114 / 214 may be received as a data transfer in operation 361 from the interaction history database 124 / 224 of the system 100 by the software code 110 executed by the hardware processor 104. However, it should be noted that in some embodiments, the interaction history database 124 / 224 may not be stored in the system memory 106 of the system 100, but may be a remote data storage resource accessible by the system 100 via the communications network 140 and the network communications link 142.
[0032] It should be noted that in some embodiments, the hardware processor 104 can execute the software code 110 to generate one or more new participant cohort personality profiles. By way of example, if the interaction data 114 / 214 results from a participant cohort behavior that does not reasonably match an existing one of the participant cohort personality profiles 122a, 122b, 122c, and 122d, the software code 110, when executed by the hardware processor 104, may use the new behavior to generate a new participant cohort personality profile. It should also be noted that following the generation of such a new participant cohort personality profile, the new participant cohort personality profile may be persistently stored in the participant cohort personality profile database 120, which includes the participant cohort personality profiles 122a, 122b, 122c, and 122d.
[0033] In some use cases, the interaction data 114 / 214 may further describe a context for the action identified by the interaction data 114 / 214. As described above, such context for the action may refer to actions engaged in by a particular participant cohort prior to or contemporaneously with the action identified by the interaction data 114 / 214, the goals or motivations of that particular participant cohort, environmental factors such as weather and location, and the subject of ongoing communication with the participant cohort.
[0034] Continuing to refer to FIG. 3A in combination with FIGS. 1 and 2, flowchart 360 further includes using the interaction data 114 / 214 to generate an interaction graph 232 of the participant cohort's behavior in the operation (operation 362). The interaction graph 232 generated in operation 362 may take the form of a multidimensional node-edge graph including edges representing relationships between nodes. According to one implementation, each node corresponds to an AI character behavior observed by the participant cohort, and each edge may represent a participant cohort's behavior. For example, the color or size of a node may represent how frequently a particular AI character behavior was observed, while the edge may represent how the next AI character behavior was triggered, with a thicker edge representing a more prominently displayed participant cohort's behavior. Note that different colors or edge line delineations (e.g., dotted lines vs. solid lines vs. dashed lines) may be useful for visually representing different participant cohorts.
[0035] As noted above, in some implementations, the participant cohort personality database 120 will typically store tens, hundreds, or thousands of participant cohort personality profiles, while the interaction history database 124 / 224 will typically store hundreds or thousands of interaction histories. Furthermore, as further noted above, each participant cohort may include a particular individual human, an individual AI character, or a group of humans and / or AI characters that share communication habits or other characteristic traits specific to that participant cohort. Thus, the interaction graph 232 generated in operation 362 may represent the same speech directed to each of thousands of participant cohorts, or the same event involving each of thousands of participant cohorts. Operation 362 may be performed by software code 110 executed by the hardware processor 104 of the system 100.
[0036] 3A in combination with FIG. 1 and 2, flowchart 360 further includes using AI behavioral model 130 / 230 to simulate the participation of each of the participant cohort in actions identified by interaction data 114 / 214 to provide a predicted interaction graph 234 for the participant cohort (operation 363). Similar to the interaction graph 232 generated in operation 362, predicted interaction graph 234 may take the form of a multidimensional node-edge graph including edges representing relationships between nodes. In predicted interaction graph 234, each edge may correspond to a predicted behavior of one of the participant cohort during an interaction.
[0037] Additionally, in some implementations, each or all of the edges included in predicted interaction graph 234 may identify one or more intents, e.g., goals motivating behavior, of the participant cohort represented by predicted interaction graph 234. Furthermore, in some implementations, each or all of the edges included in predicted interaction graph 234 may further identify one or more of vocalizations, facial expressions, or gestures attributed to one or more of the participant cohorts represented by predicted interaction graph 234.
[0038] It should be noted that, similar to operation 362 described above, operation 363 may include simulating the same speech directed to each of a cohort of thousands of participants, or simulating the same event involving each of a cohort of thousands of participants. Operation 363 may be performed by the hardware processor 104 of system 100 and, as described above, by software code 110 using AI behavioral model 130 / 230, which may be one or more multimodal foundation models and / or one or more large language models. It should be further noted that, while flowchart 360 lists operation 363 following operation 362, that representation is merely exemplary. In various implementations, operation 363 may precede operation 362, follow operation 362, or be performed in parallel, i.e., simultaneously, with operation 362.
[0039] 3A in combination with FIGS. 1 and 2, flowchart 360 further includes comparing the predicted interaction graph 234 with the generated interaction graph 232 to determine a similarity score of the predicted interaction graph 234 relative to the generated interaction graph 232 (operation 364). In some implementations, for example, the similarity score determined in operation 364 may be expressed as a percentage, in which case identical interaction graphs would yield a similarity score of 100%. By way of example, the comparison performed in operation 364 may compare the number of outgoing and incoming edges per node, as well as the weights applied to the edges (those weights depicted in the graph by the thickness of each edge). Furthermore, to ensure that the predicted interaction graph 234 correctly simulates cumulative interactions, a uniformity score may be assigned based on the overall validity of the paths through the predicted interaction graph 234. Operation 364 may be performed by software code 110 executed by the hardware processor 104 of system 100.
[0040] 3A in combination with FIG. 1 and FIG. 2, flowchart 360 further includes, when the similarity score identified in operation 364 satisfies a similarity criterion, which may be, for example, a predetermined criterion, using AI behavior model 130 / 230 to train an AI character, such as one of AI characters 154a or 154b, for the interaction (operation 365). Note that performance of operation 365 is contingent on the similarity score identified in operation 364 satisfying the similarity criterion. In use cases where the similarity score fails to satisfy the similarity criterion, operation 365 is not performed. However, if the similarity criterion is satisfied by the similarity score identified in operation 364, operation 365 may be performed by software code 110 and executed by hardware processor 104 of system 100 to use AI behavior model 130 / 230.
[0041] 3A in combination with Figures 1 and 2, flowchart 360 may include modifying AI behavior model 130 / 230 based on one or more differences between predicted interaction graph 234 and generated interaction graph 232 if the similarity score identified in act 364 does not meet the similarity criteria (act 366). Note that act 366 occurs only if the similarity score identified in act 364 does not meet the similarity criteria. Otherwise, the method outlined by flowchart 360 ends at act 365, described above.
[0042] In various implementations, the one or more differences between the predicted interaction graph 234 and the generated interaction graph 232 may include one or more of: edges omitted from the predicted interaction graph 234 that are present on the generated interaction graph 232; edges present on the predicted interaction graph 234 that are not present in the generated interaction graph 232; and a discrepancy in weights assigned to edges common to both the predicted interaction graph 234 and the generated interaction graph 232. Operation 366, when included in the method outlined by flowchart 360, may be performed by software code 110 executed by hardware processor 104 of system 100.
[0043] As mentioned above, Figure 3B illustrates additional operations for extending the method outlined in Figure 3A according to one implementation. Referring to Figure 3B in combination with Figures 1, 2, and 3A, in some implementations, flowchart 360 may further include, if the similarity score identified in operation 364 does not satisfy the similarity criterion, repeating the simulation and comparison of operations 363 and 364 until another similarity score for another predicted interaction graph satisfies the similarity criterion (operation 367). Operation 367, when included in the method outlined by flowchart 360, may be performed by software code 110 and executed by hardware processor 104 to use the iteratively modified AI behavioral model.
[0044] 1, 2, 3A, and 3B, in implementations in which operation 367 is performed, flowchart 360 further includes training an AI character, such as one of AI characters 154a or 154b, for interaction using the modified AI behavior model that provides another predicted interaction graph that meets the similarity criteria (operation 368). Operation 368, when included in the method outlined by flowchart 360, may be performed by software code 110 and executed by hardware processor 104 to use the modified AI behavior model that provides another predicted interaction graph that meets the similarity criteria.
[0045] With respect to the method outlined by Figures 3A and 3B, it is emphasized that operations 361, 362, 363 and 364 (hereinafter referred to as "operations 361-364"), as well as operation 365, or operations 361-364 and 366, or operations 361-364, 366, 367 and 368, may be performed in an automated process in which human involvement may be omitted.
[0046] FIG. 4 shows an exemplary diagram 400 illustrating training and deployment of a multi-source ML model-based AI character after deployment of the trained AI character, according to one embodiment. Note that diagram 400 represents an iterative coding process that uses qualitative feedback 416, such as language-based feedback from human users who interacted with an AI character trained and deployed according to the process described above with reference to FIGS. 3A and 3B. That iterative coding process may be performed automatically by the trained ML model 128 of FIG. 1, such as one or more multimodal foundation models and / or one or more large language models. The qualitative feedback 416 from the human users may then be combined with quantitative feedback data 418 from human users rating their interactions with the AI character, and may further be combined with internal measures of system success, such as unrecognized conversational turns, taken path length, detection of unintended conversational loops, and participation in parts of interactions designated as significant or emotionally powerful.
[0047] To the extent that human users often provide different emotional valences in their responses for which qualitative feedback 416 is provided, the full response may be parsed into segments 470a and 470b, with distinct segments defined by linguistic information such as syntax and pause length (for discussion responses), such that each segment 470a and 470b contains only one interaction-specific topic, as shown by section (a) of Figure 4. One or more of the trained ML models 128 may then be used for sentiment analysis to classify the sentiment 472 of each segment of the individual human user response as positive, negative, or neutral, as shown by section (b) of Figure 4. Here, "neutral" refers to a statement without any detected emotional valence.
[0048] Positive and negative responses may then be separately classified into topic codes 474, as shown by section (c) of Figure 4. These topic codes 474, the hierarchy of topic codes 474, and the definitions of individual topic codes 474 may be defined in an iterative process based on one or both of data observed within the system and input from human system experts. Note that automated system-based monitoring of AI character interaction quality is described in detail in U.S. Patent Application No. 17 / 325,676, filed May 20, 2021, and the invention therein entitled "Automated Social Agent Interaction Quality Monitoring and Improvement," which is incorporated herein by reference in its entirety. Furthermore, note that, as discussed above, the emotion classifications as either positive, neutral, or negative are provided merely as examples. In other implementations, any general emotion classification of interest (e.g., frustration, confusion) may be used, and topic codes may be applied to those classifications.
[0049] The response classification step may involve analyzing a graph of human users' interactions with the AI character to determine which parts of the interaction the human user participated in and in what order. ML models included among the trained ML models 128, such as multimodal foundation models or large language models, can also be used to generate automatic summaries of all responses currently sorted into a particular class, which can be used for further levels of detailed classification if a human systems expert or algorithm determines that the summary contains too many diverse topics. The output of the trained ML model that creates the summaries can be used to quantify how frequently certain shortcomings in the AI character's interactions were mentioned and how frequently parts of the interactive experience with the AI character were praised.
[0050] Along with quantitative feedback data 418 of the AI character interactions, optionally provided by human users, weights 476 can be applied to each topic code 474, such that topic codes that often coincide with very negative or very positive ratings are weighted more strongly compared to topic codes that were mentioned more frequently but had little or no influence on the ratings provided by human users who interacted with the AI character. This aspect of the process is illustrated by section (d) of Figure 4.
[0051] In the final step, as shown in section (e) of Figure 4, observed system behavior can be matched against system-generated topic codes 474. For example, for every human user who complains of not being understood by the AI character, the AI character's system log can be analyzed to extract an internal measure of the success of the AI character's language understanding unit. The system can then search for similar patterns within all logs of the AI character's interactions with other human users who do not mention a specific defect in their interactions with the AI character. The weights 476 applied to the topic codes 474 can then be tuned so that problems with a high overlap between their occurrence and the problem addressed by human users are weighted more strongly than interaction defects that are present in many interactions but rarely identified as defects by human users. Similarly, if an interaction feature occurs infrequently but is consistently identified as a human user's preference, the corresponding topic code can receive a stronger positive weight. The outcome of the process depicted in Figure 4 is a weighted set of the most and least preferred aspects of human users' interactions with the deployed AI character.
[0052] Note that qualitative feedback 416 generally corresponds to a single instance of qualitative feedback 116 in Figure 1, while quantitative feedback data 418 generally corresponds to a single instance of quantitative feedback data 118 in the previous figure. As a result, qualitative feedback 116 and quantitative feedback data 118 can share any of the characteristics attributed to qualitative feedback 416 and quantitative feedback data 418, and vice versa, in accordance with the present invention.
[0053] The process depicted by diagram 400 of Figure 4, as well as the functionality of software code 110, as further executed by hardware processor 104 of system 100 of Figure 1, is described below with reference to Figure 5. Figure 5 shows a flowchart 580 presenting an exemplary method for performing multi-source ML model-based AI character training and development after deployment of the trained AI character, according to one implementation. Note that certain details and features regarding the method outlined in Figure 5 have been left out in flowchart 580 so as not to obscure the discussion of inventive features herein.
[0054] 5, with further reference to FIGS. 1 and 4, a flowchart 580 includes receiving qualitative feedback 116 / 416 describing an interaction experience by a human user 112 with a trained AI character, such as one of AI characters 154a or 154b (operation 581). As illustrated by section (a) of FIG. 4, the qualitative feedback 116 / 416 may be or include a language-based qualitative description of the human user's 112's respective interaction experience with the AI character. For example, as further illustrated in FIG. 4, the qualitative feedback 116 / 416 may include a descriptive statement such as, "I didn't really know what to say at first, but once I got around to it, it was amazing." The qualitative feedback 116 / 416 may be received from the human user 112 via the communications network 140 and the network communications link 142 by the software code 110 executed by the hardware processor 104 of the system 100 in operation 581.
[0055] 5 in combination with FIG. 1 and FIG. 4, flowchart 580 further includes using a first one of trained ML models 128 to segment qualitative feedback 116 / 416 into segments 470a and 470b, each corresponding to a single interaction-specific topic (operation 582). As shown in FIG. 4, qualitative feedback 116 / 416 including the phrase "I didn't really know what to say at first, but once I warmed to it, it was amazing" can be segmented into segment 470a, which corresponds to the interaction-specific topic of "I didn't really know what to say at first," and segment 470b, which corresponds to the interaction-specific topic of "Once I warmed to it, it was amazing." The segmentation of the qualitative feedback 116 / 416 into segments 470a and 470b in operation 582 may be performed by the hardware processor 104 of the system 100 and implemented by the software code 110 using an ML model included among the trained ML models 128.
[0056] Continuing to refer to FIG. 5 in combination with FIGS. 1 and 4, flowchart 580 further includes analyzing one or more respective emotions 72 of human users 112 with respect to each interaction-specific topic using qualitative feedback 116 / 416 and at least one of the first ML models used in operation 582 or a second ML model in trained ML models 128 (operation 583). For example, as shown in FIG. 4, the emotion 472 for the interaction-specific topic “I liked it, but it was surprising” is determined to be positive as a result of the analysis performed in operation 583, whereas the emotion for the interaction-specific topic “I didn’t really know what to say at first” is determined to be negative. As described above, a “neutral” emotion may be determined for an interaction-specific pick that simply states a fact without signaling a specific emotional valence. The analysis of operation 583 may be performed by software code 110 executed by hardware processor 104 of system 100 and using an ML model included in trained ML models 128.
[0057] 5 in combination with FIG. 1 and FIG. 4, flowchart 580 further includes obtaining interaction quality evaluation data 117 generated by the control system for the AI character (operation 584). Interaction quality evaluation data 117 may include data contained in the AI character's system log and internal measurements of the success of the AI character's language understanding unit, and in some implementations may include internal measurements of the success of the AI character's physical interaction understanding unit. Interaction quality evaluation data 117 is obtained in operation 584 by software code 110 executed by hardware processor 104 of system 100.
[0058] 5 in combination with FIGS. 1 and 4, flowchart 580 further includes identifying tuning data for improving interaction performance by the AI character using the interaction quality assessment data 117 and based on the analysis performed in act 583 (act 585). In some implementations, as described above with reference to FIGS. 1 and 4, the feedback provided by the human users 112 may include quantitative feedback data 118 / 418 in the form of numerical ratings of interactions with the AI character provided by at least some of the human users 112. Such numerical ratings may be based on a numerical range from 1 to 5, with 5 being, for example, the best rating, and from 1 to 10, with 10 being the best rating. Thus, in some embodiments, hardware processor 104 may further execute software code 110 to receive quantitative feedback data 118 / 418 from human users 112, e.g., via communications network 140 and network communications link 142. In those implementations, identifying tuning data in act 585 may further utilize quantitative feedback data 118 / 418. The operation 585 may be performed by software code 110 executed by the hardware processor 104 of the system 100 .
[0059] With continued reference to FIG. 5 in combination with FIGs. 1 and 4, and with further reference to FIG. 2, flowchart 580 includes tuning the AI behavioral model 130 / 230 trained at operation 365 of flowchart 360 or tuning the modified AI behavioral model trained at operation 368 (operation 586) based on the tuning data. The tuning of the AI behavioral model at operation 586 may be performed by software code 110 executed by hardware processor 104 of system 100. With regard to the tuning data on which the tuning of the AI behavioral model 130 / 230 is based at operation 586, as mentioned above, in some implementations, the tuning data may be identified in an automated process by software code 110 executed by hardware processor 104 of system 100. However, in other implementations, the tuning data may be generated by a human or modified by a human following preliminary identification of the tuning data by software code 110 at operation 585.
[0060] Thus, system 100 can directly use quantitative and / or qualitative feedback provided by human user 112 to update the interaction content and dialogue performance of the AI character by periodically updating and improving the AI behavioral models used to train the AI character for the interactions. Additionally, new interactions can be fed back into the development of simulated behavior by the human user to favor testing of the AI behavioral models to become more reliable over time. It is emphasized that with respect to the method outlined by flowchart 580, operations 581 through 586 can be performed in an automated process, from which human involvement, other than the provision of quantitative and / or qualitative feedback by the human user interacting with the AI character, can be omitted.
[0061] Accordingly, this application discloses systems and methods for performing multi-source ML model-based AI character training and development. While the multi-source ML model-based training and development solution disclosed herein has been described above with reference to the specific use case of training, deploying, and providing ongoing development of an interactive AI character, it should be noted that the concepts may also be applied more generally. For example, the multi-source ML model-based training and development solution may be applied to the evaluation and improvement of virtually any kind of interactive user experience, generating an internal log of interaction successes and interaction failures, as well as a graph-based representation of the evolution of user interactions, respectively.
[0062] It is apparent from the foregoing description that various techniques may be used to implement the concepts described herein without departing from the scope of those concepts. Moreover, while the concepts have been specifically described with reference to particular implementations, those skilled in the art will recognize that changes can be made in form and detail without departing from the scope of those concepts. As such, the described implementations are to be considered in all respects as illustrative and not restrictive. Also, the present application is not limited to the particular implementations described herein, but it will be understood that many rearrangements, modifications, and substitutions are possible without departing from the scope of the invention.
Claims
1. A system comprising a hardware processor and memory for storing software code, The aforementioned hardware processor is Interaction data is received, and the interaction data identifies an action, and identifies multiple personality profiles corresponding to multiple participant cohorts in that action. Using the aforementioned interaction data, an interaction graph of the behaviors of the multiple participant cohorts in the aforementioned action is generated. Using a behavioral model, the participation of each of the multiple participant cohorts in the aforementioned action is simulated, and a predicted interaction graph for the multiple participant cohorts is provided. The predicted interaction graph and the generated interaction graph are compared to identify the similarity score of the predicted interaction graph to the generated interaction graph. If the similarity score satisfies the similarity criteria, the character's behavior model is used for interaction. A system configured to execute software code that modifies the behavioral model based on one or more differences between the predicted interaction graph and the generated interaction graph if the similarity score does not meet the similarity criteria.
2. The aforementioned hardware processor is If the similarity score does not meet the similarity criteria, the simulation and comparison are repeated until another similarity score for another predicted interaction graph meets the similarity criteria. The system according to claim 1, further configured to execute the software code which uses a modified behavior model that provides the other predicted interaction graph for the interaction.
3. The system according to claim 1, wherein the action includes at least one of the same speech directed to each of the multiple participant cohorts, or the same event in which each of the multiple participant cohorts is involved, and the predicted interaction graph includes a plurality of edges that each identify one or more intentions of the multiple participant cohorts.
4. The system according to claim 3, wherein each of the multiple edges further identifies at least one of the vocalizations, facial expressions, or gestures originating from one or more of the multiple participant cohorts.
5. The system according to claim 1, wherein the behavioral model includes at least one of a multimodal foundational model or a large-scale language model.
6. The system according to claim 1, further comprising multiple trained machine learning (ML) models, The aforementioned hardware processor is We receive qualitative feedback describing multiple interaction experiences from multiple human users. Using a first ML model of multiple trained ML models, the qualitative feedback is segmented into multiple segments, each corresponding to a single interaction specific topic. Using the qualitative feedback and at least one of the first or second ML models from the plurality of trained ML models, the respective emotions of one or more of the plurality of human users regarding each interaction on a specific topic are analyzed. We acquire interaction quality evaluation data generated by the control system. Using the aforementioned interaction quality evaluation data and based on the aforementioned analysis, adjustment data to improve the interaction performance by the character is identified. The system according to claim 1, further configured to execute the software code that adjusts the behavior model or one of the modified behavior models based on the adjustment data.
7. The aforementioned hardware processor is Quantitative feedback data is received to evaluate the multiple interaction experiences of the multiple human users having the aforementioned character. Identifying the adjustment data is further configured to execute the software code that further uses the quantitative feedback data, according to claim 6.
8. A method used by a system having a hardware processor and memory for storing software code, wherein the method is The steps include receiving interaction data that identifies the behavior and multiple personality profiles corresponding to each of the multiple participant cohorts in the behavior, via the software code executed by the hardware processor, Using the interaction data, the software code executed by the hardware processor generates an interaction graph of the behaviors of the multiple participant cohorts in the operation; The steps include: using the software code executed by the hardware processor to simulate the participation of each of the multiple participant cohorts in the action using a behavioral model, thereby providing a predicted interaction graph for the multiple participant cohorts; The software code executed by the hardware processor compares the predicted interaction graph with the generated interaction graph to determine a similarity score between the generated interaction graph and the predicted interaction graph. If the similarity score satisfies the similarity criteria, the software code executed by the hardware processor uses a character behavior model for interaction. A method comprising the step of modifying the behavioral model by software code executed by the hardware processor based on one or more differences between the predicted interaction graph and the generated interaction graph, if the similarity score does not satisfy the similarity criteria.
9. If the similarity score does not satisfy the similarity criteria, the software code executed by the hardware processor repeats the simulation step and the comparison step until another similarity score of another predicted interaction graph satisfies the similarity criteria. The method of claim 8, further comprising the step of using a modified behavioral model for the interaction, which provides the other predicted interaction graph, by software code executed by the hardware processor.
10. The method according to claim 8, wherein the action includes at least one of the same speech directed to each of the multiple participant cohorts, or the same event in which each of the multiple participant cohorts is involved, and the predicted interaction graph includes a plurality of edges that each identify one or more intentions of the multiple participant cohorts.
11. The method according to claim 10, wherein each of the multiple edges further identifies at least one of the vocalizations, facial expressions, or gestures attributable to one or more of the multiple participant cohorts.
12. The method according to claim 8, wherein the behavioral model comprises at least one of a multimodal foundational model or a large-scale language model.
13. The software code, executed by the hardware processor, receives qualitative feedback describing multiple interaction experiences from multiple human users. The software code executed by the hardware processor and using a first machine learning (ML) model of multiple trained ML models segments the qualitative feedback into multiple segments, each corresponding to a single interaction specific topic. The steps include analyzing the emotions of each of the one or more human users regarding each interaction-specific topic using the software code executed by the hardware processor, and using the qualitative feedback and at least one of the first or second ML models of the multiple trained ML models, The software code executed by the hardware processor acquires interaction quality evaluation data generated by the control system, The software code executed by the hardware processor, using the interaction quality evaluation data, and based on the analysis step, identifies adjustment data to improve the interaction performance of the character; The method according to claim 8, further comprising the step of tuning the behavior model or one of the modified behavior models by software code executed by the hardware processor based on the adjustment data.
14. The software code executed by the hardware processor further includes the step of receiving quantitative feedback data that evaluates the multiple interaction experiences between the multiple human users and the character, The method according to claim 13, wherein the step of identifying the adjustment data further uses the quantitative feedback data.
15. A computer-readable non-temporary medium for storing instructions, which is executed by a hardware processor. A step of receiving interaction data, wherein the interaction data identifies an action and a plurality of personality profiles corresponding to a plurality of participant cohorts in the action, The steps include generating an interaction graph of the behaviors of the multiple participant cohorts in the operation using the aforementioned interaction data, To provide a predicted interaction graph for the multiple participant cohorts, the steps include simulating the participation of each of the multiple participant cohorts in the operation using a behavioral model, The steps include comparing the predicted interaction graph with the generated interaction graph to determine the similarity score of the predicted interaction graph to the generated interaction graph, If the similarity score satisfies the similarity criteria, the step is to use a character behavior model for interaction, A computer-readable non-temporary medium that instantiates a method having the step of modifying the behavioral model based on one or more differences between the predicted interaction graph and the generated interaction graph if the similarity score does not satisfy the similarity criteria.
16. If the similarity score does not satisfy the similarity criteria, the simulation step and the comparison step are repeated until another similarity score for another predicted interaction graph satisfies the similarity criteria. The computer-readable non-temporary medium according to claim 15, further comprising the step of using a modified behavioral model for the interaction that provides the other predicted interaction graph.
17. The computer-readable non-temporary medium according to claim 15, wherein the operation includes at least one of the same speech directed to each of the multiple participant cohorts, or the same event in which each of the multiple participant cohorts is involved, and the predicted interaction graph includes a plurality of edges that each identify one or more intentions of the multiple participant cohorts.
18. The computer-readable non-temporary medium according to claim 17, wherein each of the multiple edges further identifies at least one of the vocalizations, facial expressions, or gestures originating from one or more of the multiple participant cohorts.
19. A step of receiving qualitative feedback that describes multiple interaction experiences from multiple human users, A first machine learning (ML) model of multiple trained ML models segments the qualitative feedback into multiple segments, each corresponding to a single interaction specific topic. The steps include analyzing the emotions of each of the one or more human users regarding each interaction-specific topic by using the qualitative feedback and at least one of the first ML model or the second ML model, The steps include acquiring interaction quality evaluation data generated by the control system, Based on the analysis step described above, the step of identifying adjustment data to improve the interaction performance of the character, The computer-readable non-temporary medium according to claim 15, further comprising the step of adjusting one of the behavioral models or a modified version of the behavioral model based on the adjustment data.
20. The method further includes the step of receiving quantitative feedback data that evaluates the multiple interaction experiences between the multiple human users and the character, The step of identifying the adjustment data further uses the quantitative feedback data, according to claim 19, in a computer-readable non-temporary medium.