A
system includes a hardware processor configured to execute a
machine learning (ML) model training pipeline to
train an ML model using data relevant to a world of a digital persona to provide a dialogue model, generate, using the dialogue model, first conversational outputs,
train the dialogue model, based on the first conversational outputs, to avoid hallucinations and / or undesirable expressions to provide a guardrailed dialogue model, generate, using the guardrailed dialogue model, second conversational outputs,
train the guardrailed dialogue model, based on the second conversational outputs and persona data identifying interaction characteristics of the digital persona to provide a persona-
specific model, generate, using the persona-
specific model, a response to a scripted question, determine a
quality score for the response, and further train the persona-
specific model or validate the persona-specific model for
human interaction, depending upon whether the
quality score fails to satisfy or satisfies a quality criterion.