Directional management of interactive elements in interactive environments utilizing machine learning

The director service facilitates the integration of a generative ML model with interactive environments, addressing the limitations of existing ML models by generating contextually appropriate outputs, enhancing user interaction and reducing manual development efforts.

JP2026509706APending Publication Date: 2026-03-25MICROSOFT TECHNOLOGY LICENSING LLC
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-02-20
Publication Date
2026-03-25

AI Technical Summary

Technical Problem

Existing machine learning models struggle to effectively interact with users to generate relevant, reproducible, and consistent results within environmental guidelines independently of the designer or system management, limiting the creativity of content designers in interactive environments.

Method used

A director service integrates a generative machine learning model with interactive environments, processing user and developer inputs to generate prompts that conform to environmental guidelines, ensuring the model output aligns with the interactive environment's context and intent, thereby reducing the need for manual development of dialogues and action trees.

Benefits of technology

This approach allows the ML model to adapt interactively and respond to a wider range of inputs in a more realistic and creative way, reducing the manual software development burden on developers while ensuring consistent and contextually appropriate interactions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026509706000001_ABST
    Figure 2026509706000001_ABST
Patent Text Reader

Abstract

This disclosure relates to a system and method for using a director service as an intermediate management system for integrating interactive elements between a developer, a user, a generative machine learning (ML) model, and / or an interactive environment. In some examples, the director service may receive input from a user or developer device related to an interactive element from the interactive environment. The director service can process input from one or more of the developer, user, and interactive environment to recognize the semantic context and intent objectives associated with the input. The director service may generate one or more prompts based on such input, which are processed by the ML model to produce an output. In some examples, prompts may be provided to the ML model to instruct it to provide an output that is conforming to the input and one or more environmental guidelines. The input and / or output may be multimodal.
Need to check novelty before this filing date? Find Prior Art

Description

Background Art

[0001] Background In recent years, as the Internet has spread and network functions have improved, more emphasis has been placed on the improvement of content provision (e.g., video games, virtual reality, multimodal entertainment experiences, etc.) with continuous and realistic interactive elements that simulate real-world interactions within a larger context of content. However, managing these interactive elements of content is a difficult task, and often developers need to write extensive turn-by-turn dialogues and action trees for interactive content to be effective in user-content interaction scenarios within the content. Without consistently managing and updating the dialogues and action trees, there is a risk that the interactive content will not function properly within the context of the content. To address this problem, some developers are turning to machine learning (ML). However, such ML models may not be able to effectively interact with users to generate relevant, reproducible, and consistent results within environmental guidelines independently of the designer or system management. Therefore, without further technological progress, the creativity of content designers is limited by the inability to efficiently apply the advantages of LLMs.

[0002] Regarding these and other general considerations, some examples will be described. Also, although relatively specific problems are discussed, it should be understood that these examples are not limited to solving the specific problems identified in this background section.

Summary of the Invention

Means for Solving the Problems

[0003] Summary Aspects of this disclosure relate to systems and methods for using a director service as an intermediate management system for integrating interactive elements between a developer, a user, a generative machine learning (ML) model, and / or an interactive environment. In some examples, the director service may receive input from a user or developer device related to an interactive element from the interactive environment. The director service can process input from one or more of the developer, user, and interactive environment to recognize the semantic context and intent objectives associated with the input. The director service may generate one or more prompts based on such input, which are processed by the ML model to produce an output. In some examples, the ML model may be provided with prompts to instruct it to provide an output that conforms to the input and one or more environmental guidelines. The input and / or output may be multimodal. In some examples, the director service may evaluate and modify the ML model output to ensure that it conforms to the input and environmental guidelines, and then provide that ML model output for use in influencing the operation of the interactive environment.

[0004] Therefore, as described above, the director service, in combination with the ML model, can replace or otherwise complement the use of comprehensive dialogs and action trees for effectively controlling interactive elements. This disclosure enables the ML model to be contextually integrated with interactive elements in substantially real time, allowing it to perform multiple tasks without requiring expensive and time-consuming learning to fine-tune the ML model for direct interaction. As a result of the model output facilitated by the director service, aspects of the interactive environment can be adapted, thereby reducing or eliminating the burden on developers to manually develop / create aspects of the environment development. Ultimately, this disclosure provides users with a richer experience, enabling interactive elements to adapt and respond to a wider range of inputs and / or in a more realistic and creative way within the interactive environment, while simultaneously reducing the amount of manual software development required of developers of the interactive environment.

[0005] This summary section is provided to provide a simplified introduction to various concepts that will be further discussed below in the detailed description section. This summary section is not intended to identify any major or essential features of the claimed subject matter, nor is it intended to be used to limit the scope of the claimed subject matter. Additional aspects, features, and / or advantages of some examples are described below, some will become apparent from the description, and some may be understood through the practice of this disclosure.

[0006] Brief explanation of the drawing Refer to the following diagrams to describe non-exclusive and non-exclusive examples. [Brief explanation of the drawing]

[0007] [Figure 1] This figure shows a system for using a director service to integrate ML models in an interactive environment, as described herein. [Figure 2]This block diagram shows a method for integrating an ML model with an interactive environment according to the embodiments described herein. [Figure 3] This figure shows an exemplary conversational agent evaluation portal according to the embodiments described herein. [Figure 4] This block diagram shows a method, according to the embodiments described herein, for evaluating the output of an ML model using an ML model and providing an evaluation of the output of the ML model itself. [Figure 5A] This figure shows an overview of an exemplary generative machine learning model that may be used according to the embodiments described herein. [Figure 5B] This figure shows an overview of an exemplary generative machine learning model that may be used according to the embodiments described herein. [Figure 6] A block diagram showing exemplary physical components of a computing device in which aspects of this disclosure may be implemented. [Figure 7] This is a simplified block diagram of a computing device in which the embodiments of this disclosure may be implemented. [Figure 8] This is a simplified block diagram of a distributed computing system in which the embodiments of this disclosure may be implemented. [Modes for carrying out the invention]

[0008] Detailed explanation The following detailed description will illustrate specific examples or embodiments with reference to the accompanying drawings that form part of this specification. These embodiments may be combined, other embodiments may be utilized, and structural modifications may be made without departing from the disclosure. Examples may be implemented as methods, systems, or devices. Thus, examples may take the form of hardware implementations, complete software implementations, or implementations combining software and hardware embodiments. Accordingly, the following detailed description should not be constrained to be restrictive, and the scope of this disclosure is defined by the appended claims and equivalents.

[0009] Figure 1 shows a system for using Director Services to integrate ML models in an interactive environment, according to the embodiments described herein. In some examples, System 100 includes one or more user devices 102, one or more developer devices 108, a data store 112, Director Services (DS) 120, an interactive environment 140, and a network 150. User devices 102 may include one or more applications 104 and an integration manager 106. Developer devices 108 may include one or more applications 105. DS 120 may include a Director Services Manager (DSM) 122, a scenario processor 124, an intent goal processor 126, a prompt generator 128, a model repository 130, and an output evaluator 132. User devices 102, developer devices 108, applications 104, applications 105, and data store 112 are referred to as plural because in some examples it is preferable to include multiple of these elements to accommodate various types and quantities of use. However, for the sake of clarity, each element will be referred to in the singular form in the description herein. However, the characteristics and examples of each element are applicable to multiple instances. Furthermore, while DS120 is shown as containing a single instance of elements 122-132, it should be understood that any other number of such elements may be used in other examples. Moreover, such elements may be implemented in various other computing devices, such as computing device 102, as additions or substitutions.

[0010] As shown in the figure, the user device 102, application 104, integration manager 106, developer device 108, application 105, data store 112, DS 120, DSM 122, scenario processor 124, intent goal processor 126, prompt generator 128, model repository 130, output evaluator 132, and interactive environment 140 communicate via network 150. Network 150 may include one or more networks such as a local area network (LAN), wide area network (WAN), enterprise network, or the Internet, and may include one or more of the wired, wireless, and / or optical components.

[0011] The interactive environment 140 may consist of multiple environments (both physical and / or virtual environments), and users and / or developers may be able to interact with the interactive elements of the interactive environment verbally or nonverbally. In this sense, the director service 120 functions to coordinate and integrate the interactive elements of system 100 within the interactive environment 140, according to the embodiments described herein. An interactive element is an embodiment of the interactive environment 140 in which one or more of the user, developer, and / or DS120 can interact. Those skilled in the art will understand that there are multiple different types of interactive environments containing the relevant interactive elements, and that this specification provides non-exclusive and non-limiting examples.

[0012] Inputs can be states relating to the interactive environment 140, as well as implicit and / or explicit requests for the system to perform operations and / or processes. Explicit requests can be direct input from users and / or developers. For example, among many examples, in a gaming environment, a user might ask an NL question to an NPC expecting some response from the NPC; in a video environment, a developer might request the generation of a scene relating to a 19th-century American West town; and in a manufacturing environment, a combination of images and NL inputs might be requested to design a schematic diagram of a new part. Implicit inputs can be aspects of the interactive environment 140 that are collected and referenced as triggers for potential outputs. For example, among many examples, in a gaming environment, a user achieving some goal might be collected as a trigger for some kind of celebration; in a shipping and handling environment, the state of an interactive element, such as a truck being empty and / or full, might trigger an action; and in a contact center environment, input might be that a caller is internally routed to a specific department, triggering one or more disclosure statements. As further examples, input may include keyboard or controller input, speech and / or other audio recognition including intonation features, and visual input including gestures and / or facial expressions from the user, developer, and / or other interactive elements of the interactive environment 140.

[0013] The output may also be a multimodal output relating to interactive elements within the interactive environment 140, such as NL output associated with NPCs and / or other artificial intelligence (AI) agents, program code executed or otherwise parsed to modify any interactive element of the interactive environment 140, textures, schematics, part designs, scheduling and / or ordering of events and / or actions, action and / or generation of scenes or characters, text NL output, text-to-speech output, and several other examples. Thus, the output may take on several different forms or embodiments based on the type of the interactive environment 140, the interactive elements involved, the type of input provided, and / or one or more of the additional processing performed by the DS120 described herein.

[0014] For example, the interactive environment 140 may be a gaming environment, such as a video game, online game, MMORPG, virtual reality gaming environment, and / or other game-like experiences that can be played or otherwise experienced via a user device 102 and developed and / or modified on a developer device 108. The director service 120 may be used to integrate gameplay with multiple interactive elements. Interactive elements within the game environment may be non-player characters (NPCs), animated infographics, videos, images, quizzes, game objects (e.g., hammers, vehicles, buildings, weapons, etc.), and / or any other form of the interactive environment that a user / developer may access and interact with.

[0015] In another example, the interactive environment 140 may be an industrial or manufacturing environment, and in aspects of this disclosure, a user can develop a manufacturing process by managing the industrial or manufacturing process via a user device 102 or a developer device 108. The interactive elements in this environment may be one or more automated manufacturing components or machines used in the process (e.g., 3D printers, CNC machines, automated handling machines, etc.), and in some cases include aspects of machine learning and / or artificial intelligence. In this example, the director service 120 can be used to control and integrate various manufacturing components to produce a desired product.

[0016] In a further example, the interactive environment 140 may be a processing facility for shipping and handling goods, equipped with an automated management system for organizing the facility. The interactive elements in this environment may be one or more automated machines used in the shipping and handling process, which in some cases include aspects of machine learning and / or artificial intelligence. In this example, the director service 120 can be used to integrate the shipping and handling process.

[0017] In another example, the interactive environment 140 could be a scene generator for multimedia entertainment developers. Examples of multimedia include, but are not limited to, television programs, movies, content created for online social networking platforms, websites, and / or applications for mobile devices. Interactive elements may be, alone or in combination, characters, background settings, animated or computer-generated objects within multimodal entertainment (e.g., backgrounds, filters, overlay stickers), audio files, video files, text files, and / or any other multimodal elements. In one case, as in a social networking interactive environment, there may be an input / output flow between the user device 102, DS120, and the interactive environment 140, from the user on the user device 102 who is browsing the interactive environment 140 on the application 104. In other cases, an input / output flow occurs between the developer device 108 and the application 105, DS120, and the interactive environment 140 when a developer creates the interactive environment 140 and / or scenes within it. For example, multimodal entertainment in the interactive environment 140 could be a television program, in which case one or more animated characters and background settings are interactive elements controlled, developed, and / or otherwise adapted using the director service 120. It can receive input and generate a scene containing both audio and visual elements related to the characters and background. The output may be a combination of the generated video, subtitles, and script code to run the video.

[0018] In a further example, the interactive environment 140 can be, for example, a customer service center and / or a contact center for providing customer support. The interactive elements in this interactive environment 140 can be one or more conversation agents that interact with the user via the user device 102 using the disclosed aspects. The director service 120 can process inputs (received, for example, via application 104 and / or application 105), generate prompts for an ML model (such as model repository 130), and generate response outputs for the inputs.

[0019] In another example, the interactive environment 140 can be an automated scheduling system and / or a digital assistant for individuals and / or any larger entity (such as an airline, railway company, shipping company, university, social organization, non-profit organization, etc.) that needs to coordinate the schedules, deadlines, and requirements of multiple individuals and organizations to create a unified workflow. The interactive elements in this interactive environment 140 can be scheduling components that control one or more scheduling entities and choreograph them virtually.

[0020] In some examples, the user device 102 and the developer device 108 can be any of a variety of computing devices including, but not limited to, mobile computing devices, laptop computing devices, tablet computing devices, desktop computing devices, video game computing devices, virtual reality computing devices, and / or any device capable of interacting with the interactive environment 140 and the director service 120.

[0021] The user device 102 and / or the developer device 108 may be configured to execute one or more applications, such as application 104 and / or application 105, to interact with the interactive environment 140 and / or services, and / or to manage hardware resources (such as processors and memories). Application 104 and / or application 105 may be native applications and / or web-based applications. Application 104 and / or application 105 may operate substantially locally with respect to the user device 102 and / or the developer device 108, and / or may operate in accordance with a server / client paradigm in cooperation with one or more servers. Application 104 and / or application 105 may be used for communication via the network 150 for the user / developer to provide input and receive and view output from the interactive environment 140.

[0022] Furthermore, the user device 102 may have an integration manager 106, which manages the flow of information between the user device 102, the director service 120, and the interactive environment 140. The integration manager 106 may hold user-specific data (e.g., related to aspects of the interactive environment 140), which may be stored locally for later use by the director service 120. For example, in a gaming context, the integration manager 106 may hold user-specific data related, among other things, to previous conversation history and / or gameplay states within aspects of the game. Examples of user-specific data may be one or more of the inputs, intents, prompts, prompt templates, and / or model outputs described herein. User data may be substantially local on the user device 102 and / or stored in a data store 112 for access by the integration manager 106 and / or the director service 120. As an example, according to the embodiments described herein, user-specific data such as conversation history can be used as context by the director service 120 during subsequent conversations between the NPC and the user to improve the accuracy and responsiveness of the ML model output regarding the NPC dialogue. Furthermore, in on-device scenarios, the integration manager 106 may manage updates to the application 104 if applicable.

[0023] The developer device 108 is substantially the same as the user device 102. The differences between the developer device 108 and the user device 102 may include how the developer, administrator, or creator of the interactive environment 140 uses the developer device 108 as a platform to access and / or modify the interactive elements, at the same time that the user accesses and uses the interactive environment 140 and / or the interactive elements within the interactive environment 140.

[0024] As an example, a user and / or developer may access applications 104 and / or 105 on computing device 102 and / or developer device 108 to provide input. While system 100 is described in an example where input is obtained via applications 104 and / or 105, it should be understood that input can also be obtained from various other sources. For example, input may be programmatically generated by applications 104 and / or 105, or it may be based on file contents or electronic communications, among other examples.

[0025] Users and / or developers may provide input (for example, to user device 102 and developer device 108, respectively) as language input, text input, program code, and / or any of various other inputs. For example, input may be provided via one or more input devices of computing device 102 not shown in Figure 1 (e.g., a microphone, camera, keyboard, upload of images or videos from local storage or data store 106).

[0026] User / developer input may also refer to previously created and / or known interactive elements (e.g., those stored in datastore 112). In some examples, user / developer input may include indications of applications, data formats, and / or other specification types that may generate output accordingly. Thus, input may invoke outputs of varying complexity based on the specificities and details provided by the user / developer.

[0027] DS120 is an interface that facilitates integration between the user device 102, the developer device 108, the ML model, the interactive elements, and the interactive environment 140. For example, DS120 is an interface between the interactive environment 140 and the ML model (e.g., in the model repository 130), thereby allowing the state of the interactive environment 140 to be processed by the ML model to generate an output, which can then be used to influence the environment accordingly (e.g., by the integration manager 106). As a result, DS120 can mitigate the constraints / barriers associated with the use of machine learning to adapt aspects of the interactive environment. Furthermore, the disclosed aspects reduce the technical burden on developers because they do not need to focus on the technical aspects of the interactive environment (e.g., rules, boundaries, and / or specific mechanisms), but instead only need to describe aspects of the interactive environment, thus allowing DS120 to generate model output, which in turn allows the interactive environment 140 to adapt accordingly.

[0028] DSM122 can receive user / developer input. DSM122 receives the input, performs a systematic contextual analysis, and accordingly provides environmental guidelines to other elements of DSM120. In some examples, DSM122 can associate the input with one or more environmental guidelines created by the developer and / or DSM122 for the interactive environment 140. Environmental guidelines provide a systematic context for noticing, organizing, and / or constraining the available options with respect to both prompts and ML model outputs, within the scope of the developer's intent for the interactive environment 140. In some examples, environmental guidelines may be rules that alter the applicability of the output, such as hard rules and / or soft rules. Hard rules have greater applicability as explicit rules that can prohibit or exclude certain types of output. For example, in a gaming context, there might be a hard rule that information about the next quest cannot be provided until a specific task is completed in the game. In this example, if the user input is a recurring question to an NPC about the next quest, the environmental guidelines would include a hard rule preventing the NPC from repeatedly providing information about the next quest in the output. A soft rule is a variable, enforceable rule intended to provide a certain level of attention, organization, and / or constraint on the output. For example, in a gaming context, a soft rule might be designed to guide the user to a specific activity or quest based on the gameplay state. In this case, repeated user input such as "What should I do next?" could generate an output based on a soft rule that hints or suggests to the user that a particular activity is what they should do next.

[0029] The environmental guidelines may contain multiple pieces of information, such as one or more policies for the interactive environment 140 (e.g., toxicity policy, diversity policy, standard operating procedures, and / or other general policies that may relate to input), a list of available interactive elements, and the status of various interactive elements with respect to input. The environmental guidelines may also include one or more NL rules for use in processing natural language input and for use in providing NL output. For example, in a gaming environment, NL rules can be used to prevent potentially offensive, discriminatory, and / or hateful output from the ML model. In another example, NL rules may exist to define appropriate responses to offensive, discriminatory, and / or hateful input. The environmental guidelines may also include a restricted list of information and / or items that are restricted from being included in prompts to the ML model and / or included as output to users and / or developers. The restricted list may serve to protect intellectual property, protect certain aspects of the storyline to avoid spoilers, protect trade secrets, and / or any other information related to the information that developers and / or organizations do not want to disclose beyond a certain level of access. In some cases, DSM122 can use one or more NL processing tools, rule-based analysis, and / or ML models, as described herein, to analyze one or more parts of the input and define appropriate environmental guidelines and scheduling functions.

[0030] In some examples, environmental guidelines may be a list and / or systematic information applicable to each input. In other examples, environmental guidelines may be structured as a tree containing nodes and branching logic, where a node may have a set of related environmental guidelines that determine what should be provided with the input (e.g., whether the objective has been achieved for the user to proceed) and / or whether prompt generation should be "prompted" and / or directed in a particular way.

[0031] Environmental guidelines may differ based on the type of interactive environment 140 in which the director service 120 is used. For example, if the interactive environment 140 is a gaming environment and the inputs correspond to NPCs, the environmental guidelines identified by DSM122 to generate outputs may include toxicity policies that limit the outputs, and thus, for example, one or more acceptable responses may be defined for a particular type of input. Environmental guidelines may also include systematic gameplay parameters, for example, particularly concerning gaming environments beyond the user's scope, which may relate, among other things, to notifying prompts to ML models.

[0032] The scenario processor 124 collects and analyzes specific contexts that can be used for prompt generation, for example, aspects of the specific context relating, in particular, to one or more inputs concerning the user / developer and / or one or more interactive elements within the interactive environment 140. Therefore, before, during, and / or after receiving input, the scenario processor 124 can access one or more contextual indicators associated with the interactive environment 140 to provide contextual indicators for prompt generation to the intent / goal processor 126 and prompt generator 128. Thus, the scenario processor 124 can use one or more tools (e.g., NL processing tools, extended focus tracking including sentiment, video / image analysis and processing tools, ML models, etc.) to generate or otherwise obtain contextual information about various interactive elements and contextual indicators related to the interactive environment 140. The specific contexts collected will differ based on the interactive environment 140 utilizing the DS 120. In some examples, the scenario processor may perform tracking of user / developer and interactive element functions and / or actions within the input scenarios of the interactive environment 140. In this sense, an action can be either an input to the DS120 and / or an output from an interactive element (for example, the start of a process by a manufacturing component needs to be tracked, the departure of a vehicle from a shipping location may be an action that needs to be tracked, or an NPC in a video game may throw an object into the sea in response to an input, which needs to be tracked).

[0033] For example, in a gaming context, the scenario processor 124 performs enhanced focus tracking with sentiment analysis of the user character, as well as other players, NPCs, and / or other objects and interactive elements in the input scene. These specific contextual indicators may include, among other things, audio and / or visual indicators from the player's character (e.g., player's facial expressions, tone of voice, posture, hand movements), from the interactive environment 140 (e.g., combat sounds, laughter, or silence within the environment, facial features of the user's companions in the game), and / or from the game state (e.g., in a sports game, time remaining in the match, position on the field, penalties, etc.). Furthermore, the scenario processor 124 can access or otherwise obtain specific context (e.g., from the integration manager 106 and / or data store 112) regarding the user's game state (e.g., game progress, goal tracking and achievement, user game level, user experience, available user tools and / or equipment), and / or any other information related to the user's gameplay. If a user is participating in a team scenario (e.g., a team quest, sports game, racing game, etc.) or interacting with an interactive element associated with a specific game state, the scenario processor may also collect similar specific context related to the game states of the user's teammates.

[0034] Continuing with the gaming example, the scenario processor 124 can also leverage recent and past conversation / dialogue history between the user, interactive elements, and / or optionally teammates to provide prompts and advance the interaction between the user and the interactive elements. For example, over a multi-turn dialogue, an NPC might be instructed by DS120, partly based on a specific context associated with the conversation history, so that within the game context, it responds to repeated inputs by saying, "I've told you 100 times already." This specific context can be used for various purposes during gameplay, such as providing hints and / or guidance in other conversations to guide the user toward the next goal and / or task completion. In another example, the user's game state and player-level specific context can be leveraged to prevent the user from receiving information in the output that could provide insights into gameplay the user has not yet accessed.

[0035] As another example, in an interactive environment of a contact center, input may include audio information about the user's tone of voice and / or, if available, the user's facial features indicating an emotional state (e.g., anger, frustration, joy, understanding, confusion, etc.). In addition to the input provided directly, certain contextual indicators that may be collected by application 104 and / or application 105 and / or director service 120 are useful for providing additional contextual information that can be used to refine and improve prompt generation.

[0036] Furthermore, the scenario processor 124 can store one or more of the user / developer inputs, intents, prompts, prompt templates, and model outputs in the data store 106 as semantic contexts and / or known interactive elements (for example, when model outputs generate new ones and / or modify existing interactive elements), which can then be used for later input analysis. In some examples, the scenario processor 124 may utilize one or more ML models trained to identify specific contexts or intents, rule-based processes for identifying contexts based on received inputs, or any other type of application or process capable of parsing and analyzing inputs to determine contexts and / or intents based on user / developer inputs.

[0037] The intent-goal processor 126 receives one or more of the input, systematic context, and environmental guidelines from the DSM 122, and / or specific context from the scenario processor, and uses them to determine one or more intent goals associated with the input. Once determined, the intent goals can be used to encapsulate the overall intent of the input, the requested task, and / or specific meaning, and to help generate one or more prompts related to those intent goals. To generate intent goals, the intent-goal processor 126 analyzes one or more of the input, systematic context, specific context, and / or environmental guidelines. In one example, the intent-goal processor 126 may use a rule-based approach, where the input, systematic context, specific context, and / or environmental guidelines are analyzed based on a set of rules to determine the intent goals. In another embodiment, a semantic encoding model can be used to determine the semantic context associated with the input, systematic context, specific context, and / or environmental guidelines to determine the intent goals. A semantic encoding model may determine one or more semantic components of input, systematic context, specific context, and / or environmental guidelines, and process these semantic components to generate one or more intent goals that describe the underlying intent of the user / developer input.

[0038] In a further example, inputs, a systematic context, a specific context, and / or environmental guidelines are processed, and intent goals are determined based on the language used therein. In one example, intent goals may be determined by a program that processes the inputs, the systematic context, the specific context, and / or environmental guidelines. In an additional embodiment, intent goals can be determined using one or more embeddings. Embeddings may be generated with respect to the inputs, the systematic context, the specific context, and / or environmental guidelines, with a single embedding describing each element. Alternatively, embeddings may be generated individually and / or for one or more parts of each of the inputs, the systematic context, the specific context, and / or environmental guidelines based on the desired granularity within the system. The embeddings can then be used to identify one or more intent goals associated with the semantics from the data store 112, which in this case may be configured as an embedded object memory. The intent goals associated with the semantics can then be analyzed and refined to determine one or more intent goals with respect to the input.

[0039] Furthermore, the intent goal processor 126 may be able to query data structures (e.g., those held in the data store 112 and / or integration manager 106) to determine and / or obtain further context to the elements of the input. This enables request-and-generate type input completion by querying the database 112 of known information that can be used to create the intent goal, rather than relying on potentially incorrect information. For example, the intent goal processor 126 may be able to look up recipes associated with the input, verify facts, refer to known objects and / or other stored information related to the input. The type of information that the intent goal processor 126 can access and / or look for will vary depending on the interactive environment 140. For example, in a shipping and handling environment and / or scheduling environment, the intent goal processor 126 may need to process an input that asks, "How much would it cost to move 5,000 SUVs from Mexico to Canada?" by referring to the storage capacity of a particular vehicle (e.g., a truck, airplane, storage container, ship, etc.). In a game environment, if the input is a question about an NPC asking "Where are my teammates?", the intent goal processor 126 may need to define the intent goal by referring to the game state information maintained in the integration manager 106 to identify who the user's teammates are and their current locations in the data store 112. Numerous possible examples based on the type of interactive environment 140 will be understood by those skilled in the art.

[0040] The prompt generator 128 can receive one or more inputs, systematic contexts, specific contexts, environmental guidelines, and / or intent goals, and use them to generate one or more prompts regarding the ML model. The prompt generator 128 can generate one or more prompts, which, when processed by the ML model, cause the ML model to generate outputs that are appropriate to one or more intent goals associated with the input, thereby allowing the interactive environment to be adapted according to the embodiments described herein. That is, the generated prompts enable the ML model to process contexts associated with the input that would normally be unknown. Thus, one or more prompts are used by the ML model to generate outputs that are appropriate to the input without requiring additional training or fine-tuning of the ML model before generating model outputs. In this sense, the output may be one or more of the following types of outputs: text files, audio files, images, videos, NL outputs (e.g., language and / or non-language), programming languages ​​(e.g., code), and / or any other types of outputs that can operate the interactive elements of the interactive environment 140 in a manner appropriate to the user / developer's input within the context provided by DS120.

[0041] Furthermore, the prompt generator 128 may evaluate the input, systematic context, specific context, and / or environmental guidelines to determine whether some or all of the input is a known intent objective. A known intent objective may have associated prompts stored in the data store 112, and / or the general ML model may be able to process the intent objective directly without requiring new prompt generation. This enables request-generate type input completion by querying the database 112 for known prompts, rather than generating potentially incorrect prompts. If the input contains one or more known prompts, the prompt generator 128 may retrieve the known prompts from the data store 112 and / or the integration manager 106, and if no additional prompt generation by the prompt generator 128 is required, the known prompts can be passed to the ML model for processing. If additional prompt generation is required, the prompt generator 128 may generate additional prompts according to the embodiments described herein and pass both the generated prompts and the known prompts to the ML model for processing.

[0042] In some examples, a prompt may consist of multiple prompt templates. Prompt templates can include, but are not limited to, any of a variety of data, including, in some examples, natural language, image data, audio data, video data, and / or binary data. In some examples, the type of data may depend on the type of ML model used to respond to the received input. One or more fields, regions, and / or other parts of a prompt can be populated with one or more prompt templates containing input and / or context, thereby generating a prompt that can be processed by the ML model in the model repository 130, in accordance with the embodiments described herein. In additional examples, prompt templates may include known entities, previously stored objects, and / or previously created or inputted into system 100, thereby allowing the user to refer to and further use any of the previously created model outputs and / or various other contents. For example, the data store 112 and / or integration manager 106 may contain one or more embeddings associated with previously generated model outputs and / or previously processed inputs, thereby enabling semantic retrieval of prompt templates and associated contexts (e.g., thereby allowing iteration of previously generated model outputs). In some embodiments, this involves selecting the most relevant portion of previously generated prompts and using it as input to an ML model. In some cases, only a small portion of a previously known prompt template may be sufficient to generate a suitable model output, and complete prompt generation may not be required.

[0043] In some embodiments, the prompt generator 128 may utilize one or more prompts instead of inputs to the ML model. In another example, the prompt generator 128 may provide one or more prompts in addition to inputs. Prompts can be generated in various ways. In one example, an application and / or ML model from the model repository 130 may analyze the intent and input to select from one or more prompt templates stored in the data store 106 and / or integration manager 106 and input them as prompts. In another example, an NL processing tool may be used to analyze the input and intent and goal to determine one or more relevant prompt templates and input them as prompts.

[0044] In another embodiment, the prompt generator 128 may associate one or more prompt templates with at least a portion of the input and / or intent objectives, and take each prompt template as input to generate one or more prompts accordingly. The prompt templates may contain semantic information, which, when combined with a prompt, encapsulates the specific context used by the ML model to generate model outputs in response to the input. One or more prompt templates may be retrieved by the prompt generator 128 from the data store 112 and / or the integration manager 106. In another example, embeddings may be generated individually and / or collectively with respect to the input and / or intent objectives. One or more embeddings may be stored in the data store 112 and / or the integration manager 106 for future use. The embeddings can be used to identify the prompt templates associated with the semantics from the data store 106, which is designed as an embedding object memory. The prompt templates associated with one or more semantics can be used to generate one or more prompts to be processed by one or more ML models from the model repository 130. In further embodiments, prompts can be generated by training or otherwise using ML models stored in the model repository 130. In this example, the trained ML model can be used to process one or more inputs, systematic contexts, specific contexts, environmental guidelines, and / or intent goals, and output one or more prompts that are appropriate to the input.

[0045] In some examples, the prompt generator 128 may include a summarization function to control the length of the prompt. In certain examples, the ML model may have tokens or other input restrictions that can prevent excessively long prompts from being input to the ML model. In this case, the prompt generator 128 may modify the prompt to shorten its length, prevent overflow, and reduce token costs. In some examples, this may involve summarizing one or more parts of the prompt.

[0046] In some examples, the prompt generator 128 can also be used to determine whether a prompt complies with environmental guidelines, including protecting certain aspects of intellectual property from being exposed to the ML model. If the aspects of a prompt do not comply with environmental guidelines, the prompt generator 128 may modify the prompt and / or data mask certain information by subsequent analysis before the ML model input and / or after the ML model output. Data masking may be performed in accordance with environmental guidelines, including the restriction list discussed above. Data masking may be performed so that certain information is not permitted to be included in the prompt and, in some cases, not to be exposed to the background ML model used to generate the response output.

[0047] As illustrated, the director service 120 includes a model repository 130 which may contain any of the various ML models. Generative models used in accordance with the embodiments described herein (also generally referred to herein as a type of ML model) can generate any of the various output types (and thus, in some examples, may be multimodal generative models), and may, among other examples, be generative transformer models, large language models (LLMs), and / or generative image models. Exemplary ML models include, but are not limited to, GPT-3 (Generative Pre-trained Transformer 3), BigScience BLOOM (Large Open-science Open-access Multilingual Language Model), DALL-E, DALL-E2, Stable Diffusion, or Jukebox. Additional examples of such embodiments are discussed below with respect to the generative ML models shown in Figures 5A-5B. Additional or alternative, one or more recognition models (or any of the various other types of ML models) may generate their own outputs which are processed as part of the skill chain in accordance with the embodiments described herein.

[0048] The output evaluator 132 receives the output from the ML model, evaluates the output for its suitability to the input, and can further ensure that the output meets the environmental guidelines from DSM122. In this case, the suitability evaluation determines whether the output produces an interactive element of the interactive environment 140 that satisfies the input. In some examples, the output may be suitable and meet the environmental guidelines if it is returned to the integration manager 106, application 104, and / or application 105 for use in the interactive environment 140. In some examples, before returning the output to the user, the output evaluator 132 may determine that the output is unsuitable for the input and / or unsuitable for the input. In some examples, this may be, among other things, a result of the output failing to meet a predetermined confidence threshold, or an indication of an error or other problem received (for example, as a result of processing at least part of the output, such as if the output contains code or other output that is syntactically incorrect or otherwise malformed). If a confidence threshold is used, the output evaluator 132 can receive the confidence threshold from the developer, score the output based on one or more evaluation metrics, and then compare the output score to the threshold.

[0049] In some examples, the evaluation may generate a single confidence score for the entire output and / or confidence scores for individual components of the output, where the confidence scores represent the suitability of the output or output components to the input and ensure that the output or output components meet environmental guidelines. In some examples, this may be done by individually generating one or more confidence scores for one or more components of the output and comparing one or more confidence scores to a threshold. The confidence scores may be determined using evaluation metrics developed for the ML model and the output. Component confidence scores may be a measure of the suitability of the output to the input and the degree to which the output meets environmental guidelines based on one or more evaluation metrics. For example, in a scene generation scenario, the output may include elements of the generated video, subtitles corresponding to the generated video, audio components corresponding to the video, and / or script code to execute the video, subtitles, and / or audio components. A confidence score may be generated for each component of the output (e.g., video, audio, subtitles, and / or script code) and evaluated by comparison with a threshold for suitability to the input and the degree to which the output meets environmental guidelines based on one or more evaluation metrics. In a further example, the confidence scores of individual components can be combined to create a composite confidence score, which can then be compared to a threshold to further evaluate the output.

[0050] If output fails, the output evaluator 132 may restart the output generation process so that an alternative output is created. In other examples, the output evaluator 132 may provide failure notifications to applications 104 and / or 105, displaying them to the user and / or developer, for example, that the user / developer may retry or reconfigure the input, that the input was not correctly understood, or that the requested functionality may not be available. While illustrative problems and associated problem-solving techniques have been described, it should be understood that in other examples, any of the various other problems and / or problem-solving techniques may arise / be used.

[0051] In some cases, the output evaluator 132 may determine that the output does not meet environmental guidelines. This may result from the output containing too much information that should not be conveyed to the user (for example, in a gaming scenario, this may result in gameplay spoilers; in several scenarios, the output may contain aspects of intellectual property that should be excluded; and there may also be aspects of the output that do not meet the toxicity policy of the interactive environment 140). To make a determination, one or more NL processing and / or other techniques may be used to analyze the output. In such cases, the output evaluator 132 may, if possible, modify the output to remove or data-mask elements within the output that do not meet environmental guidelines. As another example, the output evaluator 132 may restart the process of generating prompts and outputs in accordance with the aspects described herein.

[0052] For example, the output evaluator 132 may include an option to receive feedback from the user device 102 and / or the developer device 108 regarding the output received and executed by the interactive element. In some cases, this feedback may be received via a conversational agent evaluation portal, such as a chat window and / or other user interface, which allows the user / developer to provide feedback on the experience of the interactive element and the suitability of the output provided. In this case, the user / developer may interact with a human representative and provide feedback by answering a series of questions. In other cases, the user / developer may interact with a conversational agent that may utilize one or more aspects of an ML model to participate in a feedback session. As an addition or alternative, feedback may be obtained via implicit signals associated with the interaction between the user and the environment.

[0053] In further cases, the output evaluator 132 can directly query the ML model that generated the output, asking it about the suitability of the output to the input and / or whether the output meets environmental guidelines. In this case, the output evaluator 132 can instruct the prompt generator 128 to generate and provide one or more prompts to the ML model, process one or more prompts, and provide an output, where the model explains why it provided the output and why it considered the output to be suitable and satisfactory. This feedback conversation history may be stored in the data store 112 and / or fed back to the ML model as one or more prompts, thereby allowing the ML model to improve and refine its process by evaluating its own behavior and output. In some cases, one or more metrics may be developed and provided to the output evaluator 132 and / or the ML model to further refine and improve the feedback and evaluation process.

[0054] In another example, the effectiveness of the model can be evaluated by utilizing a conversational agent evaluation portal outside the interactive environment 140 to grant the user / developer access to one or more interactive elements connected to DS120. In this case, to evaluate the ML model, the user / developer can provide input in the form of a chat and interact with the interactive elements (e.g., NPCs or scheduling systems, depending on the type of interactive environment). In some examples, the user / developer may fill out a questionnaire answering questions to evaluate the responses of the interactive elements. The feedback can be used to update prompt generation, ML model output, and / or other aspects of DS120 with implicit feedback from the user / developer.

[0055] For example, input may be provided to user device 102 and / or developer device 108 via application 104 and / or application 105 and sent to DS120 for the purpose of generating a model output suitable for the input. In some examples, input may be received, for example, in the chat function of application 104 and / or application 105 used to interact with ML models in model repository 130. In some embodiments, input may be natural language (NL) input provided as voice input, text input, and / or any of various other inputs (e.g., text-based, image, video, etc.) via input devices of user device 102 and / or developer device 108 not shown in Figure 1 (e.g., microphone, camera, keyboard, upload of images or videos from local storage or data store 112). Additional or alternative, input may be programmatically generated by application 104 and / or application 105, may be based on the content of a file or electronic communication, and may include images, other data types, and / or several other examples as understood by those skilled in the art. In some cases, the input may refer to previously created or known entities (for example, those stored in datastore 112). It should be understood that the input does not need to be in a specific format, contain proper grammar or syntax, or include a complete description of the model output that the user intends the ML model to generate. While the amount of detail provided in the input may improve the resulting model output, sparse input is usually sufficient.

[0056] In some embodiments, user device 102 and / or developer device 108 may be a mobile computing device, a desktop computing device, a virtual reality device, a gaming device, and / or a vehicle computer. User device 102 and / or developer device 108 may be configured to run one or more design applications (or “Applications”) such as application 104 and / or application 105 and / or services, and / or to manage hardware resources (e.g., processors and memory) that may be available to users of user device 102 and / or developer device 108. Application 104 and / or application 105 may be native applications or web-based applications. Application 104 and / or application 105 may operate substantially locally with respect to user device 102 and / or developer device 108, or may operate in conjunction with one or more servers (not shown) according to a server / client paradigm. Application 104 and / or application 105 may be used for communication over network 150 for users to provide input and receive and view model output from DS120.

[0057] User device 102 and / or developer device 108 can send and receive content data as input or output, which may be from, for example, a microphone, camera, or Global Positioning System (GPS) that transmits content data, from a computer executable program that generates content data, and / or from memory where data corresponding to content data is stored. Content data may include visual content data, audio content data (e.g., voice or ambient noise), user input such as voice queries or text queries, images, actions performed by the user and / or device, computer commands, programmatically evaluated gaze content data, calendar entries, emails, document data (e.g., virtual documents), weather data, news data, blog data, encyclopedia data, and / or other types of private and / or public data that can be recognized by those skilled in the art. In some examples, content data may include text, source code, commands, skills, or programmatic evaluations.

[0058] The user device 102, the developer device 108, and / or DS120 may each include at least one processor that runs software and / or firmware stored in memory. The software / firmware code includes instructions, which, when executed by the processor, cause the control logic to perform the functions described herein. As used herein, the terms “logic” or “control logic” may include software and / or firmware running on one or more programmable processors, application-specific integrated circuits (ASICs), field-programmable gate arrays (PPGAs), digital signal processors (DSPs), hardwired logic, or combinations thereof. Accordingly, various logics may be implemented in any suitable form according to those examples and maintained according to the examples disclosed herein.

[0059] In some embodiments, the user device 102 and / or the developer devices 108 and DS120 can access the data contained in the data store 112 and may also have the ability to store data in the data store 112. The data store 112 may contain multiple contents related to output generation and providing data to ML models. The data store 112 may be a network server, cloud server, network-attached storage ("NAS") device, or other suitable computing device. The data store 112 may contain one or more storage mechanisms or memories of any type, including memory devices such as magnetic disks (e.g., in hard disk drives), optical disks (e.g., in optical disk drives), magnetic tapes (e.g., in tape drives), random access memory (RAM) devices and read-only memory (ROM) devices, and / or other suitable types of storage media. In some cases, the data store 112 may be configured as embedded object memory. Although only one instance of the data store 112 is shown in Figure 1, the system 100 may contain two, three, or four or more similar instances of the data store 112. Furthermore, in some examples, network 150 may provide access to other data stores similar to data store 112, which are located outside of system 100.

[0060] In some examples, network 150 may be any suitable communication network or combination of communication networks. For example, network 150 may include a Wi-Fi network (which may include one or more wireless routers or one or more switches), a peer-to-peer network (e.g., a Bluetooth network), a cellular network (e.g., a 3G network, 4G network, 5G network, etc., compliant with any suitable standard), a wired network, etc. In some examples, network 150 may be a local area network (LAN), a wide area network (WAN), a public network (e.g., the Internet), a private or semi-private network (e.g., a corporate or university intranet), any other suitable type of network, or any suitable combination of networks. Each communication link (arrow) shown in Figure 1 may be any suitable communication link or combination of communication links, such as a wired link, fiber optic link, Wi-Fi link, Bluetooth link, or cellular link.

[0061] To be understood, the various methods, devices, applications, nodes, functions, etc. described with respect to Figure 1 or any of the figures described herein are not intended to be limited to systems performed by the specific applications and functions described herein. Accordingly, additional configurations may be used to put the methods and systems disclosed herein into practice, and / or the functions and applications described herein may be excluded, without departing from the methods and systems disclosed herein.

[0062] Figure 2 is a block diagram illustrating a method for integrating an ML model with an interactive environment according to the embodiments described herein. The general sequence of operations relating to Method 200 is shown in Figure 2. Generally, Method 200 begins with operation 202 and ends with operation 216. Method 200 may include more or fewer steps than those shown in Figure 2, or the order of the steps may be changed. Method 200 is executed as a computer executable instruction performed by a computer system and may be encoded or stored on a computer-readable medium or other non-temporary computer storage medium. Furthermore, Method 200 may be executed by gates or circuits associated with a processor, ASIC, FPGA, SOC, or other hardware device. Hereafter, Method 200 will be described with reference to the systems, components, devices, modules, software, data structures, characteristic representations of data, signaling diagrams, methods, etc., described in relation to Figures 1, 3, 4, 5A, 5B, 6, 7, and 8.

[0063] In operation 202, input is received. Input may be received from applications (e.g., applications 104 and / or applications 105) and / or an integration manager (e.g., integration manager 106) on a computing device (e.g., user device 102 and / or developer device 105). Input may indicate a request regarding the model output. Input may also include information collected by the DS (e.g., DS120) and / or integration manager (e.g., integration manager 106), including systematic context and / or environmental guidelines from the DSM (e.g., DSM122), and specific context from the scenario processor (e.g., scenario processor 124).

[0064] In operation 204, an intent goal may be generated by, for example, an intent goal processor (e.g., intent goal processor 126) based on one or more of the input, specific context, systematic context, and / or environmental guidelines. The intent goal can encapsulate the overall intent or specific meaning of the input and may be used to assist in generating one or more prompts that are particularly relevant to that goal.

[0065] In operation 206, a prompt generator (e.g., prompt generator 128) generates one or more prompts. The prompts are generated based on one or more of the following: the input corresponding to the interactive environment (e.g., interactive environment 140) in which the input is received, environmental guidelines, systematic context, specific context, and / or intent objectives. In some cases, the prompts may be known prompts, in which case the prompt generator can access those prompts from the data store (e.g., data store 112) and / or integration manager (e.g., and / or integration manager 106). The prompts are used to provide sufficient context to a generic ML model so that the ML model can produce model outputs that are appropriate to the input.

[0066] In operation 208, the prompt generator (e.g., prompt generator 128) compares one or more prompts to the environmental guidelines and, if a prompt does not meet the guidelines, can modify one or more parts of the prompt and / or data mask parts of the prompt to make the prompt meet the environmental guidelines. Operation 208 is shown as a dashed box to indicate that it is optional and may be omitted in certain examples.

[0067] In operation 210, one or more prompts are processed by an ML model, such as a Generative Large-Scale Language Model (LLM) that is part of a model repository (e.g., model repository 130).

[0068] In operation 212, the model output may be evaluated by an output evaluator (e.g., output evaluator 132) to determine whether the output is suitable for the input and meets environmental guidelines. In some cases, the evaluation may be performed by generating confidence scores for one or more components of the output and comparing one or more confidence scores to a threshold. The confidence scores may be determined using evaluation metrics developed for the ML model and output. The confidence scores may be a measure of the suitability of the output to the input and the degree to which the output meets environmental guidelines based on one or more evaluation metrics. In operation 214, the output evaluator may attempt to modify one or more aspects of the model output and / or data mask parts of the output as necessary to make the output suitable for the input and / or guidelines. If the model output is unsuitable and cannot be modified, the flow proceeds to operation 206, where a new prompt may be generated. The new prompt may be refined by a new method of generating prompts or by expanding or contracting the parameters of a previously used prompt generation method, as described above. A new prompt is input to the ML model, generating a new model output. This loop continues until the model output is determined to be compliant with the input and environmental guidelines. If the model output is compliant, the flow proceeds to operation 216, where the model output is provided to the user and / or developer by the DSM (e.g., DSM1221). Operation 214 is shown as a dashed box to indicate that this step is optional and may be omitted in certain cases, such as when the output is compliant in step 212.

[0069] In operation 218, the DS (e.g., DS120) monitors the interactive environment for user and / or developer feedback and additional input from users and / or developers that may be received through the applications (e.g., applications 104 and / or 105) and / or the conversational agent evaluation portal. If additional feedback and / or input is received, the flow proceeds to operation 202 for further processing of the input described above. The lines for operation 218 and operations 218-202 are shown with dashed lines to indicate that they are optional and may be omitted in certain examples, such as a gaming context.

[0070] In operation 220, one or more of the inputs, intent goals, one or more prompts, and / or model outputs are stored in the data store (e.g., data store 112) and / or the integration manager (e.g., integration manager 106). Operation 220 is shown as a dashed box to indicate that this step is optional and may be omitted in certain examples.

[0071] Figure 3 shows an exemplary conversational agent evaluation portal according to the embodiments described herein. The conversational agent evaluation portal may be provided by an output evaluator (e.g., output evaluator 132) for the purpose of receiving input from a user / developer to an interactive element located just outside the interactive environment (e.g., interactive environment 140). In some cases, the portal is a chat window, as shown in portal 300, but it may also be a smaller version of the interactive environment, where non-text input (e.g., audio, visual, status of interactive elements, etc.) may be received and processed by a DS (e.g., DS 120) according to the embodiments described herein. In some examples, the portal may have an introductory message 302 that provides some context about the user / developer in the situation. Furthermore, there may be an advising and / or warning message 304 that informs the user / developer that they are interacting with an ML model in a learning environment and that potentially aggressive and / or inappropriate responses may occur during the learning of the ML model. In this case, portal 300 is a chat window having a text input field 312 for displaying messages 306, 308, and 310. Messages 306 and 308 are generated in relation to the interactive elements of the interactive environment 140, in this case the NPC. The portal 300 may have one or more selectable control functions, such as an exit button 314, for further interaction with the portal and / or with the interactive elements, and such interactions may be further based on selectable control functions available to the user / developer within the interactive environment.

[0072] Figure 4 is a block diagram illustrating a method, according to the embodiments described herein, for evaluating the output of an ML model using an ML model and providing an evaluation of the output of the ML model itself. The general sequence of operations relating to Method 400 is shown in Figure 4. Generally, Method 400 begins with operation 402 and ends with operation 412. Method 400 may include more or fewer steps than those shown in Figure 4, or the order of the steps may be changed. Method 400 is executed as a computer executable instruction performed by a computer system and may be encoded or stored on a computer-readable medium or other non-temporary computer storage medium. Furthermore, Method 400 may be executed by gates or circuits associated with a processor, ASIC, FPGA, SOC, or other hardware device. Hereafter in this specification, Method 400 will be described with reference to the systems, components, devices, modules, software, data structures, characteristic representations of data, signaling diagrams, methods, etc., described in relation to Figures 1, 2, 3, 5A, 5B, 6, 7, and 8.

[0073] In operation 402, an output evaluator (e.g., output evaluator 132) may retrieve user / developer feedback and / or content logs from a datastore (e.g., datastore 112) and / or an integration manager (e.g., integration manager 106). In operation 404, the output evaluator may associate one or more ML model outputs with user / developer feedback and content logs from the interactive environment 140. In some examples, in operation 404, one or more ML model outputs and users may be pre-associated with user / developer feedback and content logs and stored in a datastore (e.g., datastore 106). In operation 406, a prompt generator (e.g., prompt generator 128) may generate one or more prompts about the ML model based on the associated outputs, user / developer feedback, and / or content logs. In operation 408, the ML model may process one or more prompts and generate descriptive outputs of previous behavior as a means for developers to understand the ML model's previous outputs and actions. In operation 410, the performance of the ML model may be evaluated by an output evaluator (e.g., output evaluator 132) which generates confidence scores for one or more components of the output in some cases, and by comparing one or more confidence scores to a threshold. The confidence scores may be determined using evaluation metrics developed for the ML model and its output. The confidence scores may be a measure of the fit of the output to the input and the degree to which the output meets environmental guidelines based on one or more evaluation metrics. In optional operation 412, the DS (e.g., DS120) and / or ML model may be updated based on the model evaluation results from operation 410. Operation 412 is indicated by a dashed box to show that it is optional.

[0074] Figures 5A and 5B illustrate an exemplary generative machine learning model that may be used according to the embodiments described herein. Referring first to Figure 5A, the conceptual diagram 500 illustrates a pre-trained generative model package 504 that processes inputs and prompts 502 to produce an embodiment of the model output 506 described herein.

[0075] In some examples, the generative model package 504 is pre-trained on a variety of inputs (e.g., various human languages, various programming languages, and / or various content types) and therefore does not need to be fine-tuned or trained for a specific scenario. Rather, the generative model package 504 can be pre-trained more generally, and the input 502 includes prompts that are generated, selected, or otherwise designed to prompt the generative model package 504 to produce a specific generative model output 506. It should be understood that the input 502 and the generative model output 506 can each contain any of the various content types, including, but not limited to, text output, image output, audio output, video output, program output, and / or binary output, among other examples. In some examples, the input 502 and the generative model output 506 may have different content types, such as when the generative model package 504 contains a generative multimodal machine learning model.

[0076] Therefore, the generative model package 504 can be used in any of the various scenarios, and furthermore, a different generative model package can be used instead of the generative model package 504 without substantially modifying other related aspects (e.g., similar to those described herein with respect to Figures 1, 2, 3, and 4). Thus, the generative model package 504 acts as a tool on which machine learning processing is performed, and a specific input 502 to the generative model package 504 is programmed or otherwise determined, thereby the generative model package 504 generates a model output 506 that can later be used for further processing.

[0077] The generative model package 504 may be provided or otherwise used according to any of the various paradigms. For example, the generative model package 504 may be used locally for a computing device (e.g., user device 102 in Figure 1) or it may be accessed remotely from a machine learning service (e.g., director service 120). In other examples, embodiments of the generative model package 504 are distributed across multiple computing devices. In some cases, the generative model package 504 may be accessible via an application programming interface (API), such as one provided by the operating system of a computing device and / or by a machine learning service, among other examples.

[0078] Referring here to the illustrated embodiment of the generative model package 504, the generative model package 504 includes an input tokenizer 508, an input embedding 510, a model layer 512, an output layer 514, and an output decoder 516. In some examples, the input tokenizer 508 processes the input 502 to generate the input embedding 510, which contains a sequence of symbolic representations corresponding to the input 502. Thus, the input embedding 510 is processed by the model layer 512, the output layer 514, and the output decoder 516 to generate the model output 506. An exemplary architecture corresponding to the generative model package 504 is shown in Figure 5B, which will be discussed in more detail below. Nevertheless, it should be understood that the architectures illustrated and described herein should not be interpreted in an restrictive sense, and in other examples, any of various other architectures may be used.

[0079] Figure 5B is a conceptual diagram showing an exemplary architecture 550 of a pre-trained generative machine learning model that may be used in accordance with the embodiments described herein. As stated above, in other examples, any of the various alternative architectures and corresponding ML models may be used without departing from the embodiments described herein.

[0080] As illustrated, architecture 550 processes input 502 to produce a generated model output 506. Its embodiments were discussed above with respect to Figure 5A. Architecture 550 is represented as a transformer model including encoder 552 and decoder 554. Encoder 552 processes input embedding 558 (its embodiments may be similar to input embedding 510 in Figure 5A), which contains a sequence of symbolic representations corresponding to input 556. In some examples, input 556 contains inputs and prompts related to the generation 502 (e.g., corresponding to skills in a skill chain). Furthermore, position encoding 560 can introduce information about the relative and / or absolute positions of the tokens in input embedding 558. Similarly, output embedding 574 contains a sequence of symbolic representations corresponding to output 572, and position encoding 576 can similarly introduce information about the relative and / or absolute positions of the tokens in output embedding 574.

[0081] As shown in the illustration, encoder 552 includes an exemplary layer 570. It should be understood that any number of such layers may be used, and the illustrated architecture is simplified for illustrative purposes. The exemplary layer 570 includes two sub-layers: a multi-head attention layer 562 and a feedforward layer 566. In some examples, residual connections are included across each layer 562, 566, and normalization layers 564 and 568 are included after layers 562, 566, respectively.

[0082] The decoder 554 includes an exemplary layer 590. As with the encoder 552, any number of such layers may be used in other examples, and the architecture of the illustrated decoder 554 is simplified for illustrative purposes. As shown, the exemplary layer 590 includes three sub-layers: a masked multi-head attention layer 578, a multi-head attention layer 582, and a feedforward layer 586. Embodiments of the multi-head attention layer 582 and the feedforward layer 586 may be similar to the embodiments described above with respect to the multi-head attention layer 562 and the feedforward layer 566, respectively. Furthermore, the masked multi-head attention layer 578 performs multi-head attention on the output of the encoder 552 (e.g., output 572). In some examples, the masked multi-head attention layer 578 prevents one position from paying attention to a subsequent position. Such masking, combined with embedding offsets (e.g., one position at a time, as shown by the multi-head attention layer 582), can ensure that predictions for a given position depend on known outputs for one or more positions smaller than that position. As shown, residual connections are included across layers 578, 582, and 586, followed by normalization layers 580, 584, and 588, respectively. The multi-head attention layers 562, 578, and 582 can each linearly project queries, keys, and values ​​into their corresponding dimensions using a set of linear projections. Each linear projection can be processed using an attention function (e.g., dot product or additive attention) to generate n-dimensional output values ​​for each linear projection. The resulting values ​​can be concatenated and reprojected, and the values ​​are then processed as shown in Figure 5B (e.g., by the corresponding normalization layers 564, 580, or 584).

[0083] Feedforward layers 566 and 586 may each be a fully connected feedforward network applied at each location. In some examples, feedforward layers 566 and 586 each contain multiple linear transformations, including rectifying linear activations in between. In some examples, each linear transformation may be the same across different locations, but with different parameters compared to other linear transformations in the feedforward network.

[0084] Furthermore, embodiments of the linear transformation 592 may be similar to the linear transformations discussed above with respect to the multi-head attention layers 562, 578, and 582, and the feedforward layers 566 and 586. The softmax 594 may further transform the output of the linear transformation 592 into the predicted next token probability indicated by the output probability 596. The illustrated architecture is provided as an example, and it should be understood that in other examples, any of the various other model architectures may be used according to the embodiments disclosed. Thus, the output probability 596 can generate a model output 506 according to the embodiments described herein, and the output of the generated ML model defines the output corresponding to the input. For example, the model output 506 may be associated with an interactive element of an interactive environment 140, among other examples.

[0085] Figures 6-8 and related descriptions provide a discussion of various operating environments in which embodiments of this disclosure may be implemented. However, the devices and systems illustrated and discussed with respect to Figures 6-8 are for illustrative and explanatory purposes only and do not limit the numerous computing device configurations that may be used to implement embodiments of this disclosure described herein.

[0086] Figure 6 is a block diagram showing the physical components (e.g., hardware) of a computing device 600 in which embodiments of this disclosure may be implemented. The computing device components described below may be suitable for the computing device described above, including the user device 102 in Figure 1. In a basic configuration, the computing device 600 may include at least one processing unit 602 and system memory 604. Depending on the configuration and type of the computing device, the system memory 604 may comprise, but is not limited to, volatile storage (e.g., random access memory), non-volatile storage (e.g., read-only memory), flash memory, or any combination thereof.

[0087] System memory 604 may include an operating system 605 and one or more program modules 606 suitable for running software applications 620, such as one or more components supported by the system described herein. For example, system memory 604 may store an integration manager 624 and / or director services 626. The operating system 605 may be suitable, for example, for controlling the operation of a computing device 600.

[0088] Furthermore, embodiments of this disclosure may be implemented in conjunction with graphics libraries, other operating systems, or any other application programs, and are not limited to any particular application or system. This basic configuration is shown in Figure 6 by the component within the dashed line 608. The computing device 600 may have additional features or functions. For example, the computing device 600 may include additional data storage devices (removable and / or non-removable), such as magnetic disks, optical disks, or tapes. Such additional storage is shown in Figure 6 by the removable storage device 609 and the non-removable storage device 610.

[0089] As described above, the system memory 604 can store a number of program modules and data files. While running on the processing unit 602, the program module 606 (e.g., application 620) can execute processes including, but not limited to, the embodiments described herein. Other program modules that may be used in accordance with embodiments of this disclosure include email and contacts applications, word processing applications, spreadsheet applications, database applications, slide presentation applications, drawing or computer-aided application programs, and the like.

[0090] Furthermore, embodiments of the present disclosure may be implemented in electrical circuits including discrete electronic elements, packaged or integrated electronic chips including logic gates, circuits utilizing microprocessors, or single chips including electronic elements or microprocessors. For example, embodiments of the present disclosure may be implemented by a system-on-a-chip (SOC), in which each or many of the components shown in Figure 6 may be integrated on a single integrated circuit. Such an SOC device may include one or more processing units, graphics units, communication units, system virtualization units, and various application functions, all of which are integrated (or "burned") onto a chip substrate as a single integrated circuit. When operating by an SOC, the functions described herein with respect to the client's ability to switch protocols may be operated by application-specific logic integrated on a single integrated circuit (chip) together with other components of the computing device 600. Some embodiments of the present disclosure may also be implemented using other techniques capable of performing logical operations such as AND, OR, and NOT, including, but not limited to, mechanical, optical, fluidic, and quantum techniques. Furthermore, some aspects of this disclosure may be implemented within a general-purpose computer or within any other circuit or system.

[0091] The computing device 600 may also have one or more input devices 612, such as a keyboard, mouse, pen, sound or voice input device, or touch or swipe input device. Output devices 614, such as a display, speaker, or printer, may also be included. The devices described above are examples, and other devices may be used. The computing device 600 may include one or more communication connections 616 that enable communication with other computing devices 650. Suitable communication connections 616 include, but are not limited to, radio frequency (RF) transmitters, receivers, and / or transceiver circuit configurations; Universal Serial Bus (USB), parallel ports, and / or serial ports.

[0092] As used herein, the term "computer-readable medium" may include computer storage medium. Computer storage mediums may include volatile and non-volatile, removable and non-removable media implemented in any method or technique for storing information such as computer-readable instructions, data structures, or program modules. System memory 604, removable storage device 609, and non-removable storage device 610 are all examples of computer storage mediums (e.g., memory storage). Computer storage mediums may include RAM, ROM, electrically erasable read-only memory (EEPROM), flash memory or other memory technologies, CD-ROM, digital versatile disk (DVD) or other optical storage, magnetic cassette, magnetic tape, magnetic disk storage or other magnetic storage devices, or any other manufactured article that can be used to store information and can be accessed by computing device 600. Any such computer storage medium may be part of computing device 600. Computer storage mediums do not include carrier waves or other propagated or modulated data signals.

[0093] Communication media may be embodied by computer-readable instructions, data structures, program modules, or other data within a modulated data signal, such as a carrier wave or other transmission mechanism, and include any information delivery medium. The term “modulated data signal” may refer to a signal having one or more characteristics that are set or modified to encode information within the signal. Examples of communication media, without limitation, include wired media such as wired networks and direct wired connections, and wireless media such as acoustic, radio frequency (RF), and infrared.

[0094] Figure 7 is a block diagram showing the architecture of one embodiment of a computing device. That is, the computing device can incorporate a system (e.g., architecture) 702 to implement several embodiments. In some examples, system 702 is implemented as a “smartphone” capable of running one or more applications (e.g., a browser, email, calendar, contact manager, messaging client, game, and media client / player). In some embodiments, system 702 is integrated as a computing device such as an integrated personal digital assistant (PDA) and a wireless phone.

[0095] One or more application programs 766 may be loaded into memory 762 and executed on or in connection with the operating system 764. Examples of application programs include telephone dialer programs, email programs, personal information management (PIM) programs, word processing programs, spreadsheet programs, internet browser programs, and messaging programs. System 702 also includes a non-volatile storage area 768 within memory 762. The non-volatile storage area 768 may be used to store persistent information that must not be lost when System 702 is powered off. Application programs 766 may use and store information such as email or other messages used by email applications in the non-volatile storage area 768. A synchronization application (not shown) also resides in System 702 and is programmed to interact with a corresponding synchronization application residing on the host computer to keep the information stored in the non-volatile storage area 768 synchronized with the corresponding information stored on the host computer. As can be understood, other applications can also be loaded into memory 762 and run on the mobile computing device 700 described herein (e.g., an embedded object memory insertion engine, an embedded object memory search engine, etc.).

[0096] System 702 has a power supply 770, which may be implemented as one or more batteries. The power supply 770 may further include an external power source such as an AC adapter or a powered docking cradle to replenish or recharge the batteries.

[0097] System 702 may also include a radio interface layer 772 that performs the function of sending and receiving radio frequency communications. The radio interface layer 772 facilitates a radio connection between System 702 and the "outside world" via a communication carrier or service provider. Transmission to and from the radio interface layer 772 is under the control of the operating system 764. In other words, communications received by the radio interface layer 772 can be transmitted to an application program 766 via the operating system 764, and vice versa.

[0098] A visual indicator 720 may be used to provide visual notifications, and / or an audio interface 774 may be used to generate audible notifications via an audio transducer 725. In the illustrated example, the visual indicator 720 is a light-emitting diode (LED), and the audio transducer 725 is a speaker. These devices may be directly coupled to a power supply 770 and, once activated, remain on for a period specified by the notification mechanism, even if the processor 760 and / or dedicated processor 761 and other components are shut down to conserve battery power. The LED may be programmed to remain on indefinitely until a user takes action to indicate that the device is powered on. The audio interface 774 is used to provide and receive audible signals from the user. In addition to being coupled to, for example, an audio transducer 725, the audio interface 774 may also be coupled to a microphone for receiving audible input, for example, to facilitate a telephone conversation. According to aspects of this disclosure, the microphone may also function as an audio sensor to facilitate notification control, as described below. System 702 may further include a video interface 776 that enables the operation of the onboard camera 730 to record still images, video streams, and the like.

[0099] A computing device implementing System 702 may have additional features or functions. For example, the computing device may include additional data storage devices (removable and / or non-removable) such as magnetic disks, optical disks, or tapes. Such additional storage is shown in Figure 7 by a non-volatile storage area 768.

[0100] Data / information generated or captured by a computing device and stored via system 702 may be stored locally on the computing device as described above, or the data may be stored in a variety of storage media that the device can access via the wireless interface layer 772 or via a wired connection between the computing device and another computing device associated with it (e.g., a server computer in a distributed computing network such as the Internet). As understood, such data / information may be accessed by the computing device via the wireless interface layer 772 or via a distributed computing network.

[0101] Similarly, such data / information can be easily transferred, stored, and used between computing devices according to well-known data / information transfer and storage means, including email and collaborative data / information sharing systems.

[0102] Figure 8 shows one aspect of the architecture of a system for processing data received by a computing system from a remote source, such as a personal computer 804, a tablet computing device 806, or a mobile computing device 808, as described above. The content displayed on the server device 802 may be stored in different communication channels or other types of storage. For example, various documents may be stored using a directory service 824, a web portal 825, a mailbox service 826, an instant messaging store 828, or a social networking site 830.

[0103] Application 820 (for example, similar to application 620) may be employed by a client communicating with server device 802. Additionally or alternatively, an embedded object memory insertion engine 821 and / or an embedded object memory search engine 822 may be employed by server device 802. Server device 802 can exchange data with client computing devices such as personal computer 804, tablet computing device 806, and / or mobile computing device 808 (e.g., a smartphone) via network 815. As an example, the computer system described above may be embodied in personal computer 804, tablet computing device 806, and / or mobile computing device 808 (e.g., a smartphone). Any of these examples of computing devices can retrieve content from store 816, in addition to receiving graphic data that can be used for pre-processing in a graphics sending system or post-processing in a receiving computing system.

[0104] As can be understood from the foregoing disclosure, one aspect of the technology relates to a system comprising at least one processor and memory storing instructions that, when executed by the at least one processor, cause the system to perform a set of operations. The set of operations includes: receiving inputs by a director service to modify the interactive elements of an interactive environment; analyzing the interactive environment with respect to a specific context based on the inputs by the director service; receiving one or more environment guidelines; relating the inputs to one or more environment guidelines that provide a systematic context for the interactive environment by the director service; determining intent goals by the director service based on the inputs, the specific context, and one or more of the one or more environment guidelines; generating prompts for a generative machine learning model based on the intent goals by the director service; running the generative machine learning model using the prompts to produce model outputs; evaluating the model outputs for their suitability to the inputs and environment guidelines by the director service; and, if the model outputs are suitable, modifying the interactive elements of the interactive environment based on the model outputs. In one example, a set of operations further includes monitoring the interactive environment for subsequent inputs based on the provided model output;, when there is a subsequent input, having the director service analyze the interactive environment for a specific context based on the subsequent input; associating the subsequent input with one or more environment guidelines that provide a systematic context about the interactive environment; determining a second intent objective based on the subsequent input, the specific context, and one or more of the one or more environment guidelines; and generating a second prompt about the generative machine learning model based on the intent objective. In another example, generating a prompt further includes associating one or more prompt templates with the intent objective and combining one or more prompt templates with a prompt.In a further example, associating one or more prompt templates further includes generating embeddings for intent goals and identifying one or more prompt templates that are semantically associated with the intent goals based on the embeddings. In yet another example, generating intent goals further includes generating embeddings for one or more of the inputs, specific contexts, and environmental guidelines and identifying intent goals that are semantically associated with the inputs, specific contexts, and environmental guidelines based on the intent goals. In yet another example, evaluating model outputs for suitability further includes receiving confidence thresholds for evaluating model outputs and generating one or more confidence scores for one or more components of the model outputs, wherein the confidence scores measure suitability to the inputs and suitability to the environmental guidelines based on one or more metrics and comparing one or more confidence scores for one or more components of the outputs with confidence thresholds. In yet another example, a set of operations further includes storing one or more of the inputs, one or more intent goals, prompts, and model outputs. In a further example, when the model output is unsuitable, a set of operations may further include generating a new prompt for the generative machine learning model based on the intent goal, and then running the generative machine learning model using the new prompt to generate a new model output.

[0105] In another embodiment, the technology relates to a method for modifying interactive elements of a gaming environment, which includes receiving an input to modify interactive elements of a gaming environment; analyzing the gaming environment for a specific context based on the input; receiving one or more environmental guidelines; relating the input to one or more environmental guidelines that provide a systematic context for the gaming environment; determining an intent goal based on the input, the specific context, and one or more of the environmental guidelines; generating prompts for a generative machine learning model based on the intent goal; running the generative machine learning model using the prompts to produce a model output; evaluating the model output for its suitability to the input and environmental guidelines; and, if the model output is suitable, modifying interactive elements of the gaming environment based on the model output. In one example, the method further includes monitoring an interactive environment for subsequent inputs based on a provided model output; analyzing a gaming environment for a specific context based on the subsequent inputs, associating the subsequent inputs with one or more environmental guidelines that provide a systematic context about the gaming environment; determining a second intent goal based on the subsequent inputs, the specific context, and one or more of the one or more environmental guidelines; and generating a second prompt for the generative machine learning model based on the intent goal. In another example, generating a prompt further includes associating one or more prompt templates with the intent goal and combining one or more prompt templates with a prompt. In yet another example, associating one or more prompt templates further includes generating an embedding about the intent goal and identifying one or more prompt templates semantically associated with the intent goal based on the embedding.In yet another example, generating intent goals further includes generating embeddings for one or more of the inputs, specific contexts, and environmental guidelines, and identifying intent goals that are semantically related to the inputs, specific contexts, and environmental guidelines based on the intent goals. In yet another example, evaluating model outputs for relevance further includes receiving confidence thresholds for evaluating model outputs, generating one or more confidence scores for one or more components of the model outputs, the confidence scores measuring the degree of relevance to the inputs and to the environmental guidelines based on one or more metrics, and comparing one or more confidence scores for one or more components of the outputs with confidence thresholds. In yet another example, this method further includes storing one or more of the inputs, intent goals, prompts, and model outputs. In yet another example, when the model outputs are not relevant, the method further includes generating new prompts for the generative machine learning model based on the intent goals, and running the generative machine learning model using the new prompts to generate new model outputs. In yet another example, interactive elements include non-player characters (NPCs), animated infographics, videos, images, quizzes, game objects, and other aspects of the gaming environment that a user may be able to access and interact with. In yet another example, the gaming environment includes video games, online games, MMORPGs, and virtual reality environments.

[0106] In a further embodiment, the technology relates to a computer storage medium containing instructions, wherein, when the instructions are executed by the processor, the processor is caused to: receive an input to modify the interactive elements of an interactive environment; analyze the interactive environment for a specific context based on the input; receive one or more environmental guidelines; associate the input with one or more environmental guidelines that provide a systematic context for the interactive environment; determine an intent objective based on the input, the specific context, and one or more of the environmental guidelines; generate a prompt for a generative machine learning model based on the intent objective; run the generative machine learning model using the prompt to produce a model output; evaluate the model output for its suitability to the input and environmental guidelines; and, if the model output is suitable, modify the interactive elements of the interactive environment based on the model output. In one example, if the model output is not suitable, the processor is further caused to generate a new prompt for a generative machine learning model based on the intent objective and run the generative machine learning model using the new prompt to produce a new model output.

[0107] For example, aspects of the present disclosure are described above with reference to block diagrams and / or operational diagrams of methods, systems, and computer program products relating to aspects of the present disclosure. The functions / operations described within a block may be performed in an order different from that shown in any flowchart. For example, two consecutively shown blocks may actually be executed substantially simultaneously, or blocks may sometimes be executed in reverse order depending on the functions / operations they relate to.

[0108] The descriptions and illustrations of one or more embodiments provided in this application are not intended in any way to limit or restrict the scope of the claimed disclosure. The embodiments, examples, and details provided in this application are deemed sufficient to adequately describe the invention and to enable others to create and use the claimed embodiments of the disclosure. The claimed disclosure should be construed as not being limited to any embodiments, examples, or details provided in this application.

[0109] Whether illustrated and described in combination or individually, various features (both structural and methodological) are intended to be selectively included or omitted to produce embodiments having a particular set of features. The description and illustrations provided in this application enable those skilled in the art to envision variations, modifications, and alternative embodiments that fall within the spirit of the broader aspects of the overall inventive concept embodied herein, and these too do not deviate from the broader scope of the claimed disclosure.

Claims

1. At least one processor, When executed by the aforementioned at least one processor, a memory and A system comprising the above set of operations, The director service receives input to modify the interactive elements of the interactive environment, The director service analyzes the interactive environment with respect to a specific context based on the input, Receiving one or more environmental guidelines, The director service associates the input with one or more environmental guidelines that provide a systematic context for the interactive environment, The Director Service determines intent objectives based on the input, the specific context, and one or more of the one or more environmental guidelines, The director service generates prompts for the generative machine learning model based on the intent objective, The generative machine learning model is executed using the aforementioned prompt to generate model output, The director service evaluates the model output for compliance with the input and the environmental guidelines, When the model output is suitable, the interactive elements of the interactive environment are modified based on the model output. A system that includes this.

2. The aforementioned set of operations Based on the provided model output, the interactive environment is monitored for subsequent inputs, When there is a subsequent input, the director service analyzes the interactive environment for a specific context based on the subsequent input, The subsequent input is associated with one or more environmental guidelines that provide a systematic context for the interactive environment, Determining a second intent objective based on the subsequent input, the specific context, and one or more of the one or more environmental guidelines, To generate a second prompt for the generative machine learning model based on the aforementioned intent objective. The system according to claim 1, further comprising:

3. Generating a prompt is Associating one or more prompt templates with the aforementioned intent objective, The one or more of the above prompt templates are combined into a prompt. The system according to claim 1, further comprising:

4. Associating one or more prompt templates is possible. To generate an embedding for the aforementioned intent objective, Identifying one or more prompt templates associated with the intent and semantics based on the aforementioned embedding. The system according to claim 3, further comprising:

5. Generating intent goals is To generate embeddings for one or more of the aforementioned input, the aforementioned specific context, and the aforementioned environmental guidelines, Based on intent objectives, identify intent objectives that are semantically associated with the input, the specific context, and the environmental guidelines. The system according to claim 1, further comprising:

6. Evaluating the model output for suitability is Receiving a confidence threshold for evaluating the model output, The method involves generating one or more confidence scores for one or more components of the model output, wherein the confidence scores measure the suitability to the input and the suitability to the environmental guidelines based on one or more indicators. Comparing the one or more confidence scores for one or more components of the model output with the confidence threshold. The system according to claim 1, further comprising:

7. When the aforementioned model output is not suitable, the set of operations is To generate new prompts for the generative machine learning model based on the aforementioned intent objectives, The generative machine learning model is executed using the new prompt to generate a new model output. The system of claim 1, further comprising:

8. Receiving input to modify the interactive elements of the gaming environment, Based on the aforementioned input, the gaming environment is analyzed for a specific context, Receiving one or more environmental guidelines, Associating the input with one or more environmental guidelines that provide a systematic context regarding the gaming environment, Determining intent objectives based on the aforementioned input, the aforementioned specific context, and one or more of the aforementioned one or more environmental guidelines, To generate prompts for a generative machine learning model based on the aforementioned intent objective, The generative machine learning model is executed using the aforementioned prompt to generate model output, The model output is evaluated for its compliance with the aforementioned inputs and environmental guidelines, When the model output is suitable, modify the interactive elements of the gaming environment based on the model output. A method that includes this.

9. Based on the provided model output, the interactive environment is monitored for subsequent inputs, When there is a subsequent input, the gaming environment is analyzed for a specific context based on the subsequent input, The subsequent input is associated with one or more environmental guidelines that provide a systematic context regarding the gaming environment, Determining a second intent objective based on the subsequent input, the specific context, and one or more of the one or more environmental guidelines, To generate a second prompt for the generative machine learning model based on the aforementioned intent objective. The method according to claim 8, further comprising:

10. Generating a prompt is Associating one or more prompt templates with the aforementioned intent objective, The one or more of the above prompt templates are combined into a prompt. The method according to claim 8, further comprising:

11. Associating one or more prompt templates is possible. To generate an embedding for the aforementioned intent objective, Identifying one or more prompt templates associated with the intent and semantics based on the aforementioned embedding. The method according to claim 10, further comprising:

12. Generating intent goals is To generate embeddings for one or more of the aforementioned input, the aforementioned specific context, and the aforementioned environmental guidelines, Based on the aforementioned intent objectives, identify intent objectives that are semantically associated with the input, the specific context, and the environmental guidelines. The method according to claim 8, further comprising:

13. Evaluating the model output for suitability is Receiving a confidence threshold for evaluating the model output, The method involves generating one or more confidence scores for one or more components of the model output, wherein the confidence scores measure the suitability to the input and the suitability to the environmental guidelines based on one or more indicators. Comparing the one or more confidence scores for one or more components of the model output with the confidence threshold. The method according to claim 8, further comprising:

14. If the aforementioned model output is not suitable, the method is: To generate new prompts for the generative machine learning model based on the aforementioned intent objectives, The generative machine learning model is executed using the new prompt to generate a new model output. The method according to claim 8, further comprising:

15. A computer storage medium containing instructions, wherein when the instructions are executed by a processor, the processor: Receiving input to modify the interactive elements of an interactive environment, Analyzing the interactive environment for a specific context based on the aforementioned input, Receiving one or more environmental guidelines, Associating the input with one or more environmental guidelines that provide a systematic context for the aforementioned interactive environment, Determining intent objectives based on the aforementioned input, the aforementioned specific context, and one or more of the aforementioned one or more environmental guidelines, To generate prompts for a generative machine learning model based on the aforementioned intent objective, The generative machine learning model is executed using the aforementioned prompt to generate model output, The model output is evaluated for its conformity to the aforementioned inputs and environmental guidelines, When the model output is suitable, the interactive elements of the interactive environment are modified based on the model output. When the model output is not suitable, a new prompt for the generative machine learning model is generated based on the intent objective. The generative machine learning model is executed using the new prompt to generate a new model output. This includes evaluating and A computer storage medium that performs this task.