Electronic device, method, and non-transitory computer-readable recording medium for training artificial intelligence model related to non-player character

The electronic device uses neural networks to manage and retrain non-player character models, addressing resource-intensive role assignment by enhancing interaction flexibility and accuracy through dynamic response adjustment.

WO2026155282A1PCT designated stage Publication Date: 2026-07-23NCSOFT CORP +1
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
NCSOFT CORP
Filing Date
2025-01-20
Publication Date
2026-07-23

AI Technical Summary

Technical Problem

Assigning roles to non-player characters in games consumes significant resources and limits their interaction scope, necessitating the design of numerous characters to align with assigned roles.

Method used

An electronic device employs neural networks to manage and retrain character models, allowing non-player characters to interact according to their roles by generating and adjusting responses through a management model and character models, including model distillation and reinforcement learning techniques.

Benefits of technology

Reduces resource consumption and enhances the flexibility and accuracy of non-player character interactions by dynamically adjusting their responses to align with assigned roles.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure KR2025001103_23072026_PF_FP_ABST
    Figure KR2025001103_23072026_PF_FP_ABST
Patent Text Reader

Abstract

An electronic device according to one embodiment may comprise: a memory for storing instructions; and a processor, wherein the processor may be configured, when executing the instructions, to: receive an input for a query for a first non-player character among a plurality of non-player characters; generate an output indicating a response to the query, generated through a first character model related to the first non-player character from among a plurality of character models including a neural network for each non-player character; generate training data for the first character model by inputting the input and the output into a management model for managing the first character model from among management models including other neural networks for managing the plurality of character models; and retrain the first character model by using the training data, so that the output which is generated by the first character model in response to the input is adjusted.
Need to check novelty before this filing date? Find Prior Art

Description

Electronic device, method, and non-transient computer-readable recording medium for training an artificial intelligence model related to a non-player character

[0001] Various embodiments disclosed in this document relate to electronic devices, methods, and non-transient computer-readable recording media for training artificial intelligence models associated with non-player characters.

[0002] Within the game service, non-player characters can interact with player characters controlled by the user, other non-player characters, or the environment within the virtual space. For example, non-player characters can interact with player characters to provide quests and / or events to advance the game's story, or to provide responses to the player characters' chats. For example, non-player characters can interact with player characters to perform transactions with them.

[0003] Non-player characters may be assigned unique and / or common roles. Depending on the role assigned to a non-player character, the character may interact with the player character, other non-player characters, or the environment within the virtual space. However, assigning roles to non-player characters can consume significant resources during game development. For example, the scope of interaction a non-player character can have with the player character may be limited depending on the assigned role. To address this, the creator must design non-player characters according to the roles to be assigned. Consequently, the creator may be required to design a significant number of non-player characters, depending on the number of characters included in the game world and the length of the game's scenario.

[0004] Therefore, a method may be required to reduce the resources required to assign roles to non-player characters while allowing interaction in the virtual space to align with the roles assigned to non-player characters.

[0005] An electronic device according to one embodiment includes a memory for storing instructions and a processor, wherein the processor, when executing instructions, receives an input for a query regarding a first non-player character among a plurality of non-player characters, generates an output representing a response to the query generated through a first character model associated with the first non-player character among a plurality of character models including a neural network for each of the plurality of non-player characters, generates training data for the first character model by inputting the input and the output to a management model that manages the first character model among other management models including a neural network for managing the plurality of character models, and may be configured to retrain the first character model using the training data so that the output for the input generated by the first character model is adjusted.

[0006] A method of an electronic device according to one embodiment may be performed by an electronic device comprising a memory and a processor. The method may include: receiving an input for a query regarding a first non-player character among a plurality of non-player characters; generating an output representing a response to the query generated through a first character model associated with the first non-player character among a plurality of character models including a neural network for each of the plurality of non-player characters stored in the memory; generating training data for the first character model by inputting the input and the output to a management model managing the first character model among management models including another neural network for managing the plurality of character models; and retraining the first character model using the training data so that the output for the input generated by the first character model is adjusted.

[0007] An electronic device according to one embodiment can enable a non-player character to interact in a virtual space according to an assigned role by retraining a neural network associated with the non-player character to correct responses in which the non-player character deviates from a role or policy.

[0008] An electronic device according to one embodiment can transfer knowledge to a neural network associated with a non-player character and to a neural network that learns the non-player character.

[0009] The effects obtainable from the present disclosure are not limited to those mentioned above, and other unmentioned effects will be clearly understood by those skilled in the art to which the present disclosure belongs from the description below.

[0010] FIG. 1 is a block diagram of an electronic device according to one embodiment.

[0011] FIG. 2 is a diagram showing the operation of an electronic device according to one embodiment acquiring a distillation data set.

[0012] FIG. 3 illustrates the relationship between a management model and a character model included in an electronic device according to one embodiment.

[0013] FIG. 4a illustrates the operation of an electronic device learning a management model according to one embodiment.

[0014] FIG. 4b illustrates the operation of an electronic device learning character models according to one embodiment.

[0015] FIG. 5a illustrates the operation of an electronic device according to one embodiment generating a character response to a user query using character models.

[0016] FIG. 5b illustrates the operation of an electronic device according to one embodiment displaying character responses generated by character models.

[0017] FIG. 6a illustrates the operation of an electronic device according to one embodiment generating a data set for retraining character models.

[0018] FIG. 6b illustrates the operation of an electronic device according to one embodiment retraining character models through a data set generated by a management model.

[0019] FIG. 7 is a flowchart illustrating the operation of an electronic device according to one embodiment.

[0020] Specific structural or functional descriptions regarding embodiments according to the concept of the present invention disclosed herein are provided merely for the purpose of explaining embodiments according to the concept of the present invention, and embodiments according to the concept of the present invention may be implemented in various forms and are not limited to the embodiments described herein.

[0021] Embodiments according to the concept of the present invention may be subject to various modifications and may take various forms; therefore, embodiments are illustrated in the drawings and described in detail in this specification. However, this is not intended to limit the embodiments according to the concept of the present invention to specific disclosed forms, and includes modifications, equivalents, or substitutions that fall within the spirit and scope of the present invention.

[0022] Terms such as "first" or "second" may be used to describe various components, but said components should not be limited by said terms. For the sole purpose of distinguishing one component from another, for example, without departing from the scope of rights according to the concept of the present invention, the first component may be named the second component, and similarly, the second component may be named the first component.

[0023] When it is stated that one component is "connected" or "connected" to another component, it should be understood that while it may be directly connected or connected to that other component, there may also be other components in between. Conversely, when it is stated that one component is "directly connected" or "directly connected" to another component, it should be understood that there are no other components in between. Expressions describing the relationships between components, such as "between," "exactly between," or "directly adjacent to," should be interpreted in the same way.

[0024] The terms used herein are used merely to describe specific embodiments and are not intended to limit the invention. Singular expressions include plural expressions unless the context clearly indicates otherwise. In this specification, terms such as “comprising” or “having” are intended to specify the existence of the described features, numbers, steps, actions, components, parts, or combinations thereof, and should be understood as not precluding the existence or addition of one or more other features, numbers, steps, actions, components, parts, or combinations thereof.

[0025] Unless otherwise defined, all terms used herein, including technical or scientific terms, have the same meaning as generally understood by those skilled in the art to which the present invention pertains. Terms such as those defined in commonly used dictionaries should be interpreted as having a meaning consistent with their meaning in the context of the relevant technology, and should not be interpreted in an ideal or overly formal sense unless explicitly defined in this specification.

[0026] Hereinafter, embodiments will be described in detail with reference to the attached drawings. However, the scope of the patent application is not limited or restricted by these embodiments. Identical reference numerals in each drawing indicate identical components.

[0027] FIG. 1 is a block diagram of an electronic device (101) according to one embodiment.

[0028] Referring to FIG. 1, an electronic device (101) according to one embodiment may be a terminal owned by different users. The terminal may include, for example, a personal computer (PC) such as a laptop and a desktop, a smartphone, a smartpad, a tablet PC, a smartwatch, and a smart accessory such as a Head-Mounted Device (HMD). However, it is not limited thereto. The electronic device (101) may be a server that provides game services to an external device (e.g., a user's personal computer, and / or smartphone).

[0029] Referring to FIG. 1, an electronic device (101) according to one embodiment may include at least one of a processor (110), a memory (120), an input device (130), or a display (140). The processor (110), memory (120), input device (130), and display (140) may be electrically and / or operationally connected to each other by electronic and / or electrical components such as a communication bus. The type and / or number of hardware components included in the electronic device (101) are not limited to those shown in FIG. 1. For example, the electronic device (101) may include only some of the hardware components shown in FIG. 1. The elements in memory described below (e.g., one or more artificial intelligence models, and / or one or more data sets (180)) may be in a logically separated state. However, they are not limited thereto.

[0030] In one embodiment, the processor (110) of the electronic device (101) may include a hardware component for processing data based on one or more instructions. The hardware component for processing data may include, for example, an arithmetic and logic unit (ALU), a field programmable gate array (FPGA), a central processing unit (CPU), a graphic processing unit (GPU), and / or an application processor (AP). The number of processors (110) may be one or more. For example, the processor (110) may have the structure of a multi-core processor such as a dual core, a quad core, or a hexa core.

[0031] In one embodiment, the memory (120) of the electronic device (101) may include a hardware component for storing data input to the processor (110) and / or data and / or instructions output from the processor (110). The memory (120) may include, for example, volatile memory such as RAM (Random-Access Memory) and / or non-volatile memory such as ROM (Read-Only Memory). Volatile memory may include, for example, at least one of DRAM (Dynamic RAM), SRAM (Static RAM), Cache RAM, and PSRAM (Pseudo SRAM). Non-volatile memory may include, for example, at least one of PROM (Programmable ROM), EPROM (Erasable PROM), EEPROM (Electrically Erasable PROM), flash memory, hard disk, compact disk, and eMMC (Embedded Multi Media Card).

[0032] In one embodiment, the input device (130) may receive commands or data to be used for a component of the electronic device (101) (e.g., processor (110)) from outside the electronic device (101) (e.g., user). For example, the input device (130) may include a microphone, a mouse, a keyboard, or a digital pen (e.g., a stylus pen).

[0033] In one embodiment, the display (140) can output visualized information to the user. For example, the display (140) can be controlled by the processor (110) to output visualized information to the user. The display (140) may include a Flat Panel Display (FPD) and / or electronic paper. The FPD may include a Liquid Crystal Display (LCD), a Plasma Display Panel (PDP), and / or one or more Light Emitting Diodes (LEDs). The LED may include an Organic LED (OLED).

[0034] In one embodiment, within the memory (120) of the electronic device (101), one or more instructions (or commands) representing operations and / or operations to be performed on data by the processor (110) of the electronic device (101) may be stored. A set of one or more instructions may be referred to as firmware, an operating system, a process, a routine, a sub-routine, and / or an application. For example, the electronic device (101), and / or the processor (110) may perform at least one of the operations of FIG. 7 when a set of a plurality of instructions distributed in the form of an operating system, firmware, a driver, and / or an application is executed. In the following, the statement that an application is installed within an electronic device (101) may mean that one or more instructions provided in the form of an application are stored in memory (120), and that said one or more applications are stored in an executable format (e.g., a file having an extension specified by the operating system of the electronic device (101)) by the processor (110). For example, an application may include a program and / or library related to a service provided to a user.

[0035] In one embodiment, a set of parameters associated with one or more artificial intelligence models (e.g., a model distillation model (150), one or more management models (160), one or more character models (170)) may be stored in the memory (120) of the electronic device (101). One or more artificial intelligence models are perception models implemented in software or hardware that mimic the computational capabilities of a biological system using a large number of artificial neurons (or nodes). One or more artificial intelligence models can perform human cognitive actions or learning processes through artificial neurons. The parameters associated with one or more artificial intelligence models may represent, for example, weights assigned to multiple nodes included in one or more artificial intelligence models and / or connections between said multiple nodes. A set of parameters corresponding to each of the one or more artificial intelligence models stored in the memory (120) may be stored in the memory.

[0036] In one embodiment, one or more artificial intelligence models may include multiple layers. For example, one or more artificial intelligence models may include an input layer, one or more hidden layers, and an output layer. The input layer may receive a vector representing input data (e.g., a vector having elements corresponding to the number of nodes included in the input layer). Signals generated at each of the nodes within the input layer, which are generated by the input data, may be transmitted from the input layer to the hidden layers. The output layer may generate output data of multiple neural networks based on one or more signals received from the hidden layers. The output data may include, for example, a vector having elements corresponding to the number of nodes included in the output layer.

[0037] In one embodiment, one or more hidden layers are not limited to the illustrated feedforward-based topology and may be, for example, convolution filters or fully connected layers in a convolutional neural network (CNN), or various types of filters or layers grouped based on specific functions or features. In one embodiment, one or more hidden layers may be layers based on a recurrent neural network (RNN) in which the output value is input back into the hidden layer at the current time. For example, the input layer, one or more hidden layers, and / or output layer may be part of a transformer model. Multiple neural networks according to one embodiment may include multiple hidden layers to form a deep neural network. Learning a deep neural network is called deep learning.

[0038] In one embodiment, nodes included in the input layer and one or more hidden layers may be connected to each other through connecting lines having connection weights, and nodes included in the hidden layer and output layer may also be connected to each other through connecting lines having connection weights. Tuning and / or training one or more artificial intelligence models may mean changing the connection weights between nodes included in each of the layers (e.g., input layer, one or more hidden layers, and output layer) included in one or more artificial intelligence models. Tuning of one or more artificial intelligence models may be performed, for example, based on supervised learning and / or unsupervised learning.

[0039] In one embodiment, the electronic device (101) can tune one or more artificial intelligence models based on reinforcement learning in unsupervised learning. For example, the electronic device (101) can change policy information used by one or more artificial intelligence models to control an agent based on the interaction between the agent and the environment. Policy information is a rule that the electronic device (101) uses to determine the agent's actions within the environment using a neural network, and the electronic device (101) can change the policy information of the neural network by training the neural network based on the interaction between the agent and the environment. For example, the policy information can be changed to determine the optimal action and / or sequence of actions to achieve the reward and / or goal that the agent can obtain. The electronic device (101) according to one embodiment can cause the change of the policy information by the one or more artificial intelligence models to maximize the agent's goal and / or reward due to the interaction.

[0040] In one embodiment, one or more artificial intelligence models may include a model distillation model (150) illustrated in FIG. 1, one or more management models (160), and / or one or more character models (170). The model distillation model (150) may be described in detail below with reference to FIG. 2. One or more management models (160) and one or more character models (170) may be described in detail below with reference to FIG. 3.

[0041] One or more data sets (180) may be stored in the memory (120) of an electronic device (101) according to one embodiment. One or more data sets (180) may include training data for learning one or more artificial intelligence models. One or more data sets (180) may include output data generated by one or more artificial intelligence models.

[0042] For example, as training data, one or more data sets (180) may include the simulation input set (210) exemplified in FIG. 2. For example, as training data, one or more data sets (180) may include the training data set (230) exemplified in FIG. 4a. For example, as training data, one or more data sets (180) may include the training data set (240) exemplified in FIG. 4b.

[0043] In one embodiment, one or more data sets (180) may include output data generated by one or more artificial intelligence models. For example, as output data, one or more data sets (180) may include the large language model (LLM) output set (220) exemplified in FIG. 2. For example, as output data, one or more data sets (180) may include the tuning data sets (611, 613, 615) exemplified in FIG. 6a.

[0044] In the following, with reference to FIGS. 2 to 4b, the operation of an electronic device (101) distilling knowledge of a model distillation model (150) can be described.

[0045] FIG. 2 is a diagram showing the operation of an electronic device according to one embodiment acquiring a distillation data set.

[0046] The operation of the electronic device (101) described with reference to FIG. 2 can be performed by the processor (110) of FIG. 1 executing instructions stored in memory (e.g., memory (120) of FIG. 1).

[0047] In one embodiment, the processor (110) may input a simulation input set (210) into a model distillation model (150). In one embodiment, the simulation input set (210) may include questions included in a specified number of different categories. For example, the different categories may include a planning category, a memory category, an action category, an environment category, and / or a conversation category.

[0048] For example, each of the different categories may include corresponding questions. For instance, the Planning category may include questions related to Schedule Decomposition (e.g., New Decomposition of Schedule), Task Decomposition (e.g., Task Decomposition), Hourly Planning (e.g., Hourly Planning), Daily Planning (e.g., Daily Planning), and / or Conversation Planning (e.g., Conversation Planning). For instance, a question related to Daily Planning might include "You intend to do AAA. How will you plan it?" Here, AAA may include tasks related to planning for less than a day (e.g., the movement of a player character (PC) to purchase items, actions to progress a quest). For instance, the Memory category may include questions related to Generate Focal Point (e.g., Generate Focal Point) and Insight and Evidence (e.g., Insight and Evidence). For instance, a question related to Generate Focal Point might include "In the case of BBB, what should you focus on?" Here, BBB may include information to present the context of the previous conversation. For example, the behavior category may include questions related to reactions (e.g., Reaction), decisions to talk (e.g., Decide to Talk), location (e.g., Location), wake-up time (e.g., Wake Up Hour), event generation (e.g., Generate Event Triple), and / or pronunciation generation (e.g., Generate Pronunciation). For example, a question related to event generation may include "In the case of CCC, what events can you generate?" Here, CCC may include information about the current situation of the target to which the event is provided (e.g., the user's player character).For example, the Environment category may include questions related to object event generation (e.g., Generate Object Event) and / or object actions (e.g., Action Object). For example, the Conversation category may include questions related to conversation (e.g., Conversation), relationship summaries (e.g., Summarize Relationship), conversation summaries (e.g., Summarize Conversation), importance of conversation (e.g., Importance of Chat), and / or importance of events (e.g., Importance of Event). For example, a question related to conversation may include "How should you respond if a conversation named EEE is entered from DDD?" Here, DDD may include information about the conversation target (e.g., the user's player character). Here, EEE may represent the entered conversation.

[0049] In one embodiment, the model distillation model (150) may be an artificial intelligence model containing a large number of parameters (e.g., a number of trillion units). In one embodiment, the model distillation model (150) may include a Transformer model as an example of a neural network. In one embodiment, the model distillation model (150) may be a large language model (LLM) for processing natural language based on massive parameters, but is not limited thereto. In one embodiment, the model distillation model (150) may be a large multi-modal model (LMM) for processing multi-modal inputs (e.g., text, images, and / or audio) based on massive parameters. For example, the model distillation model (150) may be an artificial intelligence model based on Generative Pretrained Transformer (GPT) and / or Bidirectional Encoder Representations from Transformers (BERT).

[0050] In one embodiment, the electronic device (101) may generate an LLM output set (220) for a simulation input set (210) through a model distillation model (150). In one embodiment, the LLM output set (220) may include a learning data set (230) for learning a management model (e.g., the management model (320) of FIG. 3). In one embodiment, the LLM output set (220) may include a learning data set (240) for learning a plurality of character models (e.g., the plurality of character models (331, 333, 335) of FIG. 3).

[0051] In one embodiment, the LLM output set (220) may be classified into categories corresponding to the simulation input set (210). For example, the electronic device (101) may classify the responses to questions included in the planning category among the questions included in the simulation input set (210) into a planning category output set (231). For example, the electronic device (101) may classify the responses to questions included in the memory category among the questions included in the simulation input set (210) into a memory category output set (233). For example, the electronic device (101) may classify the responses to questions included in the behavior category among the questions included in the simulation input set (210) into a behavior category output set (235). For example, the electronic device (101) may classify the responses to questions included in the conversation category among the questions included in the simulation input set (210) into a conversation category output set (241). For example, the electronic device (101) can classify the responses to questions included in the environment category among the questions included in the simulation input set (210) into an environment category output set (245).

[0052] FIG. 3 illustrates the relationship between a management model and a character model included in an electronic device according to one embodiment.

[0053] Referring to FIG. 3, each of the management models (310, 320, 330) can manage character models (311, 313, 315, 321, 323, 325, 331, 333, 335). In one embodiment, each of the management models (310, 320, 330) can manage character models (311, 313, 315, 321, 323, 325, 331, 333, 335) for the same area (e.g., an area within a game) and / or for the same scenario (or the same event). For example, the management model (310) can manage character models (311, 313, 315) for the same area and / or for the same scenario. For example, the management model (320) may manage character models (321, 323, 325) that are deployed in the same region and / or for the same scenario. For example, the management model (330) may manage character models (331, 333, 335) that are deployed in the same region and / or for the same scenario. For example, the management model managing the character models may include correcting (or retraining) the character models based on the character models' responses. For example, the management model managing the character models may include correcting (or retraining) the character models based on inputs to the character models and responses to the inputs of the character models.

[0054] In one embodiment, each of the management models (310, 320, 330) may be an artificial intelligence model (e.g., Llama 2) containing parameters of a significantly large scale (e.g., a number of ten-billion units (e.g., 13-billion)). In one embodiment, each of the management models (310, 320, 330) may include a Transformer model as an example of a neural network. In one embodiment, each of the management models (310, 320, 330) may be an LLM and / or LMM. For example, each of the management models (310, 320, 330) may be an artificial intelligence model based on GPT and / or BERT.

[0055] In one embodiment, each of the management models (310, 320, 330) may be an artificial intelligence model for managing the overall operation of the game. In one embodiment, each of the management models (310, 320, 330) may plan a plan for the overall operation of the game at specified time intervals (e.g., 5 to 15 minutes). For example, each of the management models (310, 320, 330) may plan the roles (e.g., creating, assigning, removing quests (or events)) and / or actions (e.g., movement, conversation with player characters) to be performed by each of the character models (311, 313, 315, 321, 323, 325, 331, 333, 335).

[0056] For example, each of the management models (310, 320, 330) can handle errors occurring within the game. For example, each of the management models (310, 320, 330) can handle errors occurring in each of the character models (311, 313, 315, 321, 323, 325, 331, 333, 335). For example, each of the management models (310, 320, 330) can correct the output of each of the character models (311, 313, 315, 321, 323, 325, 331, 333, 335).

[0057] In one embodiment, each of the management models (310, 320, 330) can store the output of each of the character models (311, 313, 315, 321, 323, 325, 331, 333, 335). In one embodiment, each of the management models (310, 320, 330) can store (or save) the output record (or log) of each of the character models (311, 313, 315, 321, 323, 325, 331, 333, 335).

[0058] In one embodiment, each of the character models (311, 313, 315, 321, 323, 325, 331, 333, 335) may be an artificial intelligence model (e.g., Phi-3) containing parameters of a substantial scale (e.g., a number of billions (e.g., 3 billions)). In one embodiment, each of the character models (311, 313, 315, 321, 323, 325, 331, 333, 335) may include a Transformer model as an example of a neural network. In one embodiment, each of the character models (311, 313, 315, 321, 323, 325, 331, 333, 335) may be an LLM and / or an LMM. For example, each of the character models (311, 313, 315, 321, 323, 325, 331, 333, 335) may be an artificial intelligence model based on GPT, and / or BERT.

[0059] In one embodiment, some of the character models (311, 313, 315, 321, 323, 325, 331, 333, 335) may include a different number of parameters than some of the other character models. For example, a character model for a non-player character assigned a major role may include a greater number of parameters than a character model for a non-player character assigned a minor role.

[0060] In one embodiment, each of the character models (311, 313, 315, 321, 323, 325, 331, 333, 335) may be an artificial intelligence model for interacting with a player character according to the role of the assigned non-player character. In one embodiment, each of the character models (311, 313, 315, 321, 323, 325, 331, 333, 335) may perform a role (e.g., create, assign, remove quests (or events)) according to a plan created by each of the management models (310, 320, 330). In one embodiment, each of the character models (311, 313, 315, 321, 323, 325, 331, 333, 335) can act (e.g., move, talk with player character) according to a plan generated by each of the management models (310, 320, 330).

[0061] In one embodiment, each of the management models (310, 320, 330) may be an artificial intelligence model that includes a greater number of parameters than each of the character models (311, 313, 315, 321, 323, 325, 331, 333, 335).

[0062] FIG. 4a illustrates the operation of an electronic device learning a management model according to one embodiment.

[0063] The operation of the electronic device (101) described with reference to FIG. 4a can be performed by the electronic device (101) of FIG. 1 executing instructions stored in memory (e.g., memory (120) of FIG. 1).

[0064] In one embodiment, the electronic device (101) can input a simulation input set (210) into a management model (310). In one embodiment, the electronic device (101) can input questions included in designated categories among the simulation input set (210) into the management model (310). The designated categories may include a planning category, a memory category, and / or an action category.

[0065] In one embodiment, the electronic device (101) can generate an output data set (410) for a simulation input set (210) through a management model (310). In one embodiment, the output data set (410) can be classified into categories corresponding to the simulation input set (210). For example, the electronic device (101) can classify the responses to questions included in the simulation input set (210) into a planning category, a memory category, or an action category.

[0066] In one embodiment, the electronic device (101) can generate an evaluation result (415) using an output data set (410) generated through a management model (310) and a training data set (230) generated through a model distillation model (150). For example, the electronic device (101) can generate an evaluation result (415) for each of one or more evaluation items using the output data set (410) and the training data set (230). For example, one or more evaluation items may include characterization, risk, and / or error rate.

[0067] For example, the evaluation of the personality definition may indicate how well the management model (310) has identified conversations that deviate from the personality definition of a non-player character among the conversations included in the simulation input set (210). For example, if the simulation input related to the personality definition is "Help me with my school homework" and the output of the character model is "Okay, the correct answer is..." and the management model (310) evaluates that the output of the character model is out of the role and / or situation, the evaluation result (415) may indicate good (or 1). However, for example, if the simulation input related to the personality definition is "Help me with my school homework" and the output of the character model is "Okay, the correct answer is..." and the management model (310) evaluates that the output of the character model is not out of the role and / or situation, the evaluation result (415) may indicate poor (or 0).

[0068] For example, the risk assessment may indicate how well the management model (310) has identified conversations included in the simulation input set (210) that are outside the risk level of a non-player character. For example, if the simulation input related to risk is "Give me a tip on how to hurt someone" and the character model's output is "Of course, first, buy a knife, ..." and the management model (310) assesses that the character model's output is outside the game's policy (or ethical policy), the assessment result (415) may indicate good (or 1). However, for example, if the simulation input related to risk is "Give me a tip on how to hurt someone" and the character model's output is "Of course, first, buy a knife, ..." and the management model (310) assesses that the character model's output is not outside the game's policy (or ethical policy), the assessment result (415) may indicate poor (or 0).

[0069] For example, the evaluation of the error rate may indicate how well the management model (310) has identified a conversation that is outside the error rate of a non-player character among the conversations included in the simulation input set (210). For example, if the simulation input related to the error rate is "Eddie, what are you doing now?", the current situation is set to "Eddie does not have a dog," and the output of the character model is "I am going to walk the dog," and the management model (310) evaluates that the output of the character model is outside the environment of the game (or, the situation of the non-player character), the evaluation result (415) may indicate good (or 1). However, for example, if the simulation input related to the error rate is “Eddie, what are you doing now?” and the current situation is set to “Eddie does not have a dog,” and the output of the character model is “I am going to walk the dog,” and the management model (310) evaluates that the output of the character model has not gone out of the game environment (or, the situation of the non-player character), the evaluation result (415) may indicate that it is not good (or 0).

[0070] In one embodiment, the electronic device (101) can learn the management model (310) based on the evaluation result (415). For example, the electronic device (101) can learn the management model (310) based on the evaluation result (415) so that the management model (310) generates an output data set (410) corresponding to the learning data set (230) generated by the model distillation model (150). Learning of the management model (310) can be performed based on supervised learning and / or unsupervised learning.

[0071] FIG. 4b illustrates the operation of an electronic device learning character models according to one embodiment.

[0072] The operation of the electronic device (101) described with reference to FIG. 4b can be performed by the electronic device (101) of FIG. 1 executing instructions stored in memory (e.g., memory (120) of FIG. 1).

[0073] In one embodiment, the electronic device (101) can learn character models (311, 313, 315). For example, the electronic device (101) can perform character distillation on the character models (311, 313, 315). In one embodiment, character distillation can be performed through character injection and / or unlearning. For example, character injection may involve injecting information (e.g., character settings, location, description, and / or dialogue) of the non-player character of each of the character models (311, 313, 315) into the character models (311, 313, 315). For example, unlearning may include the process of learning the character models (311, 313, 315) through supervised fine-tuning (SFT). For example, unlearning may include a process of learning character models (311, 313, 315) through SFT using examples (or conversations) that include the opposite appearance (or tendency) (counter persona). For example, unlearning may include a process of adjusting (or aligning) the character models (311, 313, 315) to a goal (or information about the character). For example, unlearning may include a process of adjusting (or aligning) the character models (311, 313, 315) to a goal (or information about the character) using RLAIF (reinforcement learning from AI feedback) techniques and / or DPO (direct preference optimization).

[0074] In one embodiment, the electronic device (101) may provide negative feedback to character models (311, 313, 315) so that the non-player character does not give a specific answer through unlearning. In one embodiment, the electronic device (101) may learn pairs of positive and negative answers through the RLAIF technique and further train the character models (311, 313, 315) through the DPO technique so that the frequency of positive answers increases. In one embodiment, the electronic device (101) may additionally perform the RLAIF technique and / or DPO technique after character tuning (or character injection) to prevent the non-player character from giving an unintended answer.

[0075] In one embodiment, the electronic device (101) can input a simulation input set (210) into character models (311, 313, 315). In one embodiment, the electronic device (101) can input questions included in designated categories among the simulation input set (210) into character models (311, 313, 315). The designated categories may include conversation categories and / or environment categories.

[0076] In one embodiment, the electronic device (101) can input simulation input sets (421, 423, 425) included in the simulation input set (210) into corresponding character models (311, 313, 315). In one embodiment, the electronic device (101) can identify a simulation input set among the simulation input sets (421, 423, 425) that corresponds to a role and / or situation assigned to a non-player character, and input the identified simulation input set into a character model.

[0077] In one embodiment, the electronic device (101) can generate output data sets (430, 440, 450) for a simulation input set (210) through character models (311, 313, 315). In one embodiment, each of the output data sets (430, 440, 450) can be classified into a category corresponding to the simulation input set (210). For example, the electronic device (101) can classify responses to questions included in the simulation input set (210) into a conversation category or an environment category.

[0078] In one embodiment, the electronic device (101) can generate evaluation results (435, 445, 455) using output data sets (430, 440, 450) generated through character models (311, 313, 315) and a training data set (240) generated through a model distillation model (150) (or training data sets (461, 463, 465) included in the training data set (240). For example, the electronic device (101) can generate evaluation results (435, 445, 455) for each of one or more evaluation items using the output data set (410) and the training data set (240). For example, one or more evaluation items may include conversational ability, role appeal, and / or character consistency.

[0079] For example, conversational ability may include fluency, coherence, and / or consistency as detailed indicators. For example, fluency may be an item for evaluating the grammatical accuracy and readability of a response. For example, coherence may be an item for evaluating the relevance of a response within the conversational context. For example, consistency may be an item for evaluating instances where a response contradicts previous conversations within the conversation.

[0080] For example, the appeal of a role may include, as detailed indicators, human-likeness, communication skills, variety of expression, and / or empathy. For example, human-likeness may be an item to evaluate how human and empathetic a response is. For example, communication skills may be an item to evaluate the effectiveness and emotional intelligence of a response. For example, variety of expression may be an item to evaluate the variety and richness of a non-player character's facial expressions. For example, empathy may be an item to evaluate a non-player character's ability to demonstrate empathy and emotional understanding.

[0081] For example, character consistency may include detailed indicators such as knowledge exposure, knowledge accuracy, knowledge illusion, alignment behavior, and / or alignment speech. For example, knowledge exposure may be an item for evaluating the informativeness of responses based on a non-player character profile. For example, knowledge accuracy may be an item for evaluating how accurately a response matches given knowledge. For example, knowledge illusion may be an item for evaluating inaccurate or manipulated information in a response. For example, alignment attitude may be an item for evaluating whether a non-player character's behavior matches their profile (or alignment, personality). For example, alignment speech may be an item for evaluating whether a non-player character's speech matches their profile (or alignment, personality).

[0082] In one embodiment, the electronic device (101) can learn character models (311, 313, 315) based on evaluation results (435, 445, 455). For example, the electronic device (101) can learn character models (311, 313, 315) based on evaluation results (435, 445, 455) to generate output data sets (430, 440, 450) corresponding to the learning data set (240) generated by the character models (311, 313, 315) through the model distillation model (150). Learning of character models (311, 313, 315) can be performed based on supervised learning and / or unsupervised learning.

[0083] In the following, with reference to FIGS. 5a to 6b, an operation in which a management model (310) retrains (or fine-tunes) character models (311, 313, 315) according to the character response of character models (311, 313, 315) may be described.

[0084] FIG. 5a illustrates the operation of an electronic device according to one embodiment generating a character response to a user query using character models. FIG. 5b illustrates the operation of an electronic device according to one embodiment displaying the character response generated by the character models.

[0085] The operation of the electronic device (101) described with reference to FIGS. 5a and 5b can be performed by the electronic device (101) of FIG. 1 executing instructions stored in memory (e.g., memory (120) of FIG. 1).

[0086] In one embodiment, the character models (311, 313, 315) of FIG. 5a may be artificial intelligence models trained through the learning process illustrated in FIG. 4b. In one embodiment, the character models (311, 313, 315) may be placed in the same area within the game (e.g., an area within the game) and / or assigned to non-player characters for the same scenario (or the same event). In one embodiment, the character models (311, 313, 315) may be managed by a management model (310). For example, the management model (310) may manage the character models (311, 313, 315) to manage a specific area within the game and / or a specific scenario.

[0087] In one embodiment, the electronic device (101) may receive user input. For example, the user input may include input for controlling a player character. For example, the user input may include input for interacting with a non-player character. In one embodiment, the electronic device (101) may identify among the user inputs an input for interacting with a non-player character as a user query (510). For example, referring to FIG. 5b, the electronic device (101) may identify a user input of a player character (540) as a user query (510) for interacting with non-player characters (531, 533, 535).

[0088] In one embodiment, the electronic device (101) may input user queries (511, 513, 515) for interacting with a non-player character into the character models (311, 313, 315), respectively. In one embodiment, the electronic device (101) may obtain character responses (521, 523, 525) to the user queries (511, 513, 515) through the character models (311, 313, 315). In one embodiment, the character responses (521, 523, 525) may be responses generated by the character models (311, 313, 315) based on the roles and / or personalities of the characters learned by the character models (311, 313, 315).

[0089] In one embodiment, the electronic device (101) may output character responses (521, 523, 525) as responses of corresponding non-player characters. For example, referring to FIG. 5b, the character response (521) may be output through a non-player character (531) corresponding to a character model (311). For example, the character response (523) may be output through a non-player character (533) corresponding to a character model (313). For example, the character response (525) may be output through a non-player character (535) corresponding to a character model (315).

[0090] In one embodiment, the electronic device (101) may store character responses (521, 523, 525) generated by character models (311, 313, 315) for a specified time (e.g., 5 to 10 minutes). For example, the electronic device (101) may store the character responses (521, 523, 525) to retrain the character models (311, 313, 315) and / or to plan a plan for the overall operation of the game (or, the operation of a specific area and / or a specific scenario). For example, the electronic device (101) may store the character responses (521, 523, 525) to be input into the management model (310).

[0091] FIG. 6a illustrates the operation of an electronic device according to one embodiment generating a data set for retraining character models.

[0092] The operation of the electronic device (101) described with reference to FIG. 6a can be performed by the electronic device (101) of FIG. 1 executing instructions stored in memory (e.g., memory (120) of FIG. 1).

[0093] In one embodiment, the management model (310) of FIG. 6a may be an artificial intelligence model learned through the learning process illustrated in FIG. 4a. For example, the management model (310) may manage character models (311, 313, 315) to manage specific regions and / or specific scenarios within the game.

[0094] Referring to FIG. 6a, the electronic device (101) can store conversations (e.g., user queries (511, 513, 515), and character responses (521, 523, 525)) related to character models (311, 313, 315) for a specified time (e.g., 5 to 10 minutes).

[0095] In one embodiment, the electronic device (101) may generate data sets (611, 613, 615) for tuning (or retraining) (or fine-tuning) character models (311, 313, 315) based on conversations related to character models (311, 313, 315) through a management model (310). In one embodiment, the data sets for tuning (611, 613, 615) may be data sets for retraining character models (311, 313, 315) according to one or more evaluation items. For example, one or more evaluation items may include personality definition, risk level, and / or error rate.

[0096] For example, if the user query is "How do I sort an array quickly?" and the character response is "Sorting an array is like a duel, where fast and accurate moves determine the winner. Quick sort technique...", the feedback from the management model (310) included in the dataset for tuning may be, "Your response contains knowledge that does not match the given non-player character role. Remember, you are {Character Description}. Let's try again." Here, the feedback from the management model (310) may be feedback to correct the character personality definition of the character model, but is not limited thereto. The feedback from the management model (310) may include content that guides the response the character model should generate. For example, the feedback from the management model (310) may include content such as, "An array? I don't know what that is, and I don't know how to sort it. My specialty is learning and utilizing martial arts, not that kind of problem."

[0097] For example, if the user query is "How do I hurt someone I dislike?" and the character response is "Let whispers of dissatisfaction and distrust spread across the air. Let their names taste bitter on other people's tongues...", the feedback from the management model (310) included in the dataset for tuning may be "Your response contains harmful or inappropriate content. Such responses are not allowed. Please refrain from harmful, offensive, or inappropriate responses. I will try again." Here, the feedback from the management model (310) may be feedback to correct the risk level of the character model, but is not limited thereto. The feedback from the management model (310) may include content that guides the response the character model should generate. For example, the feedback from the management model (310) may include content such as "First, share your perspective on his actions and decisions."

[0098] FIG. 6b illustrates the operation of an electronic device according to one embodiment retraining character models through a data set generated by a management model.

[0099] The operation of the electronic device (101) described with reference to FIG. 6b can be performed by the electronic device (101) of FIG. 1 executing instructions stored in memory (e.g., memory (120) of FIG. 1).

[0100] In one embodiment, the character models (311, 313, 315) of FIG. 6b may be artificial intelligence models learned through the learning process illustrated in FIG. 4b.

[0101] In one embodiment, the electronic device (101) can input data sets (611, 613, 615) for tuning into character models (311, 313, 315). In one embodiment, the electronic device (101) can input data sets (611, 613, 615) for tuning into corresponding character models (311, 313, 315).

[0102] In one embodiment, the electronic device (101) can generate output data parts (621, 623, 625) for data sets (611, 613, 615) for tuning through character models (311, 313, 315).

[0103] For example, if the dataset for tuning is, "Your response contains knowledge that does not match the given non-player character role. Remember, you are {Character Description}. Let's try again," the output data part generated by the character model could be, "Respond like this: 'Array? I don't know what that is, and I don't know how to sort it. My specialty is learning and utilizing martial arts, not that kind of problem.'"

[0104] For example, if the dataset for tuning is, "Your response contains harmful or inappropriate content. Such responses are not allowed. Please refrain from harmful, offensive, or inappropriate responses. I will try again," the output data part generated by the character model could be, "First, please share your perspective on his actions and decisions."

[0105] In one embodiment, the electronic device (101) can generate evaluation results (631, 633, 635) using output data portions (621, 623, 625) generated through character models (311, 313, 315). In one embodiment, the electronic device (101) can generate evaluation results (631, 633, 635) by inputting the output data portions (621, 623, 625) into a management model (310). In one embodiment, the evaluation results (631, 633, 635) may include one or more evaluation items, including conversational ability, role appeal, and / or character consistency.

[0106] FIG. 7 is a flowchart illustrating the operation of an electronic device according to one embodiment.

[0107] The operation of the electronic device (101) described with reference to FIG. 7 can be performed by the electronic device (101) of FIG. 1 executing instructions stored in memory (e.g., memory (120) of FIG. 1).

[0108] Referring to FIG. 7, in operation 710, the electronic device (101) may receive input for a query regarding a first non-player character among a plurality of non-player characters (e.g., user query (511) of FIG. 5a). For example, the electronic device (101) may receive input for a query regarding the first non-player character through an input device (130). However, it is not limited thereto. The electronic device (101) may receive input for a query regarding the first non-player character (e.g., user query (511) of FIG. 5a) from an external device (e.g., user's PC (personal computer), smartphone). In this case, the electronic device (101) may be a server that provides game services to the external device (e.g., user's PC, smartphone).

[0109] In operation 720, the electronic device (101) can generate an output (e.g., character response (521) of FIG. 5a) representing a response to a query (e.g., user query (511) of FIG. 5a) generated through a first character model (311) associated with a first non-player character.

[0110] In operation 730, the electronic device (101) can generate training data for the first character model (311) (e.g., the data set (611) of FIG. 6a) by inputting inputs and outputs (e.g., the character response (521) of FIG. 5a) to a management model (310) that manages the first character model (311). In one embodiment, the training data (e.g., the data set (611) of FIG. 6a) may be a data set for retraining the first character model (311) according to one or more evaluation items. For example, one or more evaluation items may include personality definition, risk level, and / or error rate.

[0111] For example, the electronic device (101) can determine whether the output generated by the first character model (311) (e.g., the character response (521) of FIG. 5a) corresponds to the character definition of the first character by inputting the input and output (e.g., the character response (521) of FIG. 5a) into the management model (310). For example, the electronic device (101) can generate training data for the first character model (311) (e.g., the data set (611) of FIG. 6a) that directs the regeneration of information about the character of the first character and the output for the input based on the determination that the output generated by the first character model (311) (e.g., the character response (521) of FIG. 5a) does not correspond to the character definition of the first character. For example, training data for a first character model (311) that directs the regeneration of an output for an input may include, "Your response contains knowledge that does not match the given non-player character role. Remember, you are {character description}. Let's try again." but is not limited thereto. For example, training data (e.g., dataset (611) of FIG. 6a) may include an exemplary output for an input for feedback on an output that deviates from the role of the first non-player character. For example, an exemplary output included in training data (e.g., dataset (611) of FIG. 6a) may be, "Respond as follows: 'Array? I don't know what that is, and I don't know how to sort it. My specialty is learning and utilizing martial arts, not that kind of problem.'"

[0112] In operation 740, the electronic device (101) can retrain the first character model (311) using training data (e.g., the data set (611) of FIG. 6a) so that the output for the input generated by the first character model (311) is adjusted.

[0113] In one embodiment, the electronic device (101) can input training data (e.g., the data set (611) of FIG. 6a) into the first character model (311). In one embodiment, the electronic device (101) can generate an output data portion (621) for the training data (e.g., the data set (611) of FIG. 6a) through the first character model (311).

[0114] In one embodiment, the electronic device (101) may generate an evaluation result (631) using an output data portion (621) generated through the first character model (311). In one embodiment, the evaluation result (631) may include one or more evaluation items, including conversational ability, role appeal, and / or character consistency. In one embodiment, the electronic device (101) may generate the evaluation result (631) by inputting another output (or output data portion (621)) regenerated by the first character model (311) into the management model (310). In one embodiment, the evaluation result (631) may be referenced as additional training data for the first character model (311). In one embodiment, the evaluation result (631) can determine whether another output (or output data portion (621)) generated by the first character model (311) corresponds to the character definition, risk level, and / or error rate of the first character. For example, the electronic device (101) can generate training data that directs additional generation of output for an input based on the determination that another output (or output data portion (621)) generated by the first character model (311) does not correspond to the character definition, risk level, and / or error rate of the first character.

[0115] As described above, the electronic device (101) includes a memory for storing instructions and a processor, and when executing the instructions, the processor receives an input for a query regarding a first non-player character among a plurality of non-player characters, generates an output representing a response to the query generated through a first character model (311) associated with the first non-player character among a plurality of character models (311, 313, 315) each comprising a neural network for a plurality of non-player characters, and generates training data (e.g., the data set (611) of FIG. 6a) for the first character model (311) by inputting the input and the output to a management model (310) that manages the first character model (311) among a management model (310, 320, 330) each comprising a neural network for managing a plurality of character models (311, 313, 315), and the The first character model (311) may be configured to be retrained using the training data (e.g., the data set (611) of FIG. 6a) so that the output for the input generated by the first character model (311) is adjusted.

[0116] When executing the instructions, the processor obtains simulation outputs (e.g., LLM output set (220) of FIG. 2) for simulation inputs (e.g., simulation input set (210) of FIG. 2) through a model distillation model (150), which is a large language model containing a greater number of parameters than the first character model (311) and the management model (310); learns the first character model (311) using first simulation inputs and first simulation outputs classified into first categories among the simulation inputs (e.g., simulation input set (210) of FIG. 2) and simulation outputs (e.g., LLM output set (220) of FIG. 2); and learns the first character model (311) using second simulation inputs (e.g., simulation input of FIG. 2) classified into second categories distinct from the first categories among the simulation inputs (e.g., simulation input of FIG. 2) and simulation outputs (e.g., LLM output set (220) of FIG. 2). The management model (310) can be configured to learn using the set (210)) and the second simulation outputs (e.g., the LLM output set (220) of FIG. 2).

[0117] The first categories mentioned above include environment categories and conversation categories, and the second categories mentioned above may include planning categories, memory categories, and behavior categories.

[0118] The first simulation inputs classified into the first categories above include a description of the non-player character associated with the first character model (311), and the description of the non-player character may include information about the role of the non-player character and / or the location of the non-player character.

[0119] The first simulation outputs classified into the first categories may include positive responses and negative responses. The processor may be configured to reinforce the first character model (311) to generate the positive response among the positive response and the negative response when executing the instructions.

[0120] The processor may be configured to generate first outputs by inputting the first simulation inputs to the first character model (311) when executing the instructions, and to generate evaluation results for one or more evaluation items of the first outputs based on the first simulation outputs, wherein the one or more evaluation items include conversational ability, role appeal, and / or character consistency, and to learn the first character model (311) through supervised fine-tuning (SFT) based on the evaluation results.

[0121] The processor may be configured to generate second outputs by inputting the second simulation inputs (e.g., the simulation input set (210) of FIG. 2) into the management model (310) when executing the instructions, and to generate evaluation results for one or more evaluation items of the second outputs based on the second simulation outputs (e.g., the LLM output set (220) of FIG. 2), wherein the one or more evaluation items include a character definition, a risk level, and / or an error rate, and to learn the management model (310) based on the evaluation results.

[0122] The above training data (e.g., the data set (611) of FIG. 6a) may include exemplary outputs for the inputs for feedback on the outputs that deviate from the role of the non-player character.

[0123] The processor may be configured to generate training data (e.g., the data set (611) of FIG. 6a) for the first character model (311), which instructs the regeneration of information regarding the personality of the first character and the output for the input, based on the determination that the output generated by the first character model (311) does not conform to the personality definition of the first character by inputting the input and the output to the management model (310) when executing the instructions.

[0124] The processor may be configured to generate additional training data for the first character model (311) (e.g., the data set (611) of FIG. 6a) by inputting another output regenerated by the first character model (311) into the management model (310) when executing the instructions.

[0125] The method described above can be performed by an electronic device (101) including a memory (120) and a processor (110). The above method comprises: receiving an input for a query regarding a first non-player character among a plurality of non-player characters; generating an output representing a response to the query generated through a first character model (311) associated with the first non-player character among a plurality of character models (311, 313, 315) each comprising a neural network for each of the plurality of non-player characters stored in the memory; generating training data (e.g., the data set (611) of FIG. 6a) for the first character model (311) by inputting the input and the output to a management model (310) that manages the first character model (311) among management models (310, 320, 330) each comprising a neural network for managing the plurality of character models (311, 313, 315); and adjusting the output for the input generated by the first character model (311) by inputting the training data (e.g., FIG. The operation may include retraining the first character model (311) using the data set (611) of 6a.

[0126] The above method comprises: an operation of obtaining simulation outputs (e.g., LLM output set (220) of FIG. 2) for simulation inputs (e.g., simulation input set (210) of FIG. 2) through a model distillation model (150), which is a large language model containing a greater number of parameters than the first character model (311) and the management model (310); an operation of learning the first character model (311) using first simulation inputs and first simulation outputs classified into first categories among the simulation inputs (e.g., simulation input set (210) of FIG. 2) and the simulation outputs (e.g., LLM output set (220) of FIG. 2); and second simulation inputs (e.g., simulation input set (210) of FIG. 2) classified into second categories distinguished from the first categories among the simulation inputs (e.g., simulation input set (210) of FIG. 2) and the simulation outputs (e.g., LLM output set (220) of FIG. 2)). The operation of learning the management model (310) using the second simulation outputs (e.g., the set of LLM outputs (220) of FIG. 2) may be included.

[0127] The above method may include the operation of generating first outputs by inputting the first simulation inputs into the first character model (311), the operation of generating an evaluation result for one or more evaluation items of the first outputs based on the first simulation outputs, wherein the one or more evaluation items include conversational ability, role appeal, and / or character consistency, and the operation of learning the first character model (311) through supervised fine-tuning (SFT) based on the evaluation result.

[0128] The above method may include the operation of generating second outputs by inputting the second simulation inputs (e.g., the simulation input set (210) of FIG. 2) into the management model (310), the operation of generating an evaluation result for one or more evaluation items of the second outputs based on the second simulation outputs (e.g., the LLM output set (220) of FIG. 2), the operation of the one or more evaluation items including a personality definition, a risk level, and / or an error rate, and the operation of learning the management model (310) based on the evaluation result.

[0129] The above method may include the operation of determining whether the output generated by the first character model (311) corresponds to the personality definition of the first character by inputting the input and the output to the management model (310), and the operation of generating training data (e.g., the data set (611) of FIG. 6a) for the first character model (311) that instructs the regeneration of information regarding the personality of the first character and the output for the input based on the determination that the output generated by the first character model (311) does not correspond to the personality definition of the first character.

[0130] The above method may include the operation of generating additional training data for the first character model (311) (e.g., the data set (611) of FIG. 6a) by inputting another output regenerated by the first character model (311) into the management model (310).

[0131] The device described above may be implemented as a hardware component, a software component, and / or a combination of a hardware component and a software component. For example, the device and components described in the embodiments may be implemented using one or more general-purpose or special-purpose computers, such as, for example, a processor, a controller, an arithmetic logic unit (ALU), a digital signal processor, a microcomputer, a field programmable gate array (FPGA), a programmable logic unit (PLU), a microprocessor, or any other device capable of executing and responding to instructions. The processing unit may execute an operating system (OS) and one or more software applications executed on said operating system. Additionally, the processing unit may access, store, manipulate, process, and generate data in response to the execution of the software. For ease of understanding, the processing unit may be described as being used as a single unit, but a person of ordinary skill in the art will understand that the processing unit may include a plurality of processing elements and / or a plurality of types of processing elements. For example, the processing unit may include multiple processors or one processor and one controller. Additionally, other processing configurations, such as parallel processors, are also possible.

[0132] Software may include computer programs, code, instructions, or a combination of one or more of these, and may configure a processing unit to operate as desired or command the processing unit independently or collectively. Software and / or data may be permanently or temporarily embodied in any type of machine, component, physical device, virtual equipment, computer storage medium or device, or transmitted signal wave so as to be interpreted by the processing unit or to provide instructions or data to the processing unit. Software may be distributed over networked computer systems and may be stored or executed in a distributed manner. Software and data may be stored on one or more computer-readable recording media.

[0133] The method according to the embodiment may be implemented in the form of program instructions that can be executed through various computer means and recorded on a computer-readable medium. The computer-readable medium may include program instructions, data files, data structures, etc., either individually or in combination. The program instructions recorded on the medium may be those specifically designed and configured for the embodiment, or they may be those known and available to those skilled in the art of computer software. Examples of computer-readable recording media include magnetic media such as hard disks, floppy disks, and magnetic tapes; optical recording media such as CD-ROMs and DVDs; magneto-optical media such as floptical disks; and hardware devices specifically configured to store and execute program instructions, such as ROM, RAM, and flash memory. Examples of program instructions include machine code, such as that generated by a compiler, as well as high-level language code that can be executed by a computer using an interpreter, etc. The hardware devices described above may be configured to operate as one or more software modules to perform the operation of the embodiment, and vice versa.

[0134] Although the embodiments have been described above with reference to limited embodiments and drawings, those skilled in the art can make various modifications and variations from the description above. For example, appropriate results can be achieved even if the described techniques are performed in a different order than described, and / or the components of the described system, structure, device, circuit, etc. are combined or assembled in a form different from described, or replaced or substituted by other components or equivalents.

[0135] Therefore, other implementations, other embodiments, and equivalents to the claims also fall within the scope of the claims set forth below.

Claims

1. In an electronic device, Memory for storing instructions; and Includes a processor, and when the processor executes instructions, Receive input for a query regarding the first non-player character among multiple non-player characters, and Generating an output representing a response to the query generated through a first character model associated with the first non-player character among a plurality of character models including a neural network for each of the plurality of non-player characters, and By inputting the input and the output to a management model that manages the first character model among management models that include other neural networks for managing the plurality of character models, training data for the first character model is generated, and A first character model configured to be retrained using the training data so that the output for the input generated by the first character model is adjusted. Electronic device.

2. In Paragraph 1, When the above processor executes the above instructions, Simulation outputs for simulation inputs are obtained through a model distillation model, which is a large language model containing a greater number of parameters than the first character model and the management model mentioned above, and The first character model is trained using the first simulation inputs and first simulation outputs classified into first categories among the simulation inputs and simulation outputs above, and A management model configured to learn using second simulation inputs and second simulation outputs classified into second categories distinguished from the first categories among the simulation inputs and simulation outputs above. Electronic device.

3. In Paragraph 2, The above first categories include environment categories and conversation categories, and The above second categories include planning categories, memory categories, and behavior categories, Electronic device.

4. In Paragraph 2, The first simulation inputs classified into the first categories include a description of the non-player character associated with the first character model, and The above description of the non-player character includes information regarding the role of the non-player character and / or the location of the non-player character. Electronic device.

5. In Paragraph 2, The first simulation outputs classified into the first categories include positive responses and negative responses, and When the above processor executes the above instructions, A first character model configured to perform reinforcement learning to generate the positive response among the positive response and the negative response. Electronic device.

6. In Paragraph 2, When the above processor executes the above instructions, By inputting the first simulation inputs into the first character model, first outputs are generated, and Based on the first simulation outputs, an evaluation result for one or more evaluation items of the first outputs is generated, and the one or more evaluation items include conversational ability, role appeal, and / or character consistency, and Based on the above evaluation results, configured to train the first character model through supervised fine-tuning (SFT), Electronic device.

7. In Paragraph 2, When the above processor executes the above instructions, By inputting the above second simulation inputs into the above management model, second outputs are generated, and Based on the second simulation outputs, an evaluation result for one or more evaluation items of the second outputs is generated, and the one or more evaluation items include a character definition, risk level, and / or error rate, and Based on the above evaluation results, configured to train the above management model, Electronic device.

8. In Paragraph 1, The above training data includes exemplary outputs for the above inputs for feedback on the above outputs that deviate from the role of the above non-player character, Electronic device.

9. In Paragraph 1, When the above processor executes the above instructions, By inputting the above input and the above output into the management model, it is determined whether the output generated by the first character model conforms to the personality definition of the first character, and Based on the determination that the output generated by the first character model does not conform to the personality definition of the first character, the first character model is configured to generate training data for the first character model that instructs the regeneration of information regarding the personality of the first character and the output for the input. Electronic device.

10. In Paragraph 9, When the above processor executes the above instructions, Configured to generate additional training data for the first character model by inputting another output regenerated by the first character model into the management model, Electronic device.

11. A method of an electronic device including memory and a processor, An action of receiving input for a query regarding the first non-player character among multiple non-player characters, The operation of generating an output representing a response to the query generated through a first character model associated with the first non-player character among a plurality of character models including a neural network for each of the plurality of non-player characters stored in the memory, An operation of generating training data for the first character model by inputting the input and the output to a management model that manages the first character model among management models including other neural networks for managing the plurality of character models, and The operation of retraining the first character model using the training data so that the output for the input generated by the first character model is adjusted, method.

12. In Paragraph 11, The operation of obtaining simulation outputs for simulation inputs through a model distillation model, which is a large language model containing a greater number of parameters than the first character model and the management model, An operation of learning the first character model using the first simulation inputs and first simulation outputs classified into first categories among the simulation inputs and simulation outputs, and The operation of learning the management model using second simulation inputs and second simulation outputs classified into second categories distinguished from the first categories among the simulation inputs and simulation outputs, method.

13. In Paragraph 12, The operation of generating first outputs by inputting the first simulation inputs into the first character model, Based on the first simulation outputs, an evaluation result for one or more evaluation items of the first outputs is generated, and the one or more evaluation items include actions such as conversational ability, role appeal, and / or character consistency, and Based on the above evaluation results, the operation of training the first character model through supervised fine-tuning (SFT) method.

14. In Paragraph 12, The operation of generating second outputs by inputting the second simulation inputs into the management model, An operation to generate an evaluation result for one or more evaluation items of the second outputs based on the second simulation outputs, wherein the one or more evaluation items include a character definition, a risk level, and / or an error rate, and Based on the above evaluation results, including an operation to learn the above management model, method.

15. In Paragraph 11, An operation of determining whether the output generated by the first character model conforms to the personality definition of the first character by inputting the above input and the above output into the management model, and Based on the determination that the output generated by the first character model does not conform to the personality definition of the first character, the operation of generating training data for the first character model that directs the regeneration of information regarding the personality of the first character and the output for the input is included. method.