Metaverse generation network for providing three-dimensional lesson environment

The metaverse creation network addresses the lack of immersive teaching environments by using scenario, object, and quest generation models to create interactive three-dimensional spaces that enhance learning through realistic scenarios and quests.

WO2025159275A1PCT designated stage Publication Date: 2025-07-31RABBIT HOLE CO LTD

Patent Information

Application Number
PCT/KR2024/015637
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-07-08
Filing Date
2024-10-15
Publication Date
2025-07-31

AI Technical Summary

Technical Problem

Existing technologies lack the ability to create immersive and interactive three-dimensional teaching environments that reflect real-life scenarios, lacking motivation for learners and effective interaction with virtual environments.

Method used

A metaverse creation network utilizing a SCENE language model for generating scenario scripts, an object selection model for selecting NPCs and objects, a space layout setting model for rendering, and a quest generation model to create a three-dimensional teaching environment with interactive quests and conversations.

Benefits of technology

Enables the creation of a highly usable three-dimensional teaching environment that reflects real-life settings, providing motivation for learning through interactive quests and conversations, enhancing the educational experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure KR2024015637_31072025_PF_FP_ABST
    Figure KR2024015637_31072025_PF_FP_ABST
Patent Text Reader

Abstract

The present invention provides a metaverse generation network for providing a three-dimensional lesson environment. According to the present invention, provided may be a metaverse generation network for providing a three-dimensional lesson environment, the network allowing generation of a digital space of a three-dimensional metaverse through only input of a text related to various environment settings, and being capable of generating, as well as providing motivation to learn, a highly useful three-dimensional metaverse lesson environment in various digital spaces reflecting a real-life environment.
Need to check novelty before this filing date? Find Prior Art

Description

A metaverse-generated network to provide a 3D teaching environment

[0001] The present invention relates to a metaverse creation network for providing a three-dimensional teaching environment.

[0002] As technological advancements usher in the full-fledged metaverse era, extensive research is being conducted to provide metaverse-related services. Interest is growing in building metaverse platforms and metaverse ecosystems, and the establishment or transformation of specialized metaverse companies is underway.

[0003] One embodiment of the present invention includes a metaverse creation network for providing a three-dimensional teaching environment, which can create a digital space of a three-dimensional metaverse with only text input regarding various environmental settings, and can create a highly usable three-dimensional metaverse teaching environment in various digital spaces reflecting real-life environments, along with motivation for learning.

[0004] In order to solve the above and other problems, the metaverse creation network of the present invention,

[0005] As a metaverse creation network to provide a 3D teaching environment,

[0006] A SCENE language model for generating scenario scripts, which are text-based dialogue / action scripts arranged chronologically about conversations or actions between a tutor and an NPC (non-player character) from prompt inputs about a set situation;

[0007] An object selection language model for selecting NPCs and surrounding objects to interact with the NPCs as objects to be implemented as 3D views in the digital space of the metaverse from a number of digital assets stored in a database, using the generated scenario script as input;

[0008] A space layout setting model for setting a rendering area by inputting the object names of the selected NPC and surrounding objects and inferring the rendering location and rendering scale at which the selected NPC and surrounding objects will be rendered in the digital space of the metaverse; and

[0009] A quest generation language model is included for generating a list of quests in the form of a list assigned as a task of an instructor from a generated scenario script, and for generating a quest list of a higher concept encompassing dialogue / action scripts arranged in time series from the generated scenario script.

[0010] According to one embodiment of the present invention, a metaverse creation network generates a 3D teaching environment by generating rendering data that provides 3D and / or 2D views of the selected NPCs, surrounding objects, and background images surrounding them, including selection of NPCs (non-player characters) as counterparts of conversation / action interactions with an instructor in the digital space from text input of environment settings for a digital space in which a metaverse is to be implemented, and a quest regarding conversation / action interactions as a task assigned to the instructor in the digital space of the set environment, and generating a conversation / action script regarding a response in response to the conversation / action script displayed from the instructor, and evaluating whether the quest assigned to the instructor has been accomplished from the conversation / action script displayed from the instructor, thereby enabling the creation of a digital space of a 3D metaverse with only text input regarding various environment settings, and providing a teaching environment that is highly usable in various digital spaces that reflect real-life environments, along with motivation of the instructor for learning. Can be provided.

[0011] FIGS. 1A and 1B are schematic diagrams showing the architecture of a Generative Pre-trained Transformer (GPT) model, which is an architecture that can be applied to a large-scale language model (LLM, language model) according to one embodiment of the present invention. FIG. 1A is a diagram showing the architecture of a GPT model in which a plurality of decoder blocks are stacked, and a natural language processing that calculates each word expression as a sum of weights according to the degree of association with a previously input word according to the input order of a word sequence, and FIG. 1B is a diagram showing a more detailed architecture of the decoder block shown in FIG. 1A, and a diagram showing a natural language processing that predicts the next word of an input word sequence.

[0012] FIG. 2 is a diagram showing a natural language processing that predicts the word that will come after a special token (S), as an example of a currently input word, and a diagram showing a natural language processing that calculates a prediction probability by taking a softmax function from the transpose vector product of the output value of the special token (S) and the 50,257 token embeddings that form the vocabulary.

[0013] FIG. 3 illustrates a diagram for explaining contrastive pre-training (contrastive language image pre-training, CLIP), which trains a text encoder and an image encoder to match each other by inputting a plurality of images forming a pair and text corresponding to captioning describing the plurality of images, and training the text embeddings output from each text encoder and the image embeddings output from the image encoder to maximize the cosine similarity between pairs that match each other and minimize the cosine similarity between pairs that do not match each other.

[0014] Figure 4a is a diagram illustrating a noising process that generates isotropic Gaussian noise in which a specific pattern of the original image is collapsed by adding noise according to a noise schedule gradually over time steps to a clear original image (swiss roll) of a specific pattern as a forward process of diffusion.

[0015] Figure 4b is a diagram illustrating a denoising process, which is a reverse process of diffusion, that is, an image generation process that restores the pattern of the original image while generating a less noisy image according to the time step while removing noise predicted from a perfect Gaussian (isotropic Gaussian noise).

[0016] Figure 5 illustrates a diagram for explaining the architecture of U-net for predicting the probability distribution of a relatively less noisy image from a noisy image at the current time step by taking as input an image with added noise and the time step of the corresponding image, and predicting the added noise at the current time step.

[0017] FIG. 6 is a diagram for explaining a latent diffusion model (LDM) in which diffusion is implemented in a latent space that is lower dimensional than the pixel space as an image generation model to which diffusion is applied, and a diagram for explaining a conditioned diffusion model in which conditioning for image generation is applied through an encoder is shown.

[0018] FIG. 7 is a diagram showing the architecture of the DALL-E-2 model that can be applied as a text-to-image model in one embodiment of the present invention.

[0019] Figure 8 illustrates a diagram for explaining the diffusion of a prior that takes the text embedding in Figure 7 as input and converts the input text embedding into a matching image embedding.

[0020] FIG. 9 is a diagram illustrating diffusion of a decoder that takes image embedding and text embedding as inputs in FIG. 7 and outputs an image as an output of text-to-image.

[0021] FIG. 10 is a diagram illustrating the overall configuration of a metaverse creation network according to one embodiment of the present invention.

[0022] FIG. 11 illustrates an example of a scenario script including a dialogue / action script between an NPC and an instructor inferred by using a set situation as a prompt input, generated from a SCENE language model according to one embodiment of the present invention, a drawing exemplarily showing a plurality of quest blocks predicted from a quest generation language model by using a scenario script generated from the SCENE language model as an input, and a quest predicted as a superordinate concept or topic from each quest block.

[0023] FIG. 12 illustrates a diagram for explaining similarity analysis through contextual embedding of a GPT model as a natural language processing for NPC selection, as an object selection language model according to one embodiment of the present invention.

[0024] FIGS. 13a to 13c are drawings showing, by way of example, a rendering area of ​​an NPC (NPC) and a rendering area of ​​surrounding objects (surrounding objects) generated from a space layout setting model according to one embodiment of the present invention.

[0025] FIGS. 14A to 14C are drawings showing examples of background images surrounding NPCs and surrounding objects generated as text-to-image, as different background images generated from a background image generation model according to one embodiment of the present invention.

[0026] FIGS. 15A to 16C illustrate an action script that forms a scenario script generated from a SCENE language model using a specific situation set in a metaverse creation network of the present invention as a prompt input, and an attribute data of a plurality of digital assets stored in a database, which exemplarily show interactions of characters selected as NPCs or surrounding objects based on the similarity between a specified motion of an NPC and a specified physical / digital interaction of surrounding objects, and an exemplarily show quests inferred from a scenario script generated from a SCENE language model using each set situation as a prompt input.

[0027] The metaverse creation network of the present invention is

[0028] As a metaverse creation network to provide a 3D teaching environment,

[0029] A SCENE language model for generating scenario scripts, which are text-based dialogue / action scripts arranged chronologically about conversations or actions between a tutor and an NPC (non-player character) from prompt inputs about a set situation;

[0030] An object selection language model for selecting NPCs and surrounding objects to interact with the NPCs as objects to be implemented as 3D views in the digital space of the metaverse from a number of digital assets stored in a database, using the generated scenario script as input;

[0031] A space layout setting model for setting a rendering area by inputting the object names of the selected NPC and surrounding objects and inferring the rendering location and rendering scale at which the selected NPC and surrounding objects will be rendered in the digital space of the metaverse; and

[0032] A quest generation language model is included for generating a list of quests in the form of a list assigned as a task of an instructor from a generated scenario script, and for generating a quest list of a higher concept encompassing dialogue / action scripts arranged in time series from the generated scenario script.

[0033] For example, the above quest generation language model can extract high-level concepts from the dialogue / action scripts forming the generated scenario script and generate a quest list in which the extracted topics or quests are arranged chronologically.

[0034] For example, the above quest generation language model,

[0035] A first natural language processing for classifying the chronologically arranged dialogue / action scripts from the generated scenario script into quest blocks according to their topics; and

[0036] A second natural language processing can be performed to generate a topic or quest of a higher level concept encompassing at least one or more dialogue / action scripts classified into the same quest block.

[0037] For example, the above quest generation language model,

[0038] The first natural language processing can be performed by analyzing the similarity between neighboring dialogue / action scripts based on the order of appearance of multiple dialogue / action scripts arranged chronologically in the generated scenario script.

[0039] For example, the above quest generation language model,

[0040] In order to set the boundaries of the quest blocks according to the order of appearance of multiple dialogue / action scripts arranged chronologically in the generated scenario script, pairs of adjacent dialogue / action scripts that are the subject of similarity analysis can be set to overlap each other.

[0041] For example, the above quest generation language model,

[0042] The second natural language processing can be performed by extracting key words or clues that appear repeatedly in multiple dialogue / action scripts classified into the same quest block.

[0043] For example, the above quest generation language model,

[0044] First and second context vectors can be derived from contextual embeddings of first and second word sequences forming multiple dialogue / action scripts classified into the same quest block, and a topic or superordinate concept quest can be generated by categorizing the first and second word sequences from the distribution of multiple context vectors mapped onto the contextual embedding space.

[0045] For example, the SCENE language model can receive a style setting or a demonstration of few-shot learning for generating a scenario script in which the subject of each dialogue / action script is specified from a prompt in which the set situation is input.

[0046] For example, the above object selection language model,

[0047] Natural language processing for selecting the above NPC and natural language processing for selecting the above surrounding objects can be performed as different processes.

[0048] For example, the above object selection language model is used for natural language processing for the selection of the NPC.

[0049] Taking as input a first word sequence regarding the above scenario script, a first context vector is generated that includes association information between each word forming the first word sequence,

[0050] An attention architecture may be included to input a second word sequence regarding entity names of a plurality of digital assets stored in the above database and attribute data stored in association with the entity names, and generate a second context vector including association information between each word forming the second word sequence.

[0051] For example, the above object selection language model, in natural language processing for the selection of the NPC,

[0052] The NPC can be selected from a similarity analysis between the first and second context vectors that are contextually embedded with respect to the first word sequence regarding the above scenario script and the second word sequence regarding the digital asset stored on the above database.

[0053] For example, the above object selection language model,

[0054] In the natural language processing for selecting the above NPC, the similarity between the action script of the generated scenario script and the dynamic motion specified for each NPC as the attribute data of the NPC is analyzed,

[0055] In natural language processing for selecting the above surrounding objects, the similarity between the action script of the generated scenario script and the physical interaction or digital interaction with the NPC specified for each surrounding object as attribute data of the surrounding object can be analyzed.

[0056] For example, the above object selection language model, in natural language processing for selection of the surrounding objects,

[0057] Natural language processing for selection of the surrounding objects can be performed based on a similarity analysis between each word forming a first word sequence regarding the above scenario script and each word forming a second word sequence regarding entity names of a plurality of digital assets stored in the above database and attribute data stored in association with the entity names.

[0058] For example, the above space layout setting model,

[0059] Create text-to-image by inputting the object names of selected NPCs and surrounding objects as text; and

[0060] Region extraction can be performed to extract the image region of each NPC and the image region of surrounding objects from the generated image.

[0061] For example, the above space layout setting model,

[0062] To create the above text-to-image,

[0063] The method may include a diffusion architecture that implements a denoising process that creates a less noisy image that is relatively closer to the original image from a noisy image so that the pattern of the original image is restored, as a reverse process of a noising process that gradually collapses the pattern of the original image by adding noise defined from a noise schedule while advancing the time step from the original image.

[0064] For example, the above space layout setting model,

[0065] To create the above text-to-image,

[0066] An encoder that learns a multi-modal embedding space of text-images to encode the input text into a text embedding by inputting a text containing the object names of the above NPC and surrounding objects;

[0067] A prior for converting the above text embedding into an image embedding that matches the above text embedding; and

[0068] It may include a decoder for generating an image as an output of text-to-image from the image embedding converted from the above-mentioned prior.

[0069] For example, the prior includes an architecture of conditioning diffusion that injects the text embedding as a condition for generating the image embedding,

[0070] The above decoder may include an architecture of conditioning diffusion that injects the image embedding as a condition for generating an image as an output of the text-to-image.

[0071] For example, the above space layout setting model, for extracting the area,

[0072] Object detection or image segmentation can be performed to distinguish the image area of ​​each NPC and the image area of ​​surrounding objects from the image generated through the text-to-image generation above.

[0073] For example, in the object detection, the boundary of each NPC image area and the image area of ​​the surrounding objects is predicted in the form of a bounding box surrounding the NPC image area and the image area of ​​the surrounding objects on the generated image,

[0074] In the image segmentation, the NPC image area and the image area of ​​the surrounding objects can be predicted on a pixel basis on the generated image.

[0075] For example, the above space layout setting model,

[0076] As a rendering area in which the above NPC and surrounding objects are to be rendered, data regarding the rendering area including the center position of the bounding box surrounding the rendering area and the width x height size of the bounding box can be generated.

[0077] For example, a metaverse creation network according to one embodiment of the present invention,

[0078] It is for generating a background image to provide an environment surrounding the above NPC and surrounding objects, and may further include a background image generation model that implements text-to-image generation by inputting text.

[0079] For example, a metaverse creation network according to one embodiment of the present invention,

[0080] A transmission code or transmission data including rendering data for rendering a three-dimensional view of an NPC and surrounding objects selected from an object selection language model that takes as input a scenario script generated from the SCENE language model, data regarding a rendering area in which the NPC and surrounding objects are to be rendered generated from a space layout setting model that takes as input a text in which the entity names of the selected NPC and surrounding objects are to be rendered, a quest list generated from a quest generation language model that takes as input a scenario script generated from the SCENE language model, and two-dimensional video frame data regarding a background image that provides an environment surrounding the selected NPC and surrounding objects can be transmitted to a local area of ​​an instructor where a digital space of the metaverse is implemented.

[0081] For example, the above transmission code or transmission data may further include animation data for implementing dynamic motion expressed as an action script of a scenario script generated from a SCENE language model among dynamic motions specified as attribute data of a selected NPC.

[0082] For example, the transmission code or transmission data may further include a call function for transmitting a dialogue / action script expressed by the instructor toward the processing server.

[0083] For example, a metaverse creation network according to one embodiment of the present invention,

[0084] It may further include an interactive language model for generating a dialogue / action script as a response to a dialogue / action script expressed by the above instructor.

[0085] For example, the above conversational language model,

[0086] Natural language processing can be performed to evaluate whether a quest generated from a quest generation language model has been achieved from a dialogue / action script expressed from the above instructor.

[0087] For example, the above conversational language model,

[0088] The achievement of the quest can be evaluated by analyzing the similarity between the word sequence forming the quest and the word sequence forming the dialogue / action script expressed by the instructor.

[0089] Hereinafter, with reference to the attached drawings, a metaverse providing system for providing a three-dimensional teaching environment according to one embodiment of the present invention will be described.

[0090] Large Language Model (LLM)

[0091] The metaverse providing system according to one embodiment of the present invention may include a large language model (LLM) such as a SCENE language model, an object selection language model, a quest generation language model, a conversational language model, etc., as described below. Hereinafter, an architecture applicable to a large language model (hereinafter, “language model”) will be described, and more specific technical details of each language model will be described in more detail later.

[0092] FIGS. 1A and 1B are schematic diagrams showing the architecture of a Generative Pre-trained Transformer (GPT) model, which is an architecture that can be applied to a large-scale language model (LLM, language model) according to one embodiment of the present invention. FIG. 1A is a diagram showing the architecture of a GPT model in which a plurality of decoder blocks are stacked, and a natural language processing that calculates each word expression as a sum of weights according to the degree of association with a previously input word according to the input order of a word sequence, and FIG. 1B is a diagram showing a more detailed architecture of the decoder block shown in FIG. 1A, and a diagram showing a natural language processing that predicts the next word of an input word sequence.

[0093] FIG. 2 is a diagram showing a natural language processing that predicts the word that will come after a special token (S), as an example of a currently input word, and a diagram showing a natural language processing that calculates a prediction probability by taking a softmax function from the transpose vector product of the output value of the special token (S) and the 50,257 token embeddings that form the vocabulary.

[0094] Referring to FIGS. 1A and 2, the language model may have an architecture in which a decoder structure of a transformer including an encoder and a decoder is based on the decoder structure of the transformer and an architecture in which a plurality of decoder layers of 12 are stacked. For example, a language model according to one embodiment of the present invention may have an auto-regressive feature that sequentially predicts the next word token from a sequential input of word tokens forming a sentence sequence. In order to perform downstream tasks such as translation, summarization, and question-answering, in-context learning may be applied to perform downstream tasks such as translation or summarization by predicting a given task from the input of a prompt without fine-tuning for a separate update of the parameters of the language model, and a large web-based data set, web text, may be constructed and utilized for dictionary learning. For example, web text, which is a data set used for pre-learning, can be pre-learned by including information on various downstream tasks such as translation, summarization, and question-answering, and accordingly, it is possible to process a given task with zero-shot learning from an input sequence entered in a prompt without fine tuning for updating separate parameters for processing data of the downstream task to be processed in the inference stage.

[0095] In various embodiments of the present invention, in-context learning can be applied to the language model. For example, one or more demonstrations can be added to the input sequence input to the prompt so that the language model can make inferences by referring to information about the task, so that the language model can predict the task. Depending on the number of examples input to the prompt, processing for a given task can be possible with zero-shot learning or few-shot learning, and such demonstrations can be added in connection with the input sequence. For example, in one embodiment of the present invention, reinforcement learning can be applied to the language model in a way that differential rewards are given according to the evaluation of a labeler or the like so that misalignment (hallucination, bias, etc.) that does not match human instructions does not occur through Reinforcement learning by human feedback (RLHF).

[0096] A language model according to one embodiment of the present invention may include a pre-trained model, and more specifically, may include a Generative Pre-trained Transformer (GPT) model. For example, as a language model of the present invention, the pre-trained model is a model trained using an unlabeled data set (an unlabeled large corpus, a large-scale text such as web text) as training data, and can acquire prior knowledge about general-purpose natural language processing. In various embodiments of the present invention, the large-scale language model can be fine-tuned by updating the model's parameters using a labeled data set as training data to suit a specific downstream task.

[0097] For example, in various embodiments of the present invention, the language model may be reinforced through RLHF (reinforcement learning by human feedback) through interaction with a labeler, for example, fine tuning of the labeler through supervised learning (supervised fine tuning) and reinforcement learning (reward model) according to the evaluation and reward of the labeler. In addition, a language model according to one embodiment of the present invention, for example, a conversational language model described below, may additionally provide a conversational interface as a model fine-tuned through supervised learning to be optimized for a conversational agent (supervised fine tuning). For example, a language model according to one embodiment of the present invention may include a GPT (Generative Pre-trained Transformer) model.

[0098] Attention

[0099] Referring to FIG. 1, in the GPT model, decoder blocks including multi-head attention and a feedforward network that are stacked in a plurality of attention blocks are accumulated and stacked in a plurality of stacks, and a plurality of decoder blocks can be connected so that the result value of each decoder block is input to the next connected decoder block, and a word or token that will come out after the word or token input to the first decoder block can be predicted from the result value of the last decoder block.

[0100] Similar to the above GPT model, the BERT (bidirectional encoder representations from transformer) model, which basically takes the structure of a transformer, is a bidirectional language model including forward and backward, and may include self-attention in the encoder of the transformer and encoder-decoder cross-attention in the decoder of the transformer. However, since the GPT model does not include the encoder structure of the transformer and only includes the decoder structure of the transformer, it may not include the encoder-decoder cross-attention of the BERT model, and may include masked self-attention as a forward language model. That is, the GPT model may include an architecture in which a plurality of decoder blocks including masked multi-headed self-attention and a feedforward network are accumulated and stacked. For example, in each decoder block included in the above GPT model, the parameters of the mask multi-head self-attention (weight vectors that are operated with each word or token to produce a query, key, and value) may have different values.

[0101] In the above mask self-attention, words or tokens after the word or token to be predicted can be masked according to the input order, and in the above mask self-attention, a score is calculated from a scaled dot product between a query vector and a key vector corresponding to the word or token to be predicted, and a softmax function is applied to the calculated score to calculate the result value of the query vector from the weighted sum of the value vectors with the normalized score as a weight. In this case, for the masked word or token, the score before applying the softmax function can be replaced with an infinite negative value so that the normalized score practically converges to zero.

[0102] As a language model according to one embodiment of the present invention, in the GPT model, an input embedding that is a sum of a token embedding, a positional embedding, and a segment embedding can be input, and the GPT model can input a special token corresponding to the beginning of a sentence sequence (for example, <sos>) from which a special token (e.g., <eos>) can be repeated until the next word or token is produced, and the probability for the next word or token can be calculated, and the next word or token can be predicted based on the calculated probability. For example, the GPT model may include parameters learned to maximize the objective function with the likelihood of the next word or token appearing as the objective function.

[0103] As a language model according to one embodiment of the present invention, the GPT model can predict the probability of each token forming a token embedding matrix to appear next from an operation between a result value from the last decoder block, that is, an expression value for each word or token, and a token embedding matrix including the entire vocabulary, and can predict the token with the highest prediction probability by applying a softmax function as the word or token to appear next.

[0104] <Multimodal AI model or multimodal embedding that learns the multimodal embedding space of text and images>

[0105] In one embodiment of the present invention, a model trained to connect pairs of images and texts describing the images, i.e., image-text pairs, may be applied, and as described below, in one embodiment of the present invention, a multi-modal AI model that has learned a multi-modal embedding space of text-images may be applied to a space layout setting model or a background image generation model.

[0106] FIG. 3 illustrates a diagram for explaining contrastive pre-training (contrastive language image pre-training, CLIP), which trains a text encoder and an image encoder to match each other by inputting a plurality of images forming a pair and text corresponding to captioning describing the plurality of images, and training the text embeddings output from each text encoder and the image embeddings output from the image encoder to maximize the cosine similarity between pairs that match each other and minimize the cosine similarity between pairs that do not match each other.

[0107] For example, in one embodiment of the present invention, the multi-modal AI model can convert an image and a text into embeddings respectively, and then predict a text embedding that is closest to the image embedding, or conversely, predict a text embedding that is closest to the text embedding. For example, in one embodiment of the present invention, as a multi-modal AI model trained to connect image-text pairs, contrastive language image pre-training (CLIP) can be applied. In one embodiment of the present invention, the multi-modal AI model can use web-based raw data, that is, a large amount of raw data without human annotation, as a training data set. For example, the model can be trained using a web-based image and a captioning text describing the image as training data as image-text pairs, and can be trained using contrastive learning.

[0108] The above multi-modal AI model may include an image encoder for producing an image representation (visual representation) for an input image and a text encoder for producing a text representation (language representation) for an input text, and inputs pairs of image-text pairs forming n pairs to each of the image encoder and the text encoder, thereby generating a mini-batch (n image training data) of images and a mini-batch (n text training data) of texts, and then the text-image pairs that form a pair with each other are designated as positive pairs, and the text-image pairs that do not form a pair with each other are designated as negative pairs, so that the cosine similarity for the n positive pairs is maximized, and n 2 - The cosine similarity for n negative pairs can be learned to be minimal, and the image encoder and the text encoder can be trained together from such contrastive learning. For example, in one embodiment of the present invention, the multi-modal AI model can learn a multi-modal image-text embedding space, and for example, the image encoder and the text encoder can be trained to map the respective image embeddings and text embeddings from an image encoder that embeds an input image into an image space and a text encoder that embeds an input text into a text space to each other.

[0109] Image generation model

[0110] An image generation model is a model that has learned the distribution of data, and can predict the data distribution of an image to be generated from data learned from a large amount of training data, and can generate a new image by sampling from the region with the highest likelihood from the predicted data distribution. For example, in one embodiment of the present invention, the image generation model can be applied to a space layout setting model or a background image generation model, as described below. For example, in one embodiment of the present invention, the image generation model can include a diffusion model capable of generating diverse and high-quality images.

[0111] Figure 4a is a diagram illustrating a noising process that generates isotropic Gaussian noise in which a specific pattern of the original image is collapsed by adding noise according to a noise schedule gradually over time steps to a clear original image (swiss roll) of a specific pattern as a forward process of diffusion.

[0112] Figure 4b is a diagram illustrating a denoising process, which is a reverse process of diffusion, that is, an image generation process that restores the pattern of the original image while generating a less noisy image according to the time step while removing noise predicted from a perfect Gaussian (isotropic Gaussian noise).

[0113] Figure 5 illustrates a diagram for explaining the architecture of U-net for predicting the probability distribution of a relatively less noisy image from a noisy image at the current time step by taking as input an image with added noise and the time step of the corresponding image, and predicting the added noise at the current time step.

[0114] FIG. 6 is a diagram for explaining a latent diffusion model (LDM) in which diffusion is implemented in a latent space that is lower dimensional than the pixel space as an image generation model to which diffusion is applied, and a diagram for explaining a conditioned diffusion model in which conditioning for image generation is applied through an encoder is shown.

[0115] Referring to FIGS. 4A to 6, in the diffusion for image generation, a diffusion process may be defined as a forward process including an iterative Markov chain of a Gaussian probability distribution in which a Gaussian noise (ε~N(0,1), random noise) defined from a noise schedule is added to the image of the previous time step while advancing discrete time steps from an original image having a specific pattern, so that a specific pattern of the original image is eliminated, and a denoising process may be defined as a reverse process including an iterative Markov chain of a Gaussian probability distribution in which a specific pattern of the original image is restored while removing the Gaussian noise (ε~N(0,1), random noise) added from the previous time step as a reverse process for the forward process, and an image (Xt) to which noise is added and a time step (t=0,1,...,T) of the corresponding image are input in the forward process, so that the added noise can be learned to be predicted, and a specific pattern of the original image can be extracted from the reverse process of removing the added noise. It is possible to perform a sampling process to create a restored image.

[0116] More specifically, as the forward process (diffusion process), when the input image (X0, an original image with a clear specific pattern) is sequentially added with a predetermined noise schedule (βt, for example, noise that gradually increases along a linear scale, a fixed noise schedule) while advancing the time step (t=0,1,...,T) and performing the forward process for a preset time step (for example, 1000 time steps, T=1000), the image (XT) of the preset time step can generate a complete Gaussian noise (normal Gaussian distribution, mean 0, variance 1, N(0,1)) in which the specific pattern of the original image is eliminated while converging to an isotropic Gaussian in all directions. For example, the forward process can be expressed as a chain (joint distribution) of conditional probability distributions for the image (Xt) of the current time step conditioned on the image (Xt-1) of the previous time step, and can be expressed as a chain (joint distribution) of Markov processes in which the state at the current time step (Xt) depends only on the state at the previous time step (Xt-1). More specifically, in the forward process, the conditional probability distribution q(Xt│Xt-1) for the image (Xt) of the current time step conditioned on the image (Xt-1) of the previous time step can be defined as a conditional Gaussian probability distribution with the following mean and variance.

[0117]

[0118] In the above forward process, an image at the current time step (Xt) can be generated by adding Gaussian noise (ε~N(0,1), random noise) while decreasing the value (Xt-1, e.g., pixel value, etc.) at the previous time step according to a noise schedule (βt).

[0119]

[0120] The above reverse process (denoising process) restores the original image (X0) from complete Gaussian noise (XT), which may correspond to a learning target in an actual image generation model and may correspond to sampling for image generation. More specifically, the above reverse process (denoising process, sampling process) may restore a specific pattern of the original image (X0) while gradually removing Gaussian noise from complete Gaussian noise (XT), and this reverse process may be expressed as a chain (joint distribution) of conditional probability distributions for the image (Xt-1) of the previous time step conditioned on the image (Xt) of the current time step, and may be expressed as a chain (joint distribution) of Markov processes in which the state of the current time step (Xt) depends only on the state of the previous time step (Xt-1). In the above reverse process, the conditional probability distribution q(Xt-1│Xt) for the image (Xt) of the current time step conditioned on the image (Xt-1) of the previous time step is difficult to directly derive the conditional probability distribution whose conditional time point is reversed from the conditional probability distribution q(Xt│Xt-1) of the forward process, unlike the forward process in which noise is added according to the noise schedule (βt). Therefore, the conditional probability distribution q(Xt-1│Xt) of the reverse process is derived from the conditional probability distribution P of the reverse process through a neural network model with a parameter (θ). θ It can be approximated as (Xt-1│Xt).

[0121]

[0122] The above reverse process may correspond to an image generation (sampling process) in the inference stage, and the above reverse process may correspond to a learning target of a neural network model for image generation, and the neural network model for image generation may be learned to predict a reverse process conditional probability distribution q(Xt-1│Xt), and among the conditional probability distributions q(Xt-1│Xt) of the reverse process, the variance is fixed (for example, σ t 2 =βt) Average (μ θ (Xt,t)) can be learned to predict, for example, the conditional probability distribution q(Xt-1│Xt) of the reverse process and the conditional probability distribution P of the reverse process approximated through a neural network model. θ An objective function for training a neural network model can be derived so that the KL divergence is minimized between (Xt-1│Xt), and the error between the predicted mean from the neural network model and the actual mean (e.g., target label) produced in the forward process can be minimized, for example, by using these errors as the objective function or loss function, and updating the parameters of the neural network model that approximates the reverse process from the descent gradient of the objective function or loss function. In the following, E may mean an expectation.

[0123]

[0124] At this time, the average can be expressed as follows, and the neural network model approximating the reverse process is not trained to directly predict the average of the less noisy image (Xt-1) at the previous time step from the image (Xt) at the current time step, but rather the Gaussian noise ε added from the image (Xt-1) at the previous time step. θ By predicting (Xt,t), the average of the less noisy image (Xt) at the current time step can be predicted, and as described above, the objective function of the neural network model can be configured to minimize the error between the average predicted from the neural network model and the actual average (target label) produced in the forward process.

[0125]

[0126]

[0127] Here,

[0128]

[0129]

[0130] Here,

[0131]

[0132] In learning a neural network model that approximates the reverse process as described above, the parameters of the neural network model can be updated by taking the gradient descent as follows.

[0133]

[0134] And, from the image (Xt) at the current time step with added noise from the Gaussian noise predicted from the neural network model learned as above, an image (Xt-1) at the previous time step with less noise can be generated, and more specifically, an image (Xt) at the current time step with predicted Gaussian noise (ε) can be generated. θ From (xt,t)), we can calculate the average of the image (Xt-1) at the previous time step with less noise, and have a fixed variance (σ t 2 =βt, for example, assuming a value such as noise schedule βt) and Gaussian noise for sampling (Z~N(0.1)) can be derived as follows.

[0135]

[0136] Here,

[0137]

[0138] The neural network model for approximating the above reverse process predicts the mean and variance of the image (Xt-1) at the previous time step, which is relatively less noisy, from the image (Xt) at the current time step, which is relatively noisier (Gaussian noise ε added from the image Xt at the current time step). θ By predicting (xt,t), the average of the image Xt-1 in the previous time step can be predicted), the probability distribution of the image in the previous time step can be predicted conditionally on the image in the current time step, and a specific pattern of the original image can be restored by generating an image in the previous time step that is closer to the original image, and it can be applied as a model for image generation.

[0139] For example, a neural network model for predicting Gaussian noise added to an image in a relatively less noisy previous time step, given an image in a current time step and the current time step as inputs, or in other words, a neural network model for predicting a probability distribution of an image in a relatively less noisy previous time step, given an image in a relatively more noisy current time step, can be formed with an architecture of a U-net structure. The architecture of the U-net structure can include a contracting path for generating a low-dimensional image through downsampling to extract a feature map or feature of an image from an input high-dimensional image (e.g., a noisy image), and an expanding path for generating a high-dimensional image, such as an input image, through upsampling to increase a matrix dimension from a low-dimensional image. In one embodiment of the present invention, a conditional probability distribution P that approximates a conditional probability distribution q(Xt-1│Xt) θ A neural network model for predicting (Xt-1│Xt) can predict the conditional probability distribution of the image at the previous time step with relatively less noise by inputting the current time step together with the image at the current time step. For this purpose, the current time step (t) can be input into the neural network model implemented with the U-net architecture through embedding, and for example, can be input into multiple layers of the U-net architecture. As described below, a conditional probability distribution P that approximates the conditional probability distribution q(Xt-1│Xt) θ The neural network model for producing (Xt-1│Xt) is an image generation model, and conditioning such as text that conditions image generation can be injected, and in one embodiment of the present invention, text conditioning can be input as text embedding through a text encoder, and an attention structure can be included for performing cross-attention using data (latent expression) of a target image to be generated as a query and the input text embedding as a key and value to increase the relevance with the input text embedding.

[0140] The image generation model to which the above diffusion is applied can be implemented in the pixel space of the image, but it can also be implemented in a latent space that is lower dimensional than the pixel space (LDM, latent diffusion model, stable diffusion). For example, while it is possible to generate diverse and high-quality images, it takes a relatively long time to generate images compared to other image generation models (GAN, VAE, flow). In addition, when learning of diffusion is performed in two stages: perceptual compression, in which high-frequency details disappear but semantics are maintained, and semantic compression, in which the essence of actual data is learned, it takes a relatively long time in perceptual compression. Therefore, rather than performing diffusion (pixel space model) on a pixel-by-pixel basis of the image, diffusion can be performed in a latent space that is lower dimensional (latent space model). The forward and reverse processes described above can be implemented in a latent space that is lower dimensional. Can be.

[0141] In LDM (Latent Diffusion Model) or stable diffusion, the forward process and reverse process of diffusion may be implemented in front and back of the neural network block, and may further include an encoder for encoding the pixel space of the input image into the latent space (encoding the pixel data of the input image into a latent representation), and a decoder for decoding the generated image into the pixel space on the latent space (decoding the latent representation of the generated image into a representation of the pixel space), and the encoder and decoder may be implemented as a pair of encoders and decoders of an autoencoder. The forward process and reverse process may be implemented on the latent space between the encoder and the decoder, and for example, as described above, a conditional probability distribution P approximating the conditional probability distribution q(Xt-1│Xt) θ A neural network model of the U-net architecture for generating (Xt-1│Xt) can be implemented between the encoder and the decoder. In the neural network model for image generation as illustrated in the figure, the input to the encoder and the output from the decoder can be made in the pixel space, and the forward process and the reverse process of diffusion between the encoder and the decoder can be implemented in the latency space, and in the reverse process (denoising process), the image at the current time step and the current time step are input to predict a conditional probability distribution to generate an image at a previous time step that is less noisy (for example, Gaussian noise ε θ (Xt,t) to predict the average in the previous time step that is less noisy) A neural network model of the U-net architecture can be implemented. At this time, a text encoder for injecting conditioning on the generated image (conditioning diffusion) can be connected to the U-net structure on the reserve process (denoising process), and in one embodiment of the present invention, the text encoder can be implemented as a multi-modal AI model that has learned a multi-modal embedding space of image-text as described above, and for example, in one embodiment of the present invention, as a multi-modal AI model, it can be a text encoder for CLIP text embedding.

[0142] For example, the text encoder can generate a text embedding that can match an image embedding from an input text, and the text embedding from the text encoder can inject conditioning for image generation through cross-attention with the target image to be generated. More specifically, cross-attention can be performed using the data (latent representation) of the target image as a query and the input text embedding as a key and value. In addition, Gaussian noise ε predicted from the reverse process of the diffusion is applied between the reverse process of the diffusion and the decoder. θ Conditional probability distribution P over images at relatively less noisy previous time steps from (Xt,t) θ A sampling process can be linked to generate a less noisy image from the previous time step by sampling from the mean and variance of (Xt-1│Xt).

[0143] In the above LDM (latent diffusion model) or stable diffusion, the encoder for latent representation and the forward process (diffusion process), which may not be necessary for the inference process as a configuration for learning, can be deployed without being equipped in order to make the neural network model lightweight after learning, and the probability distribution of the image in the previous time step that is less noisy than the image in the current time step can be predicted from the reverse process including the learned parameters.

[0144] <Text-to-Image Generation Model>

[0145] FIG. 7 is a diagram showing the architecture of the DALL-E-2 model that can be applied as a text-to-image model in one embodiment of the present invention.

[0146] Figure 8 illustrates a diagram for explaining the diffusion of a prior that takes the text embedding in Figure 7 as input and converts the input text embedding into a matching image embedding.

[0147] FIG. 9 is a diagram illustrating diffusion of a decoder that takes image embedding and text embedding as inputs in FIG. 7 and outputs an image as an output of text-to-image.

[0148] Referring to FIGS. 7 to 9, in one embodiment of the present invention, a generative model for text-to-image for generating an image from an input text can be applied to a space layout setting model and a background image generation model in one embodiment of the present invention, as described below. More specifically, in one embodiment of the present invention, the space layout setting model and the background image generation model, as described below, can be implemented as a text-to-image generation model that encompasses a multi-modal AI model that has learned the multi-modal embedding space of the text-to-image described above, and an image generation model. For example, in one embodiment of the present invention, the DALL-E-2 model can be applied as the text-to-image generation model.

[0149] In one embodiment of the present invention, a text-to-image generation model may include a text encoder for encoding an input text into a text embedding, a prior for converting the text embedding into an image embedding matching the text embedding, for example, an image embedding matching the text embedding in a multi-modal image-to-text embedding space, and a decoder for generating an image as an output of the text-to-image from the image embedding converted from the prior. For example, in one embodiment of the present invention, the text encoder for encoding the text into the text embedding may be implemented as a text encoder as a multi-modal AI model that has learned the multi-modal image-to-text embedding space as described above.

[0150] In one embodiment of the present invention, the prior can convert a text embedding output from a text encoder into an image embedding matching the corresponding text embedding in a multi-modal embedding space of text and images, and can generate an image embedding matching the text embedding from a diffusion that applies the text embedding output from the text encoder as a conditioning of the generation.

[0151] For example, in one embodiment of the present invention, the prior may be trained to generate image embeddings that match image-text pairs by taking as input text embeddings output from a text encoder, for example, to generate less noisy image embeddings from relatively noisy image embeddings, and the text embeddings may be fed as conditions of a diffusion model. For example, for the training data of the prior, a text (text or captioning) that can describe an input image can be output from a multi-modal AI model that has learned the multi-modal embedding space of text-image described above, and text and images that are matched as image-text pairs in this way, and a noisy image that has noise added to the image can be input into a text encoder and an image encoder, respectively, to generate a text embedding, an image embedding, and a noisy image embedding, and the prior can be trained to generate a less noisy image embedding from the noisy image embedding with the text embedding as a condition, and an image embedding that matches the corresponding text embedding can be output for a text embedding that has not been learned through the prior trained in this way.

[0152] The decoder can generate an image as an output of text-to-image from the image embedding converted from the prior. For example, the decoder can be implemented as a neural network model (image generation model) for image generation using diffusion, and can be trained to generate a less noisy image from a relatively noisy image by inputting the text embedding output from the text encoder and the image embedding converted from the prior, and can generate an image as an output of text-to-image through a reverse process that restores a clear original image (the pattern of the original image). For example, the decoder can generate a less noisy image from a noisy image by conditionally (e.g., concatenating) the text embedding output from the text encoder and the image embedding converted from the prior, and can generate an image as an output of text-to-image that matches the input text.

[0153] Overview of the Metaverse Creation Network

[0154] FIG. 10 is a diagram illustrating the overall configuration of a metaverse creation network according to one embodiment of the present invention.

[0155] Referring to FIG. 10, a metaverse creation network according to one embodiment of the present invention is a metaverse creation network for providing a three-dimensional teaching environment.

[0156] A SCENE language model for generating scenario scripts in which text-based dialogue / action scripts for dialogues or actions between the instructor and NPCs are arranged chronologically from prompt inputs regarding a set situation, and

[0157] An object selection language model for selecting NPCs and surrounding objects to interact with NPCs as objects to be implemented as 3D views in the digital space of the metaverse from a number of digital assets stored in a database using the generated scenario script as input,

[0158] A space layout setting model for setting a rendering area by inputting the object names of the selected NPC and surrounding objects and inferring the rendering location and rendering scale at which the selected NPC and surrounding objects will be rendered in the digital space of the metaverse,

[0159] It is for generating a list of quests in the form of a list assigned as a task of an instructor from a generated scenario script, and may include a quest generation language model for generating a quest list of a higher concept encompassing dialogue / action scripts arranged in time series from the generated scenario script.

[0160] <SCENE 언어모델>

[0161] FIG. 11 illustrates an example of a scenario script including a dialogue / action script between an NPC and an instructor inferred by using a set situation as a prompt input, generated from a SCENE language model according to one embodiment of the present invention, a drawing exemplarily showing a plurality of quest blocks predicted from a quest generation language model by using a scenario script generated from the SCENE language model as an input, and a quest predicted as a superordinate concept or topic from each quest block.

[0162] In one embodiment of the present invention, the SCENE (scenario content editing natural language engine) language model can generate a scenario script in which a text-based dialogue / action script regarding a dialogue or action between an instructor and an NPC is arranged in time series from a prompt input regarding a set situation.

[0163] The above SCENE language model can generate a scenario script in which a text-based dialogue / action script regarding a dialogue or action between an NPC and an instructor in the digital space of the metaverse is arranged in time series from a situation setting given through input of a prompt.

[0164] In one embodiment of the present invention, the task assigned to the SCENE language model may include inference regarding conversational interactions and action-type interactions that can occur between an NPC and an instructor, and the SCENE language model may infer a scenario in which text-type dialogue / action scripts regarding conversations and actions that can occur instantaneously between an NPC and an instructor are arranged in time series from an input sequence regarding a virtual reality situation to be implemented in a digital space of the metaverse as a task input through a prompt, and for example, the inferred scenario may include an instantaneous arrangement of scripts regarding conversational interactions and action-type interactions that can occur instantaneously between an NPC and an instructor, and may output a scenario script in which the inferred dialogue / action scripts are arranged instantaneously.

[0165] In one embodiment of the present invention, an interaction may include a conversation or action directed at the other party between an NPC and an instructor, and in this case, a script regarding the interaction may specify the subject and the other party of each conversation or action. For example, one of the NPC and the instructor may be the subject of the conversation or action of each script, and the other party may be the other party of the conversation or action of the corresponding script. As such, in one embodiment of the present invention, a scenario script generated from the SCENE language model may include a text sequence regarding a mutually directed interaction between an NPC and an instructor, and may include, for example, a text sequence in which the subject and the other party are alternately reversed between an NPC and an instructor. In other words, in one embodiment of the present invention, a scenario script or scenario text generated from the SCENE language model may correspond to a text sequence that instantaneously arranges interactions of conversations or actions in which the subject and the other party of the conversation and action are reversed, rather than a descriptive text sequence regarding a related topic. For example, in one embodiment of the present invention, the SCENE language model can generate a scenario script in the form of a text in which the subject of each conversation and action is specified for each text script (conversation script and action script) regarding each conversation or action, and can generate a scenario script in which one of the NPC and the instructor is specified as the subject, that is, the other party among the NPC and the instructor is specified as the counterpart of the corresponding conversation or action.

[0166] As described below, the scenario script generated from the SCENE language model can be input into an object selection language model for selecting an object to appear in the digital space of the metaverse, and can be used as input data for calculating the probability that a corresponding object among a plurality of digital assets stored in a database will appear in the digital space of the metaverse from the object selection language model. From this perspective, the scenario script output from the SCENE language model and input into the object selection language model can take the form of a text in which the subject of each conversation or action is specified, and for example, the format or style (see prompt engineering) of the scenario script generated from the SCENE language model can be specified through a prompt input into the SCENE language model, and the format or style of the output format of the scenario script generated from the SCENE language model can be specified so that the subject of each conversation or action is specified so that it can be aligned with a subsequent task for selecting an object. More specifically, the SCENE language model can receive a style setting or a demonstration of few-shot learning for generating a scenario script in which the subject of each dialogue / action script is specified from an input prompt of a set situation, and can infer a given task of generating a scenario script including a dialogue script or action script for a dialogue or action in which the subject is specified in the set situation, along with a prompt input regarding the set situation.

[0167] <Object Selection Language Model>

[0168] FIG. 12 illustrates a diagram for explaining similarity analysis through contextual embedding of a GPT model as a natural language processing for NPC selection, as an object selection language model according to one embodiment of the present invention.

[0169] In one embodiment of the present invention, the object selection language model can select an object to be implemented in the digital space of the metaverse or to appear in the digital space of the metaverse from a scenario script regarding the interaction of dialogue / action output from the SCENE language model. More specifically, the object selection language model can select an NPC as an object to appear in the digital space of the metaverse from a plurality of digital assets stored in a database, and at least one or more objects that are provided with a 3D view in the digital space of the metaverse as surrounding objects that can interact (physically or digitally) with the NPC while being placed in the digital space surrounding the NPC.

[0170] In one embodiment of the present invention, the NPC may be selected as one digital asset for each set situation or each scenario script output from the SCENE language model with each set situation as input, and the surrounding object may be selected as one or more digital assets for each scenario script. In one embodiment of the present invention, the NPC may be a subject or counterpart of a conversation or action that performs a conversational or action-type interaction with the instructor, for example, the NPC and the instructor may be matched one-to-one to become the subject or counterpart of a conversation or action, and therefore the NPC may be selected as one digital asset for each scenario script. Unlike the NPC described above, the surrounding object may be selected as one or more digital assets according to a scenario script generated from a SCENE language model. For example, in one embodiment of the present invention, the surrounding object may be an object that performs physical interaction and / or digital interaction with an NPC, may mean an object that can call an interaction depending on the proximity distance between the NPC and the object, and may mean an object that can call an interaction depending on the proximity of the NPC within a specific distance.

[0171] In one embodiment of the present invention, interactions that can be invoked between an NPC and surrounding objects may include physical interactions including contact such as collision (obstruction), grabbing, sitting, lying down, etc., and digital interactions such as accessing a website, outputting a preset image or video, making a phone call, outputting a note including a preset specific sentence, etc.

[0172] In one embodiment of the present invention, on the database, as objects rendered as 3D views in the digital space of the metaverse, object names of NPCs and surrounding objects surrounding the NPC and attribute data of each object can be stored together in a linked manner. The above object selection language model may perform natural language processing for NPC selection, which takes a scenario script as input, calculates the similarity between the entity name of each NPC stored in a database and the attribute data of the NPC stored in connection with each NPC, and calculates the predictability that the corresponding NPC will appear in the digital space of the metaverse where the corresponding scenario script will be implemented based on the calculated similarity, and natural language processing for peripheral object selection, which takes a scenario script as input, calculates the similarity between the entity name of each surrounding object stored in a database and the attribute data of the surrounding object stored in connection with each surrounding object, and calculates the predictability that the corresponding surrounding object will appear in the digital space of the metaverse where the corresponding scenario script will be implemented based on the calculated similarity. In one embodiment of the present invention, the object selection language model may perform different natural language processing depending on the NPC or surrounding object corresponding to each selection target.

[0173] In one embodiment of the present invention, in the natural language processing for selecting surrounding objects performed by the object selection language model, the possibility of the corresponding surrounding object appearing in the digital space of the metaverse implementing the corresponding scenario script can be predicted based on the similarity between each word forming the first word sequence regarding the scenario script output from the SCENE language model and each word forming the second word sequence regarding the entity name of the surrounding object and the attribute data of the surrounding object stored in the database, and more specifically, each word forming the first and second word sequences can calculate the similarity between words through dense embedding or distributed representation that distributes the meaning of the word into multiple dimensions in a low dimension rather than a sparse representation in which each dimension is separated in a high dimension, such as a one-hot vector, and the similarity to each other can be calculated in the embedding space having a relatively smaller dimension than the one-hot vector.For example, in natural language processing for selecting surrounding objects, by inputting each word forming the first word sequence for each scenario script and each word forming the second word sequence for the entity name of the surrounding object and the attribute data of the surrounding object, a word representation numerically expressed in a lower-dimensional embedding space than a one-hot vector can be produced, and a one-to-one similarity can be produced between each word forming the first word sequence and each word forming the second word sequence from each word or embedded word representation forming the first and second word sequences, and among the word representations numerically expressed in the same embedding space, a surrounding object for the second word sequence including a word or word representation having a high similarity with the word or word representation forming the first word sequence can be predicted to have a high possibility of appearing in the space of the metaverse in which the scenario script for the first word sequence corresponding to the comparison target is implemented. For example, in natural language processing for selection of surrounding objects, context information (e.g., self-attention) of a first word sequence including association information between each word forming a first word sequence regarding a scenario script may not be produced, context information (e.g., self-attention) of a second word sequence including association information between each word forming a second word sequence regarding the entity name of the surrounding object and attribute data of the surrounding object may not be produced, and association information (e.g., cross-attention) between each word forming the first word sequence and each word forming the second word sequence may not be produced.

[0174] In this specification, the similarity between each word forming the first word sequence and each word forming the second word sequence is calculated one-to-one between the embedded word representations of each word forming the first word sequence and each word forming the second word sequence, without calculating self-attention or cross-attention between the first word sequence regarding the scenario script and the second word sequence regarding the entity name of the surrounding object and the attribute data of the surrounding object, and between each word forming the first word sequence and each word forming the second word sequence, for example, by using a query vector as a word or word representation to be processed (a word representation forming the first word sequence / second word sequence), and the other words or word representations (the other other word representations forming the first word sequence / second word sequence) as key vectors, calculating a score through a scaled dot product between the query vector and the key vector, and applying a softmax function to the calculated score to calculate a new expression for the word or word representation to be processed from the weighted sum of the value vectors that have the normalized score as a weight. This may mean that it does not perform any embedding.More specifically, in one embodiment of the present invention, the object selection language module calculates a word representation for each word forming the first and second word sequences through a context-independent embedding rather than a context-based embedding, which is numerically expressed on the same embedding space, and calculates a one-to-one similarity between the word representation forming the first word sequence and the word representation forming the second word sequence, for example, by calculating a Euclidean distance between the words forming the first word sequence, which are the targets of similarity calculation, and the words forming the second word sequence on the same embedding space, or by calculating a dot product between word representations (word vectors), so that it can predict a surrounding object with a high probability of appearing with respect to the second word sequence including a word representation having a high similarity to the word representation of the first word sequence with respect to the scenario script.

[0175] As such, in one embodiment of the present invention, in the metaverse digital space of the scenario script output from the SCENE language model, the probability of appearance of each corresponding surrounding object or the prediction of the probability of appearance can be predicted from a context-independent embedding, rather than a contextual embedding. As described above, the SCENE language model can infer conversational and action-type interactions between an NPC and an instructor from the input of a prompt for a virtual reality situation to be implemented in the metaverse digital space and output them in the form of text, and can infer that surrounding objects for a second word sequence including words or word expressions having a relatively high similarity with each word or word expression forming the first word sequence in the text form of such dialogue / action script are likely to appear in the scenario script for the first word sequence. As a more specific example, as a word or word expression forming the first word sequence regarding the scenario script (a word or word expression appearing in the first word sequence expressed in text form regarding the dialogue / action script of the scenario script), a word or word expression such as “sit” can produce a high similarity with respect to a peripheral object stored in a database that includes the same entity name as “chair” or a similar entity name or includes the same or similar attribute data as “sit”, and accordingly, the corresponding peripheral object (e.g., entity name: chair) can be selected as the peripheral object appearing in the metaverse digital space where the corresponding scenario script will be implemented.

[0176] In this way, in the text-type scenario script inferred from the SCENE language model, the possibility of the appearance of the corresponding peripheral object can be predicted based on the similarity between the words or word expressions appearing in each dialogue / action script in which the subject and the counterpart are specified, and without the need to extract the overall context information of the scenario script, the context information of each sentence forming the scenario script, or the contextual information (or contextual embedding) regarding the relationship between each word forming the scenario script, the possibility of the appearance of the corresponding peripheral object can be inferred based on the similarity between each word or word expression forming the scenario script itself and each word or word expression forming the entity name and attribute data stored in connection with each peripheral object stored in the database. For example, the above-described peripheral object is an object that can invoke a physical interaction or digital interaction with an NPC, and such physical interaction or digital interaction with an NPC can directly appear in the scenario script as a word or word expression that forms a scenario script inferred from the SCENE language model, and for example, peripheral objects such as entity names chair, bed, and table can be selected from words or word expressions related to physical interactions such as sit, lie down, and put down, respectively. In addition, in one embodiment of the present invention, the entity names or attributes of peripheral objects such as apples and pencils that are the objects of the transaction can be directly expressed in the text that forms the scenario script inferred from the SCENE language model from a prompt input regarding a transaction situation of buying and selling items between an NPC and an instructor, and therefore, peripheral objects such as entity names apples and pencils can be selected from each word or word expression that forms the scenario script, as described above.

[0177] Unlike the surrounding objects described above, NPCs are not directly mentioned in the text of the scenario script inferred from the SCENE language model by inputting a prompt regarding a set situation, but rather the contextual information (or contextual embedding) of the scenario script is extracted, and the possibility of appearance for each NPC stored in the database can be inferred from the contextual information of the extracted scenario script.

[0178] The above object selection language model can perform contextual embedding for a first word sequence regarding a scenario script inferred from a SCENE language model in language processing for NPC selection, can perform contextual embedding for a second word sequence regarding each NPC stored in a database, and can calculate a similarity between word expressions or sentence expressions that are numerically expressed in the same embedding space and include association information between each word within a sentence (expressed as a weighted sum calculated by adding up the weights according to the associations between each word).

[0179] More specifically, referring to FIG. 12, the context-based embedding of the first word sequence for the scenario script inferred from the SCENE language model and the second word sequence for each NPC (entity name and attribute data of each NPC) stored in the database can be implemented from a language model (large-scale language model, LLM) as described above or a GPT model as a language model, and the input embedding input to the GPT model can include a token embedding for each word and a position embedding for position information of each word, and at this time, a special token (delimit) can be interposed between the input embedding for the first word sequence and the input embedding for the second word sequence to recognize the boundary between them. For example, the input embedding for the first word sequence and the input embedding for the second word sequence can be input in different orders, for example, the first similarity calculation in which the input embedding for the first word sequence (Text 1) - special token (delimit) - input embedding for the second word sequence (Text 2) is input in the order of the first word sequence input embedding, and the second similarity calculation in which the input embedding for the second word sequence (Text 2) - special token (delimit) - input embedding for the first word sequence can be performed in parallel, and the output values ​​of the first similarity calculation and the output values ​​of the second similarity calculation can be concatenated and passed through a linear layer or a fully connected layer to produce the final similarity.For example, by inputting the input embedding for the first word sequence and the input embedding for the second word sequence in different orders into the object selection language model, the similarity is independently calculated between different orders, and the final similarity is calculated by combining these independently calculated similarities. This is because the output value of the similarity may vary depending on the input order of the first and second word sequences. For example, in the object selection language model implemented based on the GPT model as an embodiment of a large-scale language model (LLM), among the transformer models including encoders and decoders, since it is based on the decoder structure without an encoder, unlike the BERT model (a bidirectional language model including the encoder of the transformer, forward and backward) that performs attention in both directions with the previous and next words or word expressions before and after the currently processed word or word expression, it performs attention in one direction with the previous word or word expression of the currently processed word or word expression (forward language model), and applies a mask to the word or word expression after the currently processed word or word expression. Since the so-called masked attention is performed, in one embodiment of the present invention, as an embodiment of a large-scale language model (LLM), in an object selection language model implemented based on a GPT model, the first and second similarity calculations that are processed independently by changing the input order can be combined to produce the final similarity.

[0180] In one embodiment of the present invention, extracting context information of the first and second word sequences through context-based embedding for the first word sequence regarding the scenario script and the second word sequence regarding the NPC (NPC entity name and attribute data) stored in the database may comprehensively mean context information (e.g., self-attention) of the first word sequence including association information between each word in the first word sequence, context information (e.g., self-attention) of the second word sequence including association information between each word in the second word sequence, and context information (e.g., cross-attention) between the first and second word sequences including association information between each word in the first word sequence and each word in the second word sequence, for example, using any one word or word expression (any one word expression forming the first word sequence / second word sequence) to be currently processed as a query vector, and the remaining other words or word expressions (the remaining other word expressions forming the first sequence / second sequence) as key vectors, and By calculating a score through a scaled dot product between key vectors and applying a softmax function as an output function to the calculated score, a new expression for the word or word expression of the current processing target is calculated from the weighted sum of the value vectors with the normalized score as a weight, thereby performing contextual embedding for the first and second word sequences, and calculating the similarity between the first and second word sequences from the representations of the first and second word sequences numerically expressed in the same embedding space, and calculating the similarity between the first and second word sequences calculated from the object selection language model.By passing the representation for the second word sequence through a linear layer or a fully connected layer, the similarity between the first and second word sequences can be calculated. For example, the object selection language model can output a word representation that includes context information within a sentence for each word forming the first and second word sequences through context-based embedding that includes association information between each word within the sentence of the first and second word sequences, and by numerically representing each word forming the first and second word sequences in the embedding space to include context information within the sentence, the similarity between the word representations for each word included in the first and second word sequences can be calculated, or the similarity between context vectors that include context information of the entire first and second word sequences can be calculated.

[0181] For example, in one embodiment of the present invention, the object selection language model can produce an embedding representation including association information between each word within the sentence of each of the first and second word sequences, and can produce a similarity between each word from the embedding representation of each word forming the first and second word sequences (for example, the similarity between each word representation forming the first and second word sequences can be produced through a linear layer or a fully connected layer after the last hidden layer from which word representations for each word forming the first and second word sequences are output, and the similarity between the first and second word sequences can be produced by collecting the similarities between each word representation forming the first and second word sequences, and the similarity can also be produced according to the distance on the embedding space to which the first and second word sequences are mapped). In various embodiments of the present invention, the object selection language model can produce an embedding representation including context information of the entire first word sequence output at the position of a special token or the position of the last word in the last hidden layer from which the embedding representation of each word forming the first word sequence is output. The similarity between the first and second context vectors can also be calculated by using the first context vector and the second context vector containing context information of the entire second word sequence output at the position of the special token or the last position in the last hidden layer where the embedding representation of each word forming the second word sequence is output as input to the output layer after the last hidden layer.

[0182] In one embodiment of the present invention, the object selection language model can perform different processes in natural language processing for selecting an NPC and natural language processing for selecting a surrounding object. For example, in the natural language processing for selecting an NPC, an NPC can be selected based on a similarity analysis between first and second context vectors derived from contextual embeddings for first and second word sequences regarding entity names and attribute data of a plurality of digital assets stored in a scenario script and a database, respectively, and in the natural language processing for selecting a surrounding object, a surrounding object can be selected based on a similarity analysis between each word or word vector derived from word embeddings for each word forming the first and second word sequences regarding entity names and attribute data of a plurality of digital assets stored in a scenario script and a database, respectively.

[0183] In one embodiment of the present invention, the object selection language model can analyze the similarity between the action script of the generated scenario script and the dynamic motion specified for each NPC as the attribute data of the NPC in the natural language processing for selecting the NPC, and can predict that the NPC having the dynamic motion having a high similarity with the action script output from the SCENE language model is likely to appear in the digital space of the metaverse in which the corresponding scenario script is implemented.

[0184] In one embodiment of the present invention, the object selection language model can analyze the similarity between the action script of the generated scenario script and the physical interaction or digital interaction with the NPC designated for each surrounding object as attribute data of the surrounding object in the natural language processing for selection of the surrounding object, and can predict that the surrounding object designated for the physical interaction or digital interaction having a high similarity with the action script output from the SCENE language model is likely to appear in the digital space of the metaverse in which the corresponding scenario script is implemented.

[0185] Quest Generation Language Model

[0186] FIG. 11 illustrates an example of a scenario script including a dialogue / action script between an NPC and an instructor inferred by using a set situation as a prompt input, generated from a SCENE language model according to one embodiment of the present invention, a drawing exemplarily showing a plurality of quest blocks predicted from a quest generation language model by using a scenario script generated from the SCENE language model as an input, and a quest predicted as a superordinate concept or topic from each quest block.

[0187] In one embodiment of the present invention, a quest generation language model for generating a quest list as a task assigned to an instructor from a scenario script generated from the SCENE language model can infer the steps (or topics) necessary for transitioning to the next phase in each phase that unfolds chronologically according to a set situation input as a prompt of the SCENE language model to generate a corresponding scenario script, and can generate a quest list as a chronological list of dialogue / action scripts that need to be expressed by the instructor as a subject for the development of the set situation, that is, for transitioning to each phase. In one embodiment of the present invention, rather than including a specific textual expression for each dialogue / action script, the quest list can include a chronological list of tasks or quests as a higher-level concept corresponding to the topic of each specific dialogue / action script, for example, a lower-level concept that comprehensively encompasses a plurality of specific dialogue / action scripts.

[0188] In one embodiment of the present invention, the quest generation language model for generating the quest list can perform language processing for a generative summary that encapsulates a given scenario script, for example, rather than including the sentence expressions contained in the scenario script as they are, it can perform implicit summary generation that encompasses specific dialogue / action scripts contained in the scenario script so as to generate new expressions that include key words of the scenario script (e.g., words that appear repeatedly in the corresponding quest block) but are not contained in the scenario script as they are. For example, the quest generation language model can generate the quest list as an abstractive summarization rather than an extractive summarization of the scenario script.

[0189] In one embodiment of the present invention, the SCENE language model generates text regarding a specific dialogue / action script using a set situation as an input of a prompt, whereas the quest generation language model implicitly summarizes the specifically generated dialogue / action script and generates a quest list in which a quest of a higher concept that can comprehensively encompasses the specifically inferred lower-level dialogue / action script is listed in a chronological order, so that even if the instructor does not display the specific dialogue / action script itself predicted from the SCENE language model according to the development of the situation from the corresponding quest list, the instructor can be evaluated as having accomplished each quest listed on the quest list by displaying a dialogue / action script of different expressions of the same or similar topic that corresponds to each task listed in the quest list. As described below, in the teaching stage, the dialogue / action script displayed by the instructor can be input as a prompt of the conversational language model in text form, and the conversational language model can evaluate whether or not the quest has been accomplished based on the dialogue / action script input by the instructor.

[0190] In one embodiment of the present invention, a scenario script is generated from a SCENE language model that takes a set situation as input, and a more comprehensive upper-level concept of a quest or task that includes a dialogue / action script of the scenario script generated from a quest generation language model is generated and provided to an instructor in the form of a quest list, and the instructor can achieve the purpose of the instruction in the process of expressing a specific dialogue / action script for the accomplishment of each quest by taking each quest listed in the quest list as a given task, and, for example, can be provided with an educational opportunity to master expressions (for example, expressions in a foreign language such as English) suitable for an actual situation setting from the taught learning content.

[0191] In one embodiment of the present invention, the quest generation language model can generate an implicit generative summary from text expressing each dialogue / action script forming a scenario script, for example, a quest for transition or transition to the next stage (or topic) while unfolding chronologically in a set situation may correspond to a major stage (or topic) that can be inferred under the set situation, and the sentence sequence forming the scenario script is an interaction in which an NPC and an instructor alternately become a subject and an opponent and communicated to each other, and a dialogue / action script that is performed while the instructor and the NPC change subjects and opponents may include a quest clue that implicitly expresses the corresponding stage (or topic), and such a quest clue may repeatedly appear as a major key word in a dialogue / action script communicated to each other between the subject and opponent of the interaction, for example, a first word sequence regarding the interaction expressed from the subject from a dialogue / action script in which an NPC and an instructor alternately become a subject and an opponent and communicated, and a response thereto. A quest that can implicitly encompass the relevant interaction can be created from key words or clues that appear repeatedly in the second word sequence regarding the interaction in which the other party becomes the subject or responds (for example, a quest such as greeting, finding out price information, ordering, paying, etc., which can correspond to the subject of the quest block as described below).

[0192] In one embodiment of the present invention, a key word or clue that repeatedly appears from the first and second word sequences regarding an interaction in which a subject and an opponent take turns may correspond to a word representing a specific category to which the first and second context vectors are mapped when the first and second word sequences (e.g., first and second context vectors including context information of the first and second word sequences) are mapped adjacently within a specific Euclidean distance in the embedding space, rather than limiting the same or similar expressions. For example, different word sequences expressing greetings (hi, hello) can be mapped to adjacent positions in an embedding space that reflects the similarity between words or sentences, and it can be predicted that the overall meanings of the first and second word sequences are similar from the similarity between the first and second context vectors produced through context-based embedding of the first and second word sequences, and a quest that implicitly expresses the first and second word sequences can be extracted from the first and second context vectors that include similar meanings.

[0193] In one embodiment of the present invention, the quest generation language model can classify sentences having the same or similar topic into the same quest block through similarity analysis on the sentence sequence appearing in the scenario script, and can extract key words or clues that appear repeatedly from the sentences forming the same quest block from each quest block to generate the topic of the corresponding quest block as a single quest. More specifically, in one embodiment of the present invention, each dialogue / action script or each sentence forming the scenario script generated from the SCENE language model can be input from the prompt of the quest generation language model in the order of appearance, and each input script can be sequentially subjected to similarity analysis by the quest generation language model in the order of appearance in the scenario script. For example, in one embodiment of the present invention, the quest generation language model performs similarity analysis in the order of appearance, such as similarity analysis between the first script (dialogue / action script) and the second script, similarity analysis between the second script and the third script, and similarity analysis between the third script and the fourth script in the order of input, but can sequentially perform similarity analysis on two scripts that form a pair adjacent to each other in the order of appearance, and can perform similarity analysis in an overlapping manner so that the pairs overlap each other to determine the boundary of the same quest block.In one embodiment of the present invention, the quest generation language model can be implemented as a GPT model as one embodiment of a large-scale language model (LLM), and for example, a special token (delimit) capable of recognizing the boundary of the first and second word sequences can be inserted between the input embedding for the first word sequence and the input embedding for the second word sequence to perform a similarity analysis on pairs of different first and second word sequences, and can be input to the quest generation language model (see FIG. 12, the order of the input embeddings Text 1 and Text 2 of the different first and second word sequences is changed and inputted to the GPT model, and the similarity between the first and second word sequences is analyzed from the concatenate and linear layers of the output values).

[0194] The above quest generation language model collects scripts on the same topic and classifies them into the same quest block (quest block 1, quest block 2, quest block 3, quest block 4, quest block 5), but can classify scripts that are adjacent to each other in time series as the same quest block in the order of appearance in the scenario script, for example, or in other words, the same quest block can be predicted in time series by the start and end positions of the block, and a set of at least one script that is sequentially arranged in time series can be predicted as one quest block.

[0195] In one embodiment of the present invention, the quest generation language model can predict a single quest block by collecting a set of multiple scripts that appear sequentially in time series from a similarity analysis between scripts that are sequentially input in the order of appearance of a scenario script, and the multiple scripts that form the same quest block are conversation / action scripts that are alternately transmitted between an NPC and an instructor as a subject and an opponent, and can infer a generated summary from the multiple scripts that form the same quest block, and in various embodiments of the present invention, the quest generation language model can be trained to infer a generated summary from the multiple scripts that form the same quest block, and can extract key words or clues that appear repeatedly from the multiple scripts that form the same quest block and generate a quest (such as greeting, finding out price information, placing an order, making a payment, or receiving an item) as a subject of the corresponding quest block from these key words or clues.

[0196] In one embodiment of the present invention, the quest generation language model can be implemented as a large-scale language model (LLM), and as a pre-trained model, it can acquire knowledge about general natural language processing by using an unlabeled data set (unlabeled large corpus, large-scale text such as web text) as learning data, and by acquiring knowledge for general natural language processing, such as predicting the next conversation content from a part of a conversation, such as NSP (next sentence prediction), or inferring the implication or connotation relationship between a premise and a hypothesis, such as entailment, it can infer a common topic or a generated summary about a common topic from a plurality of scripts forming the same quest block, or it can extract key words or clues that appear repeatedly from a plurality of scripts forming the same quest block, and for example, it can extract key words or clues that are mapped adjacently within a certain Euclidean distance on an embedding space in which the similarity of a plurality of word expressions is reflected from word sequences forming a plurality of scripts forming the same quest block, and extract the extracted key words or clues. From words or clues, you can infer the topic or generate a summary about the topic of the corresponding quest block.

[0197] In various embodiments of the present invention, the main steps for transitioning or transitioning to each step according to the development of a set situation input as a prompt of a SCENE language model, or the topic of each step of a dialogue / action script that is alternately delivered between an NPC and an instructor as the subject and the counterpart, can be predicted from a specific scenario script generated from a SCENE language model, or can be predicted according to a quest block generated from a quest generation language model that uses a scenario script as a prompt input, or can infer the topic of the corresponding quest block or a summary generation regarding the topic from a quest generation language model that refers to a mapping setting that connects each key word or clue and each quest through a mapping of a form set in advance according to a specifically set situation. As described above, the quest generation language model is a pre-trained model that learns general natural language processing knowledge using a large-scale unlabeled text as a learning data set, thereby inferring a common topic or generating a summary about a topic for at least one or more scripts forming the same quest block. In various embodiments of the present invention, the quest generation language model may be a fine-tuned model whose model parameters are updated using a labeled data set as learning data to suit a downstream task of extracting a common topic from one or more scripts forming the same quest block.

[0198] In this way, in one embodiment of the present invention, the quest generation language model can perform a first natural language processing for classifying dialogue / action scripts arranged in time series in a scenario script generated from the SCENE language model into quest blocks according to their topics, and a second natural language processing for generating a topic or quest of a higher concept encompassing at least one or more dialogue / action scripts classified into the same quest block.

[0199] The above quest generation language model can perform the first natural language processing by analyzing the similarity between neighboring dialogue / action scripts according to the order of appearance of a plurality of dialogue / action scripts arranged chronologically in the generated scenario script.

[0200] At this time, the quest generation language model is

[0201] In order to set the boundaries of the quest blocks according to the order of appearance of multiple dialogue / action scripts arranged chronologically in the generated scenario script, pairs of adjacent dialogue / action scripts that are the subject of similarity analysis can be set to overlap each other.

[0202] The above quest generation language model is,

[0203] The second natural language processing can be performed by extracting key words or clues that appear repeatedly in multiple dialogue / action scripts classified into the same quest block.

[0204] At this time, the quest generation language model is

[0205] First and second context vectors can be derived from contextual embeddings of first and second word sequences forming multiple dialogue / action scripts classified into the same quest block, and a topic or superordinate concept quest can be generated by categorizing the first and second word sequences from the distribution of multiple context vectors mapped onto the contextual embedding space.

[0206] <Space layout setting model>

[0207] FIGS. 13a to 13c are drawings showing, by way of example, a rendering area of ​​an NPC (NPC) and a rendering area of ​​surrounding objects (surrounding objects) generated from a space layout setting model according to one embodiment of the present invention.

[0208] In one embodiment of the present invention, NPCs and peripheral objects implemented on the digital space of the metaverse are based on a scenario script inferred from a SCENE language model as described above, and an object selection language model for calculating similarity between entity names and attribute data regarding NPCs and peripheral objects stored in a database from a scenario script generated in text form can be used to select NPCs and peripheral objects that are likely to appear on the digital space of the metaverse where the corresponding scenario script is implemented, based on the similarity calculated from the similarity calculated from the object selection language model.

[0209] In one embodiment of the present invention, NPCs and surrounding objects selected from an object selection language model can be input as a prompt of a spatial layout setting model a text including the entity name, and from the input of the prompt including the entity name of the NPC and surrounding object, the spatial layout setting model can infer a rendering area including a rendering position and a rendering scale at which an object (3D object) including the NPC and surrounding object whose entity name is input is to be rendered in the digital space of the metaverse. For example, in one embodiment of the present invention, the rendering area can predict a two-dimensional bounding box having a horizontal x vertical dimension as a rendering area in which each object (3D object) including each NPC and surrounding object is to be rendered, and for example, in one embodiment of the present invention, the spatial layout setting model can infer, for each object, a center position of a bounding box surrounding a rendering area and a horizontal x vertical size of the bounding box as information regarding a rendering area in which the corresponding object is to be rendered in the digital space of the metaverse.

[0210] For example, the spatial layout setting model may include a text-to-image generation model that takes as input a prompt in the form of a text combining the entity names of NPCs and surrounding objects selected from an object selection language model and generates an image from the input of the corresponding prompt. For example, in one embodiment of the present invention, the spatial layout setting model may generate an image in which an NPC image and an surrounding object image of an entity name input from a text combining the entity names of the NPC and surrounding objects appear together, and may perform image segmentation or object detection using the generated image as input to distinguish an NPC image region and a surrounding object image region appearing in each image. For example, in the object detection, the text combining the entity names of the NPC and surrounding objects is input, and the boundaries of each NPC image area and surrounding object image area can be predicted in the form of a bounding box surrounding the NPC image area and surrounding object image area on an image generated from text-to-image, and in the image segmentation, the boundaries of the NPC image area and surrounding object image area can be predicted on the generated image.

[0211] For example, the above-described space layout setting model can predict the space layout in the digital space of the metaverse in which rendering is implemented for each object including the NPC and surrounding objects selected from the object selection language model, for example, it can predict the rendering area in which each object is rendered, and more specifically, as the position of the rendering area in which each object is rendered, it can output the center position of the bounding box surrounding the rendering area in which each object is rendered, and as the scale of the rendering area in which each object is rendered, it can output the width x height size of the bounding box surrounding the rendering area in which each object is rendered. At this time, the space layout setting model may include a text-to-image generation model for generating text-to-image using a text combining the entity names of objects (objects implemented as 3D views, NPCs, and surrounding objects) selected from an object selection language model as a prompt input, and a region extraction model for predicting a boundary surrounding each object image on the generated image using an image generated from the text-to-image generation model as an input. For example, in one embodiment of the present invention, the text-to-image generation model may include a text encoder for encoding input text, that is, text combined with entity names of objects for which a 3D view is implemented, including NPCs and surrounding objects, into a text embedding through a multi-modal AI model that has learned a multi-modal embedding space of images and texts or a multi-modal embedding (e.g., CLIP embedding), a prior for converting the text embedding into an image embedding matching the text embedding through the multi-modal embedding, and a decoder for generating an image as an output of the text-to-image from the image embedding changed from the prior.In various embodiments of the present invention, the prior and decoder may be implemented as conditioning diffusion that injects, as information available at each stage, text input as a prompt of the spatial arrangement setting model (text combined with entity names of objects selected from an object selection language model) and / or text embedding output from the text encoder and / or image embedding converted from the prior, into conditioning for generation, or may be implemented as a latent diffusion model (LDM) or stable diffusion that implements diffusion on a low-dimensional latent space. In various embodiments of the present invention, a region extraction model for predicting the boundary of each object image on an image generated from a text-to-image generation model may include a neural network of an R-CNN series architecture (Region based CNN, R-CNN, Fast R-CNN, faster R-CNN, etc.) for predicting the boundary of each object image in the form of a bounding box, and may include an architecture of a U-net for predicting the boundary of each object image on a pixel basis.

[0212] <Background Image Generation Model>

[0213] FIGS. 14A to 14C are drawings showing examples of background images surrounding NPCs and surrounding objects generated as text-to-image, as different background images generated from a background image generation model according to one embodiment of the present invention.

[0214] In one embodiment of the present invention, in the digital space of the metaverse where the scenario script inferred from the SCENE language model is implemented, an object for which a 3D view is provided may include an NPC as a subject or counterpart of a dialogue / action script included in the scenario script expressed in text form, and surrounding objects that perform physical and digital interactions with the NPC. In addition, a background image that provides an environment surrounding the NPC and surrounding objects may be implemented in the digital space of the metaverse, and the background image may be provided as a 2D view, unlike the NPC and surrounding objects that are provided as a 3D view in the digital space of the metaverse.

[0215] In one embodiment of the present invention, a background image implemented on a digital space of the metaverse can be implemented from a text-to-image generation model, and in one embodiment of the present invention, the background image generation model can include a text encoder that takes text from a prompt as input, and encodes the input text into a text embedding through a multi-modal AI model that has learned a multi-modal embedding space of images and texts from the input text or a multi-modal embedding (e.g., CLIP embedding), and, for example, a prior for converting the text embedding into an image embedding that matches the text embedding on the multi-modal embedding space, and a decoder for generating an image as an output of the text-to-image from the image embedding converted from the prior, and, for example, in one embodiment of the present invention, the prior and the decoder can be implemented based on diffusion. In various embodiments of the present invention, the prior and decoder may be implemented as a conditioning diffusion that injects, as available information at each stage, a text input as a prompt of the background image generation model and / or a text embedding output from the text encoder and / or an image embedding converted from the prior into the conditioning of the generation, or may be implemented as a latent diffusion model (LDM) or stable diffusion that implements diffusion on a low-dimensional latent space.

[0216] In one embodiment of the present invention, the background image generation model can implement text-to-image image generation from a prompt input in the form of text, and generates a background image of a 2D view that provides an environment surrounding the NPC and surrounding objects, rather than being provided as a 3D view in the digital space of the metaverse, such as an NPC performing conversational interaction and / or action interaction with an instructor in a scenario script inferred from a SCENE language model, and surrounding objects performing physical interaction and / or digital interaction with the NPC, thereby saving computational resources and implementing a digital space of the metaverse that can sufficiently describe a setting situation input as a prompt of the SCENE language model to infer a scenario script.

[0217] In one embodiment of the present invention, the text input as a prompt of the background image generation model may be based on the text of a scenario script generated from the SCENE language model, and for example, the text or word sequence from which the entity names of the subject or counterpart of the dialogue / action script forming the generated scenario script, and the NPCs and surrounding objects selected from the object selection language model as objects that will appear in the digital space of the metaverse in which the corresponding scenario script is to be implemented using the generated scenario script as input, have been deleted may be used as the input prompt of the background image generation model. For example, the text input as a prompt of the background image generation model may be based on the text from which the entity names of the subject or counterpart of the dialogue / action script forming the corresponding scenario script from the SCENE language model, and the entity names of the NPCs and surrounding objects selected from the object selection language model, or words similar to these objects, have been deleted. For example, the background image generation model can perform natural language processing to delete the names of the subjects or counterparts of the dialogue / action scripts forming the corresponding scenario scripts from the scenario scripts input from the SCENE language model, and also to delete the entity names of NPCs and surrounding objects selected from the object selection language model or words similar to the entity names using the corresponding scenario scripts as input. For example, in the natural language processing, words identical to or similar to the entity names of the NPCs and surrounding objects selected from the object selection language model can be extracted and deleted based on a similarity analysis between the names of the subjects (the subjects of the alternating dialogue / action scripts may include the names of the subjects and counterparts) from the scenario scripts in which the subjects of each dialogue / action script are specified and the text forming the scenario script (each word or word expression of the text).The above background image generation model may include text-to-image generation that generates text-to-image using text processed in a natural language as input for a prompt. In one embodiment of the present invention, in order to prevent an image that overlaps with an image provided in a 2D view as a background image and an image of an NPC and surrounding objects provided in a 3D view, the 2D background image generated from the background image generation model may not include an image of an NPC and surrounding objects provided in a 3D view.

[0218] <Transmission code or transmission data transmitted from the server area of ​​the metaverse creation network to the instructor's local area>

[0219] In one embodiment of the present invention, rendering data for rendering NPCs and surrounding objects selected from an object selection language model using the generated scenario script as input, excluding a scenario script generated from a SCENE language model using a set situation as input, on a digital space of a metaverse where a set situation is implemented, data regarding a rendering area including a rendering position and a rendering scale at which the selected NPCs and surrounding objects are to be rendered on a digital space of the metaverse using text including the object names of the selected NPCs and surrounding objects as input (including an output of a space layout setting model), a quest list generated from a quest generation language model using the generated scenario script as input, and image data regarding a background image generated from a text-to-image generation model (including an output of a background image generation model) using text through a prompt as input, may form a transmission code or transmission data transmitted from a server area (a metaverse generation network or a processing server in which a metaverse generation network is implemented) toward a local area (an instructor or a teaching location) to implement a teaching environment of a 3D metaverse, and the transmission code or transmission Data can be transmitted from the server area where rendering data, etc. are generated to the local area where instruction takes place.

[0220] In one embodiment of the present invention, the transmission code or transmission data transmitted from a server area where rendering data, etc., providing a 3D teaching environment of the metaverse are generated to a local area may include 2D image frame data regarding a background image, 3D rendering data for rendering a 3D view regarding NPCs and surrounding objects, a set of functions (e.g., camera functions, etc.) that implement a rendering image using the rendering data of these NPCs and surrounding objects as arguments or generate arguments input to another function, or a set of functions that can change the settings of the 3D view.

[0221] In one embodiment of the present invention, a transmission code or transmission data transmitted from a server area where rendering data, etc. are generated to provide a teaching environment of a 3D metaverse to a local area at a location of an instructor may include a call function, etc. that can transmit a corresponding conversation / action script to a server area (an interactive language model of the server area) for natural language processing of a conversation / action script displayed by an instructor in the local area. For example, the call function may be called according to an input of the instructor (input of a conversation / action script), and may access the server area along a connection path defined in the function to transmit the conversation / action script displayed by the instructor from the local area to the server area.

[0222] In one embodiment of the present invention, a dialogue / action script expressed by an instructor is transmitted from a local area to a server area, and a dialogue / action script as a response to the dialogue / action script expressed by the instructor can be generated through natural language processing of an interactive language model built in the server area (processing server), and whether a quest assigned as a task to the instructor has been accomplished can be evaluated based on the dialogue / action script expressed by the instructor. The processing result of this interactive language model can be transmitted from the server area to the local area, and provided to the instructor through an interface in the local area.

[0223] More specifically, the transmission code or transmission data transmitted from the server area to the local area can be driven through an application of a terminal providing a teaching environment, more specifically, a PC provided in the local area of ​​the instructor, a portable computing device, a head-mounted set or controller providing a geared metaverse environment, etc., and for example, rendering data for rendering objects provided as 3D views, such as NPCs and surrounding objects, can be rendered through webGL (web Graphic Library), which is a graphic library or rasterization engine that can be driven in a web browser, and three.js, which is a 3D graphic library of JavaScript.In one embodiment of the present invention, the transmission code or transmission data transmitted from the server area to the terminal equipped in the local area of ​​the instructor may include animation data for granting dynamic motion to the NPC, and for example, for designated actions granted as attribute data of each NPC, animation data (e.g., animation clips) for implementing each designated action may be included, and in one embodiment of the present invention, dynamic motions that can be implemented by each NPC may be designated in advance as attribute data of each NPC, and the dynamic motion designated for each NPC from attribute data stored in conjunction with the object name of the NPC is reference data referenced by an object selection language model that predicts the probability or possibility of appearing in a scenario script generated based on the similarity with the attribute data of each NPC stored in a database by inputting a scenario script generated from a SCENE language model and selects an NPC to appear in the digital space of the metaverse in which the generated scenario script is to be implemented, and the object selection language model refers to the dynamic motion designated in the attribute data of each NPC and executes an action script expressed in the scenario script generated from the SCENE language model. An NPC capable of implementing the dynamic motion required for implementation can be selected. In other words, the object selection language model can select an NPC with a dynamic motion specified to implement the action script expressed in the scenario script generated from the SCENE language model.

[0224] The processing server in the server area can transmit rendering data necessary for implementing an NPC and surrounding objects selected from an object selection language model as a 3D view in the digital space on a terminal in the local area where the instructor is located so that the digital space of the metaverse can be implemented on the terminal in the local area. At this time, the rendering data transmitted to the terminal in the local area can include rendering data for rendering the NPC as a 3D view, as well as animation data for implementing dynamic motions of the NPC. More specifically, among a plurality of animation data for implementing a plurality of dynamic motions designated in advance in connection with the selected NPC, only animation data for implementing an action script to be implemented on a scenario script generated from the SCENE language model can be selected and transmitted to the instructor's terminal in the local area.

[0225] The above NPC can be rendered using webGL (web Graphic Library), which is a graphic library or rasterization engine, and three.js, which is a 3D graphic library for JavaScript, and using the renderer of three.js and webGL, which corresponds to the rasterization engine, a 3D mesh forming a scene within the scope of a scene captured by a camera can be rendered as a 2D image according to the angle of the camera, and here, the 3D mesh can include vertex data regarding the geometric shape of the NPC or a texture that can be mapped to the surface of a geometry that includes primitive data such as points, lines, and triangles. In one embodiment of the present invention, the rendering data for implementing the rendering of the NPC may include data regarding geometry of the shape of the NPC, data regarding material of surface properties of the geometry, and data regarding texture mapped to the surface of the geometry, and may also include a renderer for implementing an NPC that forms a scene within the range of a scene captured by a camera, and a camera function for setting a camera frustum or a camera function for setting a field of view (angle of view) of the camera.

[0226] In this way, in one embodiment of the present invention, the metaverse generation network or the processing server in which the metaverse generation network is implemented can transmit transmission code or transmission data including rendering data for rendering a three-dimensional view of an NPC and surrounding objects selected from an object selection language model that takes as input a scenario script generated from a SCENE language model, data regarding a rendering area in which the NPC and surrounding objects are to be rendered generated from a space layout setting model that takes as input a text in which the entity names of the selected NPC and surrounding objects are to be rendered, a quest list generated from a quest generation language model that takes as input a scenario script generated from the SCENE language model, and two-dimensional image frame data regarding a background image that provides an environment surrounding the selected NPC and surrounding objects, to a local area of ​​an instructor in which a digital space of the metaverse is implemented.

[0227] Conversational Language Model

[0228] In one embodiment of the present invention, a dialogue / action script displayed by a tutor can be input as a prompt for a conversational language model in text form, and the conversational language model can be implemented as a GPT model, which is a forward language model based on the decoder structure of a transformer. For example, the GPT model can be pre-trained from a large-scale training data (unlabeled corpus) to acquire general-purpose natural language processing knowledge. For example, the GPT model can learn parameters from an objective function or loss function that maximizes the conditional probability (log probability) of the next word, conditioned on a set of previously input words, so as to predict the next word from a sequentially input word sequence. For example, the conversational language model can predict the probability of the next word token by taking the inner product between the output value for the input embedding and a plurality of word tokens forming the vocabulary of the conversational language model and the softmax as an output function. The next word can be inferred from the predicted probability in an auto-regressive manner, predicting the next word from all previously input words. For example, the above conversational language model can sequentially predict the next word to appear by using the conversation / action script expressed by the instructor as input for a prompt as a reaction or response to the conversation / action script expressed by the instructor, thereby generating a reaction or response to the conversation / action script expressed by the instructor.

[0229] The above conversational language model is a model pre-trained from a large amount of unlabeled learning data, and can be used for general natural language processing. It can implement in-context learning, which allows processing of tasks without separate information, through zero-shot learning, which allows inference of a given task without separate fine-tuning for a specific downstream task, or it can implement few-shot learning, which provides a few examples so that it can infer a given task without misalignment through input of a prompt.

[0230] In one embodiment of the present invention, the conversational language model can generate a conversation / action script as a response to a conversation / action script expressed by a tutor, and the generated conversation script can be output in text form through a screen-type interface in which a digital space of the metaverse is implemented, and the generated action script can be implemented through the motion of an NPC implemented as a 3D view in the digital space of the metaverse. For example, in one embodiment of the present invention, rendering data for rendering the NPC on the digital space of the metaverse may include animation data for generating NPC motion, and for example, in one embodiment of the present invention, a glb file or a glTF (graphic library Transmission Format) file exported from Blender, which provides a 3D graphic tool, may be imported by the GLTF loader provided by three.js, and the glTF file may store independent animation clips, such as individual movements or individual animations of the NPC, for example, greeting, grabbing, and walking of the NPC, in different fields, and each animation clip may be played by the animation mixer of three.js. For example, the animation mixer may implement a specified movement of the animation, together with a loop function for updating the movement of the NPC or the animation mixer over time.

[0231] In one embodiment of the present invention, the conversational language model generates a conversation / action script as a response to a conversation / action script displayed by the instructor as an input prompt, and at the same time, performs an evaluation of whether a quest provided as a quest list as a task assigned to the instructor in advance from the conversation / action script displayed by the instructor has been accomplished. For example, in one embodiment of the present invention, the quest evaluation from the conversational language model can calculate the similarity between the dialogue / action script expressed from the instructor and each quest in the quest list, for example, by calculating the similarity between the first word sequence regarding the dialogue / action script expressed from the instructor and the second word sequence regarding each quest provided in the quest list, or the similarity between the first and second context vectors including the context information of the first and second word sequences, the dialogue / action script expressed from the instructor can be evaluated as having accomplished the corresponding quest for the second word sequence or the second context vector regarding the quest, which is predicted to have a high similarity with the first word sequence or the first context vector regarding the dialogue / action script expressed from the instructor. For example, in one embodiment of the present invention, the first and second context vectors may be mapped onto a lower-dimensional embedding space than the sparse representation so as to reflect the similarity between different word expressions or different sentence expressions, and the first and second context vectors or the first and second word sequences mapped or embedded in adjacent positions, for example, adjacent positions within a Euclidean distance, may be determined to be similar to each other, and the similarity may be calculated from the inner product between the first and second context vectors.

[0232] In one embodiment of the present invention, a geared metaverse environment can be provided through gear such as a head-mounted set worn by an instructor or a controller into which the instructor's operations are input, and among the dialogue / action scripts displayed by the instructor, an action script can be input through the instructor's manipulation of the controller, and such manipulation of the controller by the instructor can be input as a prompt of the interactive language model in text form. For example, the instructor's physical manipulation of the controller can be processed into a data signal that can be expressed in text form through the controller and input as a prompt of the interactive language model, and for example, it can be input through the prompt of the interactive language model in time series together with other dialogue / action scripts displayed by the instructor. In this way, the conversation / action script displayed by the instructor can be sequentially input as a prompt of the conversational language model, and the conversational language model generates a conversation / action script as a response to the conversation / action script of the instructor sequentially input through the prompt, and at the same time, it can evaluate whether the conversation / action script of the instructor has been completed by calculating the similarity between the conversation / action script of the instructor and each quest on the quest list.

[0233] FIGS. 15A to 16C illustrate an action script that forms a scenario script generated from a SCENE language model using a specific situation set in a metaverse creation network of the present invention as a prompt input, and an attribute data of a plurality of digital assets stored in a database, which exemplarily show interactions of characters selected as NPCs or surrounding objects based on the similarity between a specified motion of an NPC and a specified physical / digital interaction of surrounding objects, and an exemplarily show quests inferred from a scenario script generated from a SCENE language model using each set situation as a prompt input.

[0234] More specifically, FIG. 15a illustrates a diagram exemplarily showing a scenario script (action script) generated from a SCENE language model with a situation set to order a cafe drink as a prompt input and an action script that forms a scenario script among a plurality of digital assets in which motions designated as attribute data are stored in a database, and a designated motion (required character interaction) required for an NPC to implement the action script. In addition, FIGS. 15b and 15c illustrate diagrams exemplarily showing a quest generated from the quest generation language model with a situation set to order a cafe drink as a prompt input and a scenario script generated from the SCENE language model as input.

[0235] Also, in Fig. 16a, a diagram is shown exemplarily showing a designated motion (required character interaction) required for an NPC to implement an action script that forms a scenario script among a number of digital assets in which a scenario script (action script) generated from a SCENE language model and a motion designated as attribute data stored in a database are stored, with a situation set as a stall as a prompt input.

[0236] And, in Fig. 16b, a diagram is shown exemplarily showing a designated physical or digital interaction required for a surrounding object to implement an action script that forms a scenario script among a number of digital assets in which a scenario script (action script) generated from a SCENE language model and a physical or digital interaction designated as attribute data in a database are stored, using a situation set as a stall as a prompt input.

[0237] In addition, FIG. 16c shows an example of a quest generated from the quest generation language model by using a situation set as a stall as a prompt input and a scenario script generated from the SCENE language model as input.

[0238] Although the present invention has been described with reference to the embodiments shown in the attached drawings, these are merely exemplary, and those skilled in the art to which the present invention pertains will understand that various modifications and equivalent other embodiments are possible therefrom.

[0239] The present invention relates to a metaverse creation network for providing a three-dimensional teaching environment, and can be applied to a system for providing various teaching environments including an LMS (Learning Management System).< / eos> < / sos>

Claims

As a metaverse creation network to provide a 1.3D teaching environment, A SCENE language model for generating scenario scripts, which are text-based dialogue / action scripts arranged chronologically about conversations or actions between a tutor and an NPC (non-player character) from prompt inputs about a set situation; An object selection language model for selecting NPCs and surrounding objects to interact with the NPCs as objects to be implemented as 3D views in the digital space of the metaverse from a number of digital assets stored in a database, using the generated scenario script as input; A space layout setting model for setting a rendering area by inputting the object names of the selected NPC and surrounding objects and inferring the rendering location and rendering scale at which the selected NPC and surrounding objects will be rendered in the digital space of the metaverse; and A metaverse generation network for providing a three-dimensional teaching environment, including a quest generation language model for generating a quest list in the form of a list assigned as a task of an instructor from a generated scenario script, and a quest generation language model for generating a quest list of a higher concept encompassing dialogue / action scripts arranged in time series from the generated scenario script.

2. In paragraph 1, The above quest generation language model is a metaverse generation network for providing a three-dimensional teaching environment, characterized in that it extracts a topic of a higher concept from a dialogue / action script forming a generated scenario script and generates a quest list in which the extracted topic or quest is arranged in chronological order.

3. In paragraph 1, The above quest generation language model is, A first natural language processing for classifying the chronologically arranged dialogue / action scripts from the generated scenario script into quest blocks according to their topics; and A metaverse generation network for providing a three-dimensional teaching environment, characterized by performing a second natural language processing to generate a topic or quest of a higher concept encompassing at least one or more dialogue / action scripts classified into the same quest block.

4. In paragraph 3, The above quest generation language model is, A metaverse generation network for providing a three-dimensional teaching environment, characterized in that the first natural language processing is performed by analyzing the similarity between neighboring dialogue / action scripts according to the order of appearance of a plurality of dialogue / action scripts arranged chronologically on the generated scenario script.

5. In paragraph 4, The above quest generation language model is, A metaverse generation network for providing a three-dimensional teaching environment, characterized in that pairs of adjacent dialogue / action scripts that are the subject of similarity analysis are set to overlap each other so that the boundaries of the quest block can be set according to the order of appearance of a plurality of dialogue / action scripts arranged chronologically on the generated scenario script.

6. In paragraph 3, The above quest generation language model is, A metaverse generation network for providing a three-dimensional teaching environment, characterized in that it performs the second natural language processing by extracting key words or clues that appear repeatedly in multiple dialogue / action scripts classified into the same quest block.

7. In paragraph 3, The above quest generation language model is, A metaverse generation network for providing a three-dimensional teaching environment, characterized in that first and second context vectors are derived from contextual embeddings of first and second word sequences forming a plurality of dialogue / action scripts classified into the same quest block, and a quest of a topic or upper concept categorizing the first and second word sequences is generated from a distribution of a plurality of context vectors mapped onto a contextual embedding space.

8. In paragraph 1, The above SCENE language model is a metaverse creation network for providing a three-dimensional teaching environment, characterized in that it receives a style setting or a demonstration of few-shot learning for generating a scenario script in which the subject of each dialogue / action script is specified from a prompt in which the above-mentioned set situation is input.

9. In paragraph 1, The above object selection language model is, A metaverse system for providing a three-dimensional teaching environment, characterized in that natural language processing for selecting the above NPC and natural language processing for selecting the above surrounding objects are performed as different processes.

10. In paragraph 1, The above object selection language model is for natural language processing for the selection of the NPC. Taking as input a first word sequence regarding the above scenario script, a first context vector is generated that includes association information between each word forming the first word sequence, A metaverse system for providing a three-dimensional teaching environment, characterized in that it includes an attention architecture, which inputs a second word sequence regarding entity names of a plurality of digital assets stored in the database and attribute data stored in association with the entity names, and generates a second context vector including association information between each word forming the second word sequence.

11. In paragraph 1, The above object selection language model, in natural language processing for the selection of the NPC, A metaverse system for providing a three-dimensional teaching environment, characterized in that the NPC is selected from a similarity analysis between first and second context vectors that are contextually embedded with respect to a first word sequence regarding the scenario script and a second word sequence regarding a digital asset stored on the database.

12. In paragraph 1, The above object selection language model is, In the natural language processing for selecting the above NPC, the similarity between the action script of the generated scenario script and the dynamic motion specified for each NPC as the attribute data of the NPC is analyzed, A metaverse generation network for providing a 3D teaching environment, characterized in that in natural language processing for selection of the above surrounding objects, the similarity between the action script of the generated scenario script and the physical interaction or digital interaction with the NPC specified for each surrounding object as attribute data of the surrounding object is analyzed.

13. In paragraph 1, The above object selection language model is, in natural language processing for selection of the surrounding objects, A metaverse generation network for providing a three-dimensional teaching environment, characterized in that natural language processing is performed for selection of the surrounding objects from a similarity analysis between each word forming a first word sequence regarding the scenario script and each word forming a second word sequence regarding entity names of a plurality of digital assets stored in the database and attribute data stored in association with the entity names.

14. In paragraph 1, The above space layout setting model is, Create text-to-image by inputting the object names of selected NPCs and surrounding objects as text; and A metaverse generation network for providing a 3D teaching environment, characterized in that it performs region extraction to extract each NPC image region and surrounding object image region from the generated image.

15. In paragraph 14, The above space layout setting model is, To create the above text-to-image, A metaverse generation network for providing a 3D teaching environment, characterized by including a diffusion architecture that implements a denoising process that generates a less noisy image that is relatively closer to the original image from a noisy image so that the pattern of the original image is restored as a reverse process of a noising process that gradually collapses the pattern of the original image by adding noise defined from a noise schedule while advancing the time step from the original image.

16. In paragraph 14, The above space layout setting model is, To create the above text-to-image, An encoder that learns a multi-modal embedding space of text-images to encode the input text into a text embedding by inputting a text containing the object names of the above NPC and surrounding objects; A prior for converting the above text embedding into an image embedding that matches the above text embedding; and A metaverse generation network for providing a three-dimensional teaching environment, characterized in that it comprises a decoder for generating an image as an output of text-to-image from an image embedding converted from the above-mentioned prior.

17. In paragraph 16, The above prior includes an architecture of conditioning diffusion that injects the text embedding as a condition for generating the image embedding, A metaverse generation network for providing a three-dimensional teaching environment, characterized in that the decoder includes an architecture of conditioning diffusion that injects the image embedding as a condition for generating an image as an output of the text-to-image.

18. In paragraph 14, The above space layout setting model is used to extract the above area. A metaverse generation network for providing a 3D teaching environment, characterized in that object detection or image segmentation is performed to distinguish each NPC image area and the image area of surrounding objects from an image generated through the text-to-image generation.

19. In paragraph 18, In the above object detection, the boundary of each NPC image area and the image area of the surrounding objects is predicted in the form of a bounding box surrounding the NPC image area and the image area of the surrounding objects on the generated image, A metaverse generation network for providing a 3D teaching environment, characterized in that in the image segmentation, the NPC image area and the image area of the surrounding objects are predicted on a pixel basis on the generated image.

20. In paragraph 14, The above space layout setting model is, A metaverse generation network for providing a 3D teaching environment, characterized in that it generates data regarding a rendering area in which the NPC and surrounding objects are to be rendered, the rendering area including the center position of a bounding box surrounding the rendering area and the width x height size of the bounding box.

21. In paragraph 1, A metaverse generation network for providing a 3D teaching environment, characterized in that it further includes a background image generation model that implements text-to-image generation by inputting text, for generating a background image to provide an environment surrounding the NPC and surrounding objects.

22. In paragraph 1, A metaverse generation network for providing a 3D teaching environment, characterized in that it transmits transmission code or transmission data including rendering data for rendering a 3D view of an NPC and surrounding objects selected from an object selection language model that uses a scenario script generated from the SCENE language model as input, data regarding a rendering area in which the NPC and surrounding objects are to be rendered generated from a space layout setting model that uses a text combining the entity names of the selected NPC and surrounding objects as input, a quest list generated from a quest generation language model that uses a scenario script generated from the SCENE language model as input, and 2D video frame data regarding a background image that provides an environment surrounding the selected NPC and surrounding objects, to a local area of an instructor in which a digital space of the metaverse is implemented.

23. In paragraph 22, A metaverse creation network for providing a three-dimensional teaching environment, characterized in that the above transmission code or transmission data further includes animation data for implementing dynamic motion expressed as an action script of a scenario script generated from a SCENE language model among dynamic motions specified as attribute data of a selected NPC.

24. In paragraph 22, A metaverse creation network for providing a three-dimensional teaching environment, characterized in that the above transmission code or transmission data further includes a call function for transmitting a dialogue / action script expressed by the instructor toward the processing server.

25. In paragraph 1, A metaverse generation network for providing a three-dimensional teaching environment, characterized in that it further includes an interactive language model for generating a dialogue / action script as a response to a dialogue / action script expressed by the instructor.

26. In paragraph 25, The above conversational language model is, A metaverse generation network for providing a three-dimensional teaching environment, characterized in that it performs natural language processing to evaluate whether a quest generated from a quest generation language model is accomplished from a dialogue / action script expressed from the above instructor.

27. In paragraph 26, The above conversational language model is, A metaverse generation network for providing a three-dimensional teaching environment, characterized in that the achievement of the quest is evaluated based on a similarity analysis between the word sequence forming the quest and the word sequence forming the dialogue / action script expressed from the instructor.

Citation Information

Patent Citations

  • System for providing metaverse based language education service

    KR102535936B1

  • Multimodal-Based Metaverse Environment Implementation System, Method and Computer-Recordable Medium

    KR102612320B1

  • C.i.c.

    KR102640880B1

  • Use of immersive real-time metaverse and avatar and 3-D hologram for medical and veterinary applications using spatially coordinated multi-imager based 3-D imaging

    US11850005B1

  • KR20230132260A

Cited By

  • Multi-source data and physics combined driven drainage basin runoff uncertainty forecasting method

    CN120596859A