Method for progressing content based on the reaction of user in a metaverse environment

KR103024689B1Active Publication Date: 2026-09-29ELECTRONICS & TELECOMM RES INST
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
KR1020220153756
Authority / Receiving Office
KR · KR
Patent Type
Patents
Current Assignee / Owner
Filing Date
2022-11-16
Publication Date
2026-09-29
Estimated Expiration
2042-11-16

Smart Images

  • Figure 112022122197910-PAT00001_ABST
    Figure 112022122197910-PAT00001_ABST
Patent Text Reader

Abstract

An embodiment of the present invention relates to a method for providing content in a metaverse environment, a content providing device, and a learning device, wherein the content providing device may include the steps of: providing content to a metaverse user; the content providing device obtaining user response information corresponding to the content; the content providing device obtaining a user response level based on the user response information using a multimodal artificial intelligence model; and the content providing device providing modified content based on the user response level to the metaverse environment.
Need to check novelty before this filing date? Find Prior Art

Description

Technology Field

[0001] The present invention relates to a method for changing content provided in a metaverse environment based on the responsiveness of a plurality of users connected to the metaverse environment. Background Technology

[0002] Recently, various companies have been producing diverse media or content targeting the general public for advertising or corporate event promotion, and providing it through websites and mass media. In addition, with the rapid rise of non-face-to-face services, metaverse-type content that enables multiple users to interact online is becoming more active.

[0003] However, since the content provided in the aforementioned metaverse environment is produced based on the author's experience and creativity, there was a lack of methodology to quantitatively verify the level of satisfaction and fondness users have for such content.

[0004] Recently, there are also examples where users' emotional states are identified through sensing of biosignals or behaviors and used as a basis for analyzing service utility.

[0005] Specifically, examples include technology that analyzes a user's emotional state by analyzing facial images captured while the user is watching content, and technology that derives the user's emotions by analyzing electroencephalograms (EEGs) detected in the frontal lobe using biosignal sensors.

[0006] However, conventional methods focused on identifying the emotional state of people encountering or experiencing content, or were used to evaluate the utility of content based on analyzed emotional states, and there was no methodology to change content provided in the metaverse environment in real time. The problem to be solved

[0007] The present disclosure is intended to use artificial intelligence to change content provided to multiple users connected to a metaverse environment based on user response. means of solving the problem

[0008] The present disclosure discloses a method for providing content in a metaverse environment, comprising the steps of: a content providing device providing content to a metaverse user; the content providing device acquiring user response information corresponding to the content; the content providing device acquiring a user response level based on the user response information using a multimodal artificial intelligence model; and the content providing device providing modified content to the metaverse environment based on the user response level.

[0009] Additionally, the user response information includes at least one of the user's facial expression information, gaze information, and motion information, and the step of the content providing device obtaining a user response based on the user response information using a multimodal artificial intelligence model may include a step of determining whether the content corresponds to a predetermined event time and a step of obtaining the user response by inputting at least one of the user's facial expression information, gaze information, and motion information into the multimodal artificial intelligence model.

[0010] In addition, the step of the content providing device providing modified content based on the user response to the metaverse environment may include the step of outputting to the metaverse environment by loading effect assets and props from a resource database.

[0011] In addition, the step of the content providing device acquiring user response based on user response information using a multimodal artificial intelligence model may include a step of determining whether the content corresponds to a predetermined content start time and a step of acquiring user response by inputting at least one of the user's facial expression information, gaze information, and motion information into the multimodal artificial intelligence model.

[0012] Additionally, the step of the content providing device providing modified content to the metaverse environment based on the user response level may include changing the content scenario to a scenario corresponding to the level of response obtained among the content scenarios and outputting it, or providing the modified content by outputting it after the currently ongoing scenario.

[0013] In addition, the multimodal artificial intelligence model includes a plurality of convolutional neural networks, and each of the plurality of convolutional neural networks can output features corresponding to the user's facial expression information, gaze information, and motion information, respectively.

[0014] In addition, the multimodal artificial intelligence model includes a recurrent neural network, and the recurrent neural network can output user responsiveness based on the user's facial expression information, gaze information, and motion information output from each of the plurality of convolutional neural networks.

[0015] A content providing device for providing content in a metaverse environment according to an embodiment of the present invention may include a playback unit for providing content to a metaverse user, a shooting unit for acquiring user response information corresponding to the content, and a processor for acquiring user response based on the user response information using a multimodal artificial intelligence model and providing modified content based on the user response to the metaverse environment. Additionally, when the content corresponds to a predetermined event time, the processor may acquire the user response by inputting at least one of the user's facial expression information, gaze information, and motion information into the multimodal artificial intelligence model.

[0016] In addition, when the content corresponds to a predetermined event time, the processor can provide modified content based on user response to the metaverse environment by loading effect assets and props from the resource database.

[0017] In addition, when the content corresponds to a predetermined content start time, the processor can input at least one of the user's facial expression information, gaze information, and motion information into the multimodal artificial intelligence model to obtain the user response.

[0018] In addition, the processor may change to a scenario corresponding to the level of responsiveness obtained among the content scenarios and output it, or provide changed content by outputting it after the currently ongoing scenario.

[0019] A learning device for providing content to a metaverse environment using an artificial intelligence model according to an embodiment of the present disclosure may include: a stimulus storage unit storing a stimulus video, a still image, or a sound designed to induce emotions related to user responsiveness; a playback unit playing the content; a shooting unit acquiring user facial expressions, gaze, posture, and motion information using at least one camera; a user input processing unit storing specific segments set by a user; and a labeling unit providing the content of the specific segments to a user, acquiring a responsiveness score based on a predetermined standard, and labeling the video sequence data of the specific segments with the responsiveness score determined by the user to generate learning data.

[0020] Additionally, the aforementioned specific section may include the onset point and the ending point among the points marked by clicking the user's mouse or remote control, with the preceding point being set as the onset point and the subsequent point as the ending point. Effects of the invention

[0021] The present disclosure can provide users of the metaverse environment with a newer content experience and higher satisfaction by using artificial intelligence to change content provided by a content provider based on user response to content received by multiple users connected to the metaverse environment.

[0022] In addition, the quality of services can be improved by providing service providers with objective result data regarding user response. Brief explanation of the drawing

[0023] FIG. 1 shows an example of a content provision environment according to one embodiment of the present invention. FIG. 2 illustrates the process of learning an artificial intelligence model according to one embodiment of the present invention. FIG. 3 relates to a process in which a user marks a highly responsive part during the learning process of an artificial intelligence model according to an embodiment of the present invention. Figure 4 shows the configuration of an artificial intelligence model according to an embodiment of the present invention. Figure 5 illustrates a content provision process based on user responsiveness according to an embodiment of the present invention. Figure 6 illustrates an example of a content provision process based on user responsiveness according to an embodiment of the present invention. Specific details for implementing the invention

[0024] The technology described below is subject to various modifications and may have various embodiments, and specific embodiments are illustrated in the drawings and described in detail. However, this is not intended to limit the technology described below to specific embodiments, and it should be understood that it includes all modifications, equivalents, and substitutions that fall within the spirit and scope of the technology described below.

[0025] Terms such as first, second, A, B, etc., may be used to describe various components, but such components are not limited by the said terms and are used solely for the purpose of distinguishing one component from another. For example, without departing from the scope of rights of the technology described below, the first component may be named the second component, and similarly, the second component may be named the first component. The term "and / or" includes a combination of multiple related described items or any of the multiple related described items.

[0026] In terms used in this specification, singular expressions should be understood to include plural expressions unless the context clearly indicates otherwise, and terms such as “includes” should be understood to mean that the described features, number, steps, actions, components, parts, or combinations thereof exist, and not to exclude the existence or addition of one or more other features, numbers, steps, actions, components, parts, or combinations thereof.

[0027] Before providing a detailed description of the drawings, it is to clarify that the classification of components in this specification is merely based on the primary function each component is responsible for. That is, two or more components described below may be combined into a single component, or a single component may be divided into two or more components based on more subdivided functions. Furthermore, each component described below may additionally perform some or all of the functions of other components in addition to its own primary function, and it goes without saying that some of the primary functions of each component may be exclusively performed by other components.

[0028] Furthermore, in performing the method or operation method, each process constituting the method may occur differently from the specified order unless a specific order is clearly indicated in the context. That is, each process may occur in the same order as specified, may be performed substantially simultaneously, or may be performed in the reverse order.

[0029] First, the definitions of the terms used in the following description are as follows.

[0030] The "metaverse" according to the embodiments of the present disclosure may refer to a virtual world connected to real life and legally recognized activities such as occupation, finance, and learning. Specifically, as a higher-level concept than virtual reality and augmented reality, it may refer to a system that extends reality into a digital-based virtual world, enabling all activities to be performed in a virtual space.

[0031] The 'content' according to an embodiment of the present disclosure may include video, document, image, VR, and AR content, and may include information that can be decoded through an electronic device or application.

[0032] The 'user response information' according to an embodiment of the present disclosure may include various information for deriving the 'user response level' to be described later. For example, the user response information may include user facial expression information, gaze information, user action information (user posture, movements, gestures, etc.).

[0033] 'User responsiveness' may be a score representing the degree based on the above-mentioned user responsiveness information. For example, if the degree of user response based on features extracted from user responsiveness information exceeds a preset level, user responsiveness may be measured as high, and if the degree of user response is lower than the preset level, user responsiveness may be measured as low.

[0034] A multimodal artificial intelligence model according to an embodiment of the present disclosure may refer to an artificial neural network model designed to derive user responsiveness based on the user responsiveness information.

[0035] A 'metaverse service provider' according to an embodiment of the present disclosure may provide the metaverse environment, and the metaverse service may be implemented through a specific terminal, a server, or a cloud.

[0036] Additionally, 'content provider' may refer to the author or provider of content executed in the aforementioned metaverse environment. A terminal or server used by the content provider may be referred to as a content providing device.

[0037] Depending on the implementation example, the metaverse service provider and the content provider may be the same.

[0038] A user according to an embodiment of the present disclosure may correspond to a user who accesses the metaverse environment, and a user terminal may refer to an electronic device used by the user. Examples of user terminals may include various devices capable of accessing the metaverse, such as mobile phones, laptops, HMDs (Head Mounted Displays), VR devices, etc.

[0039] The actions performed by the user in the metaverse environment described below may be transmitted to a service provider or a content provider device via a user terminal, and the user terminal, content provider terminal, and service provider terminal may include a device that processes input data in a specific manner and performs operations necessary for content provision according to a specific model or algorithm. For example, the content provider device may be implemented in the form of a PC, a server on a network, a smart device, a chipset with a design program embedded therein, etc.

[0040] The following describes a configuration for changing content based on the responsiveness of multiple users participating in a virtual content environment within a metaverse.

[0041] FIG. 1 is an example illustrating a user content provision environment in one embodiment of the present disclosure.

[0042] Meanwhile, Figure 1 illustrates a content provision environment for obtaining user responsiveness corresponding to content, and a system for implementing the content provision environment can be implemented similarly to an artificial intelligence model learning environment.

[0043] The above system may include an output device including a display and a speaker, at least one shooting unit (10) for receiving user operation information, and a user terminal providing a metaverse environment. Depending on the implementation example, it is also possible for an output unit and a camera mounted on the user terminal to perform the role of the output device and the camera.

[0044] From the perspective of a content provider, the content provision system of Fig. 1 above may be a system for learning an artificial intelligence model.

[0045] Specifically, it may be a system that collects training data necessary for training an artificial intelligence model to determine how users react to or respond to content provided by a content provider.

[0046] When collecting training data, data (still images or videos) regarding the user's facial expressions, gaze, and body movements can be obtained through a shooting unit (including a camera, etc.) equipped in the user terminal or the system while users are watching specific content.

[0047] Meanwhile, acquiring diverse training data is crucial for improving the performance of AI models. For example, it is very important to obtain videos of various movements, gazes, and facial expressions related to user reactions.

[0048] The training data of the above artificial intelligence model may include various emotional states related to positive and negative emotions, such as joy, emotion, and sadness, as well as stimulus information such as still images, videos, and sounds that can induce said emotional states.

[0049] The above stimulus information can be obtained by capturing the reaction patterns of each user to content provided to the user through a display or speaker using a camera unit. The acquired data can be stored in a database as training data.

[0050] The learning data according to an embodiment of the present invention can be formed based on various factors such as user preference, stimulus threshold, race, gender, age, nationality, and language.

[0051] For example, even if the stimulus information is the same video or sound content, the response of each user may differ based on various factors such as preference, race, gender, age, nationality, and language. Therefore, it is also possible to generate training data by allowing each user to freely select what they prefer regarding a specific emotional state during the preparation stage before filming.

[0052] The learning process of the artificial intelligence model of the present invention is described below.

[0053] FIG. 2 illustrates the process of learning an artificial intelligence model according to one embodiment of the present invention.

[0054] Referring to Figure 2, the process of generating training data is illustrated by providing at least one piece of content to a user to obtain a corresponding user response, and storing it in a database.

[0055] The subject of the operation shown in FIG. 2 may be a learning device, and the learning device may be equipped with separate equipment to store a learning database, and the stored database may be provided to a content provider or transmitted to a server for learning. In addition, the learning device may be a content provider.

[0056] After the artificial intelligence model is trained by the learning device, the content provider can acquire user responsiveness in the same way as the operation performed by the learning device. Therefore, the content provider can be interpreted as encompassing all the components of the learning device.

[0057] According to an embodiment of the present disclosure, a learning device may include a stimulus storage unit (20) for storing content, a selection unit (30) for selecting content, a stimulus playback unit (40) for playing content, a shooting unit (10) for photographing a user, a synchronization unit (50) for synchronizing the shooting unit (10) and the stimulus playback unit (40), a user input processing unit (60) for processing user input, a user survey unit (70) for storing user surveys, and a storage unit (80, 90) for performing data labeling and storage. The above components may be electrically or structurally connected and may be operated by at least one processor.

[0058] The selection unit (30) selects content from the stimulus storage unit (20) which stores stimulus videos, still images, or sounds designed to induce emotions related to responsiveness, and the playback unit (40) can play the selected content.

[0059] At this time, the stimulus playback unit (40) is synchronized with the shooting unit (10), and the shooting unit (10) can acquire user facial expressions, gaze, posture, and motion information using at least one camera.

[0060] The user input processing unit (60) can store specific sections set by the user watching the content. At this time, the specific section can be marked by clicking the user's mouse or remote control, and among the marked points, the earlier point is set as the Onset point (61) and the subsequent point is set as the Ending point (62), and the area between the Onset point (61) and the Ending point (62) can be set as the specific section. This will be explained in detail below in FIG. 3.

[0061] The user survey section (70) provides the content of the above-mentioned specific section to the user again and can obtain an appropriate response score from a predetermined standard through the survey.

[0062] The labeling unit (80) and the data storage unit (90) can generate training data by labeling video sequence data of a specific section with a response score determined by the user (90).

[0063] FIG. 3 relates to a process in which a user marks a highly responsive part during the learning process of an artificial intelligence model according to an embodiment of the present invention.

[0064] According to an embodiment of the present invention, training data can be generated as a responsiveness score based on user facial expressions, gaze, posture, and movement information obtained in a specific section of the entire content. In other words, the 'user facial expressions, gaze, posture, and movement information' and 'user responsiveness' (score) of a specific section can be stored as a training data pair.

[0065] Meanwhile, it is also possible to use at least one of the user's facial expressions, gaze, posture, and motion information.

[0066] The above specific section can be set by the user. For example, the learning device can set a specific section of the content by requesting the user to click a wireless mouse or remote control (Onset point) when a scene section with high likeability appears while watching video content, and by requesting the user to click once more (Ending point) when it is determined that the likeability has decreased.

[0067] When a specific section is set, the shooting unit (10) can capture the user's facial expression, gaze, posture, and movements during the specific section to obtain information on the user's facial expression, gaze, posture, and movements. Additionally, the responsiveness score of the specific section can be obtained by the user inputting it later.

[0068] Meanwhile, the aforementioned specific interval can be set in the content in advance by the content provider, and if the interval set by the content provider differs from the specific interval set by the user when actual training data is generated, user facial expressions, gaze, posture, and motion data acquired in the specific interval set by the user may be used for training.

[0069] Through this, since users are continuously filmed with a camera for a considerable amount of time, it will be possible to identify which parts show high levels of engagement and acquire only valid data for training, such as videos of user facial expressions and movements while watching those specific parts.

[0070] FIG. 4 is a diagram showing the configuration of an artificial intelligence model according to an embodiment of the present disclosure.

[0071] Referring to FIG. 4, the artificial intelligence model may be composed of a convolutional neural network (200) and a recurrent neural network (400). Specifically, training input data (100) is input into the convolutional neural network (200), and the result value (300) output from the convolutional neural network (200) is input into the recurrent neural network (400) to derive a final result (500).

[0072] As mentioned above, the input data (100) required for artificial intelligence learning consists of various facial expression information, gaze information, posture information, and motion information obtained from multiple users. This is highly correlated with the user's emotional state, such as content immersion, satisfaction, and responsiveness. However, the form of reaction expressed or displayed externally may vary from person to person for each user.

[0073] For example, the aforementioned response patterns may vary based on race, age, personality, culture, response threshold, etc.; in the case of certain individuals, the degree of response may be strongly evident in facial expressions, whereas in others, it can be easily identified through large gestures.

[0074] Therefore, it is necessary to objectively determine the level of response of people encountering or experiencing the content service by simultaneously acquiring several characteristic external reaction patterns, such as facial expressions, gaze, and body poses, that users adopt. This is explained in Fig. 2.

[0075] An artificial intelligence model according to an embodiment of the present invention inputs each of the user's facial expression information, gaze information, and motion information (which may include posture information) into a convolutional neural network (200) to detect features for each of the facial expression information, gaze information, and motion information.

[0076] For example, the convolutional neural network (200) may include a Convolutional Neural Network (CNN), and the CNN model may be composed of three convolutional neural networks corresponding to facial expression information, gaze information, and motion information, respectively.

[0077] Specifically, the convolutional neural network (200) can classify the type of facial expression from facial expression information, detect the direction of gaze from gaze information, and detect motion features from motion information.

[0078] For example, the convolutional neural network (200) can classify multiple emotional situations (310) through facial expressions that appear as a response to the content. Examples of the emotions may include Neutral, Happy, Sad, Surprised, Angry, Disgust, Fear, etc.

[0079] Additionally, the convolutional neural network (200) can detect Roll, Pitch, and Yaw information (320) through the gaze direction according to the head pose to determine the level of concentration on the content being played.

[0080] Additionally, the convolutional neural network (200) can detect full-body posture and movement or behavior features to determine the user's movement features (330). By detecting facial expressions, gaze, and movement features as described above, the artificial intelligence model may be a multi-modal model that reflects multiple feature information.

[0081] According to an embodiment of the present disclosure, the output result (300) of an individual convolutional neural network classified and identified as above can be used as an input parameter to a recurrent neural network (RNN).

[0082] Specifically, a recurrent neural network is a sequence model that processes inputs and outputs in sequence units, and it can be a multimodal model that considers facial expressions, gaze, movements, and posture information that change according to specific intervals.

[0083] According to an embodiment of the present disclosure, a recurrent neural network can output a final result of a degree of responsiveness based on facial expressions, gaze, movements, and posture information that change according to a specific section.

[0084] Meanwhile, in order to implement the convolutional neural network and recurrent neural network according to the above embodiment, learning can be performed by labeling the learning input data (100) that appears as a response to the content in the first stage of learning with the result data (300), and learning can be performed in advance by finally evaluating the degree of responsiveness on a 10-point scale through fusion learning based on multimodal features that consider facial expressions, gaze, movements and actions together in the second stage of learning.

[0086] Figure 5 illustrates a content provision process based on user responsiveness according to an embodiment of the present invention.

[0087] First, the artificial intelligence model according to an embodiment of the present invention may be provided in the form of an algorithm, embedded as an engine in a metaverse content platform, etc., used by a content provider. Through the provided artificial intelligence algorithm, users can view various metaverse-type content provided by the content provider, and the content provider may identify user response to the content and modify or change the scenario of the provided content based on said user response.

[0088] To this end, metaverse service providers may pre-produce resources, such as assets for content and effects for various scenarios based on the level of engagement, and download firmware to store them on servers managed by content providers or on each user's device. An embodiment in which a content scenario is modified or changed will be described in detail below through FIGS. 5 and 6.

[0089] FIG. 6 illustrates a metaverse content change scenario according to an embodiment of the present disclosure.

[0090] Referring to Fig. 6, when entering a predefined event (e1, e2, e3) in a situation where metaverse content transmitted from a service provider's server is unfolding over time, special effect events such as setting off fireworks, presenting heart emoticons, other emotional emoticons, and emojis can be produced.

[0091] Alternatively, it shows an example of modifying or changing a content scenario using the response obtained from multiple users watching the metaverse content at the start of the content (S1, S2).

[0092] According to an embodiment of the present invention, a plurality of users can freely view metaverse content by accessing a website or platform of a metaverse service provider using a user terminal. At this time, a content providing device can output the content to the user terminal by providing the metaverse content to a website or metaverse platform using a server or a content provider terminal.

[0093] A camera (510), such as a webcam attached to each user's device, can acquire the user's facial expressions, gaze, and motion information (520) and transmit it to a content provider (530).

[0094] The artificial intelligence algorithm embedded in the above content device can analyze the level of responsiveness expressed by the user (540).

[0095] A content providing device according to an embodiment of the present invention inputs facial expressions, gaze, and motion information obtained from each user into an artificial intelligence model when entering an event or content starting point to derive the response level identified from each of the multiple users, and can measure the response level of all multiple users to the content being provided at that time or section using the derived response level (560).

[0096] In this case, the overall responsiveness of multiple users may be the sum or average of the respective user responsivenesses.

[0097] Meanwhile, it is also possible to calculate the level of responsiveness by inputting at least one piece of information among facial expressions, gaze, and motion information obtained from the user into an artificial intelligence model.

[0098] A content providing device according to an embodiment of the present disclosure can provide modified content by reflecting various effect assets and props in the environment of the metaverse content by loading them from a resource database (570) based on the response calculated at the time of events (e1, e2, e3) (580).

[0099] Through this, it will be possible to induce users to take a greater interest in or continuously immerse themselves in the content being served.

[0100] Alternatively, the content providing device may provide modified or changed content by changing and loading a scenario that matches the level of responsiveness among several pre-set scenarios based on the responsiveness calculated at the start time of the content (S1, S2) (600), or by naturally connecting and unfolding it after the scenario that has been progressed so far (610).

[0101] Through this, interactive content can be provided that allows multiple users to determine the direction of change in the content.

[0102] Through this, in a situation where scenario-rich metaverse-type content is serviced online, an engine within the content platform can continuously analyze in real-time the reaction patterns or overall responsiveness expressed by a large number of users participating in the service environment, and by changing the flow of content or the virtual environment of the metaverse content accordingly, it will be possible to provide users with a newer content experience and higher satisfaction. Furthermore, by providing feedback on the responsiveness of large numbers of users to service creators, it can be used to improve the quality of services currently in operation, or to utilize this information when planning or producing content for new metaverse-type services in the future.

[0103] A person skilled in the art will understand that the various exemplary logic blocks, modules, processors, means, circuits, and algorithm steps described in connection with the embodiments disclosed herein may be implemented by electronic hardware, various forms of programs or design code (referred to herein as software for convenience), or a combination of all such.

[0104] The present invention described above can be implemented as computer-readable code on a medium on which a program is recorded. A computer-readable medium includes all types of recording devices in which data that can be read by a computer system is stored. Examples of computer-readable media include HDD (Hard Disk Drive), SSD (Solid State Disk), SSD (Silicon Disk Drive), ROM, RAM, CD-ROM, magnetic tape, floppy disk, optical data storage device, etc.

Claims

Claim 1 A method for providing content in a metaverse environment, comprising: a step in which a content providing device provides content to a metaverse user; a step in which the content providing device obtains user response information corresponding to the content; a step in which the content providing device obtains a user response level based on the user response information using a multimodal artificial intelligence model; and a step in which the content providing device provides modified content to the metaverse environment based on the user response level, wherein the user response information includes user facial expression information, gaze information, and motion information, the multimodal artificial intelligence model includes a plurality of convolutional neural networks, each of the plurality of convolutional neural networks outputs a feature corresponding to each of the user facial expression information, the gaze information, and the motion information, and the content providing device loads specific effect assets and props from a resource database based on the user response level to provide the modified content to the metaverse environment. Claim 2 delete Claim 3 delete Claim 4 delete Claim 5 A method for providing content according to claim 1, wherein the step of the content providing device providing modified content based on the user response level to the metaverse environment includes the step of changing and outputting a scenario corresponding to the level of response level obtained among the content scenarios, or providing modified content by outputting it after the currently ongoing scenario. Claim 6 delete Claim 7 A content provision method according to claim 1, wherein the multimodal artificial intelligence model further includes a recurrent neural network, and the recurrent neural network outputs user responsiveness based on the user's facial expression information, the gaze information, and the motion information output from each of the plurality of convolutional neural networks. Claim 8 A content providing device for providing content in a metaverse environment, comprising: a playback unit for providing content to a metaverse user; a shooting unit for acquiring user response information corresponding to the content; and a processor for acquiring user response based on the user response information using a multimodal artificial intelligence model and providing modified content to the metaverse environment based on the user response, wherein the user response information includes user facial expression information, gaze information, and motion information, the multimodal artificial intelligence model includes a plurality of convolutional neural networks, each of the plurality of convolutional neural networks outputs a feature corresponding to each of the user facial expression information, the gaze information, and the motion information, and the processor loads specific effect assets and props from a resource database based on the user response to provide the modified content to the metaverse environment. Claim 9 delete Claim 10 delete Claim 11 delete Claim 12 In claim 8, the above processor is a content providing device that outputs a scenario corresponding to a level of responsiveness obtained among content scenarios, or outputs the changed content after the currently ongoing scenario. Claim 13 delete Claim 14 A content providing device according to claim 8, wherein the multimodal artificial intelligence model further includes a recurrent neural network, and the recurrent neural network outputs user responsiveness based on the user's facial expression information, the gaze information, and the motion information output from each of the plurality of convolutional neural networks. Claim 15 delete Claim 16 delete

Citation Information

Patent Citations

  • Method for analysing feedback of virtual reality image

    KR101977258B1

  • Methods for training emotional response predictors utilizing attention in visual objects

    US20160078369A1

  • Increased user efficiency and interaction performance through dynamic adjustment of auxiliary content duration

    US20160127776A1

  • Interactive artificial intelligence analytical system

    US20200065612A1