Image generation device and method therefor

The image generating device and method efficiently convert user motion into realistic metaverse character interactions by removing the user's appearance from background data and matching the user's motion with the metaverse character's motion, addressing the limitations of conventional methods in terms of time, cost, and realism.

WO2025110328A1PCT designated stage expired Publication Date: 2025-05-30EIFENINTERACTIVE CO LTD
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
PCT/KR2023/020788
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-11-20
Filing Date
2023-12-15
Publication Date
2025-05-30

AI Technical Summary

Technical Problem

Conventional metaverse character production methods are time-consuming and costly, requiring manual creation of movements in 2D or 3D computer graphic videos, which lack realism between movements.

Method used

An image generating device and method that tracks a user's motion, converts it into a metaverse character, and generates an image where the metaverse character interacts with other objects without a sense of incongruity, by using background data to remove the user's appearance and a character implementation module to match the user's motion with the metaverse character's motion.

Benefits of technology

The method achieves natural interaction between the user and objects, maximizes the reality of the metaverse character, and improves efficiency by capturing user and object interactions in the same space and time, while maintaining image quality and avoiding incongruity.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure KR2023020788_30052025_PF_FP_ABST
    Figure KR2023020788_30052025_PF_FP_ABST
Patent Text Reader

Abstract

An image generation device is provided. The image generation device comprises: a background generation module that uses background data generated by photographing the appearance of an object interacting with a user to thereby generate a first image from which the user is removed; a character implementation module that uses user motion data related to a motion of the user so as to generate a second image including a motion of a metaverse character; a marker removal module that uses object motion data related to a motion of the object and the background data to generate a third image including a first marker removal object related to a first marker attached to the surface of the object; and an image synthesis module that uses the second image and the third image to generate a final image on the basis of the first image.
Need to check novelty before this filing date? Find Prior Art

Description

Image generation device and method thereof

[0001] The present invention relates to an image generation device and method thereof. Specifically, the present invention relates to a device and method for generating an image by tracking a user's motion, transforming the user into a metaverse character, and allowing the transformed character to interact seamlessly with objects other than the user.

[0002]

[0003] The creation of metaverse character videos is a highly complex process. Traditional metaverse character creation, using 2D or 3D computer graphics, requires creators to manually create each movement, which is both time-consuming and costly. Furthermore, the lack of realism between movements can be a drawback.

[0004] To overcome these shortcomings, recent metaverse character video generation devices are using various methods to bring out the realism of characters, and there is a recent attempt to maximize character realism by implementing the movements of metaverse characters using the movements of actual actors.

[0005]

[0006] The object of the present invention is to provide an image generation device and method that maximize the reality of an image in which a metaverse character appears.

[0007] In addition, the present invention aims to provide an image generation device and method that maximize the reality of an image including a process in which a metaverse character interacts with an object other than a user.

[0008] In addition, the present invention provides an image generation device and method that improve efficiency by capturing a user performing an act and an object other than the user in the same space and time.

[0009] In addition, the object of the present invention is to provide an image generating device and method for generating a natural image without a sense of incongruity in a part where a converted character and an object other than the user overlap each other in a process in which the motion of a user performing a performance is converted into a character.

[0010] The objectives of the present invention are not limited to those mentioned above. Other objectives and advantages of the present invention not mentioned above can be understood through the following description and will be more clearly understood through the embodiments of the present invention. Furthermore, it will be readily apparent that the objectives and advantages of the present invention can be realized by the means and combinations thereof set forth in the claims.

[0011]

[0012] According to some embodiments of the present invention for solving the above problem, an image generation device includes a background generation module that generates a first image from which the user's appearance is removed using background data including the user's appearance, a character implementation module that generates a second image including a motion of a metaverse character matching the user's motion using user motion data related to the user's motion, and an image synthesis module that generates a final image using the first image and the second image, wherein the background data includes an outline of the user and an appearance of an object interacting with the user.

[0013] In addition, the background generation module may include an object grouping module that generates user outline information from the background data received from an image capturing device, a user tracking module that generates user tracking data that tracks the user based on the outline information, and a user removal module that generates the first image using the user tracking data.

[0014] In addition, the user removal module can generate masking data corresponding to the outline information using the user tracking data, generate a user removal image by removing the outline information from the background data using the masking data, and generate the first image using the background data and the user removal image.

[0015] In addition, the user removal module can generate masking data corresponding to the outline information using the user tracking data, correct the masking data using the background data to generate a user removal image, and generate the first image using the background data and the user removal image.

[0016] In addition, the object grouping module can group pixels based on a predetermined criterion centered on each of a plurality of markers attached to the user, and connect the groups of pixels to each other to generate outline information of the user.

[0017] In addition, the character implementation module may further use the user tracking data generated by the background generation module to generate the second image, wherein the background data includes at least one portion where the appearance of the object overlaps the appearance of the user, and the second image may include an image of the metaverse character with the overlapped portion removed.

[0018] According to some embodiments of the present invention for solving the above problem, an image generation method using an image generation device that generates a metaverse character image using a user's motion is provided, the method comprising: a step of generating user outline information using background data including the user's appearance and the appearance of an object interacting with the user; a step of generating user tracking data that tracks the user based on the outline information; a step of generating the first image using the user tracking data; a step of generating a second image including a motion of a metaverse character that matches the user's motion using user motion data related to the user's motion; and a step of generating a final image using the first image and the second image.

[0019] In addition, the step of generating the first image may include a step of generating masking data corresponding to the outline information using the user tracking data, a step of generating a user-removed image by removing the outline information from the background data using the masking data, and a step of generating the first image using the background data and the user-removed image.

[0020] In addition, the step of generating the first image may include a step of generating masking data corresponding to the outline information using the user tracking data, a step of generating a user-removed image by correcting the masking data using the background data, and a step of generating the first image using the background data and the user-removed image.

[0021] Additionally, the step of generating the second image may include a step of generating the metaverse character using the user tracking data.

[0022]

[0023] The image generation device and method of the present invention can naturally express the interaction between a user and an object by removing the overlapping portion between the user and the object and generating a character.

[0024] In addition, the video generation device and method of the present invention can maximize the reality of a character by applying the motion of a user performing an act to the movement and directing effects of a metaverse character.

[0025] In addition, the image generation device and method of the present invention can secure economic feasibility by saving the time and cost required to shoot an image in a separate space and time.

[0026] In addition, the image generation device and method of the present invention can express objects other than the user without a sense of incongruity with the character even after the user has been converted into a character.

[0027] In addition, the image generation device and method of the present invention can produce an image with maximized quality by naturally expressing the overlapping portion of a user and an object other than the user.

[0028] In addition to the above-described contents, the specific effects of the present invention are described together with the specific matters for carrying out the invention below.

[0029]

[0030] FIG. 1 is a drawing schematically illustrating a process of generating a final image using an image generating device according to some embodiments of the present invention.

[0031] Figure 2 is a block diagram showing the configuration of the background generation module of Figure 1.

[0032] Figure 3 is a flowchart showing the process in which the background generation module of Figure 1 generates the first image.

[0033] Figure 4 is a drawing for specifically explaining each data of Figure 1.

[0034] FIG. 5a is a drawing showing an embodiment of generating the outline information of FIG. 2.

[0035] FIG. 5b is a drawing showing another embodiment of generating the outline information of FIG. 2.

[0036] FIG. 6 is a diagram for explaining the process by which the user tracking module of FIG. 2 generates user tracking data.

[0037] FIG. 7a is a diagram illustrating an embodiment of generating a first image in the user removal module of FIG. 2.

[0038] FIG. 7b is a diagram illustrating another embodiment of generating a first image in the user removal module of FIG. 2.

[0039] FIG. 8 is a diagram showing a process of generating a first image from background data transmitted by the image capturing device of FIG. 1.

[0040] Figure 9 is a block diagram showing the configuration of the character implementation module of Figure 1.

[0041] Figure 10 is a flowchart showing the process by which the character implementation module of Figure 8 generates a second image.

[0042] FIG. 11 is a drawing showing a process of generating a second image from user motion data transmitted by the motion capture device of FIG. 1.

[0043] Figure 12 is a diagram illustrating a process of generating a second image when there is an overlapping portion between a user and an object.

[0044] FIG. 13 and FIG. 14 are diagrams for explaining a process of generating a final image through an image synthesis module according to some embodiments of the present invention.

[0045] FIG. 15 is a diagram illustrating a process of generating a final image through an image synthesis module according to some other embodiments of the present invention.

[0046]

[0047] The terms and words used in this specification and claims should not be interpreted based on their general or dictionary meanings. In accordance with the principle that inventors can define the concepts of terms and words to best describe their inventions, they should be interpreted in a way that is consistent with the technical concept of the present invention. Furthermore, the embodiments described in this specification and the configurations depicted in the drawings are merely examples of how the present invention can be realized and do not fully represent the technical concept of the present invention. Therefore, it should be understood that various equivalents, modifications, and applicable examples may exist as of the time of filing.

[0048] The terms first, second, A, B, etc. used in this specification and claims may be used to describe various components, but the components should not be limited by these terms. These terms are used only for the purpose of distinguishing one component from another. For example, without departing from the scope of the present invention, the first component may be referred to as the second component, and similarly, the second component may also be referred to as the first component. The term "and / or" includes any combination of a plurality of related listed items or any item among a plurality of related listed items.

[0049] The terminology used in this specification and claims is for the purpose of describing specific embodiments only and is not intended to limit the present invention. Singular expressions include plural expressions unless the context clearly dictates otherwise. It should be understood that terms such as "comprise" or "have" in this application do not preclude the presence or addition of features, numbers, steps, operations, components, parts, or combinations thereof described in the specification.

[0050] Unless otherwise defined, all terms used herein, including technical or scientific terms, have the same meaning as commonly understood by one of ordinary skill in the art to which the present invention belongs.

[0051] Terms defined in commonly used dictionaries should be interpreted to have a meaning consistent with their meaning in the context of the relevant technology, and will not be interpreted in an idealized or overly formal sense unless expressly defined in this application.

[0052] In addition, each configuration, process, procedure or method included in each embodiment of the present invention may be shared within a scope that is not technically inconsistent with each other.

[0053] The present invention relates to a metaverse character image generation device, wherein 'metaverse character image generation' may mean the production of an image including the movement and directing effects of a metaverse character used in the metaverse, virtual reality, augmented reality, or extended reality. In other words, the 'metaverse character' in this specification may include a virtual digital idol character, game character, etc. implemented in 2D or 3D. In addition, the 'metaverse image' refers to an image including a metaverse character, and may mean an image including not only a metaverse character but also a metaverse background, structures, etc.

[0054] Hereinafter, with reference to FIGS. 1 to 15, an image generating device and method according to some embodiments of the present invention will be described.

[0055] FIG. 1 is a diagram schematically illustrating a process for generating a final image using an image generation device according to some embodiments of the present invention. FIG. 2 is a block diagram illustrating the configuration of the background generation module of FIG. 1. FIG. 3 is a flowchart illustrating the process for generating a first image using the background generation module of FIG. 1. FIG. 4 is a diagram specifically illustrating each piece of data of FIG. 1.

[0056] Referring to FIGS. 1 to 4, an image generation device (10) according to some embodiments of the present invention may include a background generation module (100), a character implementation module (200), and an image synthesis module (300).

[0057] The background generation module (100) can receive background data (BD) from the image capturing device (20). The image capturing device (20) can capture an image to generate background data (BD). At this time, the background data (BD) can include the user's appearance. In other words, the background data (BD) can include the state of the moving user.

[0058] Background data (BD) generated by the video capturing device (20) can be provided to the background generation module (100). The background generation module (100) can generate a first image (V1) using the background data (BD). The first image (V1) can be an image from which the user's appearance has been removed from the background data (BD). A detailed description of the first image (V1) will be described later. The background generation module (100) can provide the generated first image (V1) to the image synthesis module (300).

[0059] The character implementation module (200) can receive user motion data (MD) from a motion capture device (30). The motion capture device (30) can sense the user's movements to generate user motion data (MD). The user motion data (MD) can include data related to the user's movements. Specifically, the user motion data (MD) can include data related to the user's movement direction and movement speed.

[0060] For example, the motion capture device (30) can sense a marker attached to a user to generate user motion data (MD). The user motion data (MD) may include at least one of body motion data, facial motion data, and hand motion data. In this specification, the video recording device (20) and the motion capture device (30) are expressed as different components, but the embodiments are not limited thereto. When the motion capture device (30) uses an optical marker, the video recording device (20) and the motion capture device (30) may be the same components. However, in this specification, for the convenience of explanation, the video recording device (20) and the motion capture device (30) are expressed as being distinct from each other.

[0061] The character implementation module (200) can generate a second image (V2) containing the movement of a metaverse character using user motion data (MD). The metaverse character can correspond to a user. Specifically, the movement of the metaverse character can be matched with the movement of the user contained in the user motion data (MD). In other words, the second image (V2) can be an image of a metaverse character having a movement corresponding to the user motion data (MD).

[0062] In some embodiments, the second image (V2) may be a metaverse character image from which an area where a user (UM) and an object (OB) interact are removed. The object (OB) may refer to an entity other than the user (UM). In other words, the second image (V2) may include a metaverse character image, but the metaverse character image may have an area where the user (UM) and the object (OB) overlap each other removed. A detailed description of the second image (V2) will be described later. The character implementation module (200) may provide the second image (V2) generated using user motion data (MD) and user tracking data (TD) to the image synthesis module (300).

[0063] The image synthesis module (300) can generate a final image (VF) using the first image (V1) received from the background generation module (100) and the second image (V2) received from the character implementation module (200).

[0064] According to some embodiments, the image generation device (10), the image capture device (20), and the motion capture device (30) can exchange data with each other via a network. Data can be transmitted via the network. The network can include a network using wired Internet technology, wireless Internet technology, and short-range communication technology. The wired Internet technology can include, for example, at least one of a local area network (LAN) and a wide area network (WAN).

[0065] Wireless Internet technologies may include, for example, at least one of Wireless LAN (WLAN), Digital Living Network Alliance (DLNA), Wireless Broadband (Wibro), World Interoperability for Microwave Access (Wimax), High Speed ​​Downlink Packet Access (HSDPA), High Speed ​​Uplink Packet Access (HSUPA), IEEE 802.16, Long Term Evolution (LTE), Long Term Evolution-Advanced (LTE-A), Wireless Mobile Broadband Service (WMBS), and 5G NR (New Radio) technologies. However, the present embodiment is not limited thereto.

[0066] Short-range communication technologies may include, for example, at least one of Bluetooth, Radio Frequency Identification (RFID), Infrared Data Association (IrDA), Ultra-Wideband (UWB), ZigBee, Near Field Communication (NFC), Ultra Sound Communication (USC), Visible Light Communication (VLC), Wi-Fi, Wi-Fi Direct, and 5G NR (New Radio). However, the present embodiment is not limited thereto.

[0067] The image generating device (10), the image capturing device (20), and the motion capture device (30) communicating through a network may comply with technical standards and standard communication methods for mobile communication. For example, the standard communication method may include at least one of GSM (Global System for Mobile communication), CDMA (Code Division Multi Access), CDMA2000 (Code Division Multi Access 2000), EV-DO (Enhanced Voice-Data Optimized or Enhanced Voice-Data Only), WCDMA (Wideband CDMA), HSDPA (High Speed ​​Downlink Packet Access), HSUPA (High Speed ​​Uplink Packet Access), LTE (Long Term Evolution), LTEA (Long Term Evolution-Advanced), and 5G NR (New Radio). However, the present embodiment is not limited thereto.

[0068] The background generation module (100) may include an object grouping module (110), a user tracking module (120), and a user removal module (130). The object grouping module (110) may receive background data (BD) from an image capturing device (20).

[0069] For example, the video recording device (20) can generate background data (BD) by recording the entire background including the user (UM) performing the act, the object (OB) interacting with the user (UM), and the place where the user (UM) is being filmed.

[0070] Background data (BD) may include a user (UM). Specifically, the background data (BD) may include the appearance of the user (UM). In this case, the appearance of the user (UM) may refer to the appearance of the user (UM) captured by a video recording device (20). The user (UM) may be the subject performing the performance. Specifically, the user (UM) may be the person performing the lines and actions of the metaverse character.

[0071] Additionally, background data (BD) may include objects (OBs). Objects (OBs) may represent entities other than the user (UM). Objects (OBs) may be people or objects, and may also include tangible entities that interact with the user (UM).

[0072] Meanwhile, the object grouping module (110) can generate outline information (OI) of the user (UM) using background data (BD) (S100). The outline information (OI) may be data that distinguishes the user (UM) from other objects. In this case, the outline information (OI) may be generated from the appearance of the user (UM) included in the background data (BD). In other words, the outline information (OI) may be a boundary line that distinguishes the appearance of the user (UM) from other objects.

[0073] In other words, the object grouping module (110) can generate outline information (OI) by distinguishing users (UM) and objects (OB) from background data (BD). That is, the object grouping module (110) can find and segment objects from background data (BD). Specifically, the object grouping module (110) can find objects from background data (BD) and determine users (UM) and objects (OB) for the found objects, thereby distinguishing users (UM) and objects (OB). That is, the object grouping module (110) can find and segment users (UM) and objects (OB) from background data (BD) to generate outline information (OI).

[0074] The user tracking module (120) can generate user tracking data (TD) that tracks the user (UM) based on the outline information (OI) (S110). The user tracking data (TD) can include changes in the outline information (OI) over time. In other words, the user tracking data (TD) can be data that collects the outline information (OI) of the user (UM) by time zone. In addition, the user tracking module (120) can transmit the user tracking data (TD) to the user removal module (130). A detailed description of the user tracking module (120) will be described later.

[0075] The user removal module (130) can receive user tracking data (TD) from the user tracking module (120). The user removal module (130) can generate a first image (V1) using the user tracking data (TD) (S120). At this time, the user removal module (130) can generate the first image (V1) using the user tracking data (TD). That is, the user removal module (130) can remove the user's appearance using the user tracking data (TD). In other words, the first image (V1) can be an image from which the user's appearance (UM) has been removed and which includes the appearance of the object (OB).

[0076] In some embodiments, the user removal module (130) may generate user removal images (DUa, DUb) using background data (BD) and user tracking data (TD). Subsequently, the user removal module (130) may generate a first image (V1) using the user removal images (DUa, DUb). A detailed description of the first image (V1) will be provided below.

[0077]

[0078] Fig. 5a is a drawing showing one embodiment of generating the outline information of Fig. 2. Fig. 5b is a drawing showing another embodiment of generating the outline information of Fig. 2.

[0079] Referring to FIGS. 2, 5A, and 5B, the object grouping module (110) can group background data (BD) by object using a pre-trained artificial intelligence algorithm. In other words, the object grouping module (110) can receive background data (BD) from the image capturing device (20), and input the received background data (BD) into a pre-trained artificial intelligence algorithm, thereby segmenting and grouping objects included in the background data (BD). For example, the object grouping module (110) can perform supervised learning using a dataset including images and annotations for the images. As another example, the object grouping module (110) can perform self-supervised learning or unsupervised learning using a dataset including images. However, this is an exemplary description, and the embodiments are not limited thereto. That is, the image segmentation technology used in the object grouping module (110) can utilize various known artificial intelligence algorithms.

[0080] In some embodiments, the object grouping module (110) may generate outline information (OI) using a deep learning-based object detection algorithm. The deep learning-based object detection algorithm may be, for example, at least one of YOLO (You Only Look Once), CNN, and SSD (Single Shot Multibox Detector), but the embodiments are not limited to specific algorithms for image segmentation techniques.

[0081] In some embodiments, the object grouping module (110) may segment each object included in the background data (BD). Specifically, the object grouping module (110) may segment a user (UM), an object (OB), and other objects, such as a chair on which the user (UM) is sitting and a chair on which the object (OB) is sitting, into individual objects. In this case, the shape of each object may be stored in the dataset in the form of an image or text. However, the present embodiment is not limited thereto.

[0082] According to some embodiments, the object grouping module (110) may utilize markers. Specifically, the object grouping module (110) may group objects by utilizing markers attached to the user (UM). In this case, the markers may be attached for motion capture of the user. The object grouping module (110) may group similar pixels based on the attached markers. For example, the object grouping module (110) may determine the similarity of pixels in proximity to the first marker (MK1) and group the similar pixels. That is, the object grouping module (110) may determine whether pixels are similar from a close distance based on the first marker (MK1) attached to the head. Subsequently, the object grouping module (110) may group pixels determined to be similar and segment them into the user's head area.

[0083] In addition, the object grouping module (110) can determine the similarity of nearby pixels centered on the second marker (MK2) attached to the torso. Subsequently, the object grouping module (110) can group pixels determined to be similar and segment them into the user's torso. Similarly, the object grouping module (110) can group similar pixels centered on the third a marker (MK3a) and the third b marker (MK3b) attached to both arms of the user and segment them into the user's arm part. In addition, the object grouping module (110) can group similar pixels centered on the fourth a marker (MK4a) and the fourth b marker (MK4b) attached to both legs of the user and segment them into the user's leg part. Subsequently, the object grouping module (110) can generate the user's outline information (OI) based on the segmented head, torso, arms, and legs. Specifically, the object grouping module (110) can generate outline information (OI) by connecting the boundaries of grouped similar pixels.

[0084]

[0085] FIG. 6 is a drawing for explaining the process by which the user tracking module of FIG. 2 generates user tracking data.

[0086] Referring to FIGS. 2 and 6, the user tracking module (120) can receive outline information (OI) from the object grouping module (110). The user tracking module (120) can generate user tracking data (TD) based on the received outline information (OI). The user tracking data (TD) can include temporal changes of a moving user (UM). In other words, the user tracking data (TD) can be data obtained by tracking an area including the outline information (OI) of the user (UM). That is, the user tracking module (120) can generate user tracking data (TD) by arranging the outline information (OI) in chronological order.

[0087] In some embodiments, the user tracking module (120) may be indistinguishable from the object grouping module (110). In other words, the user tracking module (120) may generate tracking data (TD) using background data (BD) including the user's appearance. That is, the user tracking module (120) may track the user (UM) using an artificial intelligence algorithm.

[0088]

[0089] FIG. 7a is a diagram illustrating an embodiment of generating a first image in the user removal module of FIG. 2. FIG. 7b is a diagram illustrating another embodiment of generating a first image in the user removal module of FIG. 2.

[0090] Referring to FIGS. 2, 6, 7a, and 7b, the user removal module (130) can generate masking data (MD) corresponding to outline information (OI) using tracking data (TD). The masking data (MD) may include a layer in which a portion is transparent. Specifically, the masking data (MD) may include a layer in which the user's area (i.e., the area inside the user's outline) is opaque and the area outside the user's area is transparent. However, the embodiments of the present invention are not limited thereto.

[0091] Meanwhile, the user removal module (130) can generate a first image (V1a, V1b) using background data (BD) and tracking data (TD). At this time, the user removal image (DUa) may be an image in which the user's area is removed from the background data (BD), and the part in which the user's appearance is removed may be empty without data or filled with a single color.

[0092] In some embodiments, the user removal module (130) can combine masking data (MD) with background data (BD). That is, the user removal module (130) can generate a user-removed image (DUa) from which the user's area is removed. The user removal module (130) can generate the user-removed image (DUa) by removing outline information (OI) from the background data (BD) using the masking data (MD). Subsequently, the user removal module (130) can correct the portion from which the user's area is removed so that it corresponds to the surrounding background area to generate the first image (V1a).

[0093] As another example, the user removal module (130) can correct the masking data (MD) using the background data (BD). That is, the user removal image (DUb) can be an image in which the user's area included in the masking data (MD) is corrected. Specifically, the user removal module (130) can generate the user removal image (DUb) by correcting the part from which the user's area is removed so that it corresponds to the background data (BD) around the user's area using the background data (BD). Subsequently, the user removal module (130) can generate the first image (V1b) by combining the user removal image (DUb) with the background data (BD). However, the embodiments of the present invention are not limited thereto.

[0094] Finally, the user removal module (130) can generate a first image (V1a, V1b). At this time, the first image (V1a, V1b) may be an image in which the area from which the user has been removed and the surrounding area are blended together, as if the user had never existed from the beginning. In other words, the first image (V1a, V1b) may be an image in which the user area has been removed and the area from which the user has been removed has been corrected.

[0095] Next, the user removal module (130) can generate a first image (V1a, V1b) using background data (BD) and user removal images (DUa, DUb).

[0096] That is, the user-removed image (DUa, DUb) may be an image from which the user has been removed from the background data (BD). In other words, the user-removed module (130) may generate the user-removed image (DUa, DUb) by removing the entire user shape, including the user outline, from the background data (BD) based on the user tracking data (TD). The user-removed module (130) may generate the first image (V1) using the user-removed image (DUa, DUb).

[0097] According to some embodiments, the user removal module (130) may generate the user removed images (DUa, DUb) into the first images (V1a, V2b) by using image synthesis or deep learning techniques. For example, the user removal module (130) may generate the first images (V1a, V1b) by correcting the user removed images (DUa, DUb) by using Generative Adversarial Networks (GANs) or Convolutional Neural Networks (CNNs). In this case, the deep learning technique may be trained by using background information surrounding the area of ​​the removed user.

[0098] As another example, the user removal module (130) can generate image data in an environment where no user exists before background data (BD) is generated, and can use this image data to generate a first image (V1). That is, the user removal module (130) can generate first images (V1a, V1b) by synthesizing the area of ​​the removed user in the user removal images (DUa, DUb) with image data captured in advance. However, the use of deep learning technology or image synthesis technology by the user removal module (130) is merely exemplary, and the present invention is not limited thereto. The user removal module (130) can transmit the first image (V1) generated using the background data (BD) to the image synthesis module (300).

[0099]

[0100] FIG. 8 is a diagram showing a process of generating a first image from background data transmitted by the image capturing device of FIG. 1.

[0101] Referring to FIGS. 1, 4, and 8, an image capturing device (20) can capture a user (UM) and an object (OB) in a given space. In this case, the given space may be a chroma key studio. Subsequently, the image capturing device (20) can transmit background data (BD) from the captured image to a background generation module (100). The background generation module (100) can generate a first image (V1) from the background data (BD). In this case, the background of the first image (V1) may be the same as the space where the image was captured.

[0102] In addition, the background generation module (100) can utilize pre-recorded image data as a background. The background generation module (100) can remove the entire user shape including the outline of the user (UM) from the background data (BD). At this time, the background generation module (100) can also remove the background portion behind the user shape. The background generation module (100) can generate the first image (V1) by correcting the removed background portion to correspond to the surrounding background area close to the corresponding portion. Accordingly, only the user (UM) can be removed from the first image (V1), and the user (UM) and the object (OB) can remain the same.

[0103]

[0104] Fig. 9 is a block diagram showing the configuration of the character implementation module of Fig. 1. Fig. 10 is a flowchart showing the process by which the character implementation module of Fig. 9 generates a second image.

[0105] Referring to FIGS. 9 and 10, the character implementation module (200) receives user motion data (MD) from the motion capture device (30) (S200). At this time, the character implementation module (200) may include a skeleton generation module (210), a retargeting module (220), and a rendering module (230). According to some embodiments, the skeleton generation module (210) may receive user motion data (MD) from the motion capture device (30).

[0106] A motion capture device (30) can record the user's movements as data through sensors and / or markers attached to the user's body. In other words, the motion capture device (30) can sense the user's movements and generate user motion data (MD). The user motion data (MD) may be data expressing the movements of the user's skeleton and muscles. In this case, the user motion data (MD) may include the user's movement direction and movement speed.

[0107] In addition, the motion capture device (30) can record user motion data (MD) for each body part of the user. For example, the motion capture device (30) may include a body capture module that records data on the user's activity motion, a facial capture module that records data on the user's facial expression, and a hand capture module that records the user's hand movements. Specifically, the body capture module can sense the movements of the user's body, such as the head, neck, arms, body, and legs, and record data. For example, the body capture module can sense the movements of the head, neck, arms, body, and legs when the user performs a specific action (walking, running, crawling, sitting, standing, lying down, dancing, fighting, kicking, swinging arms, etc.). However, the body capture module may attach a marker to the user's body and sense it, but the present embodiment is not limited thereto.

[0108] Meanwhile, the facial capture module can sense the movement of the user's face and record data. For example, the facial capture module can sense the movement of the user's face in specific facial expressions (e.g., crying, smiling, surprised, angry, contemptuous, regretful, loving, disgusted, etc.). The facial capture module can additionally use data sensed by attaching a marker to the user's face and image processing technology for the user's facial expression, but the present embodiment is not limited thereto.

[0109] In some embodiments, the hand capture module may be data regarding the user's hand movements. For example, the hand capture module may sense the movement of the finger joints in the user's hand movements (such as curling or extending the fingers). The hand capture module may be implemented as a special glove or wearable device capable of sensing the movement of the finger joints, but the present embodiment is not limited thereto.

[0110] That is, the motion capture device (30) can generate user motion data (MD) related to at least one of the user's gestures, facial expressions, and hand movements by using at least one of a body motion module, a facial motion module, and a hand motion module.

[0111] Meanwhile, the motion capture device (30) may be implemented as an optical, magnetic, and / or inertial motion capture device. The optical motion capture device may include a marker and a camera. The optical motion capture device may be a device that captures the movement of a user who has a marker attached to them with one or more cameras, and generates user motion data (MD) by inversely calculating the three-dimensional coordinates of the marker attached to the user through triangulation. The magnetic motion capture device may be a device that generates user motion data (MD) by measuring movement by attaching a sensor capable of measuring a magnetic field to the user's joints, etc., and calculating the amount of change in the magnetic field of each sensor near a magnetic field generating device. The inertial motion capture device may be a device that generates user motion data (MD) by attaching an inertial sensor including an acceleration sensor, a gyro sensor, and a geomagnetic sensor to the user's joints, etc., and reading the user's movement, rotation, and / or direction. However, the optical, magnetic and inertial motion capture devices described above are merely examples of how the motion capture device (30) according to some embodiments of the present invention can be implemented, and the embodiments are not limited thereto. For example, the motion capture device (30) may generate user motion data (MD) by using image processing technology using artificial intelligence, or in parallel with an optical, magnetic and / or inertial motion capture device. In the following description, for convenience of explanation, it is assumed that the motion capture device (30) is an optical motion capture device including a marker and a camera.

[0112] Next, the skeleton generation module can generate skeleton data using user motion data obtained from markers attached to the user's joints, etc. (S210). At this time, the skeleton data can include data regarding the user's skeleton size, skeleton position, skeleton direction, etc.

[0113] According to some embodiments, the character implementation module (200) may receive tracking data (TD) from the background generation module (100). The character implementation module (200) may modify skeletal data using the received tracking data (S220). That is, the modified skeletal data (SD) may be data obtained by modifying skeletal data using the tracking data (TD).

[0114] Next, the retargeting module (220) can receive the modified skeleton data (SD) from the skeleton generation module (210). The retargeting module can generate rigging data (RD) through a process of rigging the modified skeleton data (S230). At this time, the rigging process may refer to a process of applying movement to the character by applying a structure called a joint when implementing a 3D character. Specifically, it may refer to a process of connecting skeletons using the modified skeleton data (SD) and assigning weights to the direction and angle at which each skeleton can move so that the character can move naturally like a human. The retargeting module (220) can transmit the rigging data (RD) generated using the modified skeleton data (SD) to the rendering module (230).

[0115] The rendering module (230) can implement a character by performing a rendering task on the character's appearance based on the rigging data (RD) (S240). Specifically, the rendering process may include a step of determining the character's appearance and texture, such as the character's skin, clothes, and hair, a step of setting the lighting surrounding the character to provide brightness and shade, and a step of calculating and reproducing various visual elements, such as shadows. The rendering module (230) can generate a second image (V2) that includes the character's movement (S250). That is, the second image (V2) relates to a metaverse character image generated based on user motion data (MD), but may mean an image in which a part of the user's appearance that is obscured by an object (OB) in the metaverse character image is removed. The character implementation module (200) can transmit the second image (V2) to the image synthesis module (300).

[0116] An image generating device according to some embodiments of the present invention can naturally express interaction between a user and an object by removing an overlapping portion between the user and the object and generating a character.

[0117]

[0118] FIG. 11 is a drawing showing a process of generating a second image from user motion data transmitted by the motion capture device of FIG. 1.

[0119] Referring to FIGS. 1 and 11, a motion capture device (30) can capture a user (UM) and an object (OB) in a given space. Subsequently, the motion capture device (30) can transmit motion data (MD) from the captured image to a character implementation module (200). In addition, the character implementation module (200) can receive tracking data (TD) from a background generation module (100). The character implementation module (200) can generate a second image (V2) from the motion data (MD) using the tracking data (TD).

[0120] As described above, the second image (V2) may be an image containing the movement of a metaverse character. In this case, the metaverse character included in the second image (V2) may have the area where the user (UM) and the object (OB) interact removed. For an exemplary explanation of this, see FIG. 12.

[0121]

[0122] Figure 12 is a diagram illustrating a process of generating a second image when there is an overlapping portion between a user and an object.

[0123] Referring to FIGS. 1, 2, and 9 to 12, the character implementation module (200) omits common processes when there is no overlapping between the user (UM) and the object (OB), and describes processes with differences below. Specifically, the background data (BD) may include an image in which the user (UM) and the object (OB) directly or indirectly contact each other. That is, the background data (BD) may include an image in which the user (UM) and the object (OB) are in contact with each other or are in a straight line with the image capturing device (20) and interact with each other. In other words, the object (OB) may interact with the user (UN), and the interaction may be divided into contact interaction and non-contact interaction.

[0124] Contact interactions can involve direct physical contact, such as shaking hands, hugging, touching a body part, or bumping fists. Non-contact interactions, on the other hand, can involve communication without physical contact, such as greetings or using a device like a phone call or text message.

[0125] At this time, the background data (BD) may include a portion where the user (UM) and the object (OB) overlap each other on the screen. Specifically, when the user (UM) is farther away from the image capturing device (20) than the object (OB), the image capturing device (20), the object (OB), and the user (UM) may be sequentially positioned in a straight line, thereby obscuring a portion of the user (UM). Since a portion of the user (UM) is obscured by the object (OB), outline information (OI_O) excluding the portion of the user (UM) obscured by the object (OB) may be generated from the background data (BD).

[0126] In other words, the object grouping module (110) can generate outline information (OI_O) excluding the portion of the user (UM) that is obscured by the object (OB) from the background data (BD). Accordingly, the user tracking module (120) can generate user tracking data (TD_O) based on the outline information (OI_O) excluding the obscured portion.

[0127] On the other hand, if the user (UM) is closer to the imaging device (20) than the object (OB), the imaging device (20), the user (UM), and the object (OB) may be sequentially aligned in a straight line, thereby obscuring a portion of the shape of the object (OB). In this case, since the background data (BD) includes the entire shape of the user, the outline information (OI_O) may include the shape of the user (UM) including the portion where the user (UM) and the object (OB) overlap.

[0128] The user tracking module (120) can transmit user tracking data (TD_O) to the character implementation module (200). At this time, the user tracking data (TD_O) can include a portion where the user (UM) and the object (OB) overlap. The character implementation module (200) can generate a second image (V2) using the user tracking data (TD_O) and user motion data (MD_O).

[0129] Specifically, the skeleton generation module (210) can extract a portion of the user's (UM) appearance that is obscured by an object (OB) from the skeleton data using the tracking data (TD_O). Thereafter, the skeleton generation module (210) can delete the portion extracted from the skeleton data to generate modified skeleton data (SD). That is, the skeleton generation module (210) can precisely express a portion of the user's (UM) appearance that is obscured by an object (OB) by reflecting the tracking data (TD_O) into the skeleton data.

[0130] Next, the skeleton generation module (210) can transmit modified skeleton data (SD) generated using the user motion data (MD_O) and tracking data (TD_O) to the retarget module (220). Consequently, the modified skeleton data (SD) transmitted from the skeleton generation module (210) to the retarget module (220) may be data in which data regarding a portion of the user's skeleton that is obscured by an object (OB) is deleted.

[0131] In some embodiments, when the user (UM) is further away from the image capturing device (20) than the object (OB), the tracking data (TD_O) may omit a portion of the user's (OB's) shape in the overlapping portion between the user (UM) and the object (OB). In this case, the character implementation module (200) generates the metaverse character by omitting the portion where the user (UM) and the object (OB) overlap. For example, when generating the final video (VF) including a screen where the metaverse character and the object (OB) shake hands, the object's (OB's) hand may be placed on top of the user's (UM's) hand. In this case, the image capturing device (20) may be positioned closer to the object's (OB's) hand than to the user's (UM's) hand in a straight line. Therefore, the image capturing device (20) may capture a portion of the user's (UM's) hand by omitting it.

[0132] Tracking data (TD_O) may reflect a portion of the hand of an omitted user (UM). The character implementation module (200) may generate a second image (V2_O) by omitting a portion of the hand of the metaverse character as much as the hand of the object (OB) based on the tracking data (TD_O).

[0133] The video generation device and method of the present invention can maximize the reality of a character by reflecting the motion of a user performing an act and applying it to the movement and directing effects of a metaverse character.

[0134]

[0135] FIG. 13 and FIG. 14 are diagrams for explaining a process of generating a final image through an image synthesis module according to some embodiments of the present invention.

[0136] Referring to FIGS. 13 and 14, the image synthesis module (300) can generate a final image (VF) when the user and the object do not overlap. The image synthesis module (300) can generate the final image (VF) using at least a portion of the first image (V1) generated by the background generation module (100) and the second image (V2) generated by the character implementation module (200). Specifically, the image synthesis module (300) can receive the first image (V1) and the second image (V2) (S300). The image synthesis module (400) can generate the final image (VF) by synthesizing the second image (V2) based on the first image (V1) as a background from which the user has been removed (S310, S320). For example, the image synthesis module (400) can place the second image (V2) on the first image (V1) with the first image (V1) as the bottommost layer.

[0137] According to some embodiments, the image synthesis module (300) can generate an image including a metaverse character on behalf of a user by synthesizing a second image (V2) based on a first image (V1).

[0138] At this time, the user (UM) may exist within the first image (V1), and the metaverse character may exist within the second image (V2). Furthermore, the metaverse character within the second image (V2) may be generated by correcting the object (OB). In other words, the final image (VF) may be generated based on the captured images of the user (UM) and the object (OB).

[0139] The video generation device and method of the present invention can simultaneously capture images of a user (UM) and an object (OB) in the same space. That is, the present invention can capture images of a person playing a metaverse character and an object interacting with the person in a single take. In other words, the present invention can capture images of a user (UM) playing a metaverse character and an object (OB) interacting with the person in separate spaces and times, and generate a final image (VF) without compositing them. Accordingly, the present invention can secure economic feasibility by saving the time and cost required to capture images in separate spaces and times.

[0140]

[0141] FIG. 15 is a diagram illustrating a process of generating a final image through an image synthesis module according to some other embodiments of the present invention.

[0142] Referring to FIGS. 1 and 15, the image synthesis module (300) can generate a final image (VF_O) when a user and an object overlap.

[0143] The image synthesis module (300) can generate a final image (VF_O) by using at least a portion of the first image (V1_O) generated by the background generation module (100) and the second image (V2_O) generated by the character implementation module (200). Hereinafter, common content with Fig. 14-1 will be omitted, and the process of generating the final image (VF_O) will be described.

[0144] In some embodiments, when there is an overlapping portion between a user and an object, a first image (V1_O) can be generated by reflecting the positional relationship between the user and the object. This allows for the generation of an image from which the user has been removed. In this case, the background data (BD) may include a portion where the object's appearance overlaps the user's appearance. Additionally, the tracking data (TD) may omit the user's appearance corresponding to the portion where the user and the object overlap. In other words, the tracking data (TD) may omit a portion of the user's shape.

[0145] The character implementation module (200) can receive tracking data (TD) from the background generation module (100). The character implementation module (200) can generate a second image (V2_O) including a metaverse character. Specifically, the second image (V2_O) can be a metaverse character image from which the overlapping portion between the user and the object is removed. In this case, when the second image (V2_O) is synthesized based on the first image (V1_O), an image in which the metaverse character and the object naturally interact with each other can be generated. In other words, the final image (VF_O) can be an image in which the metaverse character and the object interact with each other, rather than the user, from the beginning.

[0146] Accordingly, the image synthesis module (300) can generate an image in which a metaverse character and an object interact by synthesizing a second image (V2_O) based on a first image (V1_O) to generate a final image (VF_O).

[0147] In addition, the image generation device and method of the present invention can express objects other than the user without a sense of incongruity with the character even after the user has been converted into a character.

[0148] In addition, the image generation device and method according to some embodiments of the present invention can produce an image with maximized quality by naturally expressing the overlapping portion of a user and an object other than the user.

[0149]

[0150] The above description is merely an example of the technical idea of ​​the present embodiment, and those skilled in the art to which the present embodiment pertains may make various modifications and variations without departing from the essential characteristics of the present embodiment. Therefore, the present embodiment is not intended to limit the technical idea of ​​the present embodiment, but to explain it, and the scope of the technical idea of ​​the present embodiment is not limited by these embodiments. The protection scope of the present embodiment should be interpreted by the following claims, and all technical ideas within a scope equivalent thereto should be interpreted as being included in the scope of the rights of the present embodiment.

Claims

1. A background generation module that generates a first image from which the user's appearance is removed using background data including the user's appearance; A character implementation module that generates a second image including the motion of a metaverse character that matches the motion of the user by using user motion data related to the motion of the user; and An image synthesis module is included that generates a final image using the first image and the second image. The above background data includes the outline of the user and the appearance of the object interacting with the user. Image generating device.

2. In paragraph 1, The above background generation module, An object grouping module that generates outline information of the user from the background data received from the video capturing device, A user tracking module that generates tracking data that tracks the user based on the above outline information, Including a user removal module that generates the first image using the above tracking data, Image generating device.

3. In paragraph 2, The above user removal module, Using the above tracking data, masking data corresponding to the outline information is generated, Using the above masking data, a user-removed image is generated by removing the outline information from the background data, Generating the first image using the above background data and the user-removed image, Image generating device.

4. In paragraph 2, The above user removal module, Using the above tracking data, masking data corresponding to the outline information is generated, Using the above background data, the masking data is corrected to generate a user-removed image, Generating the first image using the above background data and the user-removed image, Image generating device.

5. In paragraph 2, The above object grouping module, Group pixels based on predefined criteria centered on each of the multiple markers attached to the user, Generate the outline information of the user by connecting the groups of pixels mentioned above. Image generating device.

6. In paragraph 1, The above character implementation module further uses the tracking data generated by the above background generation module to generate the second image, The above background data includes at least one portion where the appearance of the object overlaps the appearance of the user, The second image includes an image of the metaverse character with the overlapping portion removed. Image generating device.

7. A method for generating an image using an image generating device that generates a metaverse character image using the user's movements. A step of generating user outline information using background data including the user's appearance and the appearance of an object interacting with the user; A step of generating tracking data that tracks a user based on the above outline information; A step of generating a first image using the above tracking data; A step of generating a second image including a motion of a metaverse character matching the motion of the user by using user motion data related to the motion of the user; and A step of generating a final image using the first image and the second image, Video creation stage.

8. In paragraph 7, The step of generating the first image is: A step of generating masking data corresponding to the outline information using the above tracking data, A step of generating a user-removed image by removing the outline information from the background data using the above masking data, A step of generating the first image using the background data and the user-removed image, Video creation stage.

9. In paragraph 7, The step of generating the first image is: A step of generating masking data corresponding to the outline information using the above tracking data, A step of generating a user-removed image by correcting the masking data using the above background data, A step of generating the first image using the background data and the user-removed image, Video creation stage.

10. In paragraph 9, The step of generating the second image is: A step of generating the metaverse character using the above tracking data, How to create a video.

Citation Information

Patent Citations

  • Augmented reality apparatus and method for supporting interactive mode

    KR1020110084748A

  • Live social media system for using virtual human awareness and real-time synthesis technology, server for augmented synthesis

    KR1020180080783A

  • Method, apparatus and terminal for providing sign language video reflecting appearance of conversation partner

    KR102098734B1

  • Content preparation systems and methods for interactive video systems

    US20130236160A1

  • KR20210153472A