Image generation device, image generation system, and image generation method

The image generation device uses a scanner to project real-space objects into a virtual space with a 2D mask for occlusion, addressing the cost and equipment limitations of existing systems, enabling low-cost, real-time composite image generation with accurate object occlusion.

JP7745726B1Active Publication Date: 2025-09-29COVER CORP
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
JP2024171042
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Filing Date
2024-09-30
Publication Date
2025-09-29
Estimated Expiration
2044-09-30

AI Technical Summary

Technical Problem

Existing systems for creating composite images of 3D computer graphics (3DCG) and real-world footage in AR live events are costly, require expensive equipment, and are limited by camera manufacturers, making real-time generation impractical due to equipment restrictions and complex object occlusion challenges.

Method used

An image generation device and method that uses a scanner to identify real-space objects, projects them into a virtual space, and applies a 2D mask to occlude objects, allowing for real-time composite image generation without expensive equipment, using a compositing device to combine real-space images with 3DCG.

Benefits of technology

Enables low-cost, real-time generation of composite images with accurate object occlusion, reducing rendering load and equipment costs, and allowing for flexible adaptation to any video production site.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007745726000001_ABST
    Figure 0007745726000001_ABST
Patent Text Reader

Abstract

Provided are an image generation device, an image generation system, and an image generation method that enable composite images of real images and 3DCG to be generated at low cost and in real time. [Solution] An image of the real space is acquired from a camera that captures the real space, the motion of an actor playing an avatar is acquired, an avatar moving according to the acquired motion and a mask corresponding to a predetermined object generated based on the acquired image of the real space are arranged, a virtual space is generated in which an image of the predetermined object in the image of the real space can be displayed corresponding to the mask, and an image is output in which the image of the real space is synthesized with an image from a virtual camera that captures the virtual space, and the avatar is occluding the predetermined object arranged in the real space.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present invention relates to an image generation device, an image generation system, and an image generation method. [Background technology]

[0002] There is a technology called AR (Augmented Reality) that augments the real world. In recent years, virtual live performances (so-called AR live performances) using such augmented reality have been held, and there is a technology that generates a video in which a CG avatar performs a performance such as singing on a stage in real space (for example, see Non-Patent Document 1 and Patent Document 1). [Prior art documents] [Non-patent literature]

[0003] [Non-Patent Document 1] “[AR x Live Music] New Live Performances Expanded with AR,” [online], [Retrieved July 11, 2024], Internet<https: / / webar-lab.palanar.com / example / ar-music-live / > [Patent documents]

[0004] [Patent Document 1] Japanese Patent Publication No. 2023-130363 Summary of the Invention [Problem to be solved by the invention]

[0005] Traditionally, in video productions that combine real-world footage with 3D computer graphics (3DCG), the composite footage and 3D CG must be edited after filming to create a highly accurate composite image (a process known as post-production), resulting in a long wait time before the final composite image is created. Furthermore, in AR live events like those described above, it is necessary to reflect the acquired motion capture data of performers in real time in the movements of avatars composited onto the real-world stage. For this reason, AR live events require dedicated systems that generate images that combine 3D CG with real-world footage with the greatest possible precision. However, these systems require very expensive equipment and may be subject to limitations, such as the manufacturer of the camera used to capture the real-world footage used for compositing. Therefore, creating highly accurate composite images of 3D CG and real-world space is impractical due to cost and equipment limitations.

[0006] The present invention was devised in light of the above situation, and provides an image generation device, an image generation system, and an image generation method that enable composite images of real-world images and 3DCG to be generated at low cost and in real time. [Means for solving the problem]

[0007] (1) An image generating device according to an aspect of the present invention is an image generating device (e.g., a compositing device 2) that generates an image in which an avatar (e.g., an avatar C) is occluded by a predetermined object (e.g., a real microphone O) placed in real space, a means for acquiring an image of the real space from a camera (e.g., a real camera RC) that captures the real space (e.g., an input unit of the composite device 2 that acquires the image of the real camera RC); a means for acquiring the motion of the actor playing the avatar (for example, an input unit of the compositing device 2 for acquiring the actor's motion data from the capture device 5); a means for generating a virtual space (for example, rendering an image captured by a virtual camera VC in a scene construction unit 21), in which an avatar moving according to the acquired motion and a mask corresponding to the predetermined object (for example, a 2D mask OM corresponding to a real microphone O set on a polygon mesh BP) generated based on the acquired image of the real space are arranged, and an image of the predetermined object in the image of the real space can be displayed corresponding to the mask (for example, a state in which an image of the real microphone O is displayed by projection mapping an image of the real space onto a polygon mesh BP corresponding to the 2D mask OM, see steps S205 and S206 in FIG. 3 and the modified example (regarding the projection of an image of a real camera RC)); The system is provided with a means for outputting an image obtained by combining an image of the real space with an image from a virtual camera that captures the virtual space (for example, the image of the real camera RC and the image of the virtual camera VC are composited and output by the scene construction unit 21, step S208 in Figure 3).

[0008] With this configuration, a mask corresponding to a specified object is placed in a virtual space, and an image of the specified object can be displayed corresponding to the mask, making it possible to output a composite image of real-world images and 3DCG on the spot in real time at low cost.

[0009] (2) In the above (1), the image generating device a means for acquiring information (e.g., an input unit of a compositing device 2 that acquires the information identified by the scanner S) from a three-dimensional information measuring device (e.g., a scanner S having a scanner function) that measures three-dimensional information of a real space, including identification information that can identify the position of the predetermined object in the real space and the position where the avatar is to be placed in the virtual space corresponding to the real space (e.g., coordinates of a real microphone O identified by the scanner S and coordinates of a standing position marker M that is the standing position of the avatar, steps S101, S103, and S104 in FIG. 2); The generating means generates a virtual space in which the mask is placed at a position in the virtual space corresponding to the position of the predetermined object, based on the identification information, and the avatar is placed at a position where the avatar should be placed (for example, in the scene construction unit 21, a mask OM is placed in the virtual space so as to correspond to the coordinates of a real microphone O, and an avatar C is placed so as to correspond to the coordinates of a standing position marker M; steps S201 and S203 to S205 in FIG. 3).

[0010] With this configuration, it is possible to identify the positional relationship between an object placed in real space and an avatar, making it possible to express occlusion and the like based on information about the context.

[0011] (3) In (2) above, the generating means generates a virtual space in which an object (e.g., a polygon mesh BP placed in the virtual space so as to correspond to the position and range of a real microphone O in the real space) is placed at a position in the virtual space corresponding to the position and range of the specified object based on the identification information, and in which the mask is placed in association with the object (e.g., a mask OM is set as material information of the polygon mesh BP, step S205 in Figure 3).

[0012] With this configuration, the processing load of rendering is reduced, making it possible to generate an image in which a predetermined object in real space and an avatar are occluded on the spot in real time.

[0013] (4) In the above (3), an image of the real space is projected onto the object (for example, an image of a real camera RC is associated with a polygon mesh BP and projection mapped (an image of the real camera is projected onto the polygon mesh) (step S205 in FIG. 3 )), The generating means comprises: The device is provided with a means (e.g., the right diagram in Figure 7(B), steps S206 and S208 in Figure 3) for executing a process in which, when the mask is in a first state (e.g., the mask OM is in an on state), the object and the image of the real space are displayed in an area cut out in the shape of the mask corresponding to the specified object.

[0014] With this configuration, an object corresponding to the position and range of a specified object and the projected image of real space are displayed cut out in the shape of a mask, making it possible to generate an image in which a specified object in real space and an avatar are occluded.

[0015] (5) In (3) above, the object is an object placed based on the identification information measured from a position-identifying component (e.g., board B attached to real microphone O) attached to the specified object in the real space (e.g., an object corresponding to the shape and coordinates of board B detected based on information measured by scanner S (e.g., point cloud data, mesh data, etc.), step S103 in Figure 2, step S203 in Figure 3).

[0016] With this configuration, clues for identifying the position and range of a specific object can be obtained easily and quickly, enabling low-cost and rapid progress in video production.

[0017] (6) In (5) above, the position identification member has a shape corresponding to the range of the specified object perpendicular to the shooting direction when the real space is photographed from a specified position by a camera that photographs the real space (for example, a plate-shaped board B corresponding to the range of the real microphone O photographed from the real camera RC).

[0018] With this configuration, clues for identifying the position and range of a specific object can be obtained easily and quickly, enabling low-cost and rapid progress in video production.

[0019] (7) In the above (6), the three-dimensional information measuring device (for example, a scanner S equipped with a camera position transmission function) identifies information on the situation of the camera capturing the real space (for example, information on the coordinates, posture, etc. of the real camera RC, step S102 in FIG. 2), The means for generating the virtual space sets and places the virtual camera so as to match the situation of the camera based on the information on the situation of the camera acquired from the three-dimensional information measuring device (for example, based on the information on the real camera RC acquired by the scanner S, the virtual camera VC is set to the same setting as the real camera RC, step S202 in FIG. 3 ); The object is positioned so that the distance from the virtual camera to the object in the virtual space corresponds to the distance from the specified position to the specified object in the real space (see, for example, Figures 5 and 6).

[0020] With this configuration, it is possible to create a realistic composite image by ensuring consistency between the image from the camera capturing the real space and the image from the virtual camera capturing the virtual space.

[0021] (8) In the above (1), a first camera that captures an image of the predetermined object from a first direction and a second camera that captures an image of the predetermined object from a second direction are arranged in the real space (for example, real cameras RC1 to RC4 illustrated in FIG. 6(A)), In the virtual space, a first virtual camera is arranged that captures a mask (for example, a mask generated to be paired with a camera or a mask shared by multiple cameras) for displaying an image of the predetermined object from a direction corresponding to the first direction, and a second virtual camera is arranged that captures a mask (for example, a mask generated to be paired with a camera or a mask shared by multiple cameras, see the modified example (regarding a mask that projects an image)) for displaying an image of the predetermined object from a direction corresponding to the second direction (for example, virtual cameras VC1 to VC4 illustrated in FIG. 6(B)), The output means outputs a first image obtained by combining an image of the virtual space captured by the first virtual camera with an image of the real space captured by the first camera, and a second image obtained by combining an image of the virtual space captured by the second virtual camera with an image of the real space captured by the second camera (for example, output screen 4 in Figures 10(A) to (C)).

[0022] According to this configuration, an image of a predetermined object photographed from a plurality of angles is generated, so that a viewer of the composite image can feel a sense of reality from the composite image.

[0023] (9) In the above (1), the mask has a shape corresponding to the shape of the predetermined object (for example, a mask having a shape cut out to include the fine wiring of a real microphone O or the mesh of a pop filter).

[0024] With this configuration, since the mask has a shape corresponding to the predetermined object, it is possible to accurately occlude the avatar even for delicate parts of the predetermined object. Furthermore, since the rendering process burden is lighter than occluding the avatar with a 3D object corresponding to the three-dimensional shape of the predetermined object, it is possible to generate a composite image on the spot in real time at low cost.

[0025] (10) In (9) above, the mask has a shape based on the difference between an image taken from a predetermined position in the real space by a camera that photographs the real space, the image being taken before the predetermined object is placed in the real space, and the image being taken after the predetermined object is placed in the real space.

[0026] With this configuration, it is possible to quickly generate a mask for generating an image in which a delicately shaped portion of a predetermined object and an avatar are occluded at low cost.

[0027] (11) In the above (1), the image of the predetermined object displayed corresponding to the mask is an image of a first object captured in real time by a camera capturing images of the real space during a predetermined period of time; The video of the avatar is a video of the avatar moving based on the motion of the performer captured in real time during the same period as the predetermined period.

[0028] With this configuration, the real space and the actor's movements are synthesized in real time on the spot, which further enhances the sense of reality that viewers feel from the synthesized video.

[0029] (12) An image generation system according to an aspect of the present invention is an image generation device (e.g., image generation system 1) that generates an image in which an avatar (e.g., avatar C) is occluded by a predetermined object (e.g., real microphone O) placed in real space, A means for acquiring an image of the real space (for example, a real camera RC, a composite device 2), a means for acquiring the motion of the actor playing the avatar (e.g., a capture acquisition device 5, a compositing device 2); a means for generating a virtual space (for example, rendering an image captured by a virtual camera VC in a scene construction unit 21), in which an avatar moving according to the acquired motion and a mask corresponding to the predetermined object (for example, a 2D mask OM corresponding to a real microphone O set on a polygon mesh BP) generated based on the acquired image of the real space are arranged, and an image of the predetermined object in the image of the real space can be displayed corresponding to the mask (for example, a state in which an image of the real microphone O is displayed by projection mapping an image of the real space onto a polygon mesh BP corresponding to the 2D mask OM, see steps S205 and S206 in FIG. 3 and the modified example (regarding the projection of an image of a real camera RC)); The system is provided with a means for outputting an image obtained by combining an image of the real space with an image from a virtual camera that captures the virtual space (for example, a compositing device 2, a switcher 3, an output screen 4, etc., in which the image of the real camera RC and the image of the virtual camera VC are composited (synthesized) and output in the scene construction unit 21, step S208 in FIG. 3).

[0030] With this configuration, a mask corresponding to a specified object is placed in a virtual space, and an image of the specified object can be displayed corresponding to the mask, making it possible to output a composite image of real-world images and 3DCG on the spot at low cost.

[0031] (13) According to an aspect of the present invention, there is provided an image generation method (e.g., a compositing device 2) for generating an image in which an avatar (e.g., an avatar C) occludes a predetermined object (e.g., a real microphone O) arranged in a real space, the image generation method comprising: A step of acquiring an image of the real space from a camera (e.g., a real camera RC) that captures the real space (e.g., acquiring an image of the real camera RC with a composite device 2); A step of acquiring the motion of the actor playing the avatar (for example, acquiring the actor's motion data from the capture device 5 using the compositing device 2); a step of generating a virtual space (for example, rendering an image captured by a virtual camera VC in the scene construction unit 21), in which an avatar moving according to the acquired motion and a mask corresponding to the predetermined object (for example, a 2D mask OM corresponding to a real microphone O set on a polygon mesh BP) generated based on the acquired image of the real space are arranged, and an image of the predetermined object in the image of the real space can be displayed corresponding to the mask (for example, a state in which an image of the real microphone O is displayed by projection mapping the image of the real space onto the polygon mesh BP corresponding to the 2D mask OM, see steps S205 and S206 in FIG. 3 and the modified example (regarding the projection of an image of a real camera RC)); The method includes a step of outputting an image obtained by combining an image of the real space with an image from a virtual camera that captures the virtual space (for example, the image of the real camera RC and the image of the virtual camera VC are composited (synthesized) and output by the scene construction unit 21, step S208 in Figure 3).

[0032] With this configuration, a mask corresponding to a specified object is placed in a virtual space, and an image of the specified object can be displayed corresponding to the mask, making it possible to output a composite image of real-world images and 3DCG on the spot at low cost. [Brief explanation of the drawings]

[0033] [Figure 1] 1 is a diagram illustrating an example of the configuration of an image generation system including an image generation device according to an embodiment of the present invention. [Figure 2] 10 is a flowchart showing an example of an operation procedure of the scanner for matching the real space with the virtual space. [Figure 3] 10 is a flowchart showing an example of a processing order of virtual production by the image generation device according to the present embodiment. [Figure 4] FIG. 10 is a diagram for explaining a method for matching a real space with a virtual space. [Figure 5] FIG. 10 is a diagram for explaining a method for matching a real space with a virtual space. [Figure 6] FIG. 2 is a diagram for explaining the correspondence between a camera in real space and a camera in virtual space. [Figure 7] 10A and 10B are diagrams for explaining a procedure for generating a mask corresponding to an object in real space. [Figure 8] 1 is a diagram for explaining the correspondence between an image captured by a camera in real space and an image captured by a virtual camera. [Figure 9] 10 is a diagram showing an example of data used for synthesis in the image generating device of the present embodiment. [Figure 10] 10 is an example of an image of a composite image generated by the image generation device of the present embodiment.

[0034] Hereinafter, an embodiment of a communication system according to the present invention will be described with reference to the drawings. Note that the present invention is not limited to the following examples, but is defined by the claims, and all modifications within the meaning and scope equivalent to the claims are intended to be included in the present invention. In the following description, the same elements in the description of the drawings will be given the same reference numerals, and redundant explanations will not be repeated.

[0035] The present invention relates to a method, device, and system for generating video by combining virtual three-dimensional computer graphics (3DCG) with video captured by a camera placed in real space. The present invention can be applied, for example, to generating video of a 3DCG avatar reflecting an actor's motions in real space (e.g., singing, talking, etc.), where the actor's motions, the actor's voice, and real-space video must be simultaneously captured and combined in real time. For example, suppose a video of a real-space studio with only a microphone and no one in front of it is captured, and a singing performance and audio of the actor in a separate space are captured at the same time as the video is being shot. In such a scenario, the present invention can be applied to combine the real-space video with the actor's performance in real time to generate a video of a virtual avatar visiting a real studio and performing a singing performance.

[0036] According to the present invention, by placing polygon mesh data that projects real-space objects into a virtual space and placing a 2D mask that enables the display of the projected real-space objects, it is possible to inexpensively generate a composite image in which a real-space object and an avatar are occluded (hereinafter, the occluding image will be referred to as a composite image) without the need for expensive equipment as in the past. This allows for real-time composition with minimal delay, even when real-space filming, actor motion data acquisition, and the reflection of the avatar's movements are performed simultaneously. Furthermore, since there are no restrictions on the types of cameras and lenses used to film the real space, it can be flexibly adapted to any video production site. In this invention, by bringing images of real-space objects into a virtual space, it is possible to create an image in which a 3DCG character appears in the real space.

[0037] Conventional techniques for occluding real-world objects with 3DCG to create composite images that make 3DCG characters appear to exist in real space include the following: Occlusion refers to a state in which a portion of an object behind an object is not rendered when it is occluded by a foreground object as seen from the camera. A typical technique involves editing in post-production. While conventional post-production techniques can create highly accurate composite images even when complexly shaped real-world objects are occluded by 3DCG, they require extensive editing time and advanced editing techniques. Therefore, it is not possible to composite 3DCG (complete composite images) in real time on the spot where real-world footage is shot. Furthermore, conventionally, if the 3DCG to be composited is an avatar that reflects the actor's motions, the CG composite was performed after the actual filming, allowing for reshooting of the motions (multiple takes) and time-consuming adjustments to ensure consistency with the real-world footage. However, when it is required to synthesize an avatar that reflects the motions of an actor who moves simultaneously with the timing of the real-world footage with real-world footage in real time, careful editing after the fact, as in the past, is not possible. Therefore, it is desirable to devise a way to ensure consistency (for example, geometric consistency) between the real and virtual spaces while still enabling real-time synthesis.

[0038] In recent years, for example, in large-scale AR live events, dedicated systems (camera tracking systems) have been used to generate composite images in real time, occluding a real-world stage and a 3DCG avatar. For example, a tracking camera (e.g., a camera equipped with optical sensors, acceleration sensors, gyro sensors, etc.) is attached to the filming camera, enabling compositing with CG almost simultaneously as the image is being shot. However, such dedicated systems require budgets in the hundreds of millions, resulting in enormous costs and making them impractical. Furthermore, each system is restricted by the manufacturer of the real camera (lens, etc.) used to film the real space. The lenses to be used must be determined in advance, and preparation often requires several months. For example, the real-world venue is scanned in advance, or the 3DCG is created in advance based on CAD drawings of the venue, so that the images captured with the lenses to be used and the images captured by the virtual camera are aligned. Therefore, even with very expensive dedicated systems, it is difficult to match the settings of any real camera (camera used for filming) with the settings of the virtual camera (CG camera) on the day of filming. In other words, there was no system that could flexibly accommodate all lenses.

[0039] Furthermore, if the real-world object that partially obscures the 3DCG has a complex shape, synthesis becomes even more difficult. For example, a professional stand microphone used by a professional singer on stage may have a complex shape, with intricate wiring and a pop filter or other fine mesh-like components attached. To achieve occlusion, the three-dimensional shape of the real-world object must first be scanned (measured in three dimensions), and then polygon mesh data corresponding to the object must be placed in virtual space based on the scanned information, creating a front-to-back relationship with the avatar.

[0040] One method for scanning three-dimensional objects is photogrammetry. This is a measurement method that identifies the shape by analyzing parallax information from 2D images taken of an object from multiple viewpoints. However, to accurately scan the detailed shapes of complex objects to the point where the boundary between the background image and the actual object is indistinguishable even in close-up images, a huge number of shots are required, and the scanning process alone can take several hours. Furthermore, depending on the environment, there are weather and time constraints, such as whether the weather is sunny or not, making scanning not an easy task.

[0041] Without photogrammetry, methods for scanning three-dimensional objects include depth cameras equipped with time-of-flight (ToF) sensors or stereo depth cameras. While there are applications available that allow users to easily scan objects using depth cameras or the aforementioned photogrammetry on smartphones, the accuracy is limited. Even with depth cameras that can simultaneously scan and capture images, existing products offer low scanning accuracy and resolution, making it difficult to produce high-quality video. For example, even when scanning real-world objects using conventional scanners, capturing close-up footage of complex shapes can result in images in which avatars appear to be embedded in parts of the object due to the low accuracy of the object's boundaries. This can result in viewers losing their sense of immersion in the synthesized video, resulting in a lack of realism.

[0042] Even if it were possible to quickly place a polygon mesh of a real-world object's three-dimensional shape in a virtual space, rendering the 3D data into a 2D image would be a heavy load if the mesh itself were complex. This is because the large number of polygon vertices required for conversion to 2D (rasterization) would require a large amount of calculations. For example, recreating a complex object's shape in great detail could result in 100 million polygons, an extremely large amount of data that would be difficult to render in real time. This would make it difficult to perform virtual production simultaneously with real-world filming, resulting in timing issues between the actor's motions, the real-world filming, and the actor's audio, which is being recorded at the same time. When filming in real time, this delay is something that must be avoided.

[0043] Furthermore, with the diversification of content in recent years, the number of virtual talents has been increasing. Given the rising popularity of virtual talents, it is desirable to eliminate the boundary between what can be achieved by ordinary, non-virtual humans and virtual talents in content production, thereby increasing the opportunities for them to flourish. However, when virtual talents distribute videos, etc., it is not realistic to spend the high costs of tens of millions to hundreds of millions of yen described above. Generally, a method is used in which an avatar's image, projected through a green screen, is overlaid on other real images, and this method does not allow for the generation of images that occlude real objects. Therefore, in order to increase the opportunities for virtual talents to flourish, it is desirable to find a way to blend virtual talents into the real world like real humans at low cost.

[0044] Therefore, in this invention, the number of polygons corresponding to complex-shaped objects that occlude an avatar is reduced as much as possible, and a two-dimensional (2D) mask cut out of the complex-shaped object, which enables identification of the front-to-back relationship with the avatar, is placed in a virtual space, and an image of the object can be displayed corresponding to the mask. This makes it possible to easily and quickly generate a composite image of a real-space image captured in real time with an avatar that reflects the real-time motion of an actor, in which the avatar occludes a part of the object placed in the real space, without using expensive equipment. Details are explained below.

[0045] 1 is a block diagram of the configuration of a video production system according to the present invention. In this embodiment, the video production system 1 includes a compositing device 2, a switcher 3, an output screen 4, a real camera RC, a scanner S, a capture device 5, and an audio capture device 6.

[0046] In the video generation system 1 of the present invention, a compositor 2 is arranged on a video line that outputs video captured by a real camera RC placed in real space to an output screen 4 such as a monitor. Video data captured by the real camera RC and motion data acquired from a capture device 5 that acquires the actor's body tracking, facial motion, and other motions are input to the compositor 2. A composite video of the video data from the real camera RC and an avatar that reflects the actor's motion is output from the compositor 2 to a switcher 3. This causes the composite video to be displayed on the output screen 4. For example, the video is as shown in FIG. 10, and details will be described later.

[0047] The compositing device 2 is a computer (such as a personal computer) equipped with a processor such as a CPU (Central Processing Unit) or a GPU (Graphics Processing Unit), a storage unit such as a RAM (Random Access Memory), a ROM (Read Only Memory), a flash memory, or an HDD (Hard Disk Drive), an input / output interface, a display device (display unit) for displaying composited images, an operation device, etc. The storage unit stores a program executed as a rendering engine and a program that enables the construction and editing of a virtual space (virtual three-dimensional space). The compositing device 2 uses a control unit (comprised of a CPU and memory such as RAM) to construct and edit a virtual space based on the stored programs and renders 3DCG and the like placed in the virtual space. For example, software that enables CG editing, video editing, and the like to realize virtual production is installed. Note that various processes that can be realized by the compositing device 2 (e.g., each step in FIG. 3, etc.) may be realized by one or more computers.

[0048] The compositing device 2 includes a scene construction unit 21 that has the function of constructing a virtual space. The scene construction unit 21 constructs a virtual space for each camera. FIG. 1 illustrates an example in which the compositing device 2 includes multiple scene construction units 21, one for each camera. However, a compositing device 2 may be provided for each camera (each camera may be provided with a corresponding PC). Various 3DCG objects and a virtual camera VC are arranged in the virtual space. The scene construction unit 21 combines virtual space data, which will become an image captured by the virtual camera VC (described later with reference to FIG. 9), with an image captured by a real camera RC. The image combined by the scene construction unit 21 is output to the switcher 3. The operator of the compositing device 2 adjusts the combined image to be output to the switcher 3 after combination, based on the image displayed on the display unit. Specifically, the operator sets masks to be placed in the virtual space in the scene construction unit 21 and fine-tunes the placement positions of objects in the virtual space and the virtual camera VC.

[0049] The switcher 3 receives the video data (video signal) and audio data (audio signal) output from the composite device 2. The switcher 3 enables the video to be displayed on an output screen 4 or the like by switching among video data shot from multiple angles.

[0050] The output screen 4 is a display monitor or the like, and is a display device on which the composited image is displayed. For example, at a video production site, a director or the like can give instructions on the shooting angle based on the image on the output screen 4. Note that audio data (or other music data) output from the switcher 3 (or a separate audio device such as a mixer) may be output as sound data such as an audio signal from an output device such as a speaker on the same terminal as the terminal having the output screen 4 or a separate speaker.

[0051] The data output destination of the switcher 3 may be a distribution PC (not shown), and the distribution PC may be connected to the network 2. The network 2 is, for example, the Internet, and is configured from access networks such as a LAN (Local Area Network), a WAN (Wide Area Network), a mobile communication network (e.g., 5G, a wireless network, etc.), a wired telephone network, a FTTH (Fiber To The Home), and a CATV (Cable Television) network. This makes it possible to perform live distribution (live streaming) of the composite video generated by the composite device 2 via the Internet.

[0052] The video production site using the video generation system 1 in this embodiment includes a studio of real space (hereinafter also referred to as a real studio) that serves as the background for synthesis, and a capture studio where motion capture is performed to obtain the motions of actors whose movements are reflected in avatars.

[0053] In the real studio, a real microphone O and a real camera RC are placed, which captures the real studio including the real microphone O. The real microphone O is, for example, a stand microphone, and has parts with complex wiring and delicate and complicated shapes such as a pop filter mesh.

[0054] The capture studio is equipped with a capture device 5 and an audio capture device 6 that capture the motions and audio of the performer. The performer performs actions within the capture studio, and the performer's body movements and facial expressions are captured (tracking input) by the sensors of the capture device 5.

[0055] The real camera RC is a camera for video shooting that does not have a scanner function. Multiple real cameras RC are placed in the real studio. The video lines of the real cameras RC are connected to the switcher 3. The above-mentioned composite device 2 is connected between the video lines.

[0056] The scanner S is a small terminal installed with an application that has the functions of a 3D scanner and the function of transmitting information acquired by the scanner S to another device. It transmits information on the three-dimensional shape of an object, such as 3D point cloud data measured by the scanner S and mesh data converted from the point cloud data, to the compositing device 2. The scanner S may be, for example, a computer such as a smartphone that has a function (e.g., a ToF (Time of Flight) sensor) that enables identification of the three-dimensional shape by measuring the distance to the object using light, a self-position estimation function, a communication function, etc. The measurement results, such as the coordinates of the real studio measured by the scanner S and the shape and coordinates of objects placed in the real studio, are transmitted to the compositing device 2.

[0057] The scanner S also has a function for transmitting information on the position and rotation of the scanner S to the composite device 2. The scanner S is positioned in front of the lens of the real camera RC, and the measured information on the position and rotation (attitude) of the scanner S is information that is approximately equivalent to the information on the position and rotation (attitude) of the real camera RC (i.e., approximately equivalent to the position and attitude information of the camera). The information acquired by the scanner S is transmitted from the scanner S to a terminal (relay terminal) on which an application that records and sets camera information is installed. Other lens information, such as physical parameters such as the focal length and sensor size of the virtual camera VC, is input by the operator into an application (also referred to as a camera setting app) that records and sets the camera information. The camera information (camera status), including the lens information received from the scanner S and set by input from the operator, is transmitted to the composite device 2.

[0058] In the composite device 2, based on the information sent from the camera setting application, camera information such as the position and rotation of the virtual camera VC placed in the virtual space formed by the scene construction unit 21 is set so as to be consistent with the position and orientation information and lens information of the real camera RC acquired by the scanner S. This makes it possible to measure the camera information immediately on the spot at the shooting location and reflect it in the virtual camera in the virtual space, even in a situation where the information on the lens of the real camera used for shooting cannot be obtained in advance.

[0059] The capture device 5 is placed in the capture studio. The capture device 5 in this embodiment is a device (tracking device or the like) that has the function of acquiring data from sensors or the like attached to the body of the performer, such as body tracking that captures the movements of the performer, and face tracking that captures the movements of the performer's facial expressions.

[0060] The audio capture device 6 is a microphone that captures the voice of the performer. The audio data captured by the audio capture device 6 is input to the switcher 3. The audio capture device 6 installed in the capture studio may be the same as the real microphone O. This allows the performer's movements in the capture studio to take into account the presence of the real microphone O, preventing the avatar's arm from accidentally penetrating the real microphone O. It also makes it possible to capture audio quality of the same quality as when audio is input to the real microphone O installed in a real studio. Note that an instrument may be played in the real studio, and performance data of the instrument may be transmitted to the switcher 3. An audio mixer for adjusting the sound may be connected between the microphone that inputs the performance data and the switcher. Even in such cases, the video generation method of the present invention makes it possible to generate video content in which the performer's performance and real live music are combined in real time.

[0061] Note that various hardware components having the functions of configuring the image generation system 1 may be communicatively connected via the network 2, and may be capable of transmitting and receiving information (data) in both directions.

[0062] <Synthetic video generation procedure> (Overview of the procedure) A method for generating a composite image of a real space image and a 3DCG avatar by the image generation system 1 of the present invention will be described in detail with reference to FIGS.

[0063] 2 is a flowchart explaining the procedure of the operation of the scanner S in real space. In this embodiment, the scanner S is used to measure information for matching the real space (real studio) that serves as the background of the composite video with the environment of the virtual space in the scene construction unit 21 where the avatar is placed. The measured information of the scanner S is sent to the compositing device 2.

[0064] FIG. 3 is a flowchart for explaining the procedure of virtual production executed within the compositing device 2 using information acquired from the scanner S.

[0065] (Shooting environment settings match) First, a procedure for matching the coordinates of the real studio with the coordinates of the virtual space constructed in the scene construction unit 21 of the compositing device 2 will be described.

[0066] In step S101 of FIG. 2, a real-space studio in which a first object is placed is scanned using a position marker of an avatar in the real space as the origin, and the scanned information is transmitted to the compositing device 2. FIG. 4 is a diagram illustrating the correspondence between the real space and the virtual space when matching the environments of the real space and the virtual space. FIG. 4(A) is a diagram illustrating a real studio. As illustrated in the left diagram of FIG. 4(A), a real microphone O is placed in the real studio, and a position marker M is attached to the floor as a marker for the position where an avatar would stand if they were standing in front of the microphone. The relationship between the height of the microphone and the position of the position marker M is adjusted taking into account the situation where the avatar were actually standing in the real space. For example, the height of the avatar is taken into consideration. In the real studio, a scanner S is used to scan the entire real studio using the position marker M as the origin. The origin is determined by, for example, specifying the origin (scan start position) using the scanner S (a marker such as a two-dimensional barcode may be placed at the origin and read). Information on the coordinates of the scanned real studio is transmitted to the compositing device 2.

[0067] Next, in step S201 of FIG. 3, upon receiving information on the coordinates of the studio scanned by the scanner S, the compositing device 2 executes a process of matching the standing position marker M of the avatar in the real space scanned by the scanner S with the origin of the coordinates of the virtual space. The origin specified by the scanner S becomes the display start position of the avatar C, and the avatar C is placed there. FIG. 4(B) is a diagram showing the virtual space constructed by the scene construction unit 21. Note that the marker VM is the origin in the virtual space and is a virtual representation to illustrate that it corresponds to the standing position marker M in the real studio; it is not actually displayed. As shown in the left diagram of FIG. 4(B), the avatar C, which is a 3DCG character, is placed at the position of the origin. Note that the board B and polygon mesh BP are components and objects for identifying the position of the real microphone O, and will be described in detail later.

[0068] Returning to FIG. 2, in step S102, the scanner S identifies camera information of the camera used to capture real space in the real studio and transmits it to the compositing device 2. Specifically, the camera information is transmitted to the compositing device 2 from a camera setting app (installed on the relay terminal) that receives the information from the scanner S. FIGS. 5 and 6 are diagrams for explaining the correspondence between real space and virtual space in explaining the method for identifying camera information. FIG. 5 is a view of the relationship between the real space and virtual space in FIG. 4, viewed from the back of the real microphone O and from the direction looking directly at the front of avatar C when avatar C is standing in front of the microphone. FIG. 6 is a view of the real microphone O and avatar C viewed from above in the relationship between the real space and virtual space in FIG. 4, illustrating an example in which multiple cameras are arranged.

[0069] FIG. 5(A) is an example of a real studio. As shown in FIG. 5(A), a scanner S is placed in front of the lens of a real camera RC that captures the real space. Then, information on the position and rotation of the scanner S is transferred to the compositing device 2 via a camera setting application. In addition, the operator also inputs physical parameters such as the sensor size and focal length of the lens of the real camera RC into the camera setting application and sends them to the compositing device 2.

[0070] In the composite device 2, in step S202 of FIG. 3, the virtual camera is set to have the same information (situation) as the camera in real space based on the camera information, etc., identified by the scanner S transmitted from the camera setting application. As a result, as shown in the example of the virtual space in FIG. 5(B), the position and orientation information of the virtual camera VC is set to be the same as that of the real camera RC. Also, as shown in FIG. 6, when multiple cameras are installed, the same operations as step S102 of FIG. 2 and step S202 of FIG. 3 are performed for each real camera RC. As a result, the camera settings (position, orientation, lens information, etc.) of the real cameras RC1 to RC4 and the virtual cameras VC1 to VC4 become consistent. Note that whenever the camera information changes, such as when the real camera RC is moved, step S102 may be performed again to reacquire the camera information, and the virtual camera VC may be set in step S202 based on the reacquired camera information.

[0071] (Identifying objects placed in real space) Next, a method for identifying an object placed in real space will be described. In this embodiment, the object placed in real space is a real microphone O, which is an object to be occluded by creating a front-to-back relationship with avatar C. In this embodiment, in order to quickly identify the position coordinates of the real microphone O and reflect them in the virtual space, a board B, which is a simple board, is used as a position identifying member for the real microphone O. In this embodiment, the board B is attached to the real microphone O, as exemplified in the right diagram of FIG. 4(A), FIG. 5(A), and FIG. 6(A).

[0072] In step S103 of FIG. 2, a position identification member newly placed in real space is scanned with the scanner S in order to identify a first object placed in real space. In step S104, 3D data (mesh data) identification information of the scanned position identification member is selected and sent to the compositing device 2, and the operation of the scanner S for setup is completed. In this embodiment, after scanning the entire real studio in step S101, a new board B is attached to the real microphone O. Only the periphery of the board B newly placed in the real studio is scanned with the scanner S. After the scan, the scanner S still contains scan data other than that of board B scanned in step S101, so the scanner S is set to output only the scan data (mesh data) of board B. The mesh data of board B is then transferred to the compositing device 2. Note that "scan data only" does not strictly mean only board B, but may be the general range of board B.

[0073] Next, in step S203 of FIG. 3, the compositing device 2 places a polygon mesh of the shape of the position-specifying component at coordinates corresponding to the coordinates of the position-specifying component based on the 3D data (mesh data) of the position-specifying component identified by the scanner S. For example, as illustrated in the right diagram of FIG. 4(B), FIG. 5(B), and FIG. 6(B), a polygon mesh BP of the board corresponding to the position and range of the board B in real space is placed. Note that the compositing device 2 may be configured in advance to automatically generate a polygon mesh based on the mesh data acquired from the scanner S, or the polygon mesh BP may be placed by the operator entering the mesh data input to the compositing device 2 into the scene construction unit 21. Strictly speaking, there may be a slight error in the coordinates between the board B and the polygon mesh BP, but the operator may be able to fine-tune this error. As such, board B has a simple shape that requires less time to scan than a microphone with a complex shape. Therefore, the position and range of the real microphone O can be quickly and roughly identified using a sensor installed on a smartphone or the like.

[0074] (Mask generation and image projection for a given object) Next, the mask of the avatar and the object to be occluded will be described. Because it is not efficient to scan the 3D data of the avatar and the object to be occluded, in this embodiment, a board B, which serves as a position identification member, is used to roughly identify the position of the real microphone O. Then, a two-dimensional mask and an image of the real camera RC are set for the polygon mesh BP in the virtual space, which is this identified data. The shape of the two-dimensional mask is cut out to the shape of the real microphone O, and when the cut-out real microphone O is turned on (also referred to as the first state), the image of the polygon mesh BP and the real camera RC is cut out to the shape of the mask. In this way, by projection mapping the image of the real camera RC onto the polygon mesh BP, it is possible to match the 3DCG space (virtual space) with the world of the image seen from the real camera.

[0075] Specifically, in step S204 of FIG. 3, a 2D mask cut out in the shape of a first object in real space is generated from video captured by a real camera corresponding to the virtual camera. For example, a video is first captured using real camera RC1 in the real studio without a real microphone O in the real studio. Then, a video is captured again with the real microphone O placed in the same position. A difference is taken between the video before and after placing the real microphone O, creating a black-and-white mask in the shape of the real microphone O. Note that the mask itself set in the virtual space is occlusion on / off switch information, so color information is not required; it is sufficient to have at least shape information. Note that the greater the difference in brightness between the background and the object, the more clearly the details of the shape can be captured. Therefore, when generating a mask by difference, it is preferable to use a white background and a black object, but this is not limited to this.

[0076] Furthermore, as shown in FIG. 6, when multiple cameras are used to capture images from multiple angles, a mask is generated based on the difference between the images captured from each of the multiple cameras' respective angles. In other words, a mask based on the difference is generated for each camera. Referring to FIG. 8, the correspondence between the image captured by a camera in real space and the image captured by a virtual camera will be described. FIG. 8(A) is an example of real space, and FIG. 8(B) is an example of virtual space corresponding to the real space. Camera images 22a and 22b in FIG. 8(B) are diagrams illustrating the relationship between the polygon mesh BP and the microphone image. Camera image 22a in FIG. 8(B) is an illustration of the image projected onto the polygon mesh BP as viewed from virtual camera VC1. Camera image 22d is an illustration of the image projected onto the polygon mesh BP as viewed from virtual camera VC4. Note that camera images 22a and 22d are diagrams for explaining the correspondence between each camera and polygon, and therefore avatar C is not shown. In addition, the polygon mesh BP in the range outside the real microphone O is also shown in the figure for the sake of explanation, but when it is actually rendered and composited into an image, the area other than that cut out by the mask will not appear in the final composite image.

[0077] For example, in the image captured from the direction of virtual camera VC1 in the virtual space of Figure 8(B), which corresponds to real camera RC1 in the real studio of Figure 8(A), the portion of the real microphone O in the image captured from real camera RC1 and projected onto polygon mesh BP is cut out using a mask OM in the shape of the real microphone O captured from real camera RC1, and the image is rendered as the image of virtual camera VC1. Similarly, in the case of real camera RC4, the portion of the real microphone O in the image captured from real camera RC4 and projected onto polygon mesh BP is cut out using a mask OM in the shape of the real microphone O captured from real camera RC4, and the image is rendered as the image of virtual camera VC4.

[0078] Next, in step S205 of FIG. 3, the 2D mask of the first object and the video data of the real camera are associated with the material information (e.g., material) of the polygon mesh of the position identification member as image data (e.g., texture) recorded in memory, and the video of the real camera is projected onto the polygon mesh. That is, the video to be projected is set as polygon mesh data, and a mask in the shape of the real object is set for the polygon mesh data as on / off information for depth information. For example, as shown in FIG. 7(A), mesh data of a board B attached to a real microphone O scanned in a real studio is placed as a polygon mesh BP in the virtual space shown in FIG. 7(B). Thereafter, as shown in the center diagram of FIG. 7(B), a mask OM in the shape of the real microphone O generated using the above-mentioned difference and the video data of the real camera RC are set as image data recorded in memory in the material information of the polygon mesh BP, and the mask OM and the video data of the real camera RC are associated with the polygon mesh BP, and the video of the real camera RC is projected onto the polygon mesh BP. For example, as shown in the center of Fig. 7(B), an image captured of the real studio in Fig. 7(C) is projected onto the entire polygon mesh BP in the left diagram of Fig. 7(B), and a mask OM is associated with the polygon mesh BP. As a result, since the polygon mesh BP corresponds to the coordinates of the real microphone O, the mask OM and the image data of the real camera RC are set to correspond to the coordinates of the real microphone O.

[0079] In step S206, the 2D mask associated with the polygon mesh is turned on, so that the polygon mesh BP within the range corresponding to the 2D mask is displayed. For example, by turning on the mask OM, the image captured by the virtual camera VC and rendered is displayed as an image cut out within the range of the mask OM, as shown in the image on the right of FIG. 7(B). Specifically, for the polygon mesh BP onto which the image of the real camera RC is projected, a rendering process is performed on only the pixels within the turned-on mask OM. At the same time as this rendering process is performed, the depth information of the polygon mesh BP (corresponding to board B) is recorded in the virtual space data system. As a result, the polygon mesh BP cut out by the mask OM is rendered as the image of the virtual camera VC (i.e., only the portion onto which the image of the real camera RC is projected is rendered), and as shown in the right of FIG. 7(B), it appears as if only the real microphone O has been extracted from the image and appears in the virtual space. In other words, the image of the real camera RC is displayed as cut out in the shape of the real microphone O. Furthermore, at the time of step S206 (when the mask OM is turned on), the depth information of the polygon mesh BP is set only for the portion of the image from the real camera RC that is carved out by the mask OM. As a result, for example, by precisely scanning the shape of an object, it is possible to significantly reduce the number of polygons to 20,000 to 30,000 polygons across the entire screen, even though the object has a complex shape with 100 million polygons in the virtual space and is made to appear in the virtual space by associating the depth-related information in 3DCG. This facilitates real-time rendering. Note that this number of polygons is merely an example of how significantly the number of polygons can be reduced, and does not limit the number of polygons for the object or the entire screen.

[0080] (Final synthesis with a moving avatar) In step S207, the body tracking and facial tracking motions of the performer moving in the capture studio are reflected in avatar C placed in the virtual space. When the performer performs a movement, for example, in the positional relationship in virtual space between avatar C and polygon mesh BP (and associated mask OM) described with reference to FIGS. 4 to 6 in this embodiment, if the performer raises his right hand, avatar C's right hand also rises, and the position of the avatar's right hand moves to the left of polygon mesh BP in FIG. 6(B) (the positive direction in the x direction (assuming the arrow direction is positive)). In this case, from the direction of virtual camera VC1, an image of avatar C's right hand moving to the back side of the image of real microphone O is captured, as shown in the right diagram of FIG. 7(B). Therefore, when rendered, avatar C's right hand, which is behind real microphone O, is not drawn in the composite image (for example, the back hand of avatar C, which is obscured by real microphone O in FIG. 9(A)). On the other hand, if the performer raises his left hand, avatar C's left hand also rises, and the position of the avatar's left hand moves to the right of polygon mesh BP in Figure 6(B) (in the negative x direction in the figure). When this happens, an image of avatar C's left hand moving in front of the image of real microphone O is captured from the direction of virtual camera VC1, as shown in the right diagram of Figure 7(B). Therefore, when rendered, the part of real microphone O behind avatar C's right hand is not drawn in the composite image (for example, the real microphone O that is occluded by avatar C's hand in front in Figure 9(A)). In this way, occlusion between real microphone O and avatar C is possible in the virtual space.

[0081] In step S208, an image of the virtual space in which the avatar is placed and the polygon mesh hollowed out within the range of the 2D mask of the first object is placed is rendered together with the image of the real camera projecting the image, and the image is composited with the image of the real camera in the real space. The composited image is output from the compositing device 2 to the switcher 3. The data output from the compositing device 2 to the switcher 3 is image data obtained by combining data such as that shown in FIG. 9. As shown in FIG. 9, in the scene construction unit 21 corresponding to each camera, the image data of the virtual camera VC and the image data of the corresponding real camera RC are both captured and composited in real time. The image data of the virtual camera VC includes at least an avatar object (avatar C) and a position identification member polygon mesh (polygon mesh BP). The position identification member polygon mesh itself is transparent, but the image of the real camera RC is associated with a mask OM, which is a real object 2D mask corresponding to the shape of the real microphone O from the angle of the corresponding real camera RC. Furthermore, when the mask OM is turned on, the depth information of the polygon mesh BP affects the image rendered within the range of the mask OM, and the image from the real camera RC projected onto the polygon mesh BP is drawn at the position of the depth information of the polygon mesh BP. The data of these avatars C and the virtual camera that rendered the polygon mesh BP is combined with the image from the real camera RC corresponding to the virtual camera, and the combined image is returned to the image line. In this way, by cutting out the polygon onto which the image is projected within the range of the 2D mask OM, the load of rasterizing 3D data into 2D data can be reduced.

[0082] Furthermore, as shown in FIG. 7 , the virtual space does not reflect the entire shape of the real microphone O, including the legs of the stand. In this embodiment, the legs of the stand of the real microphone O are photographed at a shooting angle that does not overlap with avatar C. Therefore, it is sufficient to place a position identification member (board B) in the area that needs to be occluded and generate a polygon mesh BP. To reduce the rendering load, it is not necessary to create a polygon mesh BP and a mask OM for the portion that does not need to be occluded. For the portion of the legs that is not displayed by the mask OM, in a composite image that shows the entire real microphone O, such as the composite image displayed on the output screen 4 of FIG. 10(A), the image displayed is not the image projected onto the mask OM but the image captured by the real camera RC itself. In other words, an image of a portion of an object placed in the real studio that may overlap with avatar C is displayed at a position corresponding to the portion of the object that may overlap with avatar C in the virtual space where avatar C is placed, and the rendered data is composited with the image captured by the real camera RC. As a result, the image of the real microphone O after synthesis is a combination of two parts: a part that is the image of the real camera RC itself in the background, and a part that is the image of the real camera RC displayed on the mask OM in the virtual space.

[0083] By using the above-described virtual production method, it is possible to generate an image in which avatar C occludes a real microphone O, which is a real object, in real time while moving as shown in Fig. 10(B), as in the example of a composite image displayed on output screen 4 in Fig. 10. Also, as shown in Fig. 10(C), even if a mesh-like pop filter O2 or the like is attached to the microphone, it is possible to generate an image in which avatar C is drawn behind the mesh.

[0084] Although an example of using multiple cameras for shooting has been given, the camera angle is most preferably one that captures the widest range of board B (polygon mesh BP). For example, an angle that captures from a direction perpendicular to the widest surface of board B, such as real camera RC2 and virtual camera VC2, is preferable. The more the shooting direction deviates from the widest surface of board B, i.e., the direction perpendicular to the surface of board B and polygon mesh BP shown in FIG. 4, the more unnatural the image becomes. Therefore, it is preferable that the shape of board B corresponds to the range of real microphone O perpendicular to the shooting direction from the camera. In this invention, the 2D image onto which the microphone image is projected is set only within the range of polygon mesh BP. This uses the principle of anamorphosis, a psychological effect that creates the illusion of a three-dimensional effect when viewed from a certain direction. Therefore, the psychological effect breaks down as the direction deviates from the desired certain direction. For example, if you photograph board B (polygon mesh BP) placed as shown in Figure 4 from an angle like the ones in Figures 5 and 6, the projection surface of the microphone image on the polygon mesh BP will not be visible on the camera in the captured image.

[0085] <Examples of specific configurations and effects>

[0086] (1) In the image generation method for generating an image composite image in which an avatar C reflecting the motion of a performer is occluded by a real microphone O, which is a real object placed in a real studio (real space), in the above-described embodiment, an image captured by a real camera RC that captures the real studio is input to a compositing device 2, and motion data obtained by tracking the body and face of a singing performance in a capture studio is input to the compositing device 2. In a scene construction unit 21 that constructs a virtual space within the compositing device 2, in steps S203 to S205 of Fig. 3, a 2D mask OM corresponding to the real microphone O and the image of the real camera RC are arranged in association with a polygon mesh BP, and the image of the real camera RC is projection mapped (projected) onto the polygon mesh BP. In step S206 of FIG. 3, when the 2D mask OM is turned on (first state), the real image is cut out in the shape of the real microphone O corresponding to the 2D mask OM, and the image of the real microphone O associated with the depth information of the polygon mesh BP is drawn (it can be displayed as an image captured by the virtual camera VC). In step S207, the actor's motion data is reflected in the avatar C, and in step S208, an image captured by the virtual camera VC and featuring the mask OM and avatar C is rendered. The rendered image from the virtual camera VC is then composited with the image from the real camera RC and output to the switcher 3 or the output screen 4. In this way, a mask corresponding to a predetermined object, such as a microphone, is placed in the virtual space, and the image of the predetermined object can be displayed corresponding to the mask, making it possible to output a composite image of real image and 3DCG on the spot in real time (instantaneously, instantly) at low cost. Furthermore, the time axes of the video in real space, the motion of the performer, and the audio data (for example, the audio of the performer singing) are synchronized at desirable times.

[0087] (2) In the above-described embodiment, in steps S101, S103, and S104 of FIG. 2, the coordinates of the real microphone O in the real studio and the coordinates of the standing position marker M (e.g., the scan start position) of the avatar, which are identified by the scanner S, a three-dimensional information measuring device having a scanner function for measuring three-dimensional information of the real space, are input to the compositing device 2. In the scene construction unit 21 of the compositing device 2, in step S201 of FIG. 3, an avatar C is placed in the virtual space so as to correspond to the coordinates of the standing position marker M, and in steps S203 to S205, a mask OM is placed in the virtual space so as to correspond to the coordinates of the real camera RC. This makes it possible to identify the positional relationship between an object placed in the real space and the avatar, thereby enabling the representation of occlusion and the like based on information about the context.

[0088] (3) In the above-described embodiment, in step S203 of Fig. 3, a polygon mesh BP is placed in virtual space so as to correspond to the position and range of the real microphone O in real space based on the information identified by the scanner S in steps S103 and S104 of Fig. 2, and in step S205, a mask OM is set in the material information of the polygon mesh BP. This reduces the processing load of CG rendering, making it possible to generate, on the spot, in real time, an image in which a specific object in real space and an avatar are occluded.

[0089] (4) In the above-described embodiment, in step S205 of Fig. 3, the image of the real camera RC is associated with the polygon mesh BP, and the image is projection-mapped onto the polygon mesh BP. Furthermore, in step S206 of Fig. 3, the mask OM is turned on (first state), and the portion of the real microphone O projected onto the polygon mesh BP that is cut out within the range of the mask OM is rendered and drawn as the image of the virtual camera VC in a state in which the depth information of the polygon mesh BP is associated with it. This allows an object corresponding to the position and range of a predetermined object and the projected real-space image to be displayed cut out in the shape of the mask, making it possible to generate an image in which a predetermined object in real space and an avatar are occluded.

[0090] (5) In the above-described embodiment, the polygon mesh BP is an object corresponding to the shape and coordinates of board B, which is placed in virtual space in step S203 of Fig. 3 so as to correspond to the position and range of real microphone O in real space, based on the scan data of board B attached to real microphone O measured by scanner S in steps S103 and S104 of Fig. 2. This makes it possible to easily and quickly obtain clues to identify the position and range of a specific object, enabling low-cost and rapid progress in video production.

[0091] (6) In the above-described embodiment, as illustrated in Figures 4, 6, etc., board B has a shape that corresponds to the range of real microphone O perpendicular to the shooting direction from a predetermined position when real camera RC is placed at the predetermined position and shooting is performed, and is, for example, a board with an area that roughly corresponds to the range of real microphone O perpendicular to the shooting direction of real camera RC2. This makes it possible to easily and quickly obtain clues to identify the position and range of a predetermined object, enabling low-cost and rapid progress in video production on-site.

[0092] (7) In the above-described embodiment, in step S102 of FIG. 2, the scanner S, which has a camera position transmission function such as coordinates and attitude, transmits information specifying the camera status (position and attitude information, etc.) of the real camera RC to the composite device 2 (for example, this may be via a camera setting app, etc.), as described with reference to FIG. 5, etc. In step S202 of FIG. 3, based on the information of the real camera RC acquired by the scanner S, the information (status) of the virtual camera VC, such as the lens and position, is set to the same environment as the setting of the real camera RC. Also, as illustrated in FIGS. 5 and 6, the distance from the virtual camera VC to the polygon mesh BP and the distance from the real camera RC to the real microphone O are approximately the same. This allows consistency between the image from the camera capturing the real space and the image from the virtual camera capturing the virtual space, making it possible to generate a realistic composite image.

[0093] (8) In the above-described embodiment, as illustrated in FIG. 6A, real cameras RC1 to RC4 are placed in a real studio to capture images of a real microphone O from different directions. Furthermore, virtual cameras VC1 to VC4 are placed in a virtual space to capture images of one type or multiple types of masks OM, each of which is generated for each image of the real microphone O captured in the image captured by each real camera RC, onto which the images of the real microphone O captured by each real camera RC are projected, from directions corresponding to the image capturing directions of the real cameras RC. For example, the image of the real camera RC1 is composited with the image of the corresponding virtual camera VC1 (note that the image of the real camera RC1 projected onto the polygon mesh BP is cut out within the range of the mask OM captured by the virtual camera VC1), and the image of the corresponding virtual camera VC2 is composited with the image of the real camera RC2 (note that the image of the real camera RC2 projected onto the polygon mesh BP is cut out within the range of the mask OM captured by the virtual camera VC2). The composite image is an image in which an object in real space, such as a real microphone O, and an avatar C are occluded, as exemplified by output screen 4 in Fig. 10. Furthermore, as exemplified in Fig. 7(B), if a mask OM is generated in the portion occluded by avatar C, for example, a portion that is not occluded, such as the feet of the real microphone O, may be reflected in the composite image as shown in Fig. 7(C) included in the first real camera captured data in Fig. 9 (e.g., Fig. 10(C)). In this way, an image is generated in which a specific object is captured from multiple angles, and by switching and combining the images from multiple angles, it is possible to psychologically affect the viewer of the composite image and make them feel a sense of reality about the composite image.

[0094] (9) In the above-described embodiment, the mask OM is a mask cut out in a shape corresponding to the delicate shape of the fine wiring of the real microphone O or the mesh of a pop filter. As a result, since the mask has a shape corresponding to a predetermined object, it is possible to perform occlusion with an avatar with high accuracy even for delicately shaped parts of the predetermined object. Furthermore, since the rendering process load is lighter than occluding an avatar with a 3D object corresponding to the three-dimensional shape of the predetermined object, it is possible to generate a composite image on the spot in real time at low cost.

[0095] (10) In the above embodiment, the mask OM is generated based on the difference between the image captured by the real camera RC before and after placing the real microphone O in the real studio. This makes it possible to quickly generate a mask for generating an image in which a delicately shaped portion of a predetermined object and an avatar are occluded at low cost.

[0096] (11) In the above-described embodiment, the image of the real microphone O displayed corresponding to the mask OM is an image captured by the real camera RC in real time at the same time as it is projected onto the polygon mesh BP, and the motion of the avatar C reflects the motion data of the performer, which is captured (tracked) in real time at the same time as the image is captured by the real camera RC. This allows the real space and the performer's movements to be simultaneously combined in real time, further enhancing the sense of realism felt by the viewer in the combined image. Furthermore, because the voice data emitted by the performer, the avatar's movements, and the real image are synchronized at the desired timing, stable images can be provided to the viewer not only in recordings but also in live streaming and other similar situations.

[0097] <Modification> Modifications to the above-described embodiment are listed below.

[0098] (About the camera settings app) In the above-described embodiment, an example has been described in which mesh data of real space or real objects scanned by the scanner S is directly transmitted to the compositing device 2 with reference to steps S101, S103, S104, etc. of FIG. 2 , and camera information data is transmitted to the compositing device 2 from a terminal (relay terminal) on which an application for recording and setting camera information is installed, based on the position and orientation information of the scanner S with reference to step S102, etc. of FIG. 2 . However, this is not limiting, and mesh data of objects generated based on measurements by the scanner S may also be transmitted to the compositing device 2 via the relay terminal. For example, a manager app (which may include a camera setting app) having a function for aggregating information acquired by the scanner S and transmitting it to the compositing device 2 may be incorporated into the relay terminal, enabling the transfer of data acquired by the scanner S.

[0099] Furthermore, the camera setting app may be implemented as an application that can be installed on the scanner S or another computer, and may be a program that runs on the scanner S. For example, if the same number of real cameras RC and scanners S are available, data may be transferred directly from the scanner S to the composite device 2. Alternatively, the program of the camera setting app may be executed on the same terminal as the composite device 2. For example, information on the scanner S may be sent directly to the composite device 2, and information on physical parameters such as the sensor size and focal length of the lens of the real camera RC may be input by the operator of the composite device 2 as settings for the virtual camera VC (for example, to a camera setting app that runs in the composite device 2 or to an application that configures the scene construction unit 21).

[0100] Furthermore, for example, if there is one scanner S and multiple real cameras RC, the camera information for all of the real cameras RC aggregated in the camera setting app may be transmitted to the composite device 2 via a separate device equipped with a camera setting app. For example, the camera information may be aggregated in the manager app of the relay device described above. This allows the manager app to synchronize data such as mesh data measured by the scanner S and camera information all at once, without requiring the user to operate the scanner S screen. Specifically, this reduces the workload for scanning and synchronizing and transmitting camera information. It also allows all camera information to be recorded simultaneously, and a synchronization mechanism can be used to align the time axis of all camera information. Furthermore, multiple devices with the camera setting app installed may be used, and data may be managed by grouping cameras. For example, a camera setting app A may manage a group of real cameras RC1 to RC3, and a camera setting app B may manage a group of real cameras RC4 to RC6. Furthermore, communication between the camera setting apps A and B may enable synchronization between different groups.

[0101] (Regarding avatar position, etc.) In the above-described embodiment, as explained in step S101 of FIG. 2 and step S201 of FIG. 3, an example has been described in which the real space is scanned using the standing position marker M in the real studio as the origin, and the origin is set as the display start position of avatar C. However, the present invention is not limited to this, and any other location than the standing position marker M of the avatar may be set as the origin as long as the coordinates of the real space and the coordinates of the virtual space can be matched. Furthermore, the display start position of the avatar may be set at a position other than the origin (for example, a predetermined distance from the origin). Note that, regarding the angle at which the avatar is photographed, an example of an angle that captures the side of the avatar has been shown in the above-described embodiment, but the avatar may also be photographed from the front. Furthermore, the number of avatars (number of performers) is not limited to one, and multiple avatars (each reflecting the movements of multiple corresponding performers) may be arranged.

[0102] (Regarding location-specific components) In the above-described embodiment, a single board B is used, and one type of polygon mesh BP is obtained by scanning the single board B. However, this is not limiting. Alternatively, boards B for multiple angles may be prepared, and polygon meshes corresponding to the boards B for each angle may be arranged in the virtual space, thereby expanding the range of shooting angles even further. However, it is preferable that the camera angle be fixed. If the camera angle is fixed, steps S103 and S104 in FIG. 2 and steps S203 to S205 in FIG. 3 may be performed using multiple boards B, as long as time allows for each shooting location. Even in this case, preparation work for real-time compositing can be performed on the day, at lower cost and more quickly than with conventional virtual production methods.

[0103] The position specifying member is not limited to a flat board, but may be any member that can be easily attached to the target object and whose shape can be easily scanned to provide a hint as to the coordinates of the target object. The number of occluding objects is not limited to one, and multiple objects may be targeted. For example, in addition to a real microphone, a music stand or the like may be placed in the real space, and polygon meshes and masks for the real microphone and the music stand may be generated and placed in the virtual space to occlude the avatar.

[0104] In the above-described embodiment, because the shape of the real microphone O is complex, board B, which serves as an easily measurable position identification component, is attached to the microphone, and the coordinates of the real microphone O are identified using scan data from board B. However, this is not limiting. The polygon mesh that sets the mask onto which the real image is projected (mask OM in the above-described embodiment) may be positioned based on scanned data of a specific object, such as a microphone, onto which the image is to be projected. For example, if the shape of the real-world object itself is simple compared to the complex mesh and wiring of a laptop PC or the like, a polygon mesh may be positioned in the virtual space based on mesh data obtained by scanning the object itself. Even in this case, by setting a 2D mask of the object's shape on the polygon mesh and projecting the real image onto it, virtual production can be achieved in real time with ease and beautiful outline detail.

[0105] (About the image-projecting mask) In the above-described embodiment, an example has been described in which the mask OM, which is a mask in the shape of the real microphone O, can be generated by taking the difference between images before and after placing the real microphone O. The generation of the mask in the shape of the real microphone O using this difference may be automatically generated by an application that automatically generates a mask from the difference between images before and after the placement of the real microphone O, or, for complex shapes, may be manually generated based on the difference. Furthermore, the mask is not limited to being generated using the difference; it may be generated by automatically or manually selecting the range of the shape of the real microphone O from an image in which the real microphone O, a real object, is placed in real space, and then generating a mask cut out in the shape of the real microphone O. From the standpoint of time efficiency, it is desirable to employ a method of generating a mask using the difference; however, the latter method, which allows for more detailed cutting depending on the fineness of the object's shape, may also be employed.

[0106] The generation of the difference mask may be performed at a different timing than the order shown in Fig. 3. It may be performed after the placement position of the real microphone O has been determined, and for example, the difference mask may be generated in advance by capturing an image of the background with no microphone and an image with the microphone present at a stage before the board B is attached to the microphone in step S103 of Fig. 2.

[0107] In the above-described embodiment, an example has been described in which the masks based on the differences generated in step S204 in Fig. 3 are created for the number of cameras, as exemplified in Fig. 6 and the like. However, this is not limiting, and instead of creating all the masks for the number of cameras, it is also possible to create only masks based on angles from some of the cameras. For example, one type of mask generated from an image from an angle (e.g., real camera RC2 in Fig. 6) that captures the widest range corresponding to board B among the angles at which real microphone O is captured may be copied and placed in all scene construction units 21.

[0108] (Regarding the projection of images from Real Camera RC) In the above-described embodiment, an example has been described in which an image of a real camera RC is set onto a polygon mesh BP and projection mapping (projecting a real image onto the polygon mesh) is performed, and when the mask OM is turned on (first state), an image corresponding to the portion of the real microphone O in the image of the real camera RC associated with the depth information of the polygon mesh so as to correspond to the range of the mask OM can be drawn (displayed) as an image of the virtual camera VC. However, without being limited to this, it is also possible to project the image of the real camera RC onto the mask OM associated with the polygon mesh BP, and when the mask OM is turned on, the image of the portion of the real microphone O projected onto the mask OM associated with the depth information of the polygon mesh BP within the range of the mask OM can be drawn as an image of the virtual camera VC.

[0109] (Scanning procedure) The order of the flowcharts described in the above-described embodiment and illustrated in FIGS. 2 and 3 is not limited to this, as long as the environments of the real space and the virtual space are consistent, a mask OM corresponding to the shape of a real object is generated, an image of the real space is projected onto the polygon mesh BP, and the mask OM can be used to carve out the object. For example, after the scanning of the position-identifying member and the transmission of the 3D data identification information to the compositing apparatus 2 in steps S103 and S104 of FIG. 2, the step of transmitting camera information to the compositing apparatus 2 in step S102 may be performed. Also, for example, when scanning the entire real studio in step S101, board B may be attached to the real microphone O in advance, and mesh data of the entire studio and mesh data of board B for setting the mask may be simultaneously transmitted to the compositing apparatus 2. Specifically, board B may be attached to the real microphone O in advance, and the mesh of board B may be marked as a specific mask, so that both the mesh of the entire real space (immovable data) and the mesh of the mask of board B (movable data that can be fine-tuned) may be simultaneously transmitted to the compositing apparatus 2.

[0110] That is, the order is not limited to the above as long as the information measured in Fig. 2 (mesh data obtained by scanning the entire real space, information about the origin in the real space, mesh data identifying a first object placed in the real space, and camera information for capturing the real space) is transmitted to the compositing device 2. Using the information acquired in Fig. 2, it is only necessary to align the origin in step S201 as described in Fig. 3, and then set the virtual camera in step S202 based on the information about the origin set in step S201. Preferably, in order to utilize the difference and reduce the scanning process, after scanning the studio, it is desirable to attach board B to real microphone O and perform scanning.

[0111] (About setting up real and virtual cameras) In the above-described embodiment, the real camera RC capturing images of the real space is a video camera without a sensor for identifying the camera's position and orientation. The scanner S measures the position and orientation of the real camera RC, and the virtual camera VC in the virtual space is set based on the information measured by the scanner S (for example, via a camera setting app with lens information input). However, this is not limiting; a capturing camera equipped with a camera tracking sensor may be used and synchronized with the virtual camera. This allows capturing images while moving the real camera and virtual camera without fixing the camera position. While the angle between the plane of the 2D mask in the virtual space (a line parallel to it) and the capturing direction is orthogonal (90 degrees) for maximum efficiency, it is not limited to this, and any angle may be used as long as it does not deform the object projected onto the polygon mesh BP. Appropriate adjustments may be made depending on the shape of the object to be projected. Furthermore, the projected image may be captured by a fixed camera separate from the camera capturing the image to be combined with the virtual camera image, or it may be captured by a moving real camera. Even if camera tracking is possible, the load of rendering virtual space data is reduced.

[0112] In the above-described embodiment, an example has been described in which the scanner S measures information such as the position and attitude of a real camera, and information on the camera status including information on the position of the camera and lens information input by an operator to a camera setting app or the like is set in the virtual camera, or an example in which a camera setting app is installed in the scanner S and information on the camera status including lens information is transmitted from the scanner S to the composite device 2. However, the present invention is not limited to this, and the scanner S may be fixed to a real camera for shooting, enabling camera tracking, and, as described above, the camera position may not be fixed and the real camera and virtual camera may be moved to shoot.

[0113] (Applicable scenes of this invention) In the above-described embodiment, an example of a scene in which avatar C, reflecting the motions of a performer, sings has been described as an example of a scene in which the virtual production method of the present invention is applied. However, the present invention is not limited to this example, and may also be applied to scenes such as a video introducing avatar C's favorite items or a product explanation by avatar C. Even if the shape of a favorite item or product, which is a real object, is not limited to a simple object but is a complex object, the virtual production method of the present invention can be used to generate a mask that projects an image of the object in a virtual space, thereby creating a posterior-posterior relationship with avatar C in the virtual space. This makes it possible to occlude avatar C with a real object at low cost and in real time, thereby generating an image that makes it appear as if avatar C is actually present in real space. In addition to the real microphone mentioned above, other examples of objects with complex shapes include, but are not limited to, objects that tend to have high polygonal shapes, such as ear monitors (in-ear monitors) and round objects such as stuffed toys, as well as objects such as glass bottles that allow the laser beam to pass through and make it impossible to scan their shape, and objects such as mirrors that reflect the laser beam and make scanning difficult.

[0114] [Software implementation example] The various control blocks of the control unit of the computer of various devices, such as the composite device 2, in the above-described embodiments may be realized by logic circuits (hardware) formed on an integrated circuit (IC chip) or by software using a CPU (Central Processing Unit). When realized by software using a CPU, the computer equipped with the control unit includes a CPU that executes instructions from a program, which is software that realizes each function; a ROM (Read Only Memory) or storage device (these are referred to as "recording media") in which the program and various data are recorded so as to be readable by the computer (or CPU); and a RAM (Random Access Memory) in which the program is expanded. The object of the present invention is achieved when the computer (or CPU) reads and executes the program from the recording media. The recording media can be "non-transitory tangible media," such as tapes, disks, cards, semiconductor memories, and programmable logic circuits. The program may also be supplied to the computer via any transmission medium capable of transmitting the program (such as a communication network or broadcast waves). Note that one aspect of the present invention can also be realized in the form of a data signal embedded in a carrier wave, in which the program is embodied by electronic transmission.

[0115] A specific example of implementation by a terminal (device) on which a computer program (application software) for executing the image generation method according to the above-described embodiment is installed will be described. The program executes processing for implementing all or part of the virtual production described in FIG. 3 based on information acquired from the real camera RC, scanner S, and capture acquisition device 5. For example, a program for executing the image generation method is installed in a computer corresponding to the compositing device 2. A control unit executing the program sets a display start position for avatar C (by, for example, matching it with the origin of the virtual space) based on coordinate information in real space acquired by the scanner S, and executes processing for placing avatar C at the display start position (corresponding to step S201 in FIG. 3). Furthermore, based on camera information identified based on, for example, self-position estimation of the scanner S, a virtual camera VC is set to match the state of the real camera RC. Furthermore, based on mesh data of a position-identifying member (board B) acquired from the scanner S, a polygon mesh BP corresponding to board B is placed in the virtual space. A process is performed in which a 2D mask OM in the shape of a real microphone O, which is material data acquired and prepared by, for example, cutting out the difference from the background and storing it in advance in a storage unit, or by automatically cutting out and generating it based on real video, is associated with a polygon mesh BP (corresponding to step S205). Image data acquired from a real camera RC is projected onto the polygon mesh BP (corresponding to step S205). When it is determined that the mask OM associated with the polygon mesh BP has been turned on, a process is performed in which the polygon mesh BP (the polygon mesh BP onto which the image of the real camera RC has been projection-mapped) within the range corresponding to the mask is rendered as an image of a virtual camera VC with depth information (corresponding to step S206). Furthermore, a process is performed in which an image of the virtual space VC in which avatar C reflects the motion data is rendered, and the image is composited with the image of the real camera RC to generate image data that can be output to an output device (such as the display unit of the composite device 2 or the switcher 3) (corresponding to steps S207 and S208).Note that the operator may be able to edit and adjust some of the object data in the virtual space. For example, the operator may be able to fine-tune the position of the virtual camera VC in step S202, the position of the polygon mesh BP in step S203, and the position of the avatar C.

[0116] The embodiments disclosed herein should be considered to be illustrative in all respects and not restrictive. The scope of the present invention is defined by the claims, not by the above description, and is intended to include all modifications within the meaning and scope of the claims. [Explanation of symbols]

[0117] 1 Video generation system, 2 Composite device, 21 Scene construction unit, 3 Switcher, 4 Output screen, 5 Capture acquisition device, 6 Audio acquisition device, S Scanner, RC Real camera, VC Virtual camera, O Real microphone, OM Mask, B Board, BP Polygon mesh, C Avatar, M Standing position marker

Claims

1. An image generation device that generates an image in which an avatar is occluded by a predetermined object placed in real space, means for acquiring an image of the real space from a camera that captures the real space; means for acquiring the motion of an actor playing the avatar; a means for generating a virtual space in which an avatar that moves according to the acquired motion and a mask corresponding to the predetermined object generated based on the acquired image of the real space are arranged, and an image of the predetermined object in the image of the real space can be displayed corresponding to the mask; and outputting an image obtained by combining an image of the real space with an image from a virtual camera that captures the virtual space.

2. The image generating device is a means for acquiring information including identification information capable of identifying a position of the predetermined object in the real space and a position where the avatar is to be placed in the virtual space corresponding to the real space, from a three-dimensional information measuring device that measures three-dimensional information of the real space; 2. The image generation device according to claim 1, wherein the generating means generates a virtual space in which the mask is placed at a position in the virtual space corresponding to a position of the predetermined object, based on the identification information, and the avatar is placed at a position where the avatar is to be placed.

3. 3. The image generating device according to claim 2, wherein the generating means generates a virtual space in which the mask is placed in association with an object that is placed at a position in the virtual space corresponding to the position and range of the specified object based on the identification information.

4. an image of the real space is projected onto the object; The generating means comprises: The image generating device according to claim 3, wherein when the mask is in the first state, the object and the image of the real space are displayed in an area cut out in the shape of the mask corresponding to the specified object.

5. The image generating device according to claim 3 , wherein the object is an object arranged based on the identification information obtained by measuring a position identification member attached to the predetermined object in the real space.

6. 6. The image generating device according to claim 5, wherein the position specifying member has a shape corresponding to a range of the predetermined object perpendicular to a direction of imaging when the real space is imaged from a predetermined position by a camera that images the real space.

7. the three-dimensional information measuring device identifies information about a situation of a camera that captures the real space; the means for generating the virtual space sets and places the virtual camera so as to be consistent with the situation of the camera based on information about the situation of the camera acquired from the three-dimensional information measuring device; 7. The image generating device according to claim 6, wherein the object is positioned so that a distance from the virtual camera to the object in the virtual space corresponds to a distance from the predetermined position to the predetermined object in the real space.

8. a first camera that captures an image of the predetermined object from a first direction and a second camera that captures an image of the predetermined object from a second direction are disposed in the real space; a first virtual camera that photographs a mask on which an image of the predetermined object is to be displayed from a direction corresponding to the first direction, and a second virtual camera that photographs a mask on which an image of the predetermined object is to be displayed from a direction corresponding to the second direction, and 2. The image generating device according to claim 1, wherein the output means outputs a first image obtained by combining an image of the virtual space photographed by the first virtual camera with an image of the real space photographed by the first camera, and a second image obtained by combining an image of the virtual space photographed by the second virtual camera with an image of the real space photographed by the second camera.

9. The image generating device according to claim 1 , wherein the mask has a shape corresponding to the shape of the predetermined object.

10. 8. The image generation device according to claim 7, wherein the mask has a shape based on a difference between an image captured from a predetermined position in the real space by a camera that captures the real space, the image being taken before the predetermined object is placed in the real space, and an image being taken after the predetermined object is placed in the real space.

11. the image of the predetermined object displayed in correspondence with the mask is an image of a first object captured in real time by a camera that captures images of the real space during a predetermined period of time; The image generating device according to claim 1 , wherein the image of the avatar is an image of the avatar moving based on the motion of the performer captured in real time during the same period as the predetermined period.

12. An image generation system that generates an image in which an avatar is occluded by a predetermined object placed in real space, means for acquiring an image of the real space; means for acquiring the motion of an actor playing the avatar; a means for generating a virtual space in which an avatar that moves according to the acquired motion and a mask corresponding to the predetermined object generated based on the acquired image of the real space are arranged, and an image of the predetermined object in the image of the real space can be displayed corresponding to the mask; and outputting an image obtained by combining an image of the real space with an image from a virtual camera that captures the virtual space.

13. 1. A video generation method for generating a video in which an avatar is occluded by a predetermined object placed in a real space, comprising: acquiring an image of the real space from a camera that captures the real space; acquiring a motion of an actor playing the avatar; generating a virtual space in which an avatar moving according to the acquired motion and a mask corresponding to the predetermined object generated based on the acquired image of the real space are arranged, and an image of the predetermined object in the image of the real space can be displayed corresponding to the mask; and outputting an image obtained by combining an image of the real space with an image from a virtual camera that captures the virtual space.

Citation Information

Patent Citations

  • Spatial relations for integrating visual images of the physical environment into virtual reality

    JP2019514101A

  • Display system for hall and method for executing event using the same

    JP2023130363A

  • Method for collaboration using head mounted display

    KR1020170044318A