Video generation device, video generation system, and video generation method

The image generation device and method efficiently combine real-world footage with 3DCG in real time, addressing the high costs and impracticality of existing systems by using a camera and mask-based virtual space generation for accurate and cost-effective composite image creation.

JP2026062511APending Publication Date: 2026-04-09COVER CORP
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2025-09-16
Publication Date
2026-04-09

AI Technical Summary

Technical Problem

Existing video production systems that combine real-world footage with 3DCG require expensive equipment and lengthy post-production processes, making real-time compositing of high-precision images impractical, especially in AR live events where motion capture data needs to be reflected in real time.

Method used

An image generation device and method that uses a camera to capture real-space images, acquires performer motion data, generates a virtual space with a mask corresponding to real objects, and combines these images in real time to create composite images of real-world footage and 3DCG, reducing rendering processing load and costs.

Benefits of technology

Enables low-cost, real-time generation of composite images with accurate occlusion and realistic representation of avatars, allowing for flexible video production environments without the need for expensive specialized systems.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026062511000001_ABST
    Figure 2026062511000001_ABST
Patent Text Reader

Abstract

This invention provides a video generation device, a video generation system, and a video generation method that enable the low-cost, real-time generation of composite images combining real-world footage and 3DCG. [Solution] The system acquires images of the real space from a camera that captures the real space, acquires the motion of an actor playing an avatar, places the avatar acting with the acquired motion and a mask corresponding to a predetermined object generated based on the acquired images of the real space, generates a virtual space in which images of the predetermined object in the real space can be displayed corresponding to the mask, and outputs an image which is a composite of the images of the real space and images from a virtual camera that captures the virtual space, in which the avatar is occluded from the predetermined object placed in the real space.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a video generation device, a video generation system, and a video generation method.

Background Art

[0002] There is a technology such as AR (Augmented Reality) that expands the real world. In recent years, virtual live (so-called AR live) using such augmented reality has been carried out, and there is a technology for generating a video in which a CG avatar performs a performance such as singing on a stage in the real space (see, for example, Non-Patent Document 1 and Patent Document 1).

Prior Art Documents

Non-Patent Documents

[0003]

Non-Patent Document 1

Patent Documents

[0004]

Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0005] Traditionally, in video production that combines real-world footage with 3DCG (3-Dimensional Computer Graphics), it was necessary to edit the compositing of real-world footage and 3DCG after filming (so-called post-production), which resulted in a time-consuming process before the final composite image could be created. Furthermore, in AR live events like those described above, it is necessary to reflect the motion capture data of the performers in real time when compositing the avatar's movements onto the real-world stage. Therefore, AR live events utilize specialized systems that generate images that combine 3DCG with motion and real-world footage with the highest possible accuracy. However, such systems are extremely expensive, and there may also be constraints from the camera manufacturer used to film the real-world footage for compositing. As a result, even if one attempts to create high-precision composite images of 3DCG and real-world space, it is often impractical due to cost and equipment limitations.

[0006] This invention was conceived in view of the above circumstances and provides an image generation device, an image generation system, and an image generation method that enable the low-cost, real-time generation of composite images of real-world images and 3DCG. [Means for solving the problem]

[0007] (1) An image generation device according to a certain aspect of the present invention is an image generation device (e.g., composite device 2) that generates an image in which an avatar (e.g., avatar C) is occluded with a predetermined object (e.g., real microphone O) placed in real space, Means for acquiring images of the real space from a camera that photographs the real space (for example, a real camera RC) (for example, the input unit of the composite device 2 that acquires images from the real camera RC), Means for acquiring the motion of the performer who plays the avatar (for example, the input unit of the composite device 2 that acquires the performer's motion data from the capture acquisition device 5), Means for generating a virtual space (for example, rendering video captured by a virtual camera VC in the scene construction unit 21), where an avatar that operates with acquired motion and a mask corresponding to the predetermined object generated based on acquired video of the real space (for example, a 2D mask OM corresponding to a real microphone O set on a polygon mesh BP) are arranged, and the video of the predetermined object in the video of the real space can be displayed corresponding to the mask (for example, the video of the real microphone O is displayed by projection mapping of the video of the real space onto the polygon mesh BP corresponding to the 2D mask OM, see steps S205, S206, and the modified example (projection of video from a real camera RC) in Figure 3, The system includes means for outputting an image obtained by combining the image of the real space with an image from a virtual camera that captures the virtual space (for example, the image from the real camera RC and the image from the virtual camera VC are composited (combined) in the scene construction unit 21 and output, step S208 in Figure 3).

[0008] With this configuration, a mask corresponding to a predetermined object is placed in the virtual space, and the image of that predetermined object can be displayed corresponding to the mask, thereby enabling the output of a composite image of real-world footage and 3DCG at low cost and in real time.

[0009] (2) In the above (1), the video generation device is The system includes means for acquiring information (for example, the input unit of a composite device 2 that acquires information identified by the scanner S) that can identify the position of a predetermined object in the real space and the position where the avatar will be placed in the virtual space corresponding to the real space, from a three-dimensional information measuring device (for example, a scanner S equipped with a scanner function) that measures three-dimensional information in real space, including identification information (for example, the coordinates of a real microphone O identified by the scanner S, the coordinates of a standing position marker M which will be the standing position of the avatar, steps S101, S103, and S104 in Figure 2), The generating means generates a virtual space in which the mask is placed at a position in the virtual space corresponding to the position of the predetermined object, and the avatar is placed at a position where the avatar is to be placed, based on the identification information (for example, in the scene construction unit 21, the mask OM is placed in the virtual space so as to correspond to the coordinates of the real microphone O, and the avatar C is placed so as to correspond to the coordinates of the standing position marker M, steps S201, S203 to S205 in Figure 3).

[0010] With this configuration, the positional relationship between objects placed in real space and the avatar can be determined, making it possible to represent occlusion and other phenomena based on front-to-back information.

[0011] (3) In (2) above, the means for generating is an object that is placed in the virtual space at a position corresponding to the position and range of the predetermined object based on the identification information (for example, a polygon mesh BP placed in the virtual space to correspond to the position and range of the real microphone O in the real space), and generates a virtual space in which the mask is placed in association with the object (for example, setting the mask OM as material information for the polygon mesh BP, step S205 in Figure 3).

[0012] This configuration reduces the rendering processing load, making it possible to generate in real time images where a given object in real space and the avatar are occluded.

[0013] (4) In (3) above, the image of the real space is projected onto the object (for example, the image of a real camera RC is associated with a polygon mesh BP and projection mapping is performed (the image of the real camera is projected onto the polygon mesh). Step S205 in Figure 3), The means for generating the above is, The system includes means for performing a process in which, when the mask enters a first state (for example, when mask OM is in the ON state), the object and the image of the real space are displayed within the range cut out by the shape of the mask corresponding to the predetermined object (for example, the right diagram in Figure 7(B), and steps S206 and S208 in Figure 3).

[0014] With this configuration, the object corresponding to the position and range of a predetermined object, and the projected image of real space, are displayed with the shape of a mask cut out, making it possible to generate an image in which the predetermined object in real space and the avatar are occluded.

[0015] (5) In (3) above, the object is an object positioned based on the identification information obtained by measuring a positioning member (for example, a board B attached to a real microphone O) attached to a predetermined object in real space (for example, an object corresponding to the shape and coordinates of the board B detected based on information measured by the scanner S (for example, point cloud data, mesh data, etc.), step S103 in Figure 2, step S203 in Figure 3).

[0016] This configuration makes it possible to easily and quickly obtain clues to identify the position and range of a given object, enabling low-cost and rapid progress in video production.

[0017] (6) In (5) above, the position-determining member has a shape that corresponds to the range of the predetermined object perpendicular to the shooting direction when the real space is photographed from a predetermined position by a camera that photographs the real space (for example, a plate-shaped board B that corresponds to the range of the real microphone O photographed from the real camera RC).

[0018] This configuration makes it possible to easily and quickly obtain clues to identify the position and range of a given object, enabling low-cost and rapid progress in video production.

[0019] (7) In the above (6), the three-dimensional information measuring device (for example, the scanner S having a camera position transmission function) specifies information on the situation of the camera that photographs the real space (for example, information such as the coordinates and posture of the real camera RC, step S102 in FIG. 2), The means for generating the virtual space sets and arranges the virtual camera so as to match the situation of the camera based on the information on the situation of the camera acquired from the three-dimensional information measuring device (for example, based on the information of the real camera RC acquired by the scanner S, the virtual camera VC is set to the same settings as those of the real camera RC, step S202 in FIG. 3), The object is arranged such that the distance from the virtual camera in the virtual space to the object corresponds to the position of the distance from the predetermined position in the real space to the predetermined object (for example, refer to FIGS. 5, 6, etc.).

[0020] According to such a configuration, it is possible to make the video from the camera that photographs the real space and the video from the virtual camera that photographs the virtual space consistent, and to generate a realistic composite video.

[0021] (8) In the above (1), in the real space, a first camera that photographs the predetermined object from a first direction and a second camera that photographs the predetermined object from a second direction are arranged (for example, real cameras RC1 to RC4 illustrated in FIG. 6(A), etc.), In the virtual space, a first virtual camera that photographs a mask (for example, a mask generated to be paired with the camera, or a mask shared by a plurality of cameras) for displaying the video of the predetermined object from a direction corresponding to the first direction, and a mask (for example, a mask generated to be paired with the camera, or a mask shared by a plurality of cameras, refer to the (mask for projecting the video in the modified example), etc.) for displaying the video of the predetermined object from a direction corresponding to the second direction are arranged (for example, virtual cameras VC1 to VC4 illustrated in FIG. 6(B), etc.), The means for outputting outputs a first video obtained by synthesizing the video of the virtual space captured by the first virtual camera with the video of the real space captured by the first camera, and a second video obtained by synthesizing the video of the virtual space captured by the second virtual camera with the video of the real space captured by the second camera (for example, the output screen 4 in FIGS. 10(A) to (C)).

[0022] According to such a configuration, since videos of a predetermined object captured from a plurality of angles are generated, a viewer of the synthesized video can have a sense of reality with respect to the synthesized video.

[0023] (9) In the above (1), the mask has a shape corresponding to the shape of the predetermined object (for example, a mask having a shape cut out for each delicate shape such as the fine wiring of the real microphone O or the mesh of the pop guard).

[0024] According to such a configuration, since the mask has a shape corresponding to the predetermined object, it is possible to accurately perform occlusion with the avatar even for the delicate shape portion of the predetermined object. In addition, since the rendering processing load is lighter than performing occlusion with the avatar using a 3D object corresponding to the three-dimensional shape of the predetermined object, it is possible to generate a synthesized video at low cost and in real time on the spot.

[0025] (10) In the above (9), the mask has a shape based on the difference between the video of the real space captured by the camera for capturing the real space from a predetermined position in the real space before the predetermined object is arranged in the real space and the video after the predetermined object is arranged.

[0026] According to such a configuration, it is possible to generate a mask for generating a video in which a delicate shape portion of a predetermined object and the avatar are occluded at low cost and quickly.

[0027] (11) In (1) above, the image of a predetermined object displayed in correspondence with the mask is an image of a first object captured in real time by a camera that captures the real space over a predetermined period of time, The avatar's video is an image of an avatar that operates based on the performer's motion, which is captured in real time over the same period as the predetermined period.

[0028] With this configuration, the real space and the performers' actions are combined in real time, which can enhance the sense of realism that viewers perceive in the composite image.

[0029] (12) An image generation system according to a certain aspect of the present invention is an image generation device (e.g., image generation system 1) that generates an image in which an avatar (e.g., avatar C) is occluded with a predetermined object (e.g., real microphone O) placed in real space, The means for acquiring images of the real space (for example, Real Camera RC, Composite Device 2), Means for acquiring the motion of the performer playing the avatar (for example, a capture acquisition device 5, a composite device 2), Means for generating a virtual space (for example, rendering video captured by a virtual camera VC in the scene construction unit 21), where an avatar that operates with acquired motion and a mask corresponding to the predetermined object generated based on acquired video of the real space (for example, a 2D mask OM corresponding to a real microphone O set on a polygon mesh BP) are arranged, and the video of the predetermined object in the video of the real space can be displayed corresponding to the mask (for example, the video of the real microphone O is displayed by projection mapping of the video of the real space onto the polygon mesh BP corresponding to the 2D mask OM, see steps S205, S206, and the modified example (projection of video from a real camera RC) in Figure 3, The system includes means for outputting an image obtained by combining the image of the real space with an image from a virtual camera that captures the virtual space (for example, a composite device 2, a switcher 3, an output screen 4, etc., which composites (combines) the image from the real camera RC and the image from the virtual camera VC in the scene construction unit 21 and outputs it, step S208 in Figure 3).

[0030] With this configuration, a mask corresponding to a predetermined object is placed in the virtual space, and the image of that predetermined object can be displayed corresponding to the mask, thereby enabling the output of a composite image of real-world footage and 3DCG at low cost and on the spot.

[0031] (13) An image generation method according to a certain aspect of the present invention is an image generation method (e.g., composite device 2) that generates an image in which an avatar (e.g., avatar C) is occluded with a predetermined object (e.g., real microphone O) placed in real space, The process involves acquiring images of the real space from a camera that captures the real space (for example, a real camera RC) (for example, acquiring images from the real camera RC with a composite device 2), The steps include acquiring the motion of the performer who plays the avatar (for example, acquiring the performer's motion data from the capture acquisition device 5 using the composite device 2), The process involves creating a virtual space (for example, rendering the video captured by the virtual camera VC in the scene construction unit 21), where an avatar that operates with acquired motion and a mask corresponding to the predetermined object generated based on acquired real-space video (for example, a 2D mask OM corresponding to the real microphone O set in the polygon mesh BP) are arranged, and the video of the predetermined object in the real-space video can be displayed corresponding to the mask (for example, the video of the real microphone O is displayed by projection mapping the video of the real-space video onto the polygon mesh BP corresponding to the 2D mask OM, see steps S205, S206 in Figure 3, and the modified example (projection of video from real camera RC), etc.), The system includes a step of outputting an image obtained by combining the image of the real space with an image from a virtual camera that captures the virtual space (for example, the scene construction unit 21 composites (combines) the image from the real camera RC and the image from the virtual camera VC and outputs it, step S208 in Figure 3).

[0032] With this configuration, a mask corresponding to a predetermined object is placed in the virtual space, and the image of that predetermined object can be displayed corresponding to the mask, thereby enabling the output of a composite image of real-world footage and 3DCG at low cost and on the spot. [Brief explanation of the drawing]

[0033] [Figure 1] This figure shows an example configuration of a video generation system including the video generation device of this embodiment. [Figure 2] This flowchart shows an example of the operating procedure of a scanner used to synchronize the real and virtual spaces. [Figure 3] This flowchart shows an example of the processing sequence of virtual production by the video generation device in this embodiment. [Figure 4] This is a diagram illustrating a method for aligning real space and virtual space. [Figure 5] This is a diagram illustrating a method for aligning real space and virtual space. [Figure 6] This diagram illustrates the correspondence between cameras in the real world and cameras in the virtual world. [Figure 7] This diagram illustrates the procedure for generating masks corresponding to objects in real space. [Figure 8] This diagram illustrates the correspondence between images captured by a real-world camera and images captured by a virtual camera. [Figure 9] This is an example of data used for synthesis in the video generation device of this embodiment. [Figure 10] This is an example of a composite image generated by the video generation device of this embodiment.

[0034] Embodiments of the communication system according to the present invention will be described below with reference to the drawings. However, the present invention is not limited to the following examples, and is intended to be included in the present invention as indicated by the claims, with all modifications within the meaning and scope equivalent to the claims being included. In the following description, the same elements in the drawings will be denoted by the same reference numerals, and redundant descriptions will not be repeated.

[0035] The present invention relates to a method, apparatus, and system for generating video by combining virtual three-dimensional computer graphics (3DCG) with video acquired from a camera placed in real space. The present invention is applicable, for example, to situations where video is generated in which a 3DCG avatar reflecting the motion of a performer is operating in real space (e.g., singing, talking), requiring simultaneous acquisition of the performer's motion, acquisition of the performer's voice, and filming of the real space, and real-time synthesis of these. For example, suppose video is filmed of a real-space studio with only a microphone in place and no one actually in front of it, while simultaneously capturing and acquiring audio of a singing performance by a performer in another space at the same time as the filming and the real-world progress. In such a scenario, applying the present invention allows for real-world filming and the performer's performance to be synthesized in real time, generating video of a virtual avatar visiting a real studio and performing a singing performance.

[0036] According to the present invention, by placing polygon mesh data that projects real-world objects into a virtual space and placing a 2D mask that allows the projected real-world objects to be displayed, it is possible to generate composite images in which real-world objects and avatars are occluded (hereinafter, the occluded images will be referred to as composite images) at low cost without requiring expensive equipment as in the past. As a result, even if the filming of the real space, the acquisition of motion data of the performer, and the reflection of the movements on the avatar are performed simultaneously, the images can be composited in real time with minimal delay. Furthermore, since there are no restrictions on the type of camera or lens used to film the real space, it can be flexibly adapted to any video production environment. In this invention, by bringing images of real-world objects into a virtual space, it is possible to represent a 3DCG character as if it had appeared in the real space.

[0037] There are several methods for creating composite images that make 3DCG characters appear to exist in the real world by occluding real-world objects and 3DCG. Occlusion refers to the state where parts of an object behind it are not rendered because they are obscured by an object located in the foreground from the camera's perspective. First, a typical method is to edit the image in post-production. Traditional post-production can create highly accurate composite images even when occluding complex real-world objects and 3DCG, but it requires long editing times and advanced editing skills. Therefore, it is not possible to composite 3DCG in real time (complete the composite image) at the same time as shooting the real-world footage. Also, traditionally, if the 3DCG to be composited is an avatar that reflects the motion of an actor, the CG compositing is done after the real-world shooting, so it is possible to reshoot the motion (multiple takes) and take time to adjust it to match the real-world footage. However, when it is necessary to combine an avatar that reflects the motion of an actor acting simultaneously with the timing of filming in the real world with the actual filmed footage in real time, it is not possible to edit it later as in the past. Therefore, it is desirable to devise ways to enable real-time compositing while ensuring that consistency (for example, geometric consistency) between the real world and the virtual world is not broken.

[0038] In addition, in recent years, for example, in large-scale AR live events, specialized systems (camera tracking systems) are sometimes used to generate composite images in real time that perform occlusion between the real-world stage and 3DCG avatars. For example, a tracking camera (a camera equipped with optical sensors, accelerometers, gyroscopes, etc.) is attached to the main camera, enabling compositing with CG almost simultaneously with shooting. However, such specialized systems require budgets in the hundreds of millions of yen, making them impractical due to their enormous costs. Furthermore, each system has constraints regarding the manufacturers of the real cameras (lenses, etc.) used to shoot in the real world, and it often takes several months of preparation after the lenses to be used have been determined in advance. For example, the venue, which will be the real-world space, is scanned in advance, or 3DCG is created in advance based on CAD drawings of the venue, so that the images shot with the lenses to be used match the images shot from the virtual camera. Therefore, even with the use of an extremely expensive specialized system, it has been difficult to match the settings of the virtual camera (CG camera) on the day of shooting, regardless of the type of real camera (camera used for shooting images). In other words, there was no system that could flexibly handle all kinds of lenses.

[0039] Furthermore, if the real-world object that partially obscures the 3DCG with occlusion has a complex shape, the compositing becomes even more difficult. For example, a professional stand microphone used by a professional singer on stage may have complex wiring and fine mesh components such as pop filters, resulting in a complex shape. To achieve occlusion, it is first necessary to scan (three-dimensionally measure) the three-dimensional shape of the real-world object, and then, based on the scanned information, place polygon mesh data corresponding to the object in the virtual space to create a front-to-back relationship with the avatar.

[0040] One method for scanning three-dimensional objects is photogrammetry. This measurement method analyzes parallax information from two-dimensional images obtained by photographing an object from multiple viewpoints to determine its shape. However, to accurately scan the fine details of a complex object to the point where the boundary between the background image and the real object is indistinguishable, even when the image is zoomed in, a huge number of photographs are required, and the scanning process alone can take several hours. Furthermore, depending on the environment, there are limitations due to weather and time, such as whether it is a sunny day or not, making it not an easy scan to perform.

[0041] If photogrammetry is not used, there are methods to scan three-dimensional objects using depth cameras equipped with ToF (Time of Flight) sensors or stereo depth cameras. There are applications that allow for simple object scanning using depth cameras or the aforementioned photogrammetry on smartphones, but the accuracy is not very high. Furthermore, even when using depth cameras that can take pictures while scanning, existing products have low scanning accuracy and resolution, making it difficult to produce high-quality images. For example, even if a real object is scanned using a conventional scanner, when a close-up image of a complex shape is taken, the low scanning accuracy of the object's boundaries results in images where avatars appear to be embedded in parts of the object. This reduces the viewer's sense of immersion in the composite image, resulting in an unrealistic image.

[0042] Even if it were possible to quickly place a polygon mesh of a real-world object's 3D shape into a virtual space, rendering the 3D data into a 2D image would be computationally intensive if the mesh itself has a complex shape. This is because the large number of polygon vertices increases the computational load required to convert it to 2D (rasterize). For example, a highly detailed reproduction of a complex object's shape could result in 100 million polygons, making real-time rendering extremely difficult. This would make it difficult to perform virtual production simultaneously with real-world filming, causing the timing of the performer's motion, the real-world footage, and the simultaneously acquired audio to become mismatched. When filming in real time, it is absolutely essential to avoid this kind of delay.

[0043] Furthermore, the increasing diversification of content in recent years has led to a rise in virtual talent. Given the rising popularity of virtual talent, it is desirable to blur the lines between what can be achieved by virtual and real people in content production, thereby expanding opportunities for them to shine. However, when virtual talents engage in video streaming, it is not realistic to spend the high costs of tens of millions to hundreds of millions of yen mentioned above. Generally, the method used involves overlaying a transparent avatar image, created using a green screen, onto other real-world footage, which does not generate images that occlude with real objects. Therefore, in order to increase opportunities for virtual talent, it is desirable to devise ways to make them blend into the real world like real people at a low cost.

[0044] Therefore, in this invention, the number of polygons corresponding to complex-shaped objects that occlude with the avatar is reduced as much as possible, and a two-dimensional (2D) mask cut out from the complex-shaped object, which allows for the identification of the front-to-back relationship with the avatar, is placed in the virtual space, and the image of the object can be displayed corresponding to the mask. As a result, it is possible to easily and quickly generate a composite image of real-time footage of the real space and an avatar that reflects the motion of a performer acting in real time, in which parts of objects placed in the real space and the avatar occlude, without using expensive equipment. Details will be explained below.

[0045] Figure 1 is a block diagram of the configuration of the video generation system according to the present invention. In this embodiment, the video generation system 1 includes a composite device 2, a switcher 3, an output screen 4, a real camera RC, a scanner S, a capture device 5, and an audio acquisition device 6.

[0046] In the video generation system 1 of the present invention, a composite device 2 is positioned on the video path from the video captured by a real camera RC placed in real space to the output screen 4, such as a monitor. The composite device 2 receives video data captured by the real camera RC and motion data acquired from a capture device 5 that acquires the performer's body tracking and facial motion. A composite image of the video data from the real camera RC and an avatar reflecting the performer's motion is output from the composite device 2 to the switcher 3. As a result, the composite image is displayed on the output screen 4. For example, the image is as shown in Figure 10, but the details will be described later.

[0047] The composite device 2 is a computer (personal computer, etc.) equipped with a processing unit such as a CPU (Central Processing Unit) or GPU (Graphics Processing Unit), a storage unit such as RAM (Random Access Memory), ROM (Read Only Memory), flash memory, or HDD (Hard Disk Drive), an input / output interface, a display device (display unit) for displaying the composite image, and an operating device. The storage unit stores a program that is executed as a rendering engine and a program that enables the construction and editing of a virtual space (a virtual 3D space). The composite device 2, through its control unit (composed of the CPU and memory such as RAM), enables the construction and editing of the virtual space based on the stored program and renders 3DCG etc. placed in the virtual space. For example, software that enables CG editing and video editing to realize virtual production is installed. Note that the various processes that can be realized by the composite device 2 (for example, each step in Figure 3, etc.) may be realized by one or more computers.

[0048] The composite device 2 includes a scene construction unit 21 that has the function of constructing a virtual space. The scene construction unit 21 constructs a virtual space for each camera. Figure 1 illustrates a case where the composite device 2 has multiple scene construction units 21 for each camera, but a composite device 2 may be provided for each camera (a PC corresponding to each camera may be provided). Various 3DCG objects and virtual camera VCs are placed in the virtual space. In the scene construction unit 21, the data of the virtual space, which is the image captured by the virtual camera VC (described later with reference to Figure 9), and the image captured by the real camera RC are combined. The image combined in the scene construction unit 21 is output to the switcher 3. The operator of the composite device 2 performs work to adjust the combined image for output to the switcher 3 based on the image displayed on the display unit. Specifically, this involves setting masks to be placed in the virtual space within the scene construction unit 21, and making fine adjustments to the placement of objects and virtual camera VCs within the virtual space.

[0049] Switcher 3 receives video data (video signals) and audio data (audio signals) output from composite device 2. Switcher 3 can display video on output screen 4, etc., by switching between video data shot from multiple angles.

[0050] Output screen 4 is a display monitor or the like, which is a display device that shows the composited video. For example, in a video production setting, a director or other person can give instructions on shooting angles based on the video on output screen 4. In addition, audio data etc. (which may also be other music data etc.) output from switcher 3 (or a separate mixer or other audio equipment) may also be output as sound data such as audio signals through an output device such as the speaker of the terminal equipped with output screen 4 or a separate speaker.

[0051] The data output destination of the switcher 3 may be a distribution PC (not shown in the diagram), and this distribution PC may be connected to network 2. Network 2 is, for example, the internet and consists of access networks such as LAN (Local Area Network), WAN (Wide Area Network), mobile communication networks (e.g., 5G, wireless networks, etc.), wired telephone networks, FTTH (Fiber To The Home), and CATV (Cable Television) networks. This makes it possible to live stream (live broadcast) of the composite video generated by the composite device 2 via the internet.

[0052] The video production site using the video generation system 1 in this embodiment includes a real-world studio that serves as the background for compositing (hereinafter also referred to as the real-world studio) and a motion capture studio that performs motion capture to acquire the motion of performers so that their movements are reflected in the avatars.

[0053] The real studio is equipped with a real microphone O and a real camera RC that films the real studio, including the real microphone O. The real microphone O is, for example, a stand microphone and has delicate and complex shapes such as intricate wiring and the mesh of a pop filter.

[0054] The capture studio is equipped with a capture acquisition device 5 and an audio acquisition device 6, which acquire the motion and voice of the performer. The performer performs actions within the capture studio, and the performer's body movements and facial expressions are acquired (tracking input) by the sensors of the capture acquisition device 5.

[0055] The Real Camera RC is a video camera that does not have a scanner function. Multiple Real Camera RCs are placed within the real studio. The video cable from the Real Camera RC is connected to Switcher 3. The aforementioned Composite Device 2 is connected as an interrupt in between these video cables.

[0056] Scanner S is a small terminal with an application installed that has the function of a 3D scanner and the function of transmitting information acquired by Scanner S to other devices. Scanner S transmits information about the three-dimensional shape of an object, such as 3D point cloud data measured by Scanner S and mesh data converted from point cloud data, to Composite Device 2. Scanner S may be a computer such as a smartphone equipped with functions that enable the identification of three-dimensional shapes by measuring the distance to an object using light (e.g., equipped with a ToF (Time of Flight) sensor), self-localization function, communication function, etc. Measurement results such as the coordinates of the real studio measured by Scanner S and the shape and coordinates of objects placed in the real studio are transmitted to Composite Device 2.

[0057] Furthermore, scanner S has a function to transmit information about its position and rotation to composite device 2. When scanner S is positioned in front of the lens of real camera RC, the measured position and rotation (attitude) information of scanner S is approximately equivalent to the position and rotation (attitude) information of real camera RC (i.e., approximately equivalent to the camera's position and attitude information). The information acquired by scanner S is also transmitted from scanner S to a terminal (relay terminal) on which an application for recording and setting camera information is installed. Other lens information, such as physical parameters of the virtual camera VC, such as focal length and sensor size, is input by the operator into the application for recording and setting camera information (also called the camera setting application). Camera information (camera status), including lens information set by the reception from scanner S and input from the operator, is transmitted to composite device 2.

[0058] In the composite device 2, based on the information transmitted from the camera setting application, camera information such as the position and rotation of the virtual camera VC placed in the virtual space formed by the scene construction unit 21 is set to match the position, orientation, and lens information of the real camera RC acquired by the scanner S. This makes it possible to measure camera information immediately on-site and reflect it in the virtual camera in the virtual space, even when it is not possible to obtain information about the lens of the real camera used for shooting in advance.

[0059] The capture device 5 is located in the aforementioned capture studio. The capture device 5 in this embodiment is a device (tracking device, etc.) that has the function of acquiring body tracking data to capture the performer's movements and face tracking data to capture the performer's facial movements from sensors attached to the performer's body.

[0060] The audio acquisition device 6 is a microphone that acquires the performer's voice. The audio data acquired by the audio acquisition device 6 is input to the switcher 3. The audio acquisition device 6 placed in the capture studio may be the same as the real microphone O. This ensures that when the performer moves within the capture studio, their movements take into account the presence of the real microphone O, preventing movements such as the avatar's arm accidentally passing through the real microphone O. It also makes it possible to acquire sound quality of the same quality as when audio is input to the real microphone O placed in the real studio. In addition, musical instruments may be played in the real studio, and the performance data of the instrument may be sent to the switcher 3. Furthermore, an audio mixer for sound adjustment may be connected between the microphone inputting the performance data and the switcher. Even in such cases, the video generation method of the present invention makes it possible to generate video content in which the performer's performance and the real live performance are synthesized in real time.

[0061] Furthermore, the various hardware components that make up the video generation system 1 may be connected via network 2, enabling bidirectional transmission and reception of information (data).

[0062] <Procedure for generating composite images> (Regarding the procedure outline) Referring to Figures 2 to 10, the details of the method for generating a composite image of a real-world image and a 3DCG avatar using the image generation system 1 of the present invention will be explained.

[0063] Figure 2 is a flowchart illustrating the procedure for the operation of the scanner S in real space. In this embodiment, the scanner S is used to measure information necessary to match the environment of the virtual space within the scene construction unit 21 where the avatars are placed with the real space (real studio) that serves as the background for the composite image. The measured information from the scanner S is transmitted to the composite device 2.

[0064] Figure 3 is a flowchart illustrating the procedure for virtual production performed within the composite device 2 using information acquired from scanner S.

[0065] (Matching shooting environment settings) First, we will explain the procedure for aligning the coordinates of the real-world studio with the coordinates of the virtual space constructed in the scene construction unit 21 of the compositing device 2.

[0066] In step S101 of Figure 2, the studio in the real space where the first object is placed is scanned, using the avatar's standing position marker in the real space as the origin, and the scanned information is transmitted to the composite device 2. Figure 4 is a diagram illustrating the correspondence between the real space and the virtual space when matching the environments of the real space and the virtual space. Figure 4(A) is a diagram representing the real studio. As illustrated in the left diagram of Figure 4(A), a real microphone O is placed in the real studio, and a standing position marker M is attached to the floor as a marker for the expected standing position if the avatar were to stand in front of the microphone. The relationship between the height of the microphone and the position of the standing position marker M is adjusted considering what would happen if the avatar were actually standing in the real space. For example, the avatar's height is taken into consideration. In the real studio, the scanner S is used to scan the entire real studio, using the standing position marker M as the origin. The origin is determined, for example, by specifying the origin (scan start position) on the scanner S (this may involve placing a marker such as a two-dimensional barcode at the origin and reading it). The coordinate information of the scanned real-world studio is transmitted to the compositing device 2.

[0067] Next, in step S201 of Figure 3, when the composite device 2 receives the coordinate information of the studio scanned by the scanner S, it performs a process to match the origin of the virtual space coordinates with the standing position marker M of the avatar in the real space scanned by the scanner S. The origin specified by the scanner S becomes the display start position of avatar C, and avatar C is placed there. Figure 4(B) is a diagram showing the virtual space constructed by the scene construction unit 21. Note that the marker VM is the origin in the virtual space and is represented virtually to explain that it corresponds to the standing position marker M in the real studio, and is not actually displayed. As shown in the left diagram of Figure 4(B), avatar C, which is a 3DCG character, is placed at the position of the origin. Note that board B and polygon mesh BP are components and objects for identifying the position of real microphone O, but details will be described later.

[0068] Returning to Figure 2, in step S102, the scanner S identifies the camera information of the real-space camera in the real studio and transmits it to the composite device 2. Specifically, this camera information is transmitted to the composite device 2 from the camera setting application (installed on the relay terminal) that received the information from the scanner S. Figures 5 and 6 are diagrams used to explain the correspondence between the real space and the virtual space when describing the method of identifying the camera information. Figure 5 is a view from the direction of the rear of the real microphone O and from the direction of looking at the front of the avatar C when the avatar C is standing in front of the microphone, in the relationship between the real space and the virtual space in Figure 4 described above. Figure 6 is a view of the real microphone O and avatar C from above in the relationship between the real space and the virtual space in Figure 4 described above, and shows an example in which multiple cameras are arranged.

[0069] Figure 5(A) shows an example of a real-world studio. As shown in Figure 5(A), a scanner S is placed in front of the lens of a real camera RC that captures the real space. Then, the position and rotation information of the scanner S is transmitted to the composite device 2 via the camera settings application. Physical parameters such as the sensor size and focal length of the real camera RC's lens are also entered by the operator into the camera settings application and transmitted to the composite device 2.

[0070] In the composite device 2, in step S202 of Figure 3, the virtual camera is configured to have the same information (situation) as the camera in the real space, based on the camera information identified by the scanner S transmitted from the camera setting application. As a result, as shown in the example of the virtual space in Figure 5(B), the position and orientation information of the virtual camera VC is set to be the same as that of the real camera RC. Furthermore, as shown in Figure 6, when multiple cameras are installed, the same procedure as in step S102 of Figure 2 and step S202 of Figure 3 is performed for each real camera RC. This ensures that the settings (position, orientation, lens information, etc.) of each camera (real cameras RC1-RC4 and virtual cameras VC1-VC4) match. Alternatively, each time the camera information changes, such as by moving the real camera RC, step S102 may be repeated to reacquire the camera information, and then in step S202, the virtual camera VC may be configured based on the newly acquired camera information.

[0071] (Identifying objects placed in real space) Next, we will explain how to identify an object placed in real space. In this embodiment, the object placed in real space is the real microphone O, which is the object that creates a front-to-back relationship with the avatar C and is the target of occlusion. In this embodiment, in order to quickly identify the position coordinates of the real microphone O and reflect them in the virtual space, a board B, which is a simple-shaped plate, is used as a position identification member for the real microphone O. In this embodiment, as illustrated in the right diagram of Figure 4(A), Figure 5(A), and Figure 6(A), board B is attached to the real microphone O.

[0072] In step S103 of Figure 2, to identify the first object placed in real space, the scanner S scans the newly placed position identification member in real space. In step S104, the 3D data (mesh data) identification information of the scanned position identification member is selected and sent to the composite device 2, ending the operation of the scanner S for setup. In this embodiment, after scanning the entire real studio area in step S101, a new board B is attached to the real microphone O. The scanner S scans only the periphery of the newly placed board B in the real studio. After scanning, the scanner S still contains scan data other than board B scanned in step S101, so the scanner S is set to output only the scan data (mesh data) of board B. Then, the mesh data of board B is transferred to the composite device 2. Note that "scan data only" does not strictly mean only board B, but generally the area of ​​board B is sufficient.

[0073] Next, in the composite device 2, in step S203 of Figure 3, a polygon mesh of the shape of the position-identifying member is placed at coordinates corresponding to the coordinates of the position-identifying member, based on the 3D data (mesh data) of the position-identifying member identified by the scanner S. For example, as illustrated in the right diagram of Figure 4(B), Figure 5(B), and Figure 6(B), a polygon mesh BP of the board corresponding to the position and range of board B in real space is placed. The composite device 2 may be pre-configured to automatically generate a polygon mesh based on the mesh data acquired from the scanner S, or the mesh data input to the composite device 2 may be set in the scene construction unit 21 by the operator so that the polygon mesh BP is placed. Strictly speaking, there is a slight error in the coordinates between board B and the polygon mesh BP, but this error can be fine-tuned by the operator. In this way, board B has a simple shape that does not take time to scan compared to microphones with complex shapes. Therefore, the position and range of the real microphone O can be quickly and roughly identified using a sensor mounted on a smartphone or the like.

[0074] (Generating a mask corresponding to a given object and projecting the image) Next, we will explain the masks for the avatar and the objects to be occluded. Since scanning the 3D data of the avatar and the objects to be occluded themselves is not efficient, in this embodiment, a board B, which serves as a positioning member, is used to roughly determine the position of the real microphone O. Then, a 2D mask and the image from the real camera RC are set for the polygon mesh BP in the virtual space, which is this determined data. The shape of the 2D mask is cut out to match the shape of the real microphone O, and when the cut-out real microphone O is turned on (also called the first state), the polygon mesh BP and the image from the real camera RC are cut out to match the shape of the mask. This makes it possible to match the 3DCG space (virtual space) with the world of images seen from the real camera by projection mapping the image from the real camera onto the polygon mesh BP.

[0075] Specifically, in step S204 of Figure 3, a 2D mask is generated from images captured by the virtual camera and the corresponding real camera, with the shape of the first object in real space cut out. For example, using the real camera RC1 in the real studio, first, footage is taken without the real microphone O in the real studio. Then, footage is taken again with the real microphone O placed in the same position. The difference between the images before and after placing the real microphone O is taken to create a black and white mask of the shape of the real microphone O. Note that the mask set in the virtual space itself is switch information for occlusion on / off, so there is no color information, and it is sufficient to have at least shape information. The greater the difference in brightness between the background and the object, the cleaner the detail of the shape can be cut out. Therefore, when generating the mask by difference, it is preferable to use a white background and a black object, but it is not limited to this.

[0076] Furthermore, as illustrated in Figure 6, when using multiple cameras to capture images from multiple angles, a difference-based mask is generated from the images captured from each camera's shooting angle. In other words, a difference-based mask is generated for each camera. Refer to Figure 8 to explain the correspondence between images captured by cameras in real space and images captured by virtual cameras. Figure 8(A) is an example of real space, and Figure 8(B) is an example of virtual space corresponding to real space. Camera images 22a and 22b in Figure 8(B) are diagrams illustrating the relationship between the polygon mesh BP and the microphone image. Camera image 22a in Figure 8(B) is an image of the image projected onto the polygon mesh BP as seen from virtual camera VC1. Camera image 22d is an image of the image projected onto the polygon mesh BP as seen from virtual camera VC4. Note that camera images 22a and 22d are diagrams to explain the correspondence between each camera and polygon, so the display of Avatar C is omitted. Additionally, while polygon mesh blueprints outside the real microphone area are shown in the diagram for illustrative purposes, when the image is actually rendered and composited, areas outside the masked region are not visible in the final composite image.

[0077] For example, in the image taken from the direction of the real camera RC1 in the real studio in Figure 8(A) and the corresponding virtual camera VC1 in the virtual space in Figure 8(B), the portion of the image taken from the real camera RC1 projected onto the polygon mesh BP will be cut out with the shape of the real microphone O mask OM, and will be rendered as the image from virtual camera VC1. Similarly, in the case of real camera RC4, in the image taken from the direction of the virtual camera VC4, the portion of the image taken from the real camera RC4 projected onto the polygon mesh BP will be cut out with the shape of the real microphone O mask OM, and will be rendered as the image from virtual camera VC4.

[0078] Next, in step S205 of Figure 3, the 2D mask of the first object and the image data of the real camera, recorded in memory as image data (e.g., texture), are associated with the material information (material, etc.) of the polygon mesh of the positioning member, and the image of the real camera is projected onto the polygon mesh. In other words, the image to be projected is set in the polygon mesh data, and a mask of the shape of the real object is set for the polygon mesh data as on / off information for depth information. For example, as shown in Figure 7(A), the mesh data of board B attached to real microphone O, scanned in a real studio, is placed as a polygon mesh BP in the virtual space shown in Figure 7(B). Then, as shown in the center diagram of Figure 7(B), the mask OM of the shape of real microphone O, generated using the aforementioned difference, and the image data of real camera RC are set as image data recorded in memory in the material information of the polygon mesh BP, so that the mask OM and the image data of real camera RC are associated with the polygon mesh BP, and the image of real camera RC is projected onto the polygon mesh BP. For example, as illustrated in the center diagram of Figure 7(B), the video of the real studio in Figure 7(C) is projected onto the entire polygon mesh BP in the left diagram of Figure 7(B), and the mask OM is associated with the polygon mesh BP. As a result, since the polygon mesh BP corresponds to the coordinates of the real microphone O, the video data of the mask OM and the real camera RC are set to correspond to the coordinates of the real microphone O.

[0079] In step S206, the 2D mask associated with the polygon mesh is turned on so that the polygon mesh BP within the range corresponding to the 2D mask is displayed. For example, when mask OM is turned on, as shown in the image on the right of Figure 7(B), the image from the real camera RC is cut out within the range of mask OM and drawn as the image captured by the virtual camera VC and rendered. Specifically, for the polygon mesh BP onto which the image from the real camera RC is projected, the drawing process is performed only on the pixels in the part of the turned-on mask OM. Simultaneously with the execution of this drawing process, the depth information of the polygon mesh BP (corresponding to board B) is recorded in the virtual space data system. As a result, the polygon mesh BP cut out by mask OM is drawn as the image from the virtual camera VC (i.e., only the part onto which the image from the real camera RC is projected is drawn), and as shown in the image on the right of Figure 7(B), it appears as if only the real microphone O has been extracted from the image and appeared in the virtual space. In other words, the image from the real camera RC is displayed cut out to the shape of the real microphone O. Furthermore, at step S206 (when the mask OM is turned on), the depth information of the polygon mesh BP is set to only include the portion where the image of the real camera RC, which has been cut out by the mask OM, is rendered. As a result, for example, even if the shape of an object is scanned in detail, it becomes possible to drastically reduce the number of polygons to 20,000 to 30,000 polygons for the entire screen, even though the object's shape is made to appear in the virtual space with a complex shape that would result in 100 million polygons in the virtual space, by associating it with depth relationship information in 3DCG. This makes real-time rendering easier. Note that this number of polygons is just one example of how the number of polygons can be drastically reduced, and does not limit the number of polygons for an object or the entire screen.

[0080] (Final synthesis with a working avatar) In step S207, the avatar C, positioned in the virtual space, is made to reflect the motion of the performer's body tracking and facial tracking operating in the capture studio. When the performer performs an action, for example, if the virtual space positional relationship between the avatar C and the polygon mesh BP (and associated mask OM) is as described with reference to Figures 4 to 6 in this embodiment, then when the performer raises their right hand, the right hand of the avatar C rises, and the position of the avatar's right hand moves to the left of the polygon mesh BP in Figure 6(B) (the positive direction in the x-direction shown in the figure, assuming the arrow direction is positive). As a result, from the direction of the virtual camera VC1, the image captured is of the avatar C's right hand moving behind the image of the real microphone O, as shown in the right diagram of Figure 7(B). Therefore, when rendered, the right hand of the avatar C, which is behind the real microphone O, is not drawn in the composite image (for example, the hand of the avatar C behind the real microphone O in Figure 9(A)). On the other hand, if the performer raises their left hand, the left hand of Avatar C also rises, and the position of the Avatar's left hand moves to the right of the polygon mesh BP in Figure 6(B) (the negative direction in the x-direction shown in the figure). As a result, from the direction of the virtual camera VC1, the image is captured showing the left hand of Avatar C moving in front of the image of Real Microphone O, as shown in the right diagram of Figure 7(B). Therefore, when rendered, the part of Real Microphone O that is behind Avatar C's right hand is not drawn in the composite image (for example, Real Microphone O being obscured by the hand on the foreground side of Avatar C in Figure 9(A)). In this way, occlusion between Real Microphone O and Avatar C is made possible in the virtual space.

[0081] In step S208, an image is rendered of a virtual space where a polygon mesh cut out within the range of the 2D mask of the first object is placed, along with the image from the real camera projecting the avatar. This image is then composited with the image captured by the real camera in the real space. The composited image is output from the composite device 2 to the switcher 3. The data output from the composite device 2 to the switcher 3 is video data with data composited as illustrated in Figure 9. As illustrated in Figure 9, in the scene construction unit 21 corresponding to each camera, the shooting data of the virtual camera VC and the shooting data of the corresponding real camera RC are both captured and composited in real time. The shooting data of the virtual camera VC includes at least the avatar object (avatar C) and the position-identifying member polygon mesh (polygon mesh BP). The position-identifying member polygon mesh itself is transparent, but the image of the real camera RC is associated with the mask OM, which is a 2D mask of the real object corresponding to the shape of the real microphone O from the angle of the corresponding real camera RC. Furthermore, when the mask OM is turned on, the depth information of the polygon mesh BP affects the image rendered within the range of the mask OM, and the image from the real camera RC projected onto the polygon mesh BP is drawn at the position of the depth information of the polygon mesh BP. The data of the virtual camera that rendered these avatars C and polygon mesh BP, and the image from the real camera RC corresponding to the virtual camera are combined, and the combined image is returned to the image line. In this way, by cutting out the polygon onto which the image is projected within the range of the 2D mask OM, the load of rasterizing 3D data to 2D data can be reduced.

[0082] Furthermore, as shown in Figure 7, the virtual space does not reflect the entire shape of the real microphone O, including the base of the stand. In this embodiment, the base of the real microphone O is photographed at a shooting angle that does not overlap with the avatar C. Therefore, it is sufficient to place a positioning member (board B) in the area that needs to be occluded and generate a polygon mesh BP, and in order to reduce the rendering load, a polygon mesh BP and mask OM do not need to be created for the parts that do not need to be occluded. For the base part that is not displayed by the mask OM, in the composite image that shows the entire real microphone O, such as the composite image displayed on output screen 4 in Figure 10(A), the image shown is not the image projected onto the mask OM, but the image from the real camera RC itself. In other words, the image of the part of an object placed in the real studio that may overlap with the avatar C is displayed at the position in the virtual space where the avatar C is placed that may overlap, and the rendered data is composited with the image from the real camera RC. As a result, the resulting composite image of Real Microphone O is a combination of two parts: the image of the real camera RC itself, which serves as the background, and the image of the real camera RC displayed on the virtual space mask OM.

[0083] By using the virtual production method described above, it is possible to generate a video in which an avatar C occludes a real microphone O, which is a real object, in real time with accompanying movement, as shown in the example of the composite video displayed on output screen 4 in Figure 10 (B). Furthermore, as shown in Figure 10 (C), even if a mesh-type pop filter O2 or the like is attached to the microphone, it is possible to generate a video in which the avatar C is drawn behind the mesh.

[0084] Although an example was given using multiple cameras, the most desirable camera angle is one that captures the widest area of ​​board B (polygon mesh BP). For example, angles such as those of real camera RC2 and virtual camera VC2, which capture from a direction perpendicular to the widest surface of board B, are preferable. The further the shooting direction deviates from the direction perpendicular to the widest surface of board B, i.e., the direction perpendicular to the surface of board B and polygon mesh BP as exemplified in Figure 4, the more unnatural the image becomes. Therefore, it is preferable that the shape of board B corresponds to the range of real microphone O perpendicular to the shooting direction from the camera. In this invention, the 2D image onto which the microphone image is projected is set only within the range of polygon mesh BP. This is because it uses the principle of anamorphosis, a psychological effect in which an object is perceived as having three-dimensionality when viewed from a certain direction, and the further it deviates from the desirable direction, the more this psychological effect breaks down. For example, if board B (polygon mesh BP) arranged as shown in Figure 4 is photographed from angles as shown in Figures 5 and 6, the projection surface of the microphone image on the polygon mesh BP will not be visible to the camera in the captured video.

[0085] <Examples of LE configurations and effects>

[0086] (1) In the video generation method for generating composite video in which an avatar C reflecting the motion of the performer and a real microphone O, which is a real object placed in a real studio (real space), are occluded, the video captured by a real camera RC that films the real studio is input to the composite device 2, and motion data tracking the body and facial movements of a singing performance in the capture studio is input to the composite device 2. In the scene construction unit 21 that constructs a virtual space within the composite device 2, in steps S203 to S205 of Figure 3, a 2D mask OM corresponding to the real microphone O and the video from the real camera RC are associated with a polygon mesh BP and placed so that the video from the real camera RC is projection-mapped (projected) onto the polygon mesh BP. In step S206 of Figure 3, when the 2D mask OM is turned on (first state), the real-world image is cut out in the shape of the real microphone O corresponding to the 2D mask OM, and the image of the real microphone O, associated with the depth information of the polygon mesh BP, is rendered (making it possible to display it as an image captured by the virtual camera VC). In step S207, the performer's motion data is reflected in the avatar C, and in step S208, the image with the mask OM and avatar C positioned, captured by the virtual camera VC, is rendered. After that, the rendered image from the virtual camera VC and the image from the real camera RC are combined and output to the switcher 3 and output screen 4. This makes it possible to place a mask corresponding to a predetermined object such as a microphone in the virtual space, and to display the image of the predetermined object corresponding to that mask, thereby enabling the output of a composite image of real-world image and 3DCG at low cost and in real time (immediately, instantaneously). Furthermore, the time axes of the video in the real world, the performer's motion, and the audio data (for example, the performer's singing) will be synchronized at the desired timing.

[0087] (2) In the embodiment described above, in steps S101, S103, and S104 of Figure 2, the coordinates of the real microphone O in the real studio and the coordinates of the standing position marker M (for example, the scan start position), which will be the standing position of the avatar, are input to the composite device 2, as identified by the scanner S, which is a three-dimensional information measuring device having a scanner function for measuring three-dimensional information of the real space. In the scene construction unit 21 of the composite device 2, in step S201 of Figure 3, the avatar C is placed in the virtual space so as to correspond to the coordinates of the standing position marker M, and in steps S203 to S205, the mask OM is placed in the virtual space so as to correspond to the coordinates of the real camera RC. As a result, the positional relationship between the object placed in the real space and the avatar can be identified, making it possible to represent occlusion and the like based on front-to-back information.

[0088] (3) In the above-described embodiment, based on the information identified by the scanner S in steps S103 and S104 of Figure 2, in step S203 of Figure 3, a polygon mesh BP is placed in the virtual space to correspond to the position and range of the real microphone O in the real space, and in step S205, a mask OM is set in the material information of the polygon mesh BP. This reduces the processing load of CG rendering, making it possible to generate in real time an image in which a predetermined object in the real space and the avatar are occluded.

[0089] (4) In the embodiment described above, in step S205 of Figure 3, the image from the real camera RC is associated with the polygon mesh BP, and the image is projection-mapped onto the polygon mesh BP. Furthermore, in step S206 of Figure 3, when the mask OM is turned on (first state), the portion of the real microphone O projected onto the polygon mesh BP, which is cut out within the range of the mask OM, is rendered and drawn as the image from the virtual camera VC, with the depth information of the polygon mesh BP associated with it. As a result, the object corresponding to the position and range of a predetermined object and the projected image in real space are cut out in the shape of the mask and displayed, making it possible to generate an image in which the predetermined object in real space and the avatar are occluded.

[0090] (5) In the embodiment described above, the polygon mesh BP is an object corresponding to the shape and coordinates of board B, which is placed in the virtual space in step S203 of Figure 3 to correspond to the position and range of real microphone O in real space, based on the scan data of board B attached to real microphone O measured by scanner S in steps S103 and S104 of Figure 2. This makes it possible to easily and quickly obtain clues to identify the position and range of a given object, enabling low-cost and rapid progress in video production.

[0091] (6) In the above-described embodiment, as illustrated in Figures 4 and 6, the board B has a shape that corresponds to the range of the real microphone O perpendicular to the shooting direction from the real camera RC when it is placed in a predetermined position and takes a picture. For example, it is a board that covers a range that approximately corresponds to the range of the real microphone O perpendicular to the shooting direction of the real camera RC2. This makes it possible to easily and quickly obtain clues to identify the position and range of a predetermined object, enabling low-cost and rapid progress in the video production field.

[0092] (7) In the embodiment described above, in step S102 of Figure 2, the scanner S, which has a camera position transmission function such as coordinates and orientation, transmits information that identifies the status of the real camera RC (position and orientation information, etc.) to the composite device 2 (for example, via a camera setting application, etc.) as explained with reference to Figure 5, etc., and in step S202 of Figure 3, based on the information of the real camera RC acquired by the scanner S, the information (status) of the virtual camera VC, such as lens and position, is set to be the same environment as the settings of the real camera RC. Also, as illustrated in Figures 5 and 6, the distance from the virtual camera VC to the polygon mesh BP and the distance from the real camera RC to the real microphone O are approximately the same. This makes it possible to create a realistic composite image by ensuring consistency between the image from the camera that captures the real space and the image from the virtual camera that captures the virtual space.

[0093] (8) In the embodiment described above, as illustrated in Figure 6(A), real cameras RC1 to RC4 are arranged in the real studio to photograph the real microphone O from different directions. In addition, in the virtual space, there are multiple types of masks OM, each generated for each image of the real microphone O captured by one type of real camera RC, onto which the images of the real microphone O captured by each real camera RC are projected, and virtual cameras VC1 to VC4 are arranged to photograph from directions corresponding to the shooting directions of the real cameras RC. For example, the image from real camera RC1 is composited with the image from the corresponding virtual camera VC1 (the image from real camera RC1 projected onto the polygon mesh BP will be cut out within the range of the mask OM captured by virtual camera VC1), and the image from real camera RC2 is composited with the image from the corresponding virtual camera VC2 (the image from real camera RC2 projected onto the polygon mesh BP will be cut out within the range of the mask OM captured by virtual camera VC2). The resulting composite image, as illustrated in output screen 4 of Figure 10, will be an image in which objects such as the real microphone O in real space are occluded from the avatar C. Furthermore, as illustrated in Figure 7(B), if a mask OM is generated for the parts that are occluded from the avatar C, then parts that are not occluded, such as the feet of the real microphone O, may be shown in the composite image as the feet in Figure 7(C) included in the first real camera shooting data in Figure 9 (for example, Figure 10(C)). In this way, images of a predetermined object taken from multiple angles are generated, and by switching and combining images from multiple angles, it is possible to have a psychological effect on the viewer of the composite image and make the composite image feel more realistic.

[0094] (9) In the above embodiment, the mask OM is a mask cut out in a shape corresponding to the fine wiring of the real microphone O or the mesh of the pop filter. As a result, since the mask is shaped to correspond to a predetermined object, it is possible to perform occlusion with the avatar with high accuracy even for the delicate shape parts of the predetermined object. Furthermore, since the rendering processing burden is lighter than occlusion with the avatar using a 3D object corresponding to the three-dimensional shape of the predetermined object, it is possible to generate composite images at low cost and in real time on the spot.

[0095] (10) In the above embodiment, the mask OM is a mask generated based on the difference between the image before and after placing the real microphone O in the real studio, which is captured by the real camera RC. This makes it possible to generate a mask for producing an image in which the delicate shape parts of a given object and the avatar are occluded, at low cost and quickly.

[0096] (11) In the above embodiment, the video of the real microphone O displayed in correspondence with the mask OM is video captured in real time from the real camera RC at the same time as the projection onto the polygon mesh BP, and the motion of the avatar C reflects the motion data of the performer that is captured (tracked) in real time at the same time as the real camera RC takes the picture. As a result, the real space and the performer's movements are synthesized in real time at the same time, which can enhance the sense of realism that viewers feel towards the synthesized video. In addition, the audio data emitted by the performer, the movements of the avatar, and the real video are synchronized at a desirable timing, so stable video can be provided to viewers not only in recordings but also in live streaming and other situations.

[0097] <Variation> The following are some examples of modifications to the embodiments described above.

[0098] (Regarding camera settings apps) In the above-described embodiment, referring to steps S101, S103, S104, etc. in Figure 2, mesh data of the real space and real objects scanned by the scanner S is directly transmitted to the composite device 2, and referring to step S102, etc. in Figure 2, camera information data is transmitted to the composite device 2 from a terminal (relay terminal) on which an application for recording and setting camera information is installed, based on the position and orientation information of the scanner S, etc. However, the invention is not limited to this, and mesh data of objects generated based on measurements by the scanner S may also be transmitted to the composite device 2 via the relay terminal. For example, a manager application (which may include a camera setting application) that implements a function to aggregate information acquired by the scanner S and transmit it to the composite device 2 may be incorporated into the relay terminal, enabling the transfer of data acquired by the scanner S.

[0099] Furthermore, the camera settings application may be implemented as an application that can be installed on the scanner S or other computers, not limited to a terminal different from the scanner S, and the camera settings application program may be executed on the scanner S. For example, if the same number of real cameras RC and scanners S are available, data may be transferred directly from the scanner S to the composite device 2. Alternatively, the camera settings application program may be executed on the same terminal as the composite device 2. For example, information from the scanner S may be sent directly to the composite device 2, and information such as the sensor size and focal length of the real camera RC's lens may be input by the operator of the composite device 2 as settings for the virtual camera VC (for example, in a camera settings application launched within the composite device 2 or in an application that constitutes the scene construction unit 21).

[0100] Furthermore, for example, if there is one scanner S and multiple real cameras RC, the camera information of all real cameras RC, aggregated in the camera settings app, may be transmitted to the composite device 2 via a separate terminal equipped with the camera settings app. For example, camera information can be aggregated in one place in the manager app of the relay terminal mentioned above. This makes it possible to synchronize data such as mesh data measured by the scanner S and camera information in one place using the manager app, without having to operate the screen of the scanner S. Specifically, this reduces the workload for scanning and synchronizing / transmitting each camera's information, allows for simultaneous recording of all camera information, and enables the synchronization mechanism to align the time axis of all camera information. Alternatively, multiple terminals with the camera settings app installed may be used for operation, and data may be managed by grouping cameras. For example, camera settings app A may be used to manage the group of real cameras RC1 to RC3, and camera settings app B may be used to manage the group of real cameras RC4 to RC6. Furthermore, communication may be established between camera settings app A and camera settings app B to enable synchronization between different groups.

[0101] (Regarding Abata's position, etc.) In the embodiment described above, as explained in step S101 of Figure 2 and step S201 of Figure 3, an example was described in which the real space is scanned with the standing position marker M in the real studio as the origin, and this origin is set as the display start position of the avatar C. However, it is not limited to this, and as long as the coordinates of the real space and the coordinates of the virtual space can be aligned, the origin may be a different location from the standing position marker M of the avatar. Also, the display start position of the avatar may be set to a different location from the origin (for example, a predetermined distance from the origin). Regarding the angle from which the avatar is photographed, the embodiment above shows an example of an angle that shows the side of the avatar, but the avatar may be photographed from the front. Furthermore, the number of avatars (number of performers) is not limited to one person, and multiple avatars (each reflecting the actions of multiple corresponding performers) may be arranged.

[0102] (Regarding the location-identifying component) In the embodiment described above, there is one board B, and a single type of polygon mesh BP obtained by scanning this single board B is used. However, the invention is not limited to this, and the range of shooting angles may be further expanded by preparing multiple board Bs for each angle and arranging polygon meshes corresponding to each board B in the virtual space. However, it is preferable that the camera angle is fixed. If the camera angle is fixed, the steps S103 and S104 in Figure 2 and steps S203 to S205 in Figure 3 may be performed using multiple board Bs as time permits at each shooting location. Even in this case, the preparation work for real-time compositing can be performed on the same day at a lower cost and more quickly than using conventional virtual production methods.

[0103] Furthermore, the positioning component is not limited to a flat board; any component that is easy to attach to the target object and can scan its shape to a degree that provides a simple hint of the target object's coordinates is acceptable. Also, the object to be occluded is not limited to one; multiple objects may be targeted. For example, in addition to a real microphone, a music stand may be placed in the real space, polygon meshes and masks for the real microphone and the music stand may be generated, and these may be placed in the virtual space to be occluded with the avatar.

[0104] In the embodiment described above, because the shape of the real microphone O is complex, a board B, which serves as a positioning member that can be easily measured, was placed on the microphone, and the coordinates of the real microphone O were determined using the scan data of board B. However, this is not the only way; the polygon mesh that sets the mask (mask OM in the above embodiment) onto which the real image is projected may be positioned based on data scanned from a predetermined object, such as a microphone, that is the object to be projected onto. For example, if the shape of the object in real space is simple compared to the complex mesh, wiring, etc., of a laptop computer, a polygon mesh may be placed in the virtual space based on mesh data scanned from the object itself. Even in this case, by setting a 2D mask of the object's shape on the polygon mesh and projecting the real image, virtual production can be realized in real time in a simple manner with clean details of the outline.

[0105] (Regarding the mask used to project the image) In the embodiment described above, an example was explained in which the mask OM, which is a mask of the shape of the real microphone O, can be generated by taking the difference between the image before and after the placement of the real microphone O. The generation of the mask of the shape of the real microphone O using this difference may be automatically generated by an application that automatically generates masks from the difference between before and after images, or for complex shapes, it may be generated manually based on the difference. Furthermore, the mask is not limited to one generated using the difference; it may also be possible to automatically or manually select the range of the shape of the real microphone O from an image of the real microphone O, which is a real object, placed in real space, and generate a mask cut out with the shape of the real microphone O. From the viewpoint of time efficiency, it is desirable to adopt the method of generating a mask using the difference, but the latter method, which can cut out more finely depending on the detail of the object's shape, may also be adopted.

[0106] Note that the generation of the difference mask may be performed at a different time than the order shown in Figure 3. It is sufficient to do so after the placement position of the real microphone O has been determined. For example, the difference mask may be generated in advance by taking images of the background without the microphone and the image with the microphone before the board B is attached to the microphone in step S103 of Figure 2.

[0107] In the embodiment described above, an example was explained in which a difference-based mask generated in step S204 of Figure 3 is created for each camera, as illustrated in Figure 6. However, this is not the only option; instead of creating masks for all cameras, masks based on angles from some cameras may be used. For example, one type of mask generated from the video footage from the angle that captures the widest range corresponding to board B among the angles used to film the real microphone O (for example, real camera RC2 in Figure 6) may be copied and placed in all scene construction units 21.

[0108] (Regarding the projection of images from the Real Camera RC) In the embodiment described above, an example was explained in which, by setting the image of a real camera RC to a polygon mesh BP and performing projection mapping (projecting a real image onto the polygon mesh), when the mask OM is turned on (first state), the image corresponding to the portion of the real microphone O within the image of the real camera RC, which is associated with the depth information of the polygon mesh, can be rendered (displayed) as the image of the virtual camera VC, corresponding to the range of the mask OM. However, the embodiment is not limited to this, and the image of the real camera RC may be projected onto the mask OM associated with the polygon mesh BP, so that when the mask OM is turned on, the depth information of the polygon mesh BP is associated with the mask OM, and the image of the portion of the real microphone O projected onto the mask OM can be rendered as the image of the virtual camera VC.

[0109] (Regarding the scanning procedure) The flowchart sequence described in Figures 2 and 3, as described in the above-described embodiment, is not limited to this sequence, as long as the environments of the real space and the virtual space are consistent, a mask OM corresponding to the shape of the real object is generated, the image of the real space is projected onto the polygon mesh BP, and the mask OM can cut out the object. For example, after the scanning and 3D data identification information of the position identification member in steps S103 and S104 of Figure 2 is transmitted to the composite device 2, the step of transmitting camera information in step S102 to the composite device 2 may be performed. Alternatively, for example, when scanning the entire real studio area in step S101, a board B may be attached to the real microphone O in advance, and the mesh data of the entire studio and the mesh data of board B for setting the mask may be transmitted to the composite device 2 simultaneously. Specifically, board B may be attached to the real microphone O in advance, the mesh of board B may be marked as a specific mask, and both the mesh of the entire real space (immovable data) and the mesh of the mask of board B (data that can be moved and fine-tuned) may be transmitted to the composite device 2 simultaneously.

[0110] In other words, the information measured in Figure 2 (mesh data scanned from the entire real space, information on the origin in the real space, mesh data identifying the first object placed in the real space, and camera information for real space photography) only needs to be transmitted to the composite device 2, and is not limited to the order described above. Using the information acquired in Figure 2, it is desirable to align the origin in step S201 as explained in Figure 3, and then in step S202, set up the virtual camera based on the origin information set in step S201. Preferably, in order to reduce the scanning process by utilizing the difference, it is desirable to perform a scan of the studio first, and then attach board B to the real microphone O and perform the scan.

[0111] (Regarding the settings for real cameras and virtual cameras) In the embodiment described above, the real camera RC that captures the real space is a video camera without sensors to determine the camera's position, orientation, etc., and the position and orientation of the real camera RC are measured by a scanner S. Based on the information measured by the scanner S (for example, via a camera setting application that inputs lens information), a virtual camera VC in the virtual space is set. However, this is not limited to this, and a camera equipped with a camera tracking sensor may be used and synchronized with the virtual camera. This makes it possible to capture images while moving both the real camera and the virtual camera without fixing the camera position. Regarding the angle between the plane (or a straight line parallel to it) of the 2D mask in the virtual space and the shooting direction, orthogonal (90 degrees) provides the greatest efficiency, but this is not limited to this, and any angle within the range that does not cause deformation of the object projected onto the polygon mesh BP is acceptable. Adjustments may be made as appropriate depending on the shape of the projected object. Furthermore, the projected image may be an image captured from a fixed camera separate from the camera that captures the image to be composited with the virtual camera image, or it may be an image captured from a moving real camera. Even when camera tracking is possible, the load of rendering data in the virtual space is reduced.

[0112] In the embodiments described above, we have explained examples in which the scanner S measures information such as the position and orientation of the real camera, and the camera status information, including the camera position information and lens information input by the operator to the camera setting application, is set in the virtual camera, or in which the camera setting application is installed in the scanner S and the camera status information, including lens information, is transmitted from the scanner S to the composite device 2. However, we are not limited to these examples, and the scanner S may be fixed to the real camera used for shooting to enable camera tracking, and as mentioned above, it may be possible to move the real camera and virtual camera without fixing the camera position to enable shooting.

[0113] (Regarding the application scenarios of this invention) In the embodiments described above, an example of applying the virtual production method of the present invention was described as a scene in which avatar C, reflecting the motion of a performer, sings. However, it is not limited to this, and may also be applied to scenes such as a video introducing avatar C's favorite items or a product description by avatar C. Even if the shape of the real-world favorite item or product is not limited to a simple shape but is a complex shape, the virtual production method of the present invention can be used to generate a mask in the virtual space onto which the image of the object is projected, thereby creating a spatial relationship between avatar C and the real object in the virtual space. This makes it possible to occlude the real object and avatar C in real time and at low cost, and to generate an image that makes it appear as if avatar C is actually present in the real space. Examples of objects with complex shapes that can be scanned include, but are not limited to, the real microphones mentioned above, in-ear monitors, round objects such as stuffed animals, objects that tend to be high-polygon, objects that are difficult to scan because the laser beam passes through them, such as glass bottles, and objects that are difficult to scan because the laser light is reflected, such as mirrors.

[0114] [Examples of implementation using software] The various control blocks of the control unit in the computer of the composite device 2 and other devices described above may be implemented by logic circuits (hardware) formed on integrated circuits (IC chips), or by software using a CPU (Central Processing Unit). When implemented by software using a CPU, the computer equipped with the control unit includes a CPU that executes instructions for a program which is software that realizes each function, a ROM (Read Only Memory) or storage device (collectively referred to as a "recording medium") on which the program and various data are recorded in a readable format by the computer (or CPU), and a RAM (Random Access Memory) for loading the program. The object of the present invention is achieved when the computer (or CPU) reads the program from the recording medium and executes it. As the recording medium, a "tangible medium that is not temporary," such as tape, disk, card, semiconductor memory, or programmable logic circuit, can be used. The program may also be supplied to the computer via any transmission medium capable of transmitting the program (such as a communication network or broadcast wave). One aspect of the present invention can also be realized in the form of a data signal embedded in a carrier wave, in which the program is embodied by electronic transmission.

[0115] A specific example of implementation using a terminal (device) on which a computer program (application software) for executing the video generation method in the above-described embodiment is installed will be explained. This program executes processes to realize all or part of the virtual production described in Figure 3 based on information acquired from the real camera RC, scanner S, and capture acquisition device 5. For example, the program for executing the video generation method is installed on a computer corresponding to the composite device 2. The control unit that executes this program sets the display start position of the avatar C (by aligning it with the origin of the virtual space, etc.) based on the coordinate information of the real space acquired by the scanner S, and executes the process of placing the avatar C at the display start position (corresponding to step S201 in Figure 3). It also executes the process of setting the virtual camera VC in accordance with the state of the real camera RC based on camera information identified based on the self-position estimation of the scanner S, etc. Furthermore, it places a polygon mesh BP corresponding to board B in the virtual space based on the mesh data of the position identification member (board B) acquired from the scanner S. The process involves associating a 2D mask OM of the shape of the real microphone O, which is material data acquired and prepared by being cut out based on the difference from the background and stored in the memory unit beforehand, or by being automatically cut out and generated based on real-world video, with the polygon mesh BP (corresponding to step S205). The process involves projecting the video data acquired from the real camera RC onto the polygon mesh BP (corresponding to step S205), and when it is determined that the mask OM associated with the polygon mesh BP has been turned on, the polygon mesh BP corresponding to the mask (the polygon mesh BP with the video from the real camera RC projected onto it) is rendered as the video from the virtual camera VC with depth information (corresponding to step S206). In addition, the process involves rendering the video of the virtual space VC in which the avatar C, which reflects the motion data, is displayed, compositing it with the video from the real camera RC, and generating video data that can be output to an output device (such as the display unit of the composite device 2 or the switcher 3) (corresponding to steps S207 and S208).Furthermore, it may be possible for the operator to edit and adjust some of the object data within the virtual space. For example, the placement position of the virtual camera VC in step S202, the placement position of the polygon mesh BP in step S203, and the placement position of other avatars C may be fine-tuned by the operator.

[0116] The embodiments disclosed herein should be considered in all respects to be illustrative and not restrictive. The scope of this invention is indicated by the claims rather than by the foregoing description, and all modifications within the meaning and scope equivalent to the claims are intended to be included. [Explanation of Symbols]

[0117] 1 Video generation system, 2 Composite device, 21 Scene construction unit, 3 Switcher, 4 Output screen, 5 Capture acquisition device, 6 Audio acquisition device, S Scanner, RC Real camera, VC Virtual camera, O Real microphone, OM Mask, B Board, BP Polygon mesh, C Avatar, M Standing position marker

Claims

1. An image generation device that generates an image in which an avatar occludes with a predetermined object placed in real space, A means for acquiring images of the real space from a camera that photographs the real space, Means for acquiring the motion of the performer playing the aforementioned avatar, A means for generating a virtual space in which an avatar that operates with acquired motion and a mask corresponding to the predetermined object generated based on acquired real-world video are arranged, and the video of the predetermined object in the real-world video can be displayed in accordance with the mask, A video generation device comprising means for outputting a video obtained by combining the video of the real space with video from a virtual camera that captures the virtual space.

2. An image generation system that generates an image in which an avatar occludes with a predetermined object placed in real space, Means for acquiring images of the aforementioned real space, Means for acquiring the motion of the performer playing the aforementioned avatar, A means for generating a virtual space in which an avatar that operates with acquired motion and a mask corresponding to the predetermined object generated based on acquired real-world video are arranged, and the video of the predetermined object in the real-world video can be displayed in accordance with the mask, A video generation system comprising means for outputting a video obtained by combining the video of the real space with video from a virtual camera that captures the virtual space.

3. A method for generating images in which an avatar occludes with a predetermined object placed in real space, The steps include acquiring images of the real space from a camera that photographs the real space, The steps include obtaining the motion of the performer who plays the aforementioned avatar, The steps include: creating a virtual space in which an avatar that operates with acquired motion and a mask corresponding to the predetermined object generated based on acquired real-world video are arranged, and the video of the predetermined object in the real-world video can be displayed in accordance with the mask; A video generation method comprising the steps of outputting a video obtained by combining a video of the real space with a video from a virtual camera that captures the virtual space.

Citation Information

Patent Citations

  • Display system for hall and method for executing event using the same

    JP2023130363A