Three-dimensional video synthesis output system and three-dimensional video synthesis output method

The system integrates a shooting direction identification unit and synthesis output unit to combine 3D images from a volumetric capture system with 3D character images in real time, achieving synchronized video and audio output in a shared virtual space.

JP2025159459APending Publication Date: 2025-10-21BALUS CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
JP2024062021
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-04-08
Publication Date
2025-10-21

AI Technical Summary

Technical Problem

Existing technologies do not effectively combine 3D images generated using a volumetric capture system with images of other 3D characters in real time.

Method used

A system comprising a shooting direction identification unit, a volumetric image receiving unit, and a synthesis output unit that synthesizes a 3D image with another 3D character image in real time, utilizing a video synthesis output device, a volumetric video generation device, and a virtual character generation device connected via a network.

Benefits of technology

Enables real-time synthesis of a 3D image generated using a volumetric capture system with another 3D character image, allowing for synchronized video and audio output in a shared virtual space.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025159459000001_ABST
    Figure 2025159459000001_ABST
Patent Text Reader

Abstract

To provide a mechanism for synthesizing a volumetric video with another 3D video in real time.SOLUTION: A video synthesis output system comprises: a shooting direction specifying unit for specifying a virtual shooting direction of a 3D character generated in a virtual space; a volumetric video receiving unit for receiving a volumetric video, which is a 3D video corresponding to the virtual shooting direction; and a synthesis output unit for outputting a video obtained by synthesizing a video of the 3D character with the volumetric video.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present invention relates to a 3D video synthesis output system and a 3D video synthesis output method. [Background technology]

[0002] Japanese Patent Application Publication No. 2022-557875 (Patent Document 1) is a background technology in this technical field. This publication describes mesh-tracking-based dynamic 4D modeling for machine learning deformation training, which uses a volumetric capture system for high-quality 4D scans, establishes temporal correspondence across 4D scanned human face and whole-body mesh sequences using mesh tracking, establishes spatial correspondence between the 4D scanned human face and whole-body meshes and a 3D CG physics simulator using mesh registration, and uses machine learning to train surface deformation as deltas from the physics simulator (see Abstract). [Prior art documents] [Patent documents]

[0003] [Patent Document 1] Patent Application No. 2022-557875 Summary of the Invention [Problem to be solved by the invention]

[0004] The patent document 1 describes the generation of dynamic face and whole-body modeling including implicit deformation by machine learning (ML), which is a combination of any novel movement and natural deformation of facial expression or body language. However, the patent document does not consider the combination of 3D images generated using a volumetric capture system with images of other 3D characters in real time. Therefore, the present invention provides a mechanism for synthesizing, in real time, a three-dimensional image generated using a volumetric capture system with an image of another three-dimensional character. [Means for solving the problem]

[0005] In order to solve the above problems, for example, the configurations described in the claims are adopted. The present application includes multiple configurations and methods for solving the above problems, and one example is characterized by comprising a shooting direction identification unit that identifies a virtual shooting direction of a three-dimensional character generated in a virtual space, a volumetric image receiving unit that receives a volumetric image, which is a three-dimensional image corresponding to the virtual shooting direction, and a synthesis output unit that outputs an image that is a synthesis of the image of the three-dimensional character and the volumetric image. [Effects of the Invention]

[0006] According to the present invention, it is possible to provide a mechanism for synthesizing a three-dimensional image generated using a volumetric capture system with another three-dimensional character image in real time. Problems, configurations, and effects other than those described above will become apparent from the following description of the embodiments. [Brief explanation of the drawings]

[0007] [Figure 1] FIG. 1 is a diagram illustrating an example of the configuration of a video composition output system 10. As shown in FIG. [Figure 2] FIG. 2 shows an example of the hardware configuration of the video synthesizing output device 200. As shown in FIG. [Figure 3] FIG. 3 shows an example of the hardware configuration of a volumetric video generation device 300. [Figure 4] FIG. 4 shows an example of the hardware configuration of the virtual character generation device 400. [Figure 5] FIG. 5 is an example of a functional block diagram 500 of the video synthesizing output device 200. [Figure 6]FIG. 6 is an example of a functional block diagram 600 of the volumetric image generation device 300. [Figure 7] FIG. 7 is an example of a functional block diagram 700 of the virtual character generation device 400. [Figure 8] FIG. 8 is an example of a video synthesis output flow 800. [Figure 9] FIG. 9 is an example of a video delay processing flow 900. [Figure 10] FIG. 10 shows an example of a distribution video 1000 output from the video synthesizing output device 200. DETAILED DESCRIPTION OF THE INVENTION

[0008] Hereinafter, an embodiment will be described with reference to the drawings. [Example]

[0009] FIG. 1 is a diagram illustrating an example of the configuration of a video composition output system 10. As shown in FIG. The video synthesis output system 10 comprises a video operation terminal 250, a volumetric video generation device 300, and a virtual character generation device 400, each of which is connected to a video synthesis output device 200 via a network 100. The network 100 may be wired or wireless, and each terminal can send and receive information via the network.

[0010] Each of the image synthesis output device 200, volumetric image generation device 300, and virtual character generation device 400 of the image synthesis output system 10 may be, for example, a portable terminal (mobile terminal) such as a smartphone, tablet, mobile phone, or personal digital assistant (PDA), or a wearable terminal such as glasses, a wristwatch, or clothing. It may also be a stationary or portable computer, or a server located on the cloud or a network. It may also function as a VR (Virtual Reality) terminal, an AR (Augmented Reality) terminal, or an MR (Mixed Reality) terminal. Alternatively, it may be a combination of multiple of these terminals. For example, a combination of one smartphone and one wearable terminal can logically function as a single terminal. It may also be any other information processing terminal.

[0011] Each of the image synthesis output device 200, volumetric image generation device 300, and virtual character generation device 400 of the image synthesis output system 10 includes a processor that executes an operating system, applications, programs, etc., a main storage device such as RAM (Random Access Memory), an auxiliary storage device such as an IC card, hard disk drive, SSD (Solid State Drive), or flash memory, a communication control unit such as a network card, wireless communication module, or mobile communication module, an input device such as a touch panel, keyboard, mouse, voice input, or input based on motion detection captured by a camera unit, and an output device such as a monitor or display. Note that the output device may also be a device or terminal that transmits information to be output to an external monitor, display, printer, or other device.

[0012] The main memory stores various programs, applications, etc. (modules), and the processor executes these programs and applications to realize the various functional elements of the overall system. These modules may be implemented in hardware, such as by integration. Each module may be an independent program or application, or may be implemented as a subprogram or function within a single integrated program or application.

[0013] In this specification, each module is described as the entity (subject) that performs the processing, but in reality, the processing is carried out by a processor that processes various programs, applications, etc. (modules). Various databases (DBs) are stored in the auxiliary storage device. A "database" is a functional element (storage unit) that stores a set of data so that it can accommodate any data manipulation (e.g., extraction, addition, deletion, overwriting, etc.) from a processor or an external computer. There are no limitations on how the database is implemented; for example, it can be a database management system, spreadsheet software, or a text file such as XML or JSON.

[0014] The video synthesis output device 200 synthesizes and outputs video and audio data sent from the volumetric video generation device 300, the virtual character generation device 400, etc. In this embodiment, the video synthesis output device 200 itself may also be called a video synthesis output system. The video operation terminal 250 is a terminal device used by an operator to perform operations such as video and audio synthesis processing, rendering processing, camerawork operation, and camera switching in the video synthesis output device 200.

[0015] The volumetric video generation device 300 generates 3D images and videos based on images and videos captured by multiple cameras. As shown in Fig. 1, a volumetric capture system 350 equipped with, for example, eight cameras (A to H) is connected to the volumetric video generation device 300.

[0016] The eight cameras (A to H) installed in the volumetric capture system 350 are set up so that they can simultaneously capture an object α to be photographed. The object α may be one or more people or an object such as equipment.

[0017] 1 shows eight cameras A to H as the multiple cameras, but the number is not limited to this. The volumetric capture system 350 may be installed in a room or studio (referred to as a "volumetric studio") for performing volumetric capture. For example, in a volumetric studio, more cameras can be installed on the ceiling, walls, floor, etc.

[0018] In the volumetric capture system 350 of this embodiment, a preview monitor 360 that displays the image sent from the image synthesis output device 200 is installed at a position that is visible from the subject α.

[0019] The virtual character generation device 400 includes a motion capture system 450 . The motion capture system 450 uses, for example, three cameras to capture the image of the subject β (the performer of the virtual character's movements) and captures the movement of the subject β. At this time, the subject β is photographed wearing a suit equipped with motion capture sensors, for example. The motion capture system 450 obtains the movement of the subject β in the virtual space by capturing the position and changes of the sensors.

[0020] Virtual character generation device 400 generates (also referred to as "reconstructing") three-dimensional data based on the captured data (hereinafter referred to as "motion capture data"). Virtual character generation device 400 generates a video of a three-dimensional character (character video information) based on this three-dimensional data. The process of generating character video information will be described later.

[0021] With the above configuration, the video synthesis output system 10 can generate a video in which a three-dimensional video (volumetric video) of a real performer, generated using a volumetric capture system, and a video of a virtual character are synthesized in real time within the same virtual space.

[0022] Each component will be described in detail below. FIG. 2 shows an example of the hardware configuration of the video synthesizing output device 200. As shown in FIG. The video composition output device 200 is configured, for example, by a server located on the cloud.

[0023] The video synthesis output device 200 comprises a main memory device 201, an auxiliary memory device 202, a processor 203, an input device 204, an output device 205, and a transmission / reception unit 206, and these components are interconnected via a communication bus to communicate data, control information, etc. with each other.

[0024] The main memory device 201 stores programs and applications such as an image acquisition module 211, a shooting direction identification module 212, a camera position information transmission module 213, a volumetric image receiving module 214, an image delay processing module 215, a synthesis output module 216, an audio acquisition module 217, a background image processing module 218, a real-time image acquisition module 219, and an audio delay processing module 220, and the processor 203 executes these programs and applications to realize each functional element of the image synthesis output device 200.

[0025] The auxiliary storage device 202 includes a background data storage unit 221 that stores background data and the like, and an application storage unit 222 that stores application programs and the like executed by the processor 203. For example, the auxiliary storage device 202 is configured by a data storage device such as a hard disk drive or an SSD (Solid State Drive).

[0026] Image acquisition module 211 acquires virtual character image information and the like transmitted from virtual character generation device 400. The virtual character information includes three-dimensional data and three-dimensional images of a three-dimensional character generated in a virtual space. Details of the virtual character information will be described later.

[0027] The shooting direction identification module 212 identifies the virtual shooting direction of the three-dimensional character. Based on the image of the three-dimensional character sent from the virtual character generation device 400, the shooting direction identification module 212 identifies the shooting direction of the three-dimensional character by the virtual camera.

[0028] When a virtual camera photographs a three-dimensional character in a virtual space, the virtual camera can photograph the three-dimensional character from any angle in the virtual space. For example, the photographing direction specification module 212 specifies the position of the virtual camera relative to the three-dimensional character and the direction of the virtual camera's point of gaze relative to the three-dimensional character, based on the displayed image of the three-dimensional character.

[0029] The shooting direction specification module 212 may specify the angle from which the virtual camera is shooting the three-dimensional character, the angle of view of the virtual camera, the distance from the virtual camera, and the like. The shooting direction specification module 212 detects camerawork operations performed by an operator on an image of a three-dimensional character. Based on the detected operation information, the shooting direction specification module 212 may specify, for a virtual camera that shoots the three-dimensional character, the position of the virtual camera relative to the three-dimensional character (camera position), the direction of the virtual camera's point of gaze relative to the three-dimensional character in the virtual space (shooting direction), and the like.

[0030] For example, the operator can change the shooting angle of the three-dimensional character, i.e., the position and shooting direction of the virtual camera, by performing camerawork operations using the video operation terminal 250. The shooting direction identification module 212 identifies the shooting direction of the three-dimensional character based on information indicating the shooting direction of the virtual camera (camerawork operation information) sent from the video operation terminal 250.

[0031] The imaging direction identification module 212 identifies, from among the multiple cameras (A to H) of the volumetric capture system 350, the camera whose imaging direction corresponds to the identified imaging direction of the three-dimensional character. For example, suppose the identified shooting direction of the three-dimensional character is from 45 degrees behind and to the left of the character. On the other hand, if the front direction of the shooting target α of the volumetric capture system 350 (for example, the direction of shooting camera E) and the front direction of the three-dimensional character are the same direction, the camera corresponding to the identified shooting direction of the three-dimensional character (shooting direction from 45 degrees behind and to the left) is shooting camera B.

[0032] Furthermore, for example, if the operator uses the video operation terminal 250 to change the shooting direction of the three-dimensional character from "45 degrees behind the left" counterclockwise to "45 degrees behind the right," "45 degrees in front of the right," or "front," the corresponding camera in the volumetric capture system 350 also changes counterclockwise from "shooting camera B" to shooting cameras A, H, G, F, and E.

[0033] The camera position information transmission module 213 transmits the camera position information generated by the imaging direction identification module 212 to the volumetric image generation device 300 .

[0034] The volumetric image receiving module 214 receives volumetric image, which is a three-dimensional image corresponding to a virtual shooting direction. Specifically, the volumetric image receiving module 214 receives, from the volumetric image generation device 300, volumetric image that uses, as texture, image data captured by a shooting camera identified based on the camera position information.

[0035] The video delay processing module 215 delays and outputs the video of the character video information acquired by the video acquisition module 211. The delay processing will be described in detail later.

[0036] The synthesis output module 216 outputs an image obtained by synthesizing an image of a 3D character with a volumetric image. Furthermore, the synthesis output module 216 generates an image obtained by synthesizing a plurality of different input images and sounds. The processing details of the synthesis output module 216 will be described later.

[0037] The audio acquisition module 217 receives audio signals sent from a sound source, an audio output device, etc. to the video synthesis output device 200. The audio signals are, for example, recorded voices of virtual characters.

[0038] The background image processing module 218 generates a background image based on the background data stored in the background data storage unit 221. For example, it generates a backstage image.

[0039] The real-time image acquisition module 219 acquires preview image captured by a preview image capture camera. The preview image capture camera may be, for example, one of multiple cameras provided in the volumetric capture system 350. The image captured by the preview image capture camera is, for example, real-time image captured of the target α.

[0040] The audio delay processing module 220 delays the input audio by an audio delay time and outputs the delayed audio. For example, the audio delay processing module 220 delays the audio signal acquired by the audio acquisition module 217 and sends the delayed audio signal to the synthesis output module 216. The process of delaying the audio signal will be described later.

[0041] Next, the volumetric image generation device 300 will be described. FIG. 3 shows an example of the hardware configuration of a volumetric video generation device 300. The volumetric image generation device 300 can generate three-dimensional images based on the imaging data captured by the volumetric capture system 350.

[0042] The volumetric image generation device 300 comprises a main memory device 301, an auxiliary memory device 302, a processor 303, an input device 304, an output device 305, and a transceiver unit 306, and these components are interconnected via a communication bus to exchange data, control information, etc.

[0043] The main memory device 301 stores programs and applications such as a shooting data acquisition module 311, a camera position information acquisition module 312, a volumetric image generation module 313, and a volumetric image transmission module 314, and the processor 303 executes these programs and applications to realize each functional element of the volumetric image generation device 300.

[0044] The auxiliary storage device 302 includes an imaging data storage unit 321 that stores imaging data captured by the imaging cameras A to H of the volumetric capture system 350, and an application storage unit 322. For example, the auxiliary storage device 302 is configured by a data storage device such as a hard disk drive or an SSD (Solid State Drive).

[0045] The imaging data acquisition module 311 acquires imaging data captured by the volumetric capture system 350. For example, the imaging data acquisition module 311 acquires imaging data captured by the imaging cameras A to H and stores the data in the imaging data storage unit 321.

[0046] The camera position information acquisition module 312 acquires the camera position information transmitted from the camera position information transmission module 213 of the image synthesis output device 200. The camera position information acquisition module 312 sends the acquired camera position information to the volumetric image generation module 313.

[0047] The volumetric video generation module 313 generates a volumetric video based on the shooting data acquired by the shooting data acquisition module 311 and the camera position information acquired by the camera position information acquisition module 312 . The volumetric video generation process will be described in detail later.

[0048] The volumetric image transmission module 314 transmits the volumetric image generated by the volumetric image generation module 313 to the image synthesis output device 200 .

[0049] Next, the configuration of the virtual character generation device 400 will be described. FIG. 4 shows an example of the hardware configuration of the virtual character generation device 400.

[0050] The virtual character generation device 400 can generate a three-dimensional character image based on motion capture data captured by the motion capture system 450 (FIG. 1).

[0051] The virtual character generation device 400 comprises a main memory device 401, an auxiliary memory device 402, a processor 403, an input device 404, an output device 405, and a transmission / reception unit 406. These components are interconnected via a communication bus, and communicate data, control information, and the like with each other.

[0052] The main memory device 401 stores programs and applications such as a motion capture data acquisition module 411, a character operation information acquisition module 412, a character video information generation module 413, and a character video information transmission module 414, and the processor 403 executes these programs and applications to realize each functional element of the virtual character generation device 400.

[0053] Auxiliary storage device 402 has a character asset data storage unit 421 that stores character asset data, which is data related to the virtual character to be generated, and an application storage unit 422 that stores application programs and the like executed by processor 403. For example, auxiliary storage device 402 is configured by a data storage device such as a hard disk drive or an SSD (Solid State Drive).

[0054] The motion capture data acquisition module 411 acquires motion capture data captured by the motion capture system 450. The motion capture data includes, for example, three-dimensional data and time data that indicates time-series changes in the three-dimensional data.

[0055] The character operation information acquisition module 412 acquires operation information and the like of the virtual character input from outside the virtual character generation device 400. Details of the operation information of the virtual character will be described later.

[0056] The character video information generation module 413 generates character video information based on the motion capture data, operation information of the virtual character, and the like.

[0057] The character video information transmission module 414 transmits the character video information generated by the character video information generation module 413 to the video synthesis output device 200 . The character video information will be described in detail later.

[0058] Next, the functional blocks of the video synthesizing output device 200 will be described with reference to FIG. FIG. 5 is an example of a functional block diagram 500 of the video synthesizing output device 200.

[0059] Image acquisition module 211 acquires character image information sent from virtual character generation device 400. Image capture direction identification module 212 identifies the virtual image capture direction of the virtual character based on the character image displayed based on the character image information.

[0060] Furthermore, when the video operation terminal 250, for example, operates the camera work in the displayed character video, the video acquisition module 211 may determine the shooting direction of the virtual camera that shoots the virtual character based on the operation information from the video operation terminal 250.

[0061] The shooting direction identification module 212 generates camera position information based on the identified shooting direction. The camera position information is, for example, information specifying the shooting directions of the multiple cameras provided in the volumetric capture system 350, which corresponds to the virtual shooting direction of the virtual character identified by the shooting direction identification module 212.

[0062] The camera position information transmission module 213 may transmit the generated camera position information to the volumetric image generation device 300 .

[0063] The volumetric image receiving module 214 receives the volumetric image sent from the volumetric image generating device 300. The volumetric image receiving module 214 may input the received volumetric image to the synthesis output module 216.

[0064] The audio acquisition module 217 acquires audio data such as audio from an external sound source, character audio which is the audio of a character video, or recorded audio data. This audio data is also input to the virtual character generation device 400 at the same time. The audio acquisition module 217 inputs the acquired audio data to the audio delay processing module 220 .

[0065] The video delay processing module 215 delays the character video (video of a three-dimensional character) by a set delay time (video delay time) and outputs the delayed character video (referred to as "delayed character video") to the synthesis output module 216.

[0066] The video delay processing module 215 may calculate the delay time based on, for example, the difference between the timestamp of the character video and the timestamp of the volumetric video, which will be described later. Generating a volumetric video requires time, for example, for generating a 3D model from the image data captured by multiple cameras A to H, and for applying textures of the image data captured from specific camera directions to this 3D model. Therefore, a delay of about 3 to 4 seconds occurs between the time the subject α is captured and the time the generated volumetric video is actually sent to the video composition output device 200.

[0067] The video delay processing module 215 acquires the timestamps assigned to the real-time video (real-time video) of the subject α captured by the volumetric capture system 350 and the timestamps assigned to the real-time video of the subject β captured by the motion capture system 450, and calculates the delay time by, for example, comparing the timestamps of the same captured scene.

[0068] The video delay processing module 215 may calculate the delay time based on the difference between the time from the start of generation of the volumetric video until the volumetric video is generated (volumetric video generation time) and the character video generation time from the start of generation of the character video until the character video is generated.

[0069] The video delay processing module 215 may delay the character video by a predetermined length of time required for generating the volumetric video (for example, about 4 seconds), or may adjust the delay time according to the timestamps and video offsets of the same shooting scene.

[0070] The audio delay processing module 220 performs audio delay processing to delay the audio data input from the audio acquisition module 217 by an audio delay time. The audio delay processing module 220 outputs the delayed audio (delayed audio) to the synthesis output module 216.

[0071] The audio delay processing unit may use, as the audio delay time, the video delay time calculated by the video delay processing module 215. Furthermore, the audio delay processing module 220 may delay the audio data by, for example, a length of a preset volumetric video generation time (for example, about 4 seconds).

[0072] The background image processing module 218 generates and outputs a background image of a character image based on the background data stored in the background data storage unit 221. For example, the background image processing module 218 generates a virtual stage image or the like based on the background data and outputs it to the synthesis output module 216.

[0073] In this embodiment, the background image processing module 218 that generates the background image is provided in the image synthesis output device 200, but the present invention is not limited to this. For example, the virtual character generation device 400 may be provided with background data and a background image processing module, and may generate a virtual image including a background image.

[0074] The composite output module 216 includes a rendering output processing module 510 that acquires delayed character images, virtual stage images, and volumetric images and generates a composite image (referred to as a "virtual composite image") by performing rendering processing on these images.

[0075] The rendering output processing module 510 constructs (reconstructs) a three-dimensional space based on, for example, three-dimensional virtual character data contained in the character image (delayed character image), three-dimensional background data contained in the virtual stage image, and three-dimensional volumetric data contained in the volumetric image. Furthermore, the rendering output processing module 510 may re-render, for example, delayed character images, virtual stage images, etc., so that the lighting, brightness, etc. of the images to be rendered have a unified texture, even if the images have been rendered in advance.

[0076] The rendering output processing module 510 generates a virtual composite image by rendering the reconstructed three-dimensional space data, thereby generating an image in which a three-dimensional virtual character image, a virtual stage image, and a volumetric image are composited in the same virtual space.

[0077] Furthermore, the synthesis output module 216 outputs an image that is a composite of real-time image captured by one of multiple cameras for generating a volumetric image and an image of a non-delayed three-dimensional character. Specifically, the composite output module 216 outputs a real-time image captured by one of the multiple cameras provided in the volumetric capture system 350 to the preview monitor 360. For example, the composite output module 216 outputs a fixed-point composite image of the subject α in almost real time, which is an image different from the above-mentioned virtual composite image, with almost no delay or with less delay (low delay) than the virtual composite image, to the preview monitor 360.

[0078] This allows the subject α of the volumetric capture system 350 and the subject β of the motion capture system 450 to view a fixed-point composite image with almost no delay on the preview monitor, so subjects α and β can perform while checking each other's movements, line of sight, posture, direction, facial expression, etc. For example, subject β can perform a motion performance while checking on the preview monitor 460 the real-time movements, line of sight, posture, direction, facial expression, etc. of subject α, who is performing in a different room or location.

[0079] The synthesis output module 216 includes an audio synthesis processing module 520 that synthesizes delayed audio input from the audio delay processing module 220 with the generated virtual synthesized video. The synthesis output module 216 may distribute the video with synthesized audio over a communication network such as the Internet via an existing video distribution system, or may output the video to a video recorder or other output device. The composite output module 216 may output and distribute the composited image as 2D image data for display on a flat panel screen or the like, or may output and distribute it as 3D image data for display on VR glasses or the like.

[0080] Next, the functional blocks of the volumetric image generation device 300 will be described. FIG. 6 is an example of a functional block diagram 600 of the volumetric image generation device 300.

[0081] The volumetric video generation module 313 includes a three-dimensional data generation processing module 601 that generates three-dimensional data based on multiple pieces of imaging data acquired via the imaging data acquisition module 311. The three-dimensional data generated by the three-dimensional data generation processing module 601 is, for example, three-dimensional mesh data.

[0082] The volumetric image generation module 313 includes a shooting data processing module 602 that extracts shooting data captured by a specific camera from the shooting data acquisition module 311 . The shooting data processing module 602 identifies the shooting camera in the volumetric capture system 350 based on the camera position information acquired from the camera position information acquisition module 312 .

[0083] For example, if the shooting direction of the virtual character specified by the camera designation information is 45 degrees to the left (45 degrees behind and to the left) from the front (directly behind) of the virtual character, the shooting data processing module 602 specifies "shooting camera B" as the camera in the volumetric capture system 350 (see Figure 1) with the same shooting direction as the shooting direction of the virtual character (45 degrees behind and to the left).

[0084] Furthermore, the photographing data processing module 602 extracts photographing data photographed by the identified camera (e.g., "photographing camera B"). The photographing data processing module 602 acquires, for example, the photographing data photographed by the identified photographing camera B from the photographing data acquisition module 311.

[0085] Furthermore, the imaging data processing module 602 may perform a process of rotating the three-dimensional data, which is the generated three-dimensional mesh data, so that the identified imaging direction becomes the front.

[0086] Here, the imaging data processing module 602 attaches the acquired imaging data to the 3D mesh data as a texture. For example, the imaging data processing module 602 attaches the imaging data extracted from the imaging data acquisition module 311 to the front surface of the 3D mesh data displayed by rotation. This allows the shooting data processing module 602 to generate a volumetric image (referred to as a "volumetric image with a specific texture") in which the shooting data extracted based on the camera position information is pasted as texture onto the 3D mesh data generated by the 3D data generation processing module 601.

[0087] The volumetric video transmission module 314 transmits the volumetric video of the specific texture to the video synthesis output device 200. Note that the volumetric video of the specific texture may include a time code (time stamp) set for each frame image that constitutes the volumetric video.

[0088] Next, the functional blocks of the virtual character generation device 400 and the processing flow of the character video information generation module 413 will be described. FIG. 7 is an example of a functional block diagram 700 of the virtual character generation device 400.

[0089] First, the motion capture data acquisition module 411 acquires motion capture data, which is data captured by the motion capture system 450.

[0090] The character image information generation module 413 generates a three-dimensional character image based on the motion capture data, and the character operation information acquisition module 412 acquires operation information (character operation information) for the generated character image.

[0091] The character operation information includes, for example, audio data input from outside virtual character generation device 400, motion operation information for controlling the movement and facial expression of the character image, facial expression operation information, lighting operation information for controlling lighting, etc. The character operation information also includes character asset data stored in character asset data storage unit 421 of auxiliary storage device 402.

[0092] Here, the character information generation flow of the character video information generation module 413 will be described.

[0093] The character video information generation module 413 includes a character reconstruction processing module 711 that reconstructs the motion capture data acquired by the motion capture system 450 to reconstruct a three-dimensional character (also simply referred to as a “character”).

[0094] Furthermore, the character image information generation module 413 includes a character operation control processing module 712 that performs operation control processing on the three-dimensional character based on the character operation information.

[0095] The character operation control processing module 712 controls the movement of the character's lips and the movement of the facial expression assets based on, for example, the input lip-sync audio and facial expression operation information, thereby outputting a video in which the character speaks and changes its facial expression in time with the lip-sync audio.

[0096] Furthermore, the character operation control processing module 712 controls the lighting of the character rendered by the character reconstruction processing module 711. The character operation control processing module 712 controls the lighting of the three-dimensional character in response to, for example, the movement of the character and changes in the lighting for the character.

[0097] Furthermore, the character video information generation module 413 has a character video output processing module 713 that outputs character video information including operation information that controls the character's movements, facial expressions, and lighting, character asset data, and timestamp information corresponding to character operation control.

[0098] The character video information transmission module 414 transmits the character video information output from the character video information generation module 413 to the video synthesis output device 200 .

[0099] Next, the video output flow of the video synthesis output device 200 will be described. FIG. 8 is an example of a video synthesis output flow 800.

[0100] The image acquisition module 211 acquires character image information from the virtual character generation device 400 (S810).

[0101] The shooting direction specification module 212 detects camerawork operations performed by the video operation terminal 250 on the virtual character (VC) video (S812). The shooting direction identification module 212 identifies the shooting direction of the virtual character image based on the detected camerawork operation (S814).

[0102] In addition, even if no camerawork operation on the character image by the video operation terminal 250 is detected in S812, the shooting direction identification module 212 can identify the shooting direction in S814 based on the display state of the character in the virtual character image.

[0103] The shooting direction identification module 212 generates camera position information that identifies the shooting camera of the volumetric capture system 350 that corresponds to the identified shooting direction (S816), and transmits the camera position information to the volumetric image generation device 300 (S818).

[0104] Next, the volumetric image receiving module 214 acquires a volumetric image (referred to as a "volumetric image with a specific texture") onto which the texture of the shooting data from the identified shooting camera has been applied (S820), and the image delay processing module 215 performs delay processing of the virtual character image (S822).

[0105] The virtual character image delay process of S822 may be started at any timing after the above-mentioned S810. For example, the virtual character image delay process of S822 may be started simultaneously with the transmission of camera position information by the camera position information transmission module 213 of S818.

[0106] Next, the composite output module 216 renders and composites the volumetric image of the specific texture, the delayed virtual character image, and the virtual stage image to generate a virtual composite image (S824).

[0107] The synthesis output module 216 acquires the voice data of the virtual character (S826), and furthermore, the synthesis output module 216 acquires the voice data of the volumetric imaging target (S828).

[0108] The synthesis output module 216 delays the acquired audio data of the virtual character and the audio data of the volumetric imaging target, synthesizes the delayed audio data with the virtual synthetic image (S830), and outputs the synthesized image (S832).

[0109] The virtual character video delay process (S822) will be described below. FIG. 9 is an example of a video delay processing flow 900.

[0110] The video delay processing module 215 acquires the timestamp of the character video and the timestamp of the volumetric video (S910).

[0111] Next, the video delay processing module 215 calculates the delay time of the timestamp of the character video relative to the timestamp of the volumetric video based on the reference timestamp of the video synthesis output device 200 (S920).

[0112] The video delay processing module 215 delays the input character video by the calculated delay time and outputs the delayed character video to the synthesis output module 216 (S930). As a result, the synthesis output module 216 receives synchronized character video and volumetric video.

[0113] FIG. 10 shows an example of a distribution video 1000 output from the video synthesizing output device 200. In the distributed video 1000, a virtual character video 1010 generated by the virtual character generation device 400 and volumetric videos 1020 and 1030 of two artists generated by the volumetric video generation device 300 are displayed side by side.

[0114] In the distributed video 1000, volumetric images 1020 and 1030 of real people are loaded and composited almost in real time into a three-dimensional space including a virtual character image 1010 delayed by about four seconds and a studio image 1040 as its background image. As a result, the distributed video 1000 is a video in which the virtual character image 1010 and the image of the real person appear together in real time in the same three-dimensional space.

[0115] The present invention is not limited to the above-described embodiments and includes various modifications. For example, the above-described embodiments have been described in detail to clearly explain the present invention, and the present invention is not necessarily limited to those including all of the described configurations. Furthermore, it is possible to replace part of the configuration of one embodiment with the configuration of another embodiment, or to add the configuration of another embodiment to the configuration of one embodiment. Furthermore, it is possible to add, delete, or replace part of the configuration of each embodiment with other configurations.

[0116] Furthermore, the above-described configurations, functions, processing units, processing means, etc. may be partially or entirely implemented in hardware, for example, by designing them as integrated circuits. The above-described configurations, functions, etc. may also be implemented in software, with a processor interpreting and executing a program that implements each function. Information such as programs, tables, and files that implement each function can be stored in a memory, a recording device such as a hard disk or SSD (Solid State Drive), or a recording medium such as an IC card, SD card, or DVD. Various programs may also be stored in portable media.

[0117] In addition, the control lines and information lines shown are those that are considered necessary for the explanation, and do not necessarily show all the control lines and information lines in the product. In reality, it can be assumed that almost all components are interconnected. The above-described embodiments disclose at least the configurations described in the claims. [Explanation of symbols]

[0118] 10...Video synthesis output system, 100...Network, 200...Video synthesis output device, 250...Video operation terminal, 300...Volumetric video generation device, 350...Asset registration management module, 400...Virtual character generation device, 450...Motion capture system

Claims

1. a shooting direction specifying unit that specifies a virtual shooting direction of a three-dimensional character generated in a virtual space; a volumetric image receiving unit that receives a volumetric image that is a three-dimensional image corresponding to the virtual imaging direction; a synthesis output unit that outputs an image obtained by synthesizing the image of the three-dimensional character and the volumetric image; A video synthesis output system comprising:

2. a video delay processing unit that delays the video of the three-dimensional character by a video delay time and outputs the delayed video of the three-dimensional character; The video composition output system according to claim 1 , wherein the composition output unit outputs an image obtained by combining the delayed three-dimensional character image and the volumetric image.

3. 3. The video synthesis output system according to claim 2, wherein the video delay processing unit calculates the video delay time based on a difference between a timestamp of the video of the three-dimensional character and a timestamp of the volumetric video.

4. 3. The video synthesis output system according to claim 2, wherein the video delay processing unit calculates the video delay time based on a difference between a volumetric video generation time from the start of generation of the volumetric video until the volumetric video is generated and a three-dimensional character video generation time from the start of generation of the three-dimensional character video until the three-dimensional character video is generated.

5. 3. The video synthesis output system according to claim 2, wherein the synthesis output unit outputs a video synthesized from a real-time video captured by one of a plurality of cameras for generating the volumetric video and a non-delayed video of a three-dimensional character.

6. an audio delay processing unit that outputs the audio by delaying the input audio by an audio delay time; the audio delay processing unit uses the video delay time calculated by the video delay processing unit as the audio delay time, The video synthesis output system according to claim 2 , wherein the synthesis output unit outputs a video synthesized with the delayed audio.

7. 7. The video synthesis output system according to claim 1, wherein the synthesis output unit outputs an image in which the image of the three-dimensional character and the volumetric image are synthesized in the same virtual space.

8. specifying a virtual shooting direction of the three-dimensional character generated in the virtual space; receiving a volumetric image that is a three-dimensional image corresponding to the virtual imaging direction; outputting an image obtained by combining the image of the three-dimensional character and the volumetric image; A video synthesis output method comprising:

9. further comprising a step of delaying the image of the three-dimensional character by an image delay time to output the delayed image of the three-dimensional character; 9. The image synthesis output method according to claim 8, further comprising outputting an image obtained by synthesizing the delayed image of the three-dimensional character and the volumetric image.

10. The video synthesis output method according to claim 9 , further comprising the step of calculating the video delay time based on a difference between a timestamp of the video of the three-dimensional character and a timestamp of the volumetric video.

11. 10. The video synthesis output method according to claim 9, further comprising a step of calculating the video delay time based on a difference between a volumetric video generation time from the start of generation of the volumetric video until the volumetric video is generated and a three-dimensional character video generation time from the start of generation of the video of the three-dimensional character until the video is generated.

12. 10. The image synthesis output method according to claim 9, further comprising a step of outputting an image obtained by synthesizing a real-time image captured by one of the plurality of cameras for generating the volumetric image with an image of a non-delayed three-dimensional character.

13. a step of delaying the input audio by an audio delay time and outputting the audio; using the video delay time as the audio delay time; The video synthesis output method according to claim 9 , further comprising the step of: outputting a video synthesized with the delayed audio.

14. The image synthesis output method according to any one of claims 8 to 13, wherein the image output in the step of outputting the synthesized image is an image in which the image of the three-dimensional character and the volumetric image are synthesized in the same virtual space.

Citation Information

Patent Citations

  • Volumetric Capture and Mesh Tracking Based Machine Learning

    JP2023519846A