Image generation device, image generation method, and program

By synchronizing and coordinating 3D models from volumetric and motion capture methods within a common virtual space, the image generating device addresses the challenge of subject coordination, resulting in coherent and natural-looking virtual viewpoint images.

JP2025126679APending Publication Date: 2025-08-29CANON KK
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
JP2024023037
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-02-19
Publication Date
2025-08-29

AI Technical Summary

Technical Problem

Existing technologies struggle to generate natural-looking virtual viewpoint images when combining 3D models captured in multiple physically separate spaces due to the inability to coordinate subjects' positions and movements effectively.

Method used

An image generating device that synchronizes and coordinates 3D models generated using volumetric and motion capture methods within a common virtual space, ensuring subjects move coherently by matching spatial coordinate systems and prioritizing placement based on overlapping bounding boxes.

Benefits of technology

Enables the generation of desired virtual viewpoint images with coordinated subjects, reducing incongruity and enhancing the natural appearance of combined 3D models.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025126679000001_ABST
    Figure 2025126679000001_ABST
Patent Text Reader

Abstract

To reduce the sense of incongruity when a plurality of 3D models is generated using a plurality of methods and the 3D models are combined to generate one virtual viewpoint image.SOLUTION: An image generation device obtains a plurality of first three-dimensional (3D) models for each of a plurality of subjects, generated through a first method on the basis of shooting of a predetermined region including the plurality of subjects, and a second 3D model that is based on posture information of a specific subject, among the plurality of subjects, present in the predetermined region during the shooting, the second 3D model corresponding to the specific subject and being generated through a second method different from the first method, and outputs a virtual viewpoint image generated on the basis of the first 3D model of a subject, among the plurality of subjects, that is different from the specific subject, and the second 3D model corresponding to the specific subject.SELECTED DRAWING: Figure 4
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present disclosure relates to techniques for generating virtual viewpoint images using multiple cameras. [Background technology]

[0002] A technology called volumetric capture, which can generate a three-dimensional (3D) model of a subject from images captured using multiple imaging devices, has been attracting attention. The generated 3D model of the subject is placed in a virtual space that recreates (or simulates) the subject's shooting environment and is observed using a virtual camera that can be operated within the virtual space. This allows the virtual camera, whose viewpoint and viewing direction can be freely set in the virtual space, to be used to observe images that appear as if the subject were captured by a physical camera in a real shooting environment. The images observed by this virtual camera are called virtual viewpoint images. Since virtual viewpoint images are acquired by specifying an arbitrary viewpoint and angle of view within the virtual space, they can reproduce images that are difficult to capture in real life using a physical camera. Furthermore, a technology called motion capture is used as a shooting method for generating 3D models. In this technology, skeletal information of the photographed subject is acquired, and a predetermined CG (Computer Graphics) model is added to the skeletal information to generate a 3D model. [Prior art documents] [Patent documents]

[0003] [Patent Document 1] Japanese Patent Publication No. 2022-060058 Summary of the Invention [Problem to be solved by the invention]

[0004] Patent Document 1 describes a technology that generates a single virtual viewpoint image by combining 3D models generated using a common volumetric capture method in multiple physically separate spaces, such as a stadium and a studio. For example, a portion of a 3D model obtained by shooting in one space is placed in another space to generate a virtual viewpoint image within that space. However, with this technology, because shooting is performed in multiple different spaces, it is not possible to fully coordinate the positions and movements of the subjects present in each of the multiple spaces with each other. Therefore, when multiple 3D models are generated in different spaces and then combined to generate a single virtual viewpoint image, the generated virtual viewpoint image may appear unnatural.

[0005] The present disclosure provides a technique for generating a desired virtual viewpoint image while sufficiently coordinating subjects with each other. [Means for solving the problem]

[0006] An image generating device according to one embodiment of the present disclosure includes an acquisition means for acquiring a plurality of first three-dimensional (3D) models for each of a plurality of subjects generated by a first method based on an image of a predetermined area including the plurality of subjects, and a second 3D model for some of the plurality of subjects generated by a second method different from the first method; and an output means for outputting a virtual viewpoint image based on a virtual space in which the first 3D models generated for subjects among the plurality of subjects that are not the targets for which the second 3D models are generated are placed, and in which the second 3D models are placed without the first 3D models generated for subjects that are the targets for which the second 3D models are generated. [Effects of the Invention]

[0007] According to the present invention, a desired virtual viewpoint image can be generated while sufficiently coordinating subjects with each other. [Brief explanation of the drawings]

[0008] [Figure 1] FIG. 1 illustrates an example of the configuration of an image forming system. [Figure 2] FIG. 2 is a diagram illustrating an example of the arrangement of cameras and sensors. [Figure 3] FIG. 2 illustrates an example of the hardware configuration of a rendering device. [Figure 4] FIG. 2 is a diagram illustrating an example of the functional configuration of each device in the system. [Figure 5] FIG. 10 is a diagram illustrating an example of recorded data. [Figure 6] FIG. 10 is a diagram illustrating an example of a bounding box. [Figure 7] FIG. 10 is a diagram illustrating an example of a processing flow for generating a virtual viewpoint image. [Figure 8] FIG. 10 is a diagram illustrating an example of an output virtual viewpoint image. [Figure 9] FIG. 10 is a diagram illustrating another example of the functional configuration of each device in the system. [Figure 10] FIG. 10 is a diagram illustrating another example of data to be recorded. [Figure 11] FIG. 10 is a diagram illustrating an example of a processing flow of tracking control. [Figure 12] FIG. 10 is a diagram illustrating another example of the flow of processing for generating a virtual viewpoint image. [Figure 13] FIG. 10 is a diagram illustrating an example of an output virtual viewpoint image. DETAILED DESCRIPTION OF THE INVENTION

[0009] Hereinafter, embodiments will be described in detail with reference to the accompanying drawings. Note that the following embodiments do not limit the scope of the invention claimed. Although multiple features are described in the embodiments, not all of these multiple features are necessarily essential to the invention, and multiple features may be combined arbitrarily. Furthermore, in the accompanying drawings, the same reference numerals are used to designate the same or similar components, and redundant explanations will be omitted.

[0010] The following describes a method for generating multiple three-dimensional (3D) models using two methods, volumetric capture and motion capture, and combining the 3D models to generate a single virtual viewpoint image. In this system, a first 3D model generated using the volumetric capture method and a second 3D model created by adding a predetermined CG model to a bone model generated using motion capture can coexist in a single virtual space. CG stands for Computer Graphic. This allows for the generation of compelling content in which the first 3D model, which realistically represents a subject, and the virtual second 3D model move together. The CG model may be, for example, an avatar. The CG model is used in multiple scenes, and by linking with the bone model, the shape of the CG model is deformed to generate a second 3D model suited to the scene. While the CG model is described as being generated solely by a computer without capturing a real-world subject, the present disclosure is not limited thereto. Instead of a CG model, a model generated by photographing a real-world subject using multiple cameras using photogrammetry technology may be used, or a model generated in this way may be further processed. A second 3D model may be generated by linking such a model with a bone model.

[0011] In this embodiment, to reduce the sense of incongruity in the virtual viewpoint image generated in this manner, a single space is simultaneously captured using two methods: volumetric capture and motion capture. Two or more subjects exist in the single space, and 3D models of the two or more subjects are generated using the volumetric capture method, while 3D models of some of the subjects are generated using the motion capture method. By having multiple subjects whose 3D models are generated using multiple methods exist in a common, predetermined area, the multiple subjects can communicate with each other and work in coordination. In this case, for example, if a 3D model of a single subject is generated using both volumetric capture and motion capture, multiple 3D models will exist in spatially overlapping positions, which may create a sense of incongruity in the generated virtual viewpoint image. Therefore, this embodiment provides a technique for reducing such incongruity.

[0012] (System Configuration) 1 shows an example of the configuration of an image generation system for generating virtual viewpoint images in this embodiment. The image generation system of this embodiment is configured to include two imaging systems: a volumetric capture system 101 and a motion capture system 121. The volumetric capture system 101 performs imaging to generate a 3D model of a subject using a volumetric capture method, and the motion capture system 121 performs imaging to generate a bone model of the subject using a motion capture method. The image generation system further includes a time server 141, a storage device 142, a rendering device 143, and an information display device 144.

[0013] The time server 141 outputs a time code (time information) to the volumetric capture system 101 and the motion capture system 121. That is, the volumetric capture system 101 and the motion capture system 121 can operate at synchronized timing using the time code output from the time server 141. In each of the volumetric capture system 101 and the motion capture system 121, the same time code is associated with a 3D model generated based on data captured at the same time.

[0014] The volumetric capture system 101 includes a camera group 111 including cameras 112-1 to 112-12, a hub 113, and a first object generation device 114. The cameras 112-1 to 112-12 belonging to the camera group 111 are arranged to surround a subject. The cameras 112-1 to 112-12 capture two-dimensional (2D) images, respectively. The cameras 112-1 to 112-12 capture images at synchronized timings, thereby capturing multi-viewpoint images. The 2D images captured by each of the multiple cameras 112-1 to 112-12 are provided to the first object generation device 114 via the hub 113, along with a time code provided by a time server 141. The first object generation device 114 uses a volumetric capture method to generate a 3D model of the subject included in the capture area, combines the 3D model with the time code, and outputs the 3D model to the storage device 142. The 3D model here is a shape model in which a texture based on a captured image is applied to the surface of three-dimensional shape data formed using a volumetric capture method. Note that volumetric capture system 101 may have a configuration different from that shown in FIG. 1. For example, cameras 112-1 to 112-12 may be directly connected to first object generation device 114, and a time code may be directly provided to first object generation device 114 from time server 141. Also, cameras 112-1 to 112-12 may be daisy-chained. Also, time server 141 may input a time code to cameras 112-1 to 112-12, and cameras 112-1 to 112-12 may output 2D images together with the time code. Also, first object generation device 114 may be composed of multiple devices.

[0015] The motion capture system 121 includes a sensor group 131 including sensors 132-1 to 132-6, a hub 133, and a second object generation device 134. The sensors 132-1 to 132-6 belonging to the sensor group 131 are arranged to surround a subject and identify (calculate) the two-dimensional positions of markers attached to the subject by capturing images using infrared light. The two-dimensional positions of the markers identified by each of the multiple sensors 132-1 to 132-6 are provided to the second object generation device 134 via the hub 113, along with a time code provided by a time server 141. The second object generation device 134 calculates the three-dimensional positions of the markers using a motion capture method based on the information on the two-dimensional positions of the markers for each sensor acquired from the sensors 132-1 to 132-6, and acquires posture information of the subject. The second object generation device 134 then generates a bone model of the subject using the calculated and acquired information. The second object generation device 134 generates a 3D model by associating a predetermined CG model with the bone model, and outputs the 3D model in combination with position information and a time code to the storage device 142. Note that the motion capture system 121 may have a configuration different from that shown in FIG. 1. For example, the sensors 132-1 to 132-6 may be directly connected to the second object generation device 134, and the time code may be directly provided to the second object generation device 134 from the time server 141. Alternatively, the sensors 132-1 to 132-6 may be daisy-chained. Alternatively, the time server 141 may input the time code to the sensors 132-1 to 132-6, and the sensors 132-1 to 132-6 may output the two-dimensional position information of the markers together with the time code. Alternatively, the second object generation device 134 may be composed of multiple devices.

[0016] In this embodiment, calibration is performed in each system so that the spatial coordinate system used in the volumetric capture system 101 and the spatial coordinate system used in the motion capture system 121 match. For example, assume that the coordinates of a specific point in the volumetric capture system 101 are expressed as (Xv, Yv, Zv), and the coordinates of the specific point in the motion capture system 121 are expressed as (Xm, Ym, Zm). In this case, if Xv = Xm, Yv = Ym, and Zv = Zm are satisfied, the real-space positions of the specific point indicated by both systems match. Calibration is performed so that the coordinates of at least all points in the shooting area match. Note that, although the spatial coordinate systems are matched by calibration in both systems here, this is not essential, and only the spatial coordinates of one system may be converted to coordinate values ​​in the spatial coordinate system of the other system.

[0017] The rendering device 143 is an image generating device that generates a virtual viewpoint image using a 3D model stored in the storage device 142 according to the position and direction of a virtual viewpoint set via a UI unit (not shown) operated by a user viewing the virtual viewpoint image. The UI unit has an operation unit such as a mouse, keyboard, operation buttons, and touch panel, and accepts operations by the user. The rendering device 143 generates a virtual viewpoint image that represents the appearance from the virtual viewpoint by arranging a 3D model acquired from the storage device 142 in a virtual space and rendering the arranged 3D model according to the viewpoint position and direction of a virtual camera. In other words, a virtual viewpoint image is generated that reproduces a scene observed by a virtual camera when the virtual camera is set to face a specific direction at a specific viewpoint position in the virtual space in which the 3D model is arranged. The virtual viewpoint image in this embodiment includes an arbitrary viewpoint image (virtual viewpoint image) corresponding to a viewpoint and line of sight direction arbitrarily specified by the user. Furthermore, the virtual viewpoint image in this embodiment may include an image corresponding to a viewpoint and a viewing direction specified by a user from among a plurality of candidate viewpoints and viewing directions, or an image corresponding to a viewpoint and a viewing direction automatically specified by the device. The rendering device 143 provides the generated virtual viewpoint image to the information display device 144, and the information display device 144 acquires and displays the virtual viewpoint image generated by the rendering device 143. Note that this is just one example, and for example, the rendering device 143 may store the virtual viewpoint image in the storage device 142, and the information display device 144 may perform control to read out and display the virtual viewpoint image from the storage device 142.

[0018] Although this embodiment shows an example in which 3D models are generated using both the volumetric capture method and the motion capture method, this is not limiting. In other words, the following discussion can be applied to a system that generates 3D models using two or more of any methods capable of generating 3D models.

[0019] FIG. 2 shows an example of the arrangement of cameras 112-1 to 112-12 and sensors 132-1 to 132-6.

[0020] In FIG. 2, twelve cameras 112-1 to 112-12 are arranged to surround a shooting area 201 to be shot. The cameras 112-1 to 112-12 each capture an image of the shooting area 201 from a different direction and output the captured image to generate a virtual viewpoint image by volumetric capture. The cameras 112-1 to 112-12 are, for example, digital cameras, and may be cameras that capture still images, cameras that capture moving images, or cameras that capture both still images and moving images. In this embodiment and the appended claims, the term "image" is used to include both still images and moving images unless otherwise specified. In this embodiment, an imaging device that combines an imaging unit with an image sensor and a lens that focuses light onto the image sensor is referred to as a "camera." The volumetric capture system 101 is configured to generate a virtual viewpoint image within the shooting area 201 using images captured by multiple cameras (cameras 112-1 to 112-12) included in the camera group 111. Although this embodiment shows an example in which 12 cameras are prepared, the number of cameras may be any plural number. Also, in the example of Fig. 2, an example in which cameras 112-1 to 112-12 surround the shooting area 201 from all directions is shown, but this is not limiting. For example, the cameras may be arranged only within a certain angular range centered on a point within the shooting area 201.

[0021] In FIG. 2, six sensors 132-1 to 132-6 are arranged to surround a photographing area 201. In order to generate a bone model by motion capture, the sensors 132-1 to 132-6 photograph the photographing area 201 from different directions and acquire two-dimensional coordinates of markers attached to the subject as motion data. The motion capture system 121 is configured to generate a bone model of a specific subject present in the photographing area 201 using images photographed by multiple sensors (sensors 132-1 to 132-6) included in the sensor group 131. Note that the "photography" by the sensors here can be performed in any format as long as it can detect the positions of markers attached to the surface of the subject. For example, the sensors may be configured to detect invisible markers using infrared sensors or to detect visible markers using cameras. Note that although the present embodiment illustrates an example in which six sensors are provided to surround the subject, the number of sensors may be any multiple number. 2 shows an example in which sensors 132-1 to 132-6 surround imaging area 201 from all directions, but this is not limiting. For example, sensors may be arranged only within a certain angular range centered on a point within imaging area 201.

[0022] Note that the volumetric capture system 101 may be used to generate the bone model, instead of the motion capture system 121. For example, a known method may be used that uses silhouette data extracted by separating the foreground region, which is the subject part, from the photographed image acquired by the cameras 112-1 to 112-12 and the background, which is the part other than the foreground region, and an initial pose model using bones. For example, the initial pose model is projected onto the selected photographed image as a temporary shape model, and the similarity between the projected region and the silhouette data is evaluated. In this case, the similarity increases as the shape (silhouette) of the foreground region corresponding to the pose of the subject is more similar to the initial pose of the bones, and decreases when the shape of the foreground region is less similar to the initial pose of the bones. Next, the initial pose model is deformed to update the temporary shape model, and the similarity between the region where the shape model is projected onto the photographed image and the silhouette is evaluated again. In this way, the evaluation of the similarity between the region where the shape model is projected onto the photographed image and the silhouette is repeatedly performed while updating the temporary shape model. Then, for example, the bone corresponding to the shape model with the highest similarity may be output as the bone model corresponding to the captured image. Note that when updating the shape model and evaluating the similarity, the bone corresponding to the shape model whose similarity exceeds a predetermined value for the first time may be output as the bone model corresponding to the captured image. Note that when a continuously moving subject is being photographed, the shape model evaluated to have the highest similarity or exceed a predetermined value with respect to the captured image one frame before may be used as the initial pose model used in evaluating the similarity in a specific frame.

[0023] Alternatively, two-dimensional pose estimation may be performed from each captured image by each camera, and three-dimensional pose estimation information may be determined (calculated) by combining the two-dimensional pose estimation information of the captured images from all cameras with the camera parameters of all cameras. In this case, a bone model is generated based on the three-dimensional pose estimation information.

[0024] Furthermore, a bone model may be generated using the three-dimensional shape generated by the volumetric capture system 101. For example, AI (Artificial Intelligence) may be used to estimate 17 body parts (bones) that form the human skeleton from the three-dimensional shape. In one example, a subject for which correct answer data for a bone model exists is photographed using a volumetric capture method to generate three-dimensional shape data, and machine learning may be performed using the three-dimensional shape data as input data and the correct answer data for the bone model as training data. In this case, after a trained model is obtained by the machine learning, the three-dimensional shape of the subject obtained using the volumetric capture method is input into the trained model, thereby outputting a bone model of the subject. The above-mentioned 17 bones are composed of the waist, abdomen, chest, neck, head, right upper arm, left upper arm, right forearm, left forearm, right hand, left hand, right thigh, left thigh, right shin, left shin, right foot, and left foot. However, the human body parts are not limited to these, and bones defined in more detail or more simply may be used.

[0025] 2, subject 202 is a subject for which a 3D model is generated by motion capture system 121, and markers are attached to the surface of subject 202 for detecting its position by motion capture. Subject 203 is a subject for which a 3D model is generated by volumetric capture system 101. Here, volumetric capture system 101 also generates a 3D model for subject 202. However, in the final video after rendering, the 3D model for subject 202 generated by the motion capture method is used, and the 3D model generated by the volumetric capture method is not used. This process will be described later.

[0026] (Device configuration) 3 shows an example of the hardware configuration of the rendering device 143. The rendering device 143 includes, for example, a CPU 301, a ROM 302, a RAM 303, an auxiliary storage device 304, a communication I / F 305, and a bus 306. Note that CPU stands for Central Processing Unit, ROM stands for Read Only Memory, and RAM 303 stands for Random Access Memory. Also, I / F stands for Interface. Note that other devices in the image generation system, such as the first object generation device 114 and the second object generation device 134, may also have a hardware configuration similar to that of FIG. 3.

[0027] The CPU 301 is a processor that controls the entire rendering device 143 using computer programs and data stored in one or more memories, such as the ROM 302 and the RAM 303. The rendering device 143 may have one or more dedicated hardware components different from the CPU 301, and at least a portion of the processing performed by the CPU 301 may be executed by the dedicated hardware components. Examples of the dedicated hardware include an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), and a digital signal processor (DSP). The CPU 301 is an example of a processor, and the rendering device 143 may include one or more processors, such as a microprocessing unit (MPU). The ROM 302 stores programs and parameters that do not require modification. The RAM 303 temporarily stores programs and data supplied from the auxiliary storage device 304 and data supplied from an external device via the communication I / F 305. The ROM 302 and the RAM 303 are examples of memories, and the rendering device 143 may include one or more memories. The auxiliary storage device 304 includes, for example, a hard disk drive and stores various content data such as images and audio. The communication I / F 305 is used for communication with external devices such as the camera group 111. For example, if the rendering device 143 is connected to an external device via a wired connection, a communication cable is connected to the communication I / F 305. If the rendering device 143 has a function for wireless communication with an external device, the communication I / F 305 includes an antenna. The bus 306 connects the various components of the rendering device 143 to transmit information. Note that if the rendering device 143 has a UI unit built in, the rendering device 143 will have a display unit and an operation unit in addition to the configuration shown in FIG. 3.

[0028] 4 shows an example of the functional configuration of the first object generation device 114, the second object generation device 134, and the rendering device 143 in the image generation system. The first object generation device 114 includes, for example, an image receiving unit 401 and a foreground extraction unit 402 for each camera. The first object generation device 114 also includes a three-dimensional shape estimation unit 403 and a first 3D model generation unit 404. The second object generation device 134 includes, for example, a sensor value receiving unit 411 and a two-dimensional marker position calculation unit 412 for each sensor. The second object generation device 134 also includes a bone model generation unit 413 and a second 3D model generation unit 414. The storage device 142 includes, for example, a saving unit 421. The rendering device 143 includes, for example, a 3D model position overlap determination unit 431, a 3D model placement control unit 432, and a rendering unit 433. Each function can be realized, for example, by the CPU 301 executing a program stored in the ROM 302, the auxiliary storage device 304, etc. Furthermore, at least some of the functions may be realized by dedicated hardware.

[0029] The image receiving unit 401 receives captured images acquired by each of the cameras 112-1 to 112-12. The image receiving unit 401 outputs the received captured images to the foreground extraction unit 402. The foreground extraction unit 402 generates a foreground image from the captured images acquired from the image receiving unit 401 and outputs the generated image to the three-dimensional shape estimation unit 403. For example, the foreground extraction unit 402 stores an image for each camera when no subject is present as a background image, and extracts a foreground area using the difference between the captured image by the camera and the background image to acquire the foreground image for each camera. The three-dimensional shape estimation unit 403 generates a three-dimensional shape of the subject included in the foreground portion based on the foreground image acquired from the images captured by the multiple cameras, and outputs the generated three-dimensional shape to the first 3D model generation unit 404. The three-dimensional shape of the subject is generated by a shape estimation method such as visual hull. The three-dimensional shape of the subject is, for example, three-dimensional shape data made up of a group of points, but is not limited to this and may be expressed in other formats such as polygons.

[0030] The first 3D model generation unit 404 generates a 3D model by adding texture data extracted as a foreground region to the acquired three-dimensional shape. The first 3D model generation unit 404 then combines the generated 3D model with information indicating the position of the 3D model and a time code acquired from the time server 141, and outputs the combined data to the storage unit 421. The information indicating the position of the 3D model uses a rectangular parallelepiped circumscribing the shape of the 3D model. This rectangular parallelepiped will be referred to as a bounding box hereinafter. The bounding box defines an area corresponding to the position in virtual space where the corresponding 3D model should be placed. The coordinates of each of the eight vertices of the bounding box are determined as follows using the maximum coordinates (Xmax, Ymax, and Zmax) and minimum coordinate values ​​(Xmin, Ymin, and Zmin) of each of the X, Y, and Z axes of the 3D model shape: Vertex 1 (Xmin, Ymin, Zmin) Vertex 2 (Xmax, Ymin, Zmin) Vertex 3 (Xmin, Ymax, Zmin) Vertex 4 (Xmax, Ymax, Zmin) Vertex 5 (Xmin, Ymin, Zmax) Vertex 6 (Xmax, Ymin, Zmax) Vertex 7 (Xmin, Ymax, Zmax) Vertex 8 (Xmax, Ymax, Zmax)

[0031] The sensor value receiving unit 411 receives infrared light sensor values ​​from each of the sensors 132-1 to 132-6 and outputs them to the marker two-dimensional position calculation unit 412. The marker two-dimensional position calculation unit 412 calculates the two-dimensional position of a marker attached to the subject at the sensor position for each sensor based on the infrared light sensor value acquired from the sensor value receiving unit 411, and outputs the calculated two-dimensional position to the bone model generation unit 413. The bone model generation unit 413 calculates the three-dimensional position of the marker by combining the two-dimensional marker position information for each sensor with information on the positional relationship between the sensors obtained in advance by calibration, and generates a bone model. The bone model generation unit 413 outputs the generated bone model data to the second 3D model generation unit 414. The second 3D model generation unit 414 generates a 3D model by adding a prepared CG model to the bone model data acquired from the bone model generation unit 413. Then, the second 3D model generation unit 414 combines the generated 3D model with bounding box coordinates, which are information indicating the position of the 3D model, and the time code obtained from the time server 141, and outputs the combined data to the storage unit 421.

[0032] The storage unit 421 records the 3D model information acquired from the first 3D model generation unit 404 and the 3D model information acquired from the second 3D model generation unit 414 and manages them as a 3D model database. Here, FIG. 5 shows an example of data recorded in the storage unit 421, taking the photographing of the subjects 202 and 203 shown in FIG. 2 as an example. "Timecode" in FIG. 5 is time code information output from the time server 141 and specifies time on a frame-by-frame basis. For example, if the system operates at 60 fps (frames per second) and one second of data is recorded, separate time code information is stored as "Timecode" for each of the 60 frames corresponding to that one second. "ID" is numbered for each 3D model and is identifier information used to specify the 3D model. "3D Model" is the generated 3D model data. "Kind" indicates the method (system) used to generate the 3D model. For 3D models generated using the volumetric capture system 101, the value "Volumetric" is held as the Kind value. On the other hand, for 3D models generated using the motion capture system 121, the value "Motion" is held as the Kind value. "Coordinates" indicates the spatial coordinates of the eight vertices of the bounding box corresponding to the 3D model. In this way, the storage unit 421 associates and manages time codes and 3D model information on a frame-by-frame basis.

[0033] FIG. 6 schematically illustrates the position (and range) in the spatial coordinate system of the bounding box of a 3D model corresponding to Timecode 12:15:30.010, among the data managed as in FIG. 5. Coordinate 601 is the coordinate of the center of the shooting area 201 on the floor surface, and has world coordinate values ​​(X, Y, Z) = (0, 0, 0), which is the reference position of the spatial coordinate system handled by the system. Bounding box 602 is a bounding box corresponding to the 3D model generated by shooting the subject 202 using the volumetric capture system 101. This bounding box is identified using the Coordinates value corresponding to ID = 1 in FIG. 5. Bounding box 603 is a bounding box corresponding to the 3D model generated by shooting the subject 203 using the volumetric capture system 101. This bounding box is identified using the Coordinates value corresponding to ID = 2 in FIG. 5. Bounding box 604 is a bounding box corresponding to a 3D model generated by capturing an image of subject 202 using motion capture system 121. This bounding box is identified using the value of Coordinates corresponding to ID=3 in FIG. 5. As shown in FIG. 6, a 3D model of subject 202 is generated using both volumetric capture system 101 and motion capture system 121. Because these 3D models are generated based on the same subject 202, the corresponding bounding boxes (bounding boxes corresponding to ID=1 and ID=3) overlap.

[0034] Returning to FIG. 4 , the 3D model position overlap determination unit 431 acquires 3D model information from the storage unit 421. Then, the 3D model position overlap determination unit 431 determines whether the 3D model generated by the motion capture system 121 overlaps with the 3D model generated by the volumetric capture system 101. Then, based on the determination result, the 3D model position overlap determination unit 431 determines whether to place each 3D model in the virtual space. The 3D model overlap determination is a process for identifying whether 3D models generated based on the same subject exist, and details thereof will be described later in the description of the process using FIG. 7 . For example, as described above, for the subject 202, the 3D model generated by the motion capture system 121 overlaps with the 3D model generated by the volumetric capture system 101. In this case, the 3D model placement control unit 432 may determine not to place either the 3D model generated by the motion capture system 121 or the 3D model generated by the volumetric capture system 101 in the virtual space. For example, the 3D model placement control unit 432 may determine not to place the 3D model of the subject 202 generated by the volumetric capture system 101 in the virtual space.

[0035] The 3D model placement control unit 432 places in the virtual space 3D models that the 3D model position overlap determination unit 431 has determined should be placed. On the other hand, the 3D model placement control unit 432 does not place in the virtual space 3D models that the 3D model position overlap determination unit 431 has determined not to place. For each time code, the 3D model placement control unit 432 places all of the 3D models that are associated with that time code and stored in the storage unit 421, thereby configuring a virtual space for generating a virtual viewpoint image. The rendering unit 433 renders the virtual viewpoint image in the virtual space configured by the 3D model placement control unit 432, according to the position and direction of the virtual viewpoint set by a user operation via a UI unit (not shown).

[0036] (Processing flow) Next, an example of the flow of processing for generating a virtual viewpoint image will be described with reference to FIG. 7. This processing is executed for each time code. In this processing, an example will be described in which a 3D model generated using the motion capture system 121 is preferentially placed in the virtual space. This is just one example, and a 3D model generated using the volumetric capture system 101 may also be preferentially placed. Which 3D model should be preferentially placed can be arbitrarily set by, for example, a user operation.

[0037] First, the 3D model position overlap determination unit 431 acquires 3D model information from the storage unit 421, checks the number of 3D models associated with the time code to be processed, and executes initialization processing for processing each 3D model (S701). Here, for example, the number of 3D models associated with the time code to be processed is N, and a counter n for counting the number of processed 3D models is set to 1.

[0038] Next, the 3D model position overlap determination unit 431 determines whether the Kind of the 3D model to be processed (3D model with ID=n) is Motion (S702). If the 3D model position overlap determination unit 431 determines that the Kind of the 3D model to be processed is Motion (YES in S702), it determines that the 3D model should be placed in the virtual space (S703). That is, in this embodiment, 3D models generated using the motion capture system 121 are placed preferentially, so if the Kind of the 3D model to be processed is Motion, it is determined that the 3D model should be placed in the virtual space. On the other hand, if the Kind of the 3D model to be processed is Volumetric (NO in S702), the 3D model position overlap determination unit 431 determines whether there is another 3D model with a Kind of Motion in a position overlapping with the 3D model (S704). Then, when the 3D model position overlap determination unit 431 determines that another 3D model whose Kind is Motion exists in a position overlapping with the 3D model to be processed (YES in S704), it determines not to place the 3D model to be processed in the virtual space (S705). In other words, when another 3D model is given priority for placement in the virtual space, a Volumetric 3D model that exists in a position overlapping with that 3D model is not placed in the virtual space. On the other hand, when the 3D model position overlap determination unit 431 determines that another 3D model whose Kind is Motion does not exist in a position overlapping with the 3D model to be processed (NO in S704), it decides to place the 3D model to be processed in the virtual space (S705).

[0039] Here, an example of a method for determining whether a 3D model to be processed overlaps with another 3D model will be described. First, it is determined whether a 3D model generated by the volumetric capture system 101 exists whose corresponding bounding box overlaps with the bounding box of a 3D model generated by the motion capture system 121. For example, it is determined whether at least one vertex of the bounding box of the 3D model generated by the volumetric capture system 101 exists within the bounding box of the 3D model generated by the motion capture system 121. If such a vertex exists, it is determined that the bounding boxes overlap. If such overlapping bounding boxes exist, the degree of overlap is then evaluated. In this evaluation, the total volume Vall of the smaller of the overlapping bounding boxes and the volume Vlap of the overlapping spatial region are identified, and the degree of overlap is indicated by a value Vlap / Vall indicating the proportion of the overlapping region. For example, if this value exceeds a predetermined level, it may be determined that the 3D model to be processed overlaps with the other 3D model. The predetermined level may be, for example, 0.8 (80%). If two bounding boxes overlap only in a small portion of the spatial region as a result of such a determination, the 3D models corresponding to those bounding boxes are not determined to be overlapping, and both 3D models may be placed in the virtual space. Furthermore, for example, since two bounding boxes based on photographic results of the same subject are expected to overlap sufficiently, it is possible to place only one of the two 3D models corresponding to those bounding boxes in the virtual space.

[0040] Even when multiple 3D models are generated based on the same subject, the degree of overlap may not be 1 (100%) due to the shape of the 3D models. For example, if a prepared CG model used to generate a 3D model using the motion capture system 121 is carrying an object on its back or holding an object in its hand, the bounding box is generated taking the object into consideration. On the other hand, if the subject photographed to generate the 3D model is not holding the object, the object is not reflected in the 3D model generated using the volumetric capture system 101. For this reason, the bounding box of the 3D model generated using the motion capture system 121 is specified to be larger by the size of the object than the bounding box corresponding to the volumetric capture system 101. On the other hand, for example, the CG model used to generate the 3D model using the motion capture system 121 may be smaller than the actual subject. In this case, it is assumed that the first bounding box corresponding to the motion capture system 121 is smaller than the second bounding box corresponding to the volumetric capture system 101. In this case, most of the first bounding box is contained in the second bounding box, but if an object such as that described above is present, the first bounding box may not overlap with the second bounding box. In such a case, by using the predetermined level described above, the discrepancy between the bounding boxes due to the shape of the 3D model is taken into consideration, and it is possible to place only one of the multiple 3D models in the virtual space.

[0041] Although the predetermined level is set to 0.8 in the above example, it is not limited to this. For example, the predetermined level may be determined based on the shape of the 3D model generated by the motion capture system 121. For example, in the example of FIG. 5, the height (length in the Z-axis direction) of the 3D model with ID=1 is 1800 as indicated by the Z value of Coordinates, while the height of the 3D model with ID=3 is 1600 (note that the unit may be, for example, millimeters). The X and Y values ​​of Coordinates vary depending on the degree of fullness or thickness of the clothing of the 3D model. In this way, the size of the bounding box varies depending on the shape of the CG model used in the motion capture system 121. Furthermore, when the 3D models generated by the volumetric capture system 101 and the motion capture system 121 are placed in exactly the same position, the degree of overlap of the bounding boxes also varies. Therefore, by determining the predetermined level according to the shape of the CG model used in the motion capture system 121, it becomes possible to appropriately determine whether multiple 3D models relate to the same subject.

[0042] When the 3D model position overlap determination unit 431 determines whether or not to place the 3D model to be processed in the virtual space, it increments the counter n to change the 3D model to be processed (S706). Then, the 3D model position overlap determination unit 431 determines whether the processes of S702 to S705 have been completed for all 3D models, that is, whether the counter n has reached N+1 (S707). If the counter n has not reached N+1 (NO in S707), the 3D model position overlap determination unit 431 returns the process to S702, since this means that unprocessed 3D models remain. On the other hand, if the counter n has reached N+1 (YES in S707), the 3D model position overlap determination unit 431 advances the process to S708. In S708, the 3D model placement control unit 432 places the 3D model determined to be placed in the virtual space. Then, the rendering unit 433 generates a virtual viewpoint image, which is an image obtained when the virtual space is observed from the position and direction of the set virtual viewpoint (S709), and the process ends.

[0043] FIG. 8 shows an example of a virtual viewpoint image generated in this manner. FIG. 8 is an example of a virtual viewpoint image generated using one frame of the data in FIG. 5 at a timecode of 12:15:30.010. Object 801 is a rendered 3D model stored corresponding to ID=2 in FIG. 5. This 3D model has a volumetric kind, but is placed in the virtual space as a result of the determination in S704 that the kind does not overlap with a 3D model with a motion kind. Object 802 is a rendered 3D model stored corresponding to ID=3 in FIG. 5. This 3D model has a motion kind, so is placed in the virtual space as a result of the determination in S702. On the other hand, the 3D model stored corresponding to ID=1 in FIG. 5 has a volumetric kind, and is not placed in the virtual space as a result of the determination in S704 that the kind of ID=3 overlaps with a 3D model with a motion kind. As a result, the object corresponding to this 3D model is not displayed in the virtual viewpoint image.

[0044] As described above, in this embodiment, multiple subjects are captured in a common capture space using different methods, and 3D models generated by each system are placed in a single virtual space to generate a virtual viewpoint image. By capturing images in a common space using different methods for generating 3D models, the multiple subjects present in the common space work in coordination while recognizing each other, resulting in corresponding 3D models also working in coordination. As a result, even if multiple 3D models are generated using independent systems, it is possible to generate a virtual viewpoint image in which the multiple 3D models are seamlessly arranged and work in coordination in the virtual space. Furthermore, for a subject whose 3D models are generated using multiple methods, only one of the multiple 3D models corresponding to that subject is displayed, and the other 3D models among the multiple 3D models are not placed in the virtual space. This prevents multiple 3D models generated from a common subject from being placed overlapping in the virtual space. This reduces the sense of incongruity in the output virtual viewpoint image.

[0045] (Variation) If an abnormality occurs when generating a 3D model using the motion capture system 121, an event may occur in which the virtual 3D model disappears and a 3D model of a realistic subject generated using the volumetric capture system 101 appears. In this modified example, a control method for preventing the output of an inappropriate virtual viewpoint image due to such an event will be described.

[0046] Fig. 9 shows an example of the functional configuration of a first object generation device 114, a second object generation device 134, and a rendering device 143 in an image generation system according to this modification. Among the configurations in Fig. 9, functions having the same functions as those in Fig. 4 are assigned the same reference numerals as those in Fig. 4, and detailed description thereof will be omitted.

[0047] The second 3D model generation unit 901 in this modification generates a 3D model based on bone model data acquired from the bone model generation unit 413, in the same manner as in FIG. 4 . When a 3D model is generated, the second 3D model generation unit 901 combines the 3D model with bounding box coordinates, which are information indicating the position of the 3D model, and a time code acquired from the time server 141, and outputs the combined data to the storage unit 421. On the other hand, when a 3D model is not generated, the second 3D model generation unit 901 outputs non-generation information indicating that the 3D model has not been generated to the tracking control unit 903. For example, a 3D model is not generated when an actor corresponding to a 3D model generated using the motion capture system 121 is no longer present in the imaging area 201. Furthermore, a 3D model is not generated when the actor is present in the imaging area 201 but information required for generating the 3D model is not input to the second 3D model generation unit 901 due to some abnormality, such as a sensor failure or a disconnection of the data transmission path.

[0048] FIG. 10 shows an example of data recorded in storage unit 421 when subjects 202 and 203 shown in FIG. 2 are photographed. At timecode 12:15:30.011, a 3D model with ID=3 and Kind=Motion is recorded. However, at timecode 12:15:30.012, the 3D model with ID=3 is no longer recorded. In this case, for example, in the process shown in FIG. 7, at timecode 12:15:30.012, the 3D model with ID=1 is placed in the virtual space instead of the 3D model with ID=3. However, if such a 3D model is displayed, the display will be changed from a virtual model to a realistic model, which may cause a sense of incongruity. For this reason, in this modified example, processing is performed to prevent 3D models corresponding to such realistic subjects from being placed in the virtual space.

[0049] When the second 3D model generation unit 901 no longer generates the 3D model, it outputs non-generation information of the 3D model to the tracking control unit 903. The 3D model position overlap determination unit 902 acquires 3D model information from the storage unit 421 and determines whether to place the 3D model in the virtual space by performing the above-described overlap determination and the like. The 3D model position overlap determination unit 902 then outputs information indicating whether to place the 3D model in the virtual space to the tracking control unit 903 and the 3D model placement control unit 904. When the tracking control unit 903 receives notification of non-generation information of the 3D model from the second 3D model generation unit 901, it determines a 3D model to be tracked and tracks the 3D model. The tracking control unit 903 outputs information indicating the tracking status, such as whether tracking is being performed, to the 3D model placement control unit 904. Here, the 3D model to be tracked is a 3D model generated by the volumetric capture system 101 for a subject whose 3D model is to be generated using the motion capture system 121. That is, when a 3D model generated using the motion capture system 121 exists, the 3D model that is not placed in the virtual space becomes the target of tracking. The 3D model placement control unit 904 performs processing to place the 3D model in the virtual space based on information determined by the 3D model position overlap determination unit 902 as to whether or not to place the 3D model and information indicating the tracking status by the tracking control unit 903.

[0050] Here, an example of the flow of processing executed by the tracking control unit 903 will be described with reference to FIG. 11. The tracking control unit 903 determines whether or not non-generation information of a 3D model has been acquired from the second 3D model generation unit 901 (S1101). If the tracking control unit 903 has not acquired non-generation information of a 3D model (NO in S1101), the tracking control unit 903 sets the tracking state to OFF (S1105) and ends the processing. On the other hand, if the tracking control unit 903 has acquired non-generation information of a 3D model (YES in S1101), the tracking control unit 903 determines whether or not a 3D model determined not to be placed by the 3D model position overlap determination unit 902 exists (S1102). For example, the tracking control unit 903 determines whether or not a 3D model determined not to be placed exists in the frame before the frame in which non-generation information of a 3D model was first acquired. The 3D model position overlap determination unit 902 may determine to place a 3D model generated by the volumetric capture system 101 when a 3D model generated by the motion capture system 121 no longer exists. Therefore, it may be determined whether or not a 3D model determined not to be placed exists in the frame prior to the frame that triggers acquisition of the 3D model non-generation information. Note that due to processing delays, etc., it is assumed that at the time when the second 3D model generation unit 901 outputs the 3D model non-generation information, the 3D model position overlap determination unit 902 is processing a frame prior to the frame in which the 3D model is not generated. In this case, the tracking control unit 903 may check whether or not a 3D model determined not to be placed by the 3D model position overlap determination unit 902 exists at the time when the non-generation information is acquired.

[0051] If the tracking control unit 903 determines that a 3D model that was determined not to be placed does not exist (NO in S1102), it sets the tracking state to OFF (S1105) and ends the processing. On the other hand, if the tracking control unit 903 determines that a 3D model that was determined not to be placed exists (YES in S1102), it sets the tracking state to ON (S1103). Then, the tracking control unit 903 performs tracking on the 3D model that was determined not to be placed (S1104). Note that the tracking control unit 903 may perform the processing of FIG. 11 for each frame, for example. Then, if the tracking control unit 903 detects that a notification of non-generation information of a 3D model has no longer been acquired from the second 3D model generation unit 901, it sets the tracking state to OFF and stops the tracking processing.

[0052] Next, an example of the flow of processing for generating a virtual viewpoint image will be described with reference to FIG. 12. In this processing, after the processing in FIG. 7 determines whether to place all 3D models corresponding to the frame to be processed in the virtual space depending on whether they overlap with other 3D models, the processing in S1201 based on tracking is executed. Steps in which processing similar to that described in FIG. 7 is performed are given the same reference numerals, and descriptions thereof will be omitted. In S1201, the 3D model placement control unit 904 determines not to place the 3D model to be tracked when the tracking state is ON. For example, in the processing in FIG. 7, if the 3D model with ID=3 in FIG. 10 does not exist, it is determined that the 3D model with ID=1 is to be placed in the virtual space. However, in this modified example, since this 3D model is the target of tracking, this 3D model is not placed in the virtual space. That is, at the time point of Timecode 12:15:30.011, the 3D model position overlap determination unit 902 determines not to place the 3D model with ID=1 because the 3D model with ID=1 overlaps with the 3D model with ID=3. Meanwhile, when the tracking control unit 903 acquires from the second 3D model generation unit 901 information indicating that a 3D model has not been generated at the time point of Timecode 12:15:30.012 corresponding to the next frame, it performs tracking on the 3D model with ID=1. The 3D model placement control unit 904 then determines not to place the 3D model with ID=1 that is the tracking target in the virtual space. In this way, if the 3D model with ID=1 exists but the 3D model with ID=3 does not, which is assumed to be a failure in the generation of the 3D model using the motion capture method, the 3D model with ID=1 is not placed in the virtual space. If a 3D model with ID=3 does not exist, and a 3D model with ID=1 does not exist either, then the subject for which these 3D models are to be generated does not exist in the shooting area 201. Therefore, no tracking is performed, and the 3D model corresponding to this subject is not placed in the virtual space, and is not displayed in the virtual viewpoint image.

[0053] Here, an example of a virtual viewpoint image generated using one frame corresponding to Timecode 12:15:30.012 from the recorded data in Fig. 10 will be described. Here, as described above, based on the non-generation information of the 3D model with ID=3, the 3D model with ID=1 is targeted for tracking, and the 3D model with ID=1 is not placed, with only the 3D model with ID=2 being placed in the virtual space. Fig. 13(A) shows a virtual viewpoint image generated based on that virtual space. In this way, it is possible to prevent a situation in which the virtual 3D model generated by the motion capture system 121 suddenly disappears and a realistic 3D model of a subject generated by the volumetric capture system 101 is displayed.

[0054] In the above example, when a 3D model with ID=3 is not generated, the 3D model with ID=1 to be tracked is not placed in the virtual space, and the object is not displayed in the virtual viewpoint image. This is just one example, and a 3D model that was correctly generated in the past may continue to be displayed in the virtual viewpoint image, as shown by object 1301 in FIG. 13(B). For example, of the 3D models that were correctly generated, the latest 3D model (among the 3D models generated in the past) may continue to be displayed. That is, object 1301 corresponding to the 3D model with ID=3 at the time of Timecode 12:15:30.011 may continue to be displayed. In addition, to visually confirm that a 3D model has not been generated, an object when a 3D model has been generated and an object when a 3D model has not been generated may be displayed in different formats, as shown by the hatching in FIG. 13(B). Also, as shown in Figure 13(C), when non-generation of a 3D model due to an abnormality in the motion capture system 121 is detected, information indicating that an abnormality has occurred to the user may be output superimposed on the virtual viewpoint, as in dialog 1302.

[0055] In this way, if an abnormality occurs in the generation of a 3D model by the motion capture system 121, it is possible to prevent the virtual 3D model from suddenly disappearing and the realistic 3D model of the volumetric capture system 101 from appearing. This makes it possible to prevent an inappropriate virtual viewpoint image from being presented.

[0056] In the above example, a case where the volumetric capture method and the motion capture method are used as the 3D model generation method has been described. This is just one example, and other 3D model generation methods may be used. For example, when a first 3D model is generated for all of a plurality of subjects using a first method and a second 3D model is generated for some of the subjects using a second method, the second 3D model may be preferentially placed in the virtual space to generate a virtual viewpoint image. Note that a 3D model may be generated for a first set of subjects present in the shooting area using the first method, and a 3D model may be generated for a second set of subjects present in the shooting area using the second method. Here, a 3D model for a third set of subjects belonging to both the first set and the second set may be generated using both the first method and the second method. In this case, a virtual viewpoint image based on a virtual space in which only 3D models generated for the subjects belonging to the third set using either the 3D models corresponding to the first method or the 3D models corresponding to the second method are placed may be output. In this case, which method a 3D model corresponding to will be placed in the virtual space may be determined in advance by the system or may be determined by a user operation. Also, although the above example describes an example in which a 3D model is generated using two methods, a 3D model may be generated using three or more methods.

[0057] The present disclosure can also be realized by supplying a program that realizes one or more functions of the above-described embodiments to a system or device via a network or a storage medium, and having one or more processors in the computer of the system or device read and execute the program. It can also be realized by a circuit (e.g., ASIC) that realizes one or more functions.

[0058] (Summary of the embodiment) At least some of the above-described embodiments can be summarized as follows. (Item 1) an acquisition means for acquiring a plurality of first three-dimensional (3D) models for each of a plurality of subjects, the first 3D models being generated by a first method based on an image of a predetermined area including the plurality of subjects, and a second 3D model corresponding to a specific subject among the plurality of subjects, the specific subject being present in the predetermined area in the image of the plurality of subjects, the second 3D model being generated by a second method different from the first method based on posture information of the specific subject; an output means for outputting a virtual viewpoint image generated based on the first 3D model of a subject different from the specific subject among the plurality of subjects and the second 3D model corresponding to the specific subject; An image generating device comprising: (Item 2) 2. The image generating device according to item 1, wherein the output means outputs the virtual viewpoint image generated without using the first 3D model of the specific subject. (Item 3) 3. The image generating device according to item 1 or 2, wherein the output means outputs the virtual viewpoint image generated using the second 3D model corresponding to the specific subject instead of the first 3D model for the specific subject. (Item 4) 4. The image generating device according to item 3, wherein the position in the virtual space of the second 3D model corresponding to the specific subject is the position in the virtual space of the first 3D model for the specific subject. (Item 5) 5. The image generating device according to any one of items 1 to 4, wherein, when there is a first area in which a degree of overlap with a second area in the virtual space in which the second 3D model exists exceeds a predetermined level among a plurality of areas in which each of the first 3D models corresponding to each of the plurality of subjects exists in the virtual space, the output means outputs the virtual viewpoint image generated by placing the second 3D model corresponding to the specific subject at the position of the first area, instead of the first 3D model corresponding to the first area. (Item 6) 6. The image generating device according to any one of items 1 to 5, wherein, when the second 3D model corresponding to the specific subject is not generated, the output means outputs the virtual viewpoint image generated using the second 3D model corresponding to the specific subject that was generated in the past. (Item 7) 7. The image generating device according to item 6, wherein the display of the specific subject in the virtual viewpoint image generated using the second 3D model corresponding to the specific subject generated in the past is different from the display of the specific subject displayed in the virtual viewpoint image when the second 3D model is generated. (Item 8) 8. The image generating device according to any one of items 1 to 7, wherein the output means outputs information indicating an abnormality when the second 3D model corresponding to the specific subject is not generated and the first 3D model for the specific subject is generated. (Item 9) 9. The image generating device according to any one of items 1 to 8, wherein the first method is a volumetric capture method and the second method is a motion capture method. (Item 10) Item 10. The image generating device according to item 9, wherein the first 3D model is generated by generating three-dimensional shapes of the plurality of subjects based on photographs of the predetermined area taken by a plurality of cameras, and applying textures obtained by the photographs to the three-dimensional shapes. (Item 11) 11. The image generating device according to item 9 or 10, wherein the second 3D model is generated by generating a bone model of the specific subject from posture information based on information obtained by a plurality of sensors, and transforming a CG (Computer Graphic) model with respect to the bone model. (Item 12) 1. An image generation method performed by an image generation device, comprising: acquiring a plurality of first three-dimensional (3D) models for each of the plurality of subjects, the first 3D models being generated by a first method based on imaging of a predetermined area including the plurality of subjects, and a second 3D model corresponding to a specific subject, the specific subject being present in the predetermined area during imaging, the second 3D model being generated by a second method different from the first method based on posture information of the specific subject among the plurality of subjects; outputting a virtual viewpoint image generated based on the first 3D model of a subject different from the specific subject among the plurality of subjects and the second 3D model corresponding to the specific subject; An image generating method comprising: (Item 13) 12. A program for causing a computer to function as each of the means included in the image generating device according to any one of items 1 to 11.

[0059] The invention is not limited to the above-described embodiments, and various changes and modifications can be made without departing from the spirit and scope of the invention. Accordingly, the following claims are appended to apprise the public of the scope of the invention. [Explanation of symbols]

[0060] 404: First 3D model generation unit, 414: Second 3D model generation unit, 431: 3D model position overlap determination unit, 432: 3D model placement control unit, 433: Rendering unit, 903: Tracking control unit

Claims

1. an acquisition means for acquiring a plurality of first three-dimensional (3D) models for each of the plurality of subjects, the first 3D models being generated by a first method based on an image of a predetermined area including the plurality of subjects, and a second 3D model corresponding to a specific subject among the plurality of subjects, the specific subject being present in the predetermined area during the image capture, the second 3D model being generated by a second method different from the first method based on posture information of the specific subject; an output means for outputting a virtual viewpoint image generated based on the first 3D model of a subject different from the specific subject among the plurality of subjects and the second 3D model corresponding to the specific subject; An image generating device comprising:

2. The image generating device according to claim 1 , wherein the output means outputs the virtual viewpoint image generated without using the first 3D model of the specific subject.

3. 2. The image generating device according to claim 1, wherein the output means outputs the virtual viewpoint image generated using the second 3D model corresponding to the specific subject instead of the first 3D model for the specific subject.

4. 4. The image generating device according to claim 3, wherein the position in the virtual space of the second 3D model corresponding to the specific subject is the position in the virtual space of the first 3D model for the specific subject.

5. 2. The image generating device according to claim 1, wherein, when there is a first area in which a degree of overlap with a second area in the virtual space in which the second 3D model exists exceeds a predetermined level among a plurality of areas in the virtual space in which each of the first 3D models corresponding to the plurality of subjects exists, the output means outputs the virtual viewpoint image generated by placing the second 3D model corresponding to the specific subject at a position of the first area, instead of the first 3D model corresponding to the first area.

6. 2. The image generating device according to claim 1, wherein, when the second 3D model corresponding to the specific subject is not generated, the output means outputs the virtual viewpoint image generated using the second 3D model corresponding to the specific subject that was generated in the past.

7. 7. The image generating device according to claim 6, wherein the display of the specific subject in the virtual viewpoint image generated using the second 3D model corresponding to the specific subject generated in the past is different from the display of the specific subject displayed in the virtual viewpoint image when the second 3D model is generated.

8. 2. The image generating device according to claim 1, wherein the output means outputs information indicating an abnormality when the second 3D model corresponding to the specific subject is not generated and the first 3D model for the specific subject is generated.

9. 2. The image generating device according to claim 1, wherein the first method is a volumetric capture method, and the second method is a motion capture method.

10. 10. The image generating device according to claim 9, wherein the first 3D model is generated by generating three-dimensional shapes of the plurality of subjects based on photographs of the predetermined area taken by a plurality of cameras, and applying textures obtained by the photographs to the three-dimensional shapes.

11. 10. The image generating device according to claim 9, wherein the second 3D model is generated by generating a bone model of the specific subject from posture information based on information obtained by a plurality of sensors, and transforming a CG (Computer Graphics) model with respect to the bone model.

12. 1. An image generation method performed by an image generation device, comprising: acquiring a plurality of first three-dimensional (3D) models for each of the plurality of subjects, which are generated by a first method based on imaging of a predetermined area including the plurality of subjects, and a second 3D model corresponding to a specific subject among the plurality of subjects, which is present in the predetermined area in the imaging, and which is generated by a second method different from the first method based on posture information of the specific subject; outputting a virtual viewpoint image generated based on the first 3D model of a subject different from the specific subject among the plurality of subjects and the second 3D model corresponding to the specific subject; An image generating method comprising:

13. A program for causing a computer to function as each of the means included in the image generating device according to any one of claims 1 to 11.

Citation Information

Patent Citations

  • Image processing apparatus, image processing system, image processing method, and program

    JP2022060058A