Information processing device, information processing method, and program
By adding a light source camera at the light source position, estimating the light source position and extracting the shadow area, the problem of unrealistic shadow rendering when the background changes in real time in virtual production is solved, and a more realistic virtual background rendering effect is achieved.
Patent Information
- Application Number
- CN202480011355.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2023-02-15
- Filing Date
- 2024-01-29
- Publication Date
- 2025-09-16
Smart Images

Figure CN120660356A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to an information processing device, an information processing method, and a program, and more particularly to an information processing device, an information processing method, and a program capable of rendering a more realistic background in virtual production. Background Art
[0002] There is a technology called virtual production (or in-camera VFX) in which a 3D graphics image based on the position and posture of the imaging camera is rendered in real time on a wall display device serving as a background. The 3D graphics image is then imaged simultaneously with the performer in the foreground. Virtual production can create a composite image that appears as if the performer's image was filmed in the location rendered in the background.
[0003] In recent years, there has been an increasing trend of placing display devices not only on walls but also on floors to composite backgrounds that include the ground. However, when rendering a virtual ground on a floor, the shadows of performers, which should be present, are not rendered, making the ground appear unrealistic and creating the impression that the rendered image is artificially synthesized. This phenomenon occurs not only on floors but also when rendering a background on a wall close to the performer.
[0004] To render the performer's shadow, it's necessary to know the performer's position relative to the light source and their shape. This can be achieved by using a Time of Flight (ToF) sensor. However, this requires specialized equipment. Furthermore, if the imaging camera's tracking system uses the same IR region as the ToF sensor, their sensing may interfere with each other.
[0005] In contrast, for example, NPL 1 discloses a method in which a camera is placed at a light source position and a portion of an image captured from the light source position that is not visible due to the foreground (a portion where light is blocked) is shadowed.
[0006] [Citation List]
[0007] [Non-patent literature]
[0008] [NPL 1]
[0009] Ken Naemura, and another person, "Creation of the image space of the film", [Online], [Retrieved February 1, 2023], Internet <URL:https: / / www.iu-tokyo.ac.jp / coe / report / H14 / 21COE-ISTSC-H14_3_3_3_4.pdf > Summary of the Invention
[0010] [Technical Issues]
[0011] The technique in NPL 1 is easy to introduce because it simply requires adding an ordinary camera as a device. However, since the foreground and background are separated by image differences in images taken from the light source position, it is necessary to assume that the background is known and does not change, and that the light source position is also known and does not change.
[0012] In virtual production where the background changes in real time according to the position and posture of the imaging camera, the technology in NPL 1 cannot be applied as it is.
[0013] The present disclosure has been made in view of such circumstances, and aims to render more realistic backgrounds in virtual production.
[0014] [Solution to the problem]
[0015] An information processing device according to the present disclosure is an information processing device including: a rendering processing unit that renders a 3DCG image on a display device based on a position and posture of a first camera, the first camera imaging the display device as a background together with an object as a foreground; a position estimation unit that estimates a light source position based on a captured image captured in a second camera and based on the 3DCG image rendered on the display device, the second camera imaging the display device together with the object from a light source position where the light source is placed; and a difference extraction unit that extracts a difference area between the 3DCG image that has been projectively transformed based on the estimated light source position and the captured image from the light source position as a shadow area of the object.
[0016] The information processing method according to the present disclosure is an information processing method performed by an information processing device, the method including: rendering a 3DCG image on a display device based on the position and posture of a first camera, the first camera imaging the display device as a background together with the object as a foreground; estimating a light source position based on a captured image captured in a second camera and based on the 3DCG image rendered on the display device, the second camera imaging the display device together with the object from a light source position where the light source is placed; and extracting a difference area between the 3DCG image that has been projectively transformed based on the estimated light source position and the captured image from the light source position as a shadow area of the object.
[0017] The program according to the present disclosure is a program for causing a computer to perform the following processing: rendering a 3DCG image on a display device based on the position and posture of a first camera, which first camera images the display device as a background together with the object as a foreground; estimating the light source position based on a captured image captured in a second camera and based on the 3DCG image rendered on the display device, which second camera images the display device together with the object from the light source position where the light source is placed; and extracting the difference area between the 3DCG image that has been projectively transformed based on the estimated light source position and the captured image from the light source position as a shadow area of the object.
[0018] In the present invention, a 3DCG image is rendered on a display device based on the position and posture of a first camera, which images the display device as a background together with an object as a foreground; a light source position is estimated based on a captured image captured in a second camera and based on the 3DCG image rendered on the display device, which images the display device together with the object from a light source position where the light source is placed; and a difference area between the 3DCG image that has been projectively transformed based on the estimated light source position and the captured image from the light source position is extracted as a shadow area of the object. BRIEF DESCRIPTION OF THE DRAWINGS
[0019] Figure 1 is a diagram showing a configuration example of a virtual production system.
[0020] Figure 2 is a diagram illustrating a configuration for determining a shadow area of a foreground.
[0021] Figure 3 : is a diagram showing the flow of a process of determining a shadow area of a foreground.
[0022] Figure 4 is a block diagram illustrating a functional configuration example of an information processing device according to the present disclosure.
[0023] Figure 5 is a flowchart showing the flow of a process of rendering a background image.
[0024] Figure 6 This section shows an example of rendering a shadow image when multiple light sources are arranged.
[0025] Figure 7 is a diagram illustrating another application example of the technology according to the present disclosure.
[0026] Figure 8 A block diagram of an example configuration of a computer. DETAILED DESCRIPTION
[0027] Modes for carrying out the present disclosure (hereinafter referred to as embodiments) will be described below. The description will be given in the following order.
[0028] 1. Configuration Example of Virtual Production System
[0029] 2. Configuration for determining the shadow area of the foreground
[0030] 3. Configuration and Operation of the Information Processing Device According to the Present Disclosure
[0031] 4. Modifications and Other Applications
[0032] 5. Computer Configuration Examples
[0033] <1. Virtual Production System Configuration Example>
[0034] Figure 1 is a diagram illustrating a configuration example of a virtual production system to which the technology according to the present disclosure can be applied.
[0035] exist Figure 1 In the illustrated virtual production system 1, performer P10, acting as a foreground performers, performs in front of a display device 20, which renders a background. Display device 20 may be configured as a light-emitting diode (LED) panel having display surfaces that form wall and floor surfaces. A three-dimensional computer graphics (3DCG) image based on the position and orientation of imaging camera 30 is rendered on display device 20, and by simultaneously imaging display device 20, acting as a background, and performer P10, a composite image can be obtained that appears as if the image of performer P10 were taken at a location rendered in the background. In addition to the wall and floor surfaces, display device 20 may also have a display surface that forms a ceiling.
[0036] To obtain the position and attitude of the imaging camera 30, there is a technique that uses a tracking camera 40 rigidly connected to the imaging camera 30. The tracking camera 40 can capture the environment surrounding the imaging camera 30 and estimate the position and attitude of the imaging camera 30 based on changes in the appearance of the environment. To this end, there are techniques that place markers in the surrounding environment as landmarks, and techniques that use natural textures in the environment instead of markers. There is also a technique in which the tracking camera 40 is placed on the side of the surrounding environment rather than the imaging camera 30, and a marker is placed on the imaging camera 30 side to capture changes in the marker's position. In the technology according to the present disclosure, any technique can be used to estimate the position and attitude of the imaging camera 30.
[0037] The information processing device 50, which is composed of a computer such as a personal computer (PC), renders a 3DCG image on the display device 20 according to the position and posture of the imaging camera 30. A three-dimensional model of the virtual world for the background is prepared in advance, and the information processing device 50 renders the image of the three-dimensional model as seen from the position and posture of the imaging camera 30 as a two-dimensional image on the display device 20. As a result, a composite image C60 is obtained as an image captured by the imaging camera 30, in which the real performer P10 appears to exist in the virtual world.
[0038] In the virtual production system 1, a background that does not exist in reality must appear realistic. To achieve this, it is essential that the rendering of the background (including the foreground and background relationships of objects in the background) is consistent with the movement of the imaging camera 30, and that the colors of the background and foreground match. Furthermore, the "shadow" of the foreground cast on the background where it should exist is also considered an essential element.
[0039] <2. Configuration for determining the shadow area of the foreground>
[0040] In the technology according to the present disclosure, in order to determine the area of the "shadow" of the foreground (hereinafter referred to as the shadow area), as shown in FIG. Figure 2 As shown, in the virtual production system 1, a light source camera 110 is added at the light source position where the light source is placed. The light source camera 110 can be placed adjacent to the actual light source, or the light source camera 110 can be placed as a non-real, virtual light source. In the virtual production system 1, the light source is placed so as to be movable.
[0041] In image P120 captured by light source camera 110 at the light source position, the area (occluded area) that overlaps with performer P10 in the foreground and is therefore invisible is the area where light from the light source does not reach. Therefore, by rendering the occluded area that overlaps with the foreground as a shadow area on display device 20, the shadow of the foreground can be rendered on the background. If the position and posture of the light source and the positional relationship of display device 20 are known, the range of display device 20 corresponding to the area overlapping with the foreground in the image captured from the light source position can be calculated.
[0042] To achieve this, it is necessary to be able to separate the foreground and background in an image captured from the light source location and to determine the position and orientation of the light source. In the technique of NPL 1, a camera is also placed at the light source location to determine the area obscured by the foreground. However, to separate the foreground and background, it is assumed that the background is known and unchanged, and that the light source location is also known and unchanged.
[0043] However, in the virtual production system 1 , the background changes according to the position and posture of the imaging camera 30 , and therefore, the premise that the background is known and does not change does not hold true.
[0044] On the other hand, although the background changes in real time, the system itself renders the background based on the position and posture of imaging camera 30, so the details of the background being rendered are known. Therefore, if the position of the light source can be determined, the background image that should be seen from the light source position can be determined by projecting the two-dimensional image rendered on display device 20 into the image as seen from the light source position. The difference between the background image that should be seen from the light source position and the image actually captured from the light source position is the area obscured by the foreground, that is, the shadow area.
[0045] Will refer to Figure 3 The flow of a process for determining a shadow area of a foreground in the virtual production system 1 will be described.
[0046] First, the background image (3DCG image) rendered on the display device 20 is projectively transformed into an image viewed from the estimated light source position.
[0047] Next, an image captured by the light source camera 110 (an image actually captured from the light source position) is acquired. The image captured by the light source camera 110 includes an object (performer P10) as a foreground.
[0048] Next, a shadow area is determined based on the difference between the 3DCG image that has been projectively transformed into an image viewed from the light source position and the image captured by the light source camera 110 .
[0049] Then, the shadow image is rendered in the area corresponding to the shadow area in the 3DCG image without projection transformation. At this time, since the shape and position of the 3D object included in the 3DCG image are known, the shadow image can be further rendered in the shadow area of the 3D object based on the estimated light source position.
[0050] <3. Configuration and Operation of Information Processing Device According to the Present Disclosure>
[0051] The configuration and operation of the information processing apparatus 50 according to the present disclosure in the virtual production system 1 will be described below.
[0052] (Configuration of Information Processing Device)
[0053] Figure 4 : is a block diagram showing a functional configuration example of the information processing device 50 according to the present disclosure.
[0054] like Figure 4 As shown, in addition to the display device 20 and the tracking camera 40 , a light source camera 110 is connected to the information processing device 50 .
[0055] The information processing apparatus 50 is configured to include a three-dimensional model storage unit 211 , a position and posture estimation unit 212 , a rendering processing unit 213 , a position estimation unit 214 , an image transformation unit 215 , and a difference extraction unit 216 .
[0056] The three-dimensional model storage unit 211 stores a three-dimensional model of a virtual world prepared in advance for use as a background. The three-dimensional model stored in the three-dimensional model storage unit 211 is read out by the rendering processing unit 213 as appropriate.
[0057] The position and attitude estimation unit 212 estimates the position and attitude of the imaging camera 30 (first camera) that images the subject (performer P10) as the foreground together with the display device 20 as the background, based on the image of the surrounding environment captured by the tracking camera 40. Information indicating the estimated position and attitude of the imaging camera 30 is provided to the rendering processing unit 213.
[0058] The rendering processing unit 213 renders a 3DCG image on the display device 20 based on the position and attitude of the imaging camera 30 indicated by the information from the position and attitude estimation unit 212. Specifically, the rendering processing unit 213 renders an image of the three-dimensional model read out from the three-dimensional model storage unit 211 and seen from the position and attitude of the imaging camera 30 as a two-dimensional image on the display device 20.
[0059] The position estimation unit 214 estimates the light source position based on the captured image captured by the light source camera 110 (second camera), which images the display device 20 together with the subject (performer P10) from the light source position where the light source is placed, and based on the 3DCG image rendered on the display device 20 by the rendering processing unit 213. Information indicating the estimated light source position is provided to the image conversion unit 215.
[0060] The image conversion unit 215 performs projective conversion on the 3DCG image rendered on the display device 20 by the rendering processing unit 213 based on the light source position indicated by the information from the position estimation unit 214 , and supplies the resultant image to the difference extraction unit 216 .
[0061] The difference extraction unit 216 extracts the difference area between the 3DCG image projectively transformed by the image transformation unit 215 and the captured image captured from the light source position by the light source camera 110 as the shadow area of the subject (performer P10). Information indicating the extracted shadow area (coordinate position information) is provided to the rendering processing unit 213.
[0062] The rendering processing unit 213 renders a shadow image representing the shadow of the object (performer P10 ) in an area corresponding to the shadow area indicated by the information from the difference extraction unit 216 in the 3DCG image not projectively transformed by the image transformation unit 215 .
[0063] With the above-described configuration, the information processing apparatus 50 is able to determine the shadow area of the performer P10 which is the foreground, and render the shadow image in the area corresponding to the shadow area in the 3DCG image.
[0064] (Operation of Information Processing Device)
[0065] Will refer to Figure 5 The flowchart of FIG. 5 describes the process of rendering the background image on the display device 20 by the information processing device 50. When the imaging camera 30 in the virtual production system 1 starts to capture the image of the performer P10, Figure 5 The processing starts.
[0066] In step S1 , the position and attitude estimation unit 212 estimates the position and attitude of the imaging camera 30 based on the image of the surrounding environment captured by the tracking camera 40 .
[0067] In step S2 , the rendering processing unit 213 renders a background image (a 3DCG image according to the position and attitude of the imaging camera 30 ) on the display device 20 based on the estimated position and attitude of the imaging camera 30 .
[0068] The imaging camera 30 images the display device 20 displaying such a background image and the performer P10 in the foreground, thereby realizing virtual production.
[0069] Next, in step S3 , the position estimation unit 214 estimates the light source position based on the captured image captured in the light source camera 110 that images the display device 20 together with the performer P10 from the light source position and based on the background image rendered on the display device 20 .
[0070] At this point, in the virtual production system 1, the physical shape and size of the display device 20 are known, and the background image rendered on the display device 20 is also known. Thus, similar to the technique of estimating the position and orientation of a camera that captures a known marker image, the background image rendered on the display device 20 and appearing in the captured image can be used as a known marker image to estimate the position of the light source relative to the display device 20 (the position of the light source camera 110).
[0071] Specifically, the position estimation unit 214 estimates the light source position by minimizing the difference between the background image from an arbitrary viewpoint and the background image included in the captured image captured by the light source camera 110. Here, the light source position can be estimated by solving a perspective n-point (PnP) problem that minimizes the difference between the position of each pixel in the background image from an arbitrary viewpoint and the position of the corresponding pixel on the background image included in the captured image. The light source position can be estimated by minimizing the photometric error, which is the difference between the pixel value of each pixel in the background image from an arbitrary viewpoint and the corresponding pixel value on the background image included in the captured image.
[0072] As described above, when estimating the position of the light source camera 110 by minimizing the error (difference) on the image, the foreground (performer P10) that does not exist in the original background image is considered an outlier. When minimizing the error, outliers hinder correct convergence. Countermeasures for this can be used, such as a technique that excludes outliers through majority voting or a weighted optimization method that reduces the weight of pixels with large errors.
[0073] Due to the influence of the light source in the imaging environment and the characteristics of the camera, there is a possibility that the color tone may be different between the (known) background image rendered in the virtual production system 1 and the captured image actually captured. To address such color tone differences, a technique can be used in which a pattern image such as a Macbeth chart (color palette) is pre-displayed on the display device 20 to adjust the white balance, or a technique can be used in which the background image and the captured image are converted into brightness images and the difference between the images is calculated to mitigate the above-mentioned influence. In addition, to address optical distortion caused by the camera lens, a technique can be used in which a known pattern image is pre-displayed on the display device 20 to perform lens calibration, and a technique can be used in which the camera distortion parameters are optimized while optimizing the position and posture.
[0074] In step S4, the image transformation unit 215 projectively transforms the background image based on the estimated light source position into the coordinate system of the light source camera 110. Thus, in the process of estimating the light source position as described above, the background image is converted into an image viewed from the light source position (from a viewpoint that minimizes the error with the actually captured image).
[0075] In step S5, the difference extraction unit 216 extracts, as a shadow area, the difference area between the projectively transformed background image (from a viewpoint that minimizes the error with the actually captured captured image) and the captured image captured by the light source camera 110. Thus, the area of the foreground (performer P10) that does not exist in the projectively transformed background image is extracted as the shadow area.
[0076] In step S6 , the rendering processing unit 213 renders the shadow image in the area corresponding to the shadow area in the background image that has not been projectively transformed. As a result, the shadow of the performer P10 is rendered on the background image, which is rendered on the display device 20 .
[0077] Since the light source position is known, in step S7, the rendering processing unit 213 can further render the shadow image in the shadow area of the 3D object based on the estimated light source position. The three-dimensional shape and position of the 3D object are known and included in the background image. In this way, the shadow of the 3D object is rendered in the same direction as the shadow of the foreground, making the object in the background look more realistic.
[0078] According to the above processing, in the virtual production system 1, by determining the light source position, it is possible to extract the difference between the background image that should be seen from the light source position and the image actually captured from the light source position as the area obscured by the foreground, that is, the shadow area. As a result, even in virtual productions where the background changes depending on the position and posture of the imaging camera 30, a more realistic background can be rendered.
[0079] <4. Modifications and Other Applications>
[0080] (Variation)
[0081] As described above, the light source may be a real light source or an unreal or virtual light source. The number of light sources is not limited to one and may be multiple. When multiple light sources are arranged, regardless of whether they are real light sources or unreal or virtual light sources, the rendering processing unit 213 may adjust the density of the areas of the shadow image rendered for the light source position that do not overlap with other shadow images.
[0082] For example, Figure 6 As shown, in an area that is a shadow area cast by one light source (the area obscured by performer P10 as seen from light source camera 110-1) but not a shadow area cast by another light source (the area obscured by performer P10 as seen from light source camera 110-2), an effect such as fading the shadow can be applied.
[0083] (Other application examples)
[0084] The technology according to the present disclosure can be applied to a configuration of a virtual production system in which a camera is added at a light source position where a light source is placed, and a configuration in which a camera is added at a position where other props or equipment are placed.
[0085] For example, in Figure 7In the virtual production system 251 shown, a camera 260 is added at the location where the blower is placed, and the position and posture of the blower (camera 260) are determined in the same manner as in the above embodiment. By knowing the position and posture of the blower, it is possible to express the appearance of a 3D object included in the background image being blown by the wind in the same way as the wind is blowing on the performer P10 in the foreground.
[0086] <5. Computer Configuration Example>
[0087] The above series of processing can be executed by hardware or software. When the series of processing is executed by software, the program configuring the software is installed from a program recording medium to a computer incorporated into dedicated hardware or a general-purpose personal computer.
[0088] Figure 8 : is a block diagram showing a configuration example of computer hardware that executes the above-described series of processes using a program.
[0089] The information processing device 50 to which the technology according to the present disclosure can be applied is provided with Figure 8 The information processing device 300 of the shown configuration is implemented.
[0090] A central processing unit (CPU) 301 , a read only memory (ROM) 302 , and a random access memory (RAM) 303 are connected to one another via a bus 304 .
[0091] An input / output interface 305 is also connected to the bus 304. An input unit 306 including a keyboard and a mouse, and an output unit 307 including a display device and a speaker are connected to the input / output interface 305. In addition, a storage unit 308 including a hard disk or a nonvolatile memory, a communication unit 309 including a network interface, and a drive 310 that drives a removable medium 311 are connected to the input / output interface 305.
[0092] In the computer configured as described above, for example, the CPU 301 performs the above-described series of processes by loading a program stored in the storage unit 308 into the RAM 303 via the input / output interface 305 and the bus 304 and executing the program.
[0093] For example, the program executed by the CPU 301 is recorded on the removable medium 311 or provided via a wired or wireless transmission medium (eg, a local area network, the Internet, or digital broadcasting) to be installed in the storage unit 308 .
[0094] The program executed by the computer may be a program that executes processing chronologically in the order described in this specification, or may be a program that executes processing in parallel or at a necessary timing such as a calling time.
[0095] The embodiment according to the present disclosure is not limited to the above-described embodiment, and various modifications may be made without departing from the scope and spirit of the present disclosure.
[0096] For example, an embodiment of the present disclosure adopts a configuration of cloud computing in which a plurality of devices share and collaboratively process one function through a network.
[0097] Additionally, each step described in the flowcharts discussed above may be performed by a single device, or performed by multiple devices in a distributed manner.
[0098] Furthermore, when a single step includes a plurality of types of processing, the plurality of types of processing included in the single step may be performed by a single device, or may be performed by a plurality of devices in a distributed manner.
[0099] The beneficial effects described in this specification are merely exemplary and non-limiting, and other beneficial effects may occur.
[0100] Furthermore, the present disclosure may be configured as follows. (1)
[0102] An information processing device, comprising:
[0103] a rendering processing unit that renders a 3DCG image on a display device based on a position and a posture of a first camera that images the display device as a background together with an object as a foreground;
[0104] a position estimating unit that estimates a light source position based on a captured image captured in a second camera that images the display device together with the object from the light source position where the light source is placed and based on a 3DCG image rendered on the display device; and
[0105] A difference extraction unit extracts a difference area between the 3DCG image that has been projectively transformed based on the estimated light source position and the captured image from the light source position as a shadow area of the object. (2)
[0107] The information processing device according to (1), wherein the rendering processing unit renders the shadow image in a region corresponding to the shadow region in the 3DCG image that is not projectively transformed. (3)
[0109] The information processing device according to (1), wherein the position estimating unit estimates the light source position using a 3DCG image rendered on the display device and appearing in the captured image as a marker. (4)
[0111] The information processing device according to (3), wherein the position estimating unit estimates the light source position by minimizing a difference between a 3DCG image from an arbitrary viewpoint and a 3DCG image included in the captured image. (5)
[0113] The information processing device according to (4), wherein the position estimating unit estimates the light source position by solving a PnP (Perspective n Point) problem. (6)
[0115] The information processing device according to (4), wherein the position estimating unit estimates the light source position by minimizing a photometric error. (7)
[0117] An information processing device according to any one of (4) to (6), wherein the difference extraction unit extracts the difference area using a 3DCG image from a viewpoint having the smallest difference from a 3DCG image appearing in a captured image as a 3DCG image that has been projectively transformed into a coordinate system of a second camera. (8)
[0119] The information processing device according to any one of (2) to (7), wherein the rendering processing unit further renders a shadow image in a shadow region of the 3D object included in the 3DCG image based on the estimated light source position. (9)
[0121] The information processing device according to any one of (1) to (8), wherein the light source is a real light source. (10)
[0123] The information processing apparatus according to any one of (1) to (8), wherein the light source is a virtual light source. (11)
[0125] The information processing apparatus according to any one of (1) to (10), wherein the light source is disposed so as to be movable. (12)
[0127] The information processing device according to any one of (2) to (11), wherein, when a plurality of light sources are arranged, the rendering processing unit adjusts the density of an area that does not overlap with another shadow image in the shadow images rendered for the plurality of light source positions. (13)
[0129] An information processing method performed by an information processing device, the method comprising:
[0130] rendering a 3DCG image on a display device based on a position and a posture of a first camera, the first camera imaging the display device as a background together with the object as a foreground;
[0131] estimating a light source position based on a photographic image captured in a second camera that images the display device together with the object from the light source position where the light source is placed and based on a 3DCG image rendered on the display device; and
[0132] A difference area between the 3DCG image that has been projectively transformed based on the estimated light source position and the captured image from the light source position is extracted as a shadow area of the object. (14)
[0134] A program for causing a computer to execute the following processing:
[0135] rendering a 3DCG image on a display device based on a position and a posture of a first camera, the first camera imaging the display device as a background together with the object as a foreground;
[0136] estimating a light source position based on a photographic image captured in a second camera that images the display device together with the object from the light source position where the light source is placed and based on a 3DCG image rendered on the display device; and
[0137] A difference area between the 3DCG image that has been projectively transformed based on the estimated light source position and the captured image from the light source position is extracted as a shadow area of the object.
[0138] [Reference Signs List]
[0139] 1 virtual production system, 20 display device, 30 imaging camera, 40 tracking camera, 50 information processing device, 110 light source camera, 211 three-dimensional model storage unit, 212 position and posture estimation unit, 213 rendering processing unit, 214 position estimation unit, 215 image transformation unit, 216 difference extraction unit.
Claims
1. An information processing device, comprising: a rendering processing unit that renders a 3DCG image on a display device based on a position and a posture of a first camera, the first camera imaging the display device as a background together with an object as a foreground; a position estimating unit that estimates a light source position based on a captured image captured in a second camera that images the display device together with the object from the light source position where the light source is placed and based on the 3DCG image rendered on the display device; as well as A difference extraction unit extracts a difference area between the 3DCG image that has been projectively transformed based on the estimated light source position and the captured image from the light source position as a shadow area of the object.
2. The information processing device according to claim 1, wherein The rendering processing unit renders a shadow image in an area corresponding to the shadow area in the 3DCG image that has not been projectively transformed.
3. The information processing device according to claim 1, wherein The position estimation unit estimates the light source position using the 3DCG image rendered on the display device and appearing in the captured image as a marker.
4. The information processing device according to claim 3, wherein: The position estimating unit estimates the light source position by minimizing a difference between the 3DCG image from an arbitrary viewpoint and the 3DCG image included in the captured image.
5. The information processing apparatus according to claim 4, wherein: The position estimation unit estimates the light source position by solving a PnP (Perspective n Point) problem. The information processing apparatus according to claim 4 , wherein: The position estimation unit estimates the light source position by minimizing a photometric error.
7. The information processing apparatus according to claim 4, wherein: The difference extraction unit extracts the difference area using the 3DCG image from a viewpoint having the smallest difference from the 3DCG image appearing in the captured image as the 3DCG image projectively transformed into the coordinate system of the second camera.
8. The information processing apparatus according to claim 2, wherein: The rendering processing unit further renders the shadow image in a shadow region of a 3D object included in the 3DCG image based on the estimated light source position.
9. The information processing apparatus according to claim 1, wherein: The light source is a real light source.
10. The information processing apparatus according to claim 1, wherein: The light source is a virtual light source.
11. The information processing apparatus according to claim 1, wherein: The light source is arranged to be movable.
12. The information processing apparatus according to claim 2, wherein: When a plurality of the light sources are arranged, the rendering processing unit adjusts the density of a region that does not overlap with another shadow image in the shadow images rendered for the plurality of light source positions.
13. An information processing method performed by an information processing device, the method comprising: rendering a 3DCG image on a display device based on a position and a posture of a first camera, wherein the first camera images the display device as a background and an object as a foreground; estimating a light source position based on a captured image captured in a second camera that images the display device together with the object from the light source position where the light source is placed and based on the 3DCG image rendered on the display device; and A difference area between the 3DCG image that has been projectively transformed based on the estimated light source position and the captured image from the light source position is extracted as a shadow area of the object.
14. A program for causing a computer to execute the following processing: rendering a 3DCG image on a display device based on a position and a posture of a first camera, wherein the first camera images the display device as a background and an object as a foreground; estimating a light source position based on a captured image captured in a second camera that images the display device together with the object from the light source position where the light source is placed and based on the 3DCG image rendered on the display device; and A difference area between the 3DCG image that has been projectively transformed based on the estimated light source position and the captured image from the light source position is extracted as a shadow area of the object.