Information processing device, information processing method, and program
The system addresses delays and unnaturalness in imaging systems by layer-based rendering and motion prediction, ensuring natural background images adapt to camera changes without delay.
Patent Information
- Application Number
- PCT/JP2025/003376
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-02-19
- Filing Date
- 2025-02-03
- Publication Date
- 2025-08-28
AI Technical Summary
Existing imaging systems using large display devices for background images in studios face delays and unnatural representations when camera positions or directions change, due to high computational loads from rendering processes like ray tracing, and geometric corrections can introduce incongruities.
An information processing device and method that divides objects into multiple layers based on camera position and movement, renders background images for each layer, and synthesizes them into a single image, using motion prediction and geometric correction to maintain natural representations without delay.
Enables real-time display of natural background images that adapt to camera movements, reducing delays and unnaturalness by predicting camera motion and applying layer-based rendering and geometric corrections.
Smart Images

Figure JP2025003376_28082025_PF_FP_ABST
Abstract
Description
Information processing device, information processing method, and program
[0001] The present disclosure relates to an information processing device, an information processing method, and a program, and in particular to an information processing device, an information processing method, and a program that are capable of displaying a natural, appropriate background image without delay, even if there is a change in the position or imaging direction of the camera relative to the large display device, in an imaging system in a studio where a large display device is installed, where a background image is displayed on the display device and actors perform in front of the display device to capture images of the actors and the background.
[0002] 2. Description of the Related Art A known imaging technique for producing video content such as movies is to have actors perform in front of a so-called green screen, and then later compose the performance against a background image.
[0003] However, in recent years, instead of green screen imaging, imaging systems have been developed in which a background image is displayed on a large display device in a studio, and the performers perform in front of the display device, thereby capturing images of the performers and the background, and related technologies have also been proposed (see Patent Document 1).
[0004] These imaging systems are known as In-Camera VFX (or Virtual Production, including LED Wall Virtual Production, etc.).
[0005] US Patent Application Publication No. 2020 / 0145644
[0006] Incidentally, when capturing video content such as the above-mentioned movie, the camera position and capturing direction change depending on the movements of the performers, the direction of the performance, and so on.
[0007] In such cases, when capturing images in real space, the positional relationship with the background object will change depending on the camera position and capturing direction, so naturally the image of the captured object will also be based on its positional relationship.
[0008] On the other hand, when using an imaging system such as that proposed in Patent Document 1, if the camera position or imaging direction changes relative to the background image displayed on a large display device, displaying the displayed background image as is will not express the change in the positional relationship with the object in real space, resulting in an unnatural background image being captured.
[0009] For this reason, in an imaging system such as that proposed in Patent Document 1, each object in the background image is displayed while being changed using rendering such as ray tracing, depending on the position and imaging direction of the camera relative to the large display device.
[0010] However, rendering processes such as ray tracing generally require a large computational load and take a long time to process. For example, the display may not be able to keep up with changes in the camera position or imaging direction relative to a large display device, resulting in delays.
[0011] One way to address this issue is to predict the camera's imaging position and direction and then render the background image to reduce the delay in the displayed background image. However, if unexpected movement occurs after rendering begins, the generated background image can appear unnatural.
[0012] Furthermore, if unexpected movement occurs after rendering begins, applying geometric correction can be considered to suppress unnaturalness. However, applying geometric correction can sometimes result in an unnatural representation of the front-to-back relationship of objects in the depth direction, creating a sense of incongruity.
[0013] The present disclosure has been made in consideration of such circumstances, and in particular, in the imaging system described above, even if there is a change in the position or imaging direction of the camera relative to the large display device, an appropriate background image that is natural and does not look unnatural is displayed without delay.
[0014] An information processing device and a program according to one aspect of the present disclosure are an information processing device and a program that include: a rendering processing unit that divides objects into multiple layers based on the position and movement of an imaging device that captures a scene including a background image displayed on a display device, relative to the display device that displays the background image, the objects being based on a 3D model, and renders the background image for each layer; and a background synthesis unit that synthesizes the background images for each of the multiple layers rendered by the rendering processing unit into a single background image and displays the background image on the display device.
[0015] An information processing method according to one aspect of the present disclosure is an information processing method including: a rendering process for dividing an object into a plurality of layers and rendering the background image for each layer based on the position and movement of an imaging device that captures a scene including the background image displayed on a display device that displays the background image consisting of objects based on a 3D model; and a background synthesis process for synthesizing the background images for each of the plurality of layers rendered by the rendering process into a single background image and displaying the single background image on the display device.
[0016] In one aspect of the present disclosure, based on the position and movement of an imaging device that captures a scene including a background image displayed on a display device that displays a background image consisting of objects based on a 3D model, the objects are divided into multiple layers, the background image is rendered for each layer, and the rendered background images for each of the multiple layers are composited into a single background image and displayed on the display device.
[0017] 13 is a diagram illustrating an overview of an imaging system in a typical imaging studio. FIG. 14 is a timing chart illustrating generation of a background image by the imaging system of FIG. 1. FIG. 15 is another timing chart illustrating generation of a background image by the imaging system of FIG. 1. FIG. 16 is a timing chart illustrating an overview of the present disclosure. FIG. 17 is a diagram illustrating an overview of the present disclosure. FIG. 18 is a diagram illustrating an overview of the present disclosure. FIG. 19 is a diagram illustrating an example configuration of an imaging system of the present disclosure. FIG. 19 is a diagram illustrating an example external configuration of a camera track unit. FIG. 19 is a functional block diagram illustrating functions implemented by the imaging system of FIG. 9. FIG. 11 is a functional block diagram illustrating functions implemented by the rendering engine of FIG. 11. FIG. 12 is a functional block diagram illustrating functions implemented by the rendering processing unit of FIG. 12. FIG. 19 is a diagram illustrating an example of a histogram generated based on a 3D model. FIG. 19 is a diagram illustrating example layers of a background image based on the histogram of FIG. 14. FIG. 19 is a diagram illustrating the overlap width between adjacent layers of FIG. 15. FIG. 20 is a diagram illustrating quality verification based on geometric correction. FIG. 21 is a flowchart illustrating VFX display processing by the imaging system of FIG. 11. FIG. 21 is a diagram illustrating an example configuration of a general-purpose computer.
[0018] Preferred embodiments of the present disclosure will be described in detail below with reference to the accompanying drawings. In this specification and drawings, components having substantially the same functional configurations are designated by the same reference numerals, and redundant description will be omitted.
[0019] Hereinafter, embodiments of the present technology will be described in the following order.
[0020] 1. Overview of the Disclosure 2. Preferred Embodiments 3. Examples of Implementation by Software
[0021] <<1. Overview of the Present Disclosure>> The present disclosure is particularly directed to an imaging system that displays a background image on a large display device installed in a studio, and has performers perform in front of the image to capture images of the performers and the background, such that even if there is a change in the position or imaging direction of the camera relative to the display device, a natural, comfortable, and appropriate background image is displayed without delay.
[0022] Therefore, first, we will explain the simple configuration of an imaging system in which a background image is displayed on a large display device in an imaging studio, and actors perform in front of it to capture images of the actors and the background, and also explain the outline of the present disclosure.
[0023] Fig. 1 is a schematic diagram of an imaging system 11. The imaging system 11 in Fig. 1 is a system that performs imaging as in-camera VFX (Visual Effects) (or includes virtual production, LED wall virtual production, etc.), and the diagram shows some of the equipment that is arranged in an imaging studio.
[0024] The imaging studio is provided with a performance area 50 where performers 51 perform acts or other performances. A large display device is placed at least behind the performance area 50. The type of device used for the display device is not limited, but FIG. 1 shows an example in which an LED wall 34 is used as a large display device.
[0025] A large panel is formed on one LED wall 34 by arranging a plurality of LED panels 41 connected vertically and horizontally. The LED wall 34 displays a background image supplied from the server 32. The size of the LED wall 34 is not particularly limited, but may be any size necessary or sufficient for displaying a background when capturing an image of the performer 51.
[0026] A required number of lights 35 are placed at required positions, such as above or to the sides of the performance area 50, to illuminate the performance area 50.
[0027] A camera 31 is placed near the performance area 50 for capturing, for example, movies or other video content. A cameraman 52 can move the camera 31 and control the capturing direction, angle of view, etc. Of course, the movement of the camera 31 and the angle of view may be controlled by remote control. The camera 31 may also be designed to move and change the angle of view automatically or autonomously. For this reason, the camera 31 may be mounted on a platform or a moving body.
[0028] The camera 31 captures both the performer 51 in the performance area 50 and the image displayed on the LED wall 34. For example, by displaying a landscape as the background image vB on the LED wall 34, it becomes possible to capture an image that is similar to the image of the performer 51 actually being present in the location of that landscape and performing.
[0029] An output monitor 33 is placed near the performance area 50. The video captured by the camera 31 is displayed in real time on this output monitor 33 as a monitor image vM, allowing the director and staff producing the video content to check the video being captured.
[0030] Here, the background image vB will be explained. Even if the background image vB is displayed on the LED wall 34 and a video is taken together with the performer 51, simply displaying the background image vB will result in an unnatural background in the video. This is because the background vB is a two-dimensional image, whereas in reality it is a three-dimensional background with depth.
[0031] For example, the camera 31 can capture images of the performer 51 in the performance area 50 from various directions and can also zoom. The performer 51 does not stand still in one place. Therefore, the actual appearance of the background of the performer 51 should change depending on the position, imaging direction, and angle of view of the camera 31, but such changes cannot be obtained with the background image vB as a flat image. Therefore, the server 32 changes the background image vB so that the background appears similar to how it actually appears, including parallax.
[0032] More specifically, the camera 31 is provided with a SLAM (Simultaneous Localization and Mapping) processing unit 31a that captures images of the surroundings of the camera 31 and identifies its own position from the captured image, and an INS (Inertial Navigation System) processing unit 31b that identifies its own posture corresponding to the imaging direction.
[0033] Based on the position information supplied from the SLAM processing unit 31 a of the camera 31 and the posture information, which is information on the imaging direction supplied from the INS processing unit 31 b, the server 32 uses a 3D model (3D background data) that constitutes the background image to sequentially render the background image in real time using ray tracing or the like, and displays it on the LED wall 34.
[0034] As a result, even if the camera 31 is moved back and forth or left and right, or a zoom operation is performed, the background image is captured as an image corresponding to the change in viewpoint position that accompanies the actual movement of the camera 31.
[0035] The output monitor 33 displays a monitor image vM including the performer 51 and the background, which is an image captured by the camera 31. In other words, the background included in the captured image is an image that has been rendered in real time.
[0036] In this way, the imaging system 11 does not simply display the background image vB in a two-dimensional manner, but changes it in real time so that footage similar to that captured on actual location (location shooting) can be captured.
[0037] <Rendering> Rendering is performed in real time using 3D background data based on information about the imaging position and imaging direction of the camera 31, and is performed using ray tracing, etc., which generally requires a high computational load, takes a long time to generate the background image, and there is a risk of delays in display.
[0038] Therefore, in the rendering process, as shown in the timing chart of Figure 2, it is possible to predict future movement for the processing time related to rendering based on the self-position information supplied from the SLAM processing unit 31a and the imaging direction information supplied from the INS processing unit 31b, and present a background image rendered to reflect the prediction result.
[0039] 2, from time t0 to t21, information on the camera's own position (the position of the camera 31) is generated in the SLAM processing unit 31a, and information on the imaging direction is generated in the INS processing unit 31b. Then, from time t21 to t22, the generated information on the camera's own position and imaging direction is transmitted to the server 32.
[0040] Between times t22 and t23, the server 32 predicts the movement at time t23, which is in the future, based on the position and imaging direction information of the camera 31, by the time period related to rendering, renders a background image, and between times t23 and t24, presents the generated background image on the LED wall 34.
[0041] In the case of the processing of Figure 2, the information on the self-position supplied from the SLAM processing unit 31a at times t0 to t21 and the information on the imaging direction supplied from the INS processing unit 31b are actual measured values, so motion compensation is performed appropriately.
[0042] In contrast, the transfer of information on the self-position and information on the imaging direction from time t21 to time t23, and the background image rendered at time t23 after rendering, are based on motion prediction.
[0043] Furthermore, after rendering, when the background image is presented on the LED wall 34 between times t23 and t24, the motion cannot be compensated.
[0044] As a result, if the motion prediction is appropriate, the delay related to the display of the background image will be limited to the time represented by time t23 to t24, making it possible to suppress the apparent delay.
[0045] However, if the camera 31 operates in a manner contrary to the motion prediction and the motion prediction is incorrect, the generated background image may end up being unnatural.
[0046] Therefore, as shown in the timing chart of FIG. 3, it is conceivable to reduce unnaturalness by verifying the motion prediction result and the actual motion during rendering after rendering and geometrically correcting the difference between the two.
[0047] That is, in Figure 3, the series of processes up to rendering is the same as the processes in Figure 2, but between times t31 and t32, information on the position of the camera 31 and the imaging direction is transferred, and even after rendering has started, the INS processing unit 31b continues to acquire information on the imaging direction.
[0048] Between times t33 and t34, the INS processing unit 31b transmits information on the imaging direction to the server 32 so that at time t34 when the rendering process ends, information on the imaging direction of the camera 31 during rendering corresponding to the motion prediction is transmitted to the server 32.
[0049] At time t34 when rendering ends, server 32 verifies the difference between the motion prediction result based on the position information and imaging direction information transmitted between times t31 and t32 and the actual motion result during rendering transferred between times t33 and t34.
[0050] Then, from time t34 to t35, the server 32 performs geometric correction on the rendered background image to determine the difference between the motion prediction result, which is the verification result, and the actual motion result, and from time t35 to t36, the background image after geometric correction is presented on the LED wall 34.
[0051] As a result, the delay in displaying the background image is limited to the time represented by times t34 to t36, and even if the motion prediction is incorrect, it can be limited to the difference with the actual measured value of the motion during rendering, thereby reducing the unnaturalness of the generated background image.
[0052] However, when geometric correction is applied to a background image rendered using a 3D model, the front-to-back relationship in the depth direction of the 3D objects that make up the 3D model (3D background data) may become unnatural, which may create an unnatural feeling.
[0053] Therefore, in the present disclosure, as shown in the timing chart of FIG. 4 , multiple layers are set for each depth direction of a 3D object that constitutes a 3D model (3D background data), and background images for each of the multiple layers are generated. After that, geometric correction based on the difference between the motion prediction result and the actual motion result is performed on each background image for the multiple layers, and then the background images for the multiple layers are synthesized.
[0054] That is, in the timing chart of Figure 4, information on the position of the camera 31 is generated in the SLAM processing unit 31a from time t0 to t41, and information on the imaging direction of the camera 31 is generated in the INS processing unit 31b, and this information is transmitted to the server 32 from time t41 to t42.
[0055] Between times t42 and t43, the server 32 divides the background image (constituting objects) into multiple layers in the depth direction (number of layers=n in FIG. 4) based on the 3D model (3D background data).
[0056] Between times t43 and t45, the server 32 predicts the movement at time t45, which is in the future for the time period related to rendering, for each layer from the information on the position and imaging direction of the camera 31, and renders the background image.
[0057] On the other hand, from time t41 to time t44, even after the information on the position and imaging direction of the camera 31 has been transferred, the INS processing unit 31b continues to acquire information on the imaging direction of the camera 31.
[0058] At time t44 to t45, the INS processing unit 31b transmits information on the imaging direction of the camera 31 during rendering that corresponds to the motion prediction result to the server 32 at time t45 when the rendering process ends.
[0059] At time t45 when rendering is completed, server 32 verifies the difference between the motion prediction results based on the information on the position and imaging direction of camera 31 transmitted between times t41 and t42 and the actual motion results during rendering transferred between times t44 and t45.
[0060] Between times t45 and t46, the server 32 performs geometric correction on the rendered background image corresponding to the difference between the motion prediction result, which is the verification result for each layer, and the actual motion result.
[0061] Then, from time t46 to t47, the server 32 composites the background images of the multiple layers that have been subjected to geometric correction, and from time t47 to t48, presents the composite background image of the multiple layers on the LED wall 34.
[0062] Here, layers will be described using, as an example, the background image vB1 displayed on the LED wall 34 in FIG.
[0063] 5 is composed of object B1, which is a desk located closest to camera 31, object B2, which is a fence located next closest to camera 31, and object B3, which is the farthest object and includes a building such as a sky. The relationship between the distance from camera 31 and the angle of view is as shown in FIG.
[0064] 5 and 6, when objects B1, B2, and B3 are set according to their respective depth directions based on a 3D model (3D background data), layers L1 to L3 are set by dividing the objects in the depth direction. That is, layers L1 to L3 of the background image are set according to the distance from camera 31 based on the 3D model, and in the case of FIG. 6, layers L1, L2, and L3 are set in order of proximity to camera 31.
[0065] Based on the 3D models (3D background data) of these layers L1 to L3, the server 32 generates background images R1 to R3 corresponding to each of the objects B1 to B3 according to the distance from the camera 31, as shown in Figure 7, for example.
[0066] Then, the server 32 applies geometric correction to each of the generated background images R1 to R3 based on the difference between the motion prediction result and the actual motion result to generate background images R1' to R3', and by combining these as shown in Figure 8, generates background image vB1 (= R1' + R2' + R3') and presents it on the LED wall 34.
[0067] This series of processes generates background images for each of multiple layers set in the depth direction, making it possible to suppress unnaturalness related to the front-to-back relationship of 3D objects. Furthermore, even for changes in movement that tend to cause noticeable delays, geometric correction of the difference between the predicted movement result and the actual movement result makes it possible to suppress unnaturalness.
[0068] As a result, even if the position or imaging direction of the camera relative to the display device changes, it is possible to display a natural, natural, and appropriate background image without delay.
[0069] <<2. Preferred Embodiment>> Next, an imaging system according to a preferred embodiment of the present disclosure will be described with reference to FIG.
[0070] The imaging system 101 in FIG. 9 is composed of a camera 111, a rendering server 112, an external monitor 113, an LED wall 114, and a performance area 141.
[0071] Here, the camera 111, rendering server 112, external monitor 113, LED wall 114, and performance area 141 have configurations corresponding to the camera 31, server 32, output monitor 33, LED wall 34, and performance area 50 in the imaging system 11 of Fig. 1, and have almost the same basic functions. Of these, the external monitor 113, LED wall 114, and performance area 141 in particular have completely the same configurations as the output monitor 33, LED wall 34, and performance area 50, respectively.
[0072] Therefore, the imaging system 101 in FIG. 9 will be mainly described with reference to the camera 111 and the rendering server 112.
[0073] The camera 111 has the same functions as the camera 31, but is equipped with a camera track unit 121 that integrates functions corresponding to the SLAM processing unit 31a and the INS processing unit 31b.
[0074] <Camera Tracking Unit> More specifically, the camera tracking unit 121 is configured as shown in FIG. 10, for example, and includes cameras 151-1 to 151-5, a SLAM processing unit 152, an IMU 153, and an INS processing unit 154.
[0075] Cameras 151-1 to 151-5 are monocular RGB cameras that are attached to the exterior of camera track unit 121, capture images of the surroundings of camera 111, and supply the captured images to SLAM processing unit 152. In other words, cameras 151-1 to 151-5 do not capture images of performance area 141 together with the background image displayed on LED wall 114, as camera 111 does.
[0076] 10 shows an example of a configuration in which five cameras 151 are provided in the camera track unit 121, but it is sufficient to provide at least one camera 151. Even if multiple cameras 151 are provided, such as cameras 151-1 to 151-5, it is sufficient that at least one camera is in operation in order to reduce the processing load and power consumption.
[0077] The IMU (Inertial Measurement Unit) 153 is equipped with a three-axis acceleration sensor and a three-axis angular velocity sensor, and detects the three-axis acceleration, which is the translational motion of the camera 111, and the three-axis angular velocity, which is the rotational motion, and outputs these to the SLAM processing unit 152 and the INS processing unit 154.
[0078] The SLAM (Simultaneous Localization and Mapping) processing unit 152 estimates the depth to objects and the self-position in the captured image by using feature points and brightness values of the image based on the images of the surroundings of the camera 111 captured by the cameras 151-1 to 151-5, and outputs the results to the rendering server 112.
[0079] In addition, the SLAM processing unit 152 obtains velocity by integrating the acceleration in three axes directions and angular velocity in three axes directions based on the acceleration and angular velocity in three axes directions from the IMU 153, estimates distance by integrating the velocity, and obtains attitude (angle) by integrating the angular velocity, thereby estimating the self-position of the camera 111 with higher accuracy.
[0080] The INS (Inertial Navigation System) processing unit 154 obtains the velocity by integrating the acceleration, estimates the distance by integrating the velocity, and estimates the angle (attitude) of the camera 111 by integrating the angular velocity, based on the acceleration in the three axial directions and the angular velocity of the camera 111 supplied from the IMU 153. In this way, the INS processing unit 154 detects the velocity, position, and attitude of the camera 111 as movement, and outputs them to the rendering server 112.
[0081] Although the camera 151 is described as a monocular RGB camera, other cameras may be used. For example, a stereo camera, a depth camera, or a camera that captures distance images such as LiDAR (Light Detection and Ranging, Laser Imaging Detection and Ranging) may be used. In this case, however, the SLAM processing unit 152 and the INS processing unit 154 perform processing based on the distance image. Furthermore, the SLAM processing unit 152 may use images captured by the camera 111 in addition to the images captured by the cameras 151-1 to 151-5.
[0082] <Functions Realized by Imaging System> Next, functions realized by the imaging system 101 will be described with reference to the functional block diagram of FIG.
[0083] The camera 111 includes an imaging unit 201 and a camera track unit 121. The imaging unit 201 is composed of an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor and an optical mechanism that adjusts the angle of view and focus, and captures images of the performance area 141 and the LED wall 114, and outputs the captured images to the camera track unit 121, the external monitor 113, and the rendering engine 221 of the rendering server 112.
[0084] The camera 111 supplies the rendering engine 221 with imaging information such as the angle of view, focal length (focus position), F-number, shutter speed (exposure time), and lens information.
[0085] The camera track unit 121 has the configuration described with reference to Figure 10, and supplies the position and movement of the camera 111 to the rendering engine 221 of the rendering server 112, correlating it with a timestamp indicating the timing of capturing the image supplied from the imaging unit 201.
[0086] 11 shows an example of a configuration in which the imaging unit 201 also supplies the captured images to the camera track unit 121. Therefore, the SLAM processing unit 152 of the camera track unit 121 is configured to be able to estimate its own position using images captured by the camera 111 in addition to images captured by the cameras 151-1 to 151-5. However, the SLAM processing unit 152 may also perform SLAM processing using some of the images supplied from the cameras 151-1 to 151-5 and the camera 111.
[0087] The rendering server 112 includes a rendering engine 221 , a storage unit 222 that stores a background 3D model, and a control monitor 223 .
[0088] The memory unit 222 stores 3D background data indicating the location and range of objects that make up the background image as a 3D model, functions as a database (DB) for the 3D background data, and supplies the 3D model to the rendering engine 221 as needed.
[0089] The rendering engine 221 generates a background image vB to be displayed on the LED wall 114. To this end, the rendering engine 221 reads out necessary 3D background data from the storage unit 222. Then, the rendering engine 221 generates the background image vB by rendering the 3D background data in a form viewed from pre-specified spatial coordinates.
[0090] For processing each frame, the rendering engine 221 uses the position information and attitude information (information on the imaging direction) of the camera 111 supplied from the camera track unit 121 and the imaging information supplied from the imaging unit 201 to identify the viewpoint position and the like relative to the 3D background data, divides the background image into layers according to the distance from the camera 111, and renders each layer.
[0091] The rendering engine 221 combines the generated background images for each layer to complete the background image, and supplies it to the display control unit 241 of the LED wall 114 on a frame-by-frame basis.
[0092] The detailed functions realized by the rendering engine 221 will be described later with reference to FIG.
[0093] The display control unit 241 of the LED wall 114 generates a divided video signal nD by dividing a background image consisting of one frame of video data into video portions to be displayed on each LED panel 131, and transmits the divided video signal nD to each LED panel 131.
[0094] The display control unit 241 may be configured to calibrate individual differences / manufacturing errors in color development and the like between the LED panels 131 in advance.
[0095] Note that these processes may be performed by the rendering engine 221 without providing the display control unit 241. In other words, the rendering engine 221 may generate the divided video signal nD, perform calibration, and transmit the divided video signal nD to each LED panel 131.
[0096] Each LED processor 242 drives the LED panel 131 based on the divided video signal nD that it receives, thereby displaying the entire background image vB on the LED wall 114. The background image vB is a background image that is rendered according to the position and imaging direction of the camera 111 at that time.
[0097] The camera 111 captures the performance area 141, including the background image vB displayed on the LED wall 114. The image captured by the camera 111 is recorded on a recording medium inside the camera 111 or in an external recording device (not shown), and is also supplied in real time to the external monitor 113 and displayed as a monitor image vM.
[0098] The control monitor 223 displays an operation image vOP for controlling the rendering engine 221. A user (e.g., an engineer) operating the rendering server 112 can perform necessary settings and operations related to the rendering of the background image vB while viewing the operation image vOP. Furthermore, if an error occurs in the background image vB rendered by the rendering engine 221, such as an object being placed in an inappropriate position due to geometric correction, the control monitor 223 displays information indicating that an error has occurred in the rendering.
[0099] <Functions Realized by Rendering Engine> Next, functions realized by the rendering engine 221 will be described with reference to the functional block diagram of FIG.
[0100] The rendering engine 221 includes a motion prediction unit 271 , a rendering processing unit 272 , a motion prediction verification unit 273 , a geometric correction unit 274 , a background synthesis unit 275 , and a quality verification unit 276 .
[0101] The motion prediction unit 271 linearly predicts the future motion of the camera 111 for the processing time related to rendering in the rendering processing unit 272 based on the position information of the camera 111 supplied from the SLAM processing unit 152 of the camera track unit 121 and the attitude information of the camera 111 corresponding to the imaging direction of the camera 111 supplied from the INS processing unit 154, and supplies the motion prediction result to the rendering processing unit 272 and the motion prediction verification unit 273.
[0102] The rendering processing unit 272 divides the objects constituting the background image into a plurality of layers 1 to n based on the background 3D model stored in the storage unit 222. Furthermore, the rendering processing unit 272 renders the objects by projective transformation for each of the plurality of layers 1 to n resulting from the division based on the motion prediction result of the camera 111 supplied from the motion prediction unit 271, thereby generating corresponding first to nth background images and supplying them to the geometric correction unit 274.
[0103] The detailed configuration of the rendering processing unit 272 will be described later with reference to FIG.
[0104] As explained with reference to the timing chart of Figure 4, after the motion prediction unit 271 has supplied motion prediction information based on the position information of the camera 111 up to that point and the posture information that serves as motion compensation, and even while rendering is being performed in the rendering processing unit 272, the motion prediction verification unit 273 continues to acquire the posture information of the camera 111 that serves as motion compensation information supplied from the INS processing unit 154 of the camera track unit 121, and acquires the movement results of the camera 111 up to the time when rendering in the rendering processing unit 272 is completed (to be precise, the time before transfer, since there is a time required to transfer the motion compensation information).
[0105] The motion prediction verification unit 273 then verifies the motion prediction result by calculating the difference between the motion prediction result at the time when the rendering process supplied from the motion prediction unit 271 ends and the motion result at the time when the actual rendering process is completed (actually, a time in the past by the transfer time), and supplies the difference between the motion prediction result that serves as the verification result and the actual motion result to the geometric correction unit 274.
[0106] The geometric correction unit 274 performs geometric correction on each of the first background image to the nth background image of each of layers 1 to n supplied from the rendering processing unit 272 based on the difference between the motion prediction result, which is the verification result supplied from the motion prediction verification unit 273, and the actual motion result, and supplies the corrected images to the background synthesis unit 275.
[0107] At this time, the geometric correction unit 274 also supplies the first to n-th background images before and after the geometric correction to the quality verification unit 276, respectively.
[0108] The background synthesis unit 275 synthesizes the first background image to the nth background image of each of the geometrically corrected layers 1 to n supplied by the geometric correction unit 274 to generate a single background image, which is output to the LED wall 114 for display.
[0109] The quality verification unit 276 verifies the quality of the background images by calculating the distance between the coordinates of the representative positions of each object in the first background image through the nth background image before and after the geometric correction and comparing it with a predetermined threshold. For example, if the distance between the coordinates of the representative positions of each object in the first background image through the nth background image of each layer before and after the geometric correction is longer (farther apart) than a predetermined value for a predetermined number of objects, the quality verification unit 276 determines that there is a quality abnormality in the background image due to the geometric correction (an error has occurred). In such a case, the quality verification unit 276 controls, for example, the control monitor 223 to present information indicating that there is an abnormality in the quality of the background image due to the geometric correction and that an error has occurred.
[0110] For example, if the quality verification unit 276 detects an abnormality in the quality of the background image due to geometric correction and presents information indicating that an error has occurred on the control monitor 223, the background synthesis unit 275 may display a background image in which no error had occurred immediately before.
[0111] 13, functions realized by the rendering processing unit 272 will be described. The rendering processing unit 272 includes a motion blur calculation unit 291, a spatial blur calculation unit 292, a layer division unit 293, and a background drawing unit 294.
[0112] The motion blur calculation unit 291 calculates the amount of blur that occurs in response to the movement of the camera 111 and objects within the exposure time based on the shutter speed (exposure time) and the amount of movement of the camera 111 supplied as motion compensation information from the INS processing unit 154, and supplies this as a motion blur amount motion to the layer division unit 293 and background drawing unit 294.
[0113] The spatial blur calculation unit 292 calculates the amount of spatial blur (focus) that occurs due to focus adjustment, etc., based on the F-number and focal length (focus position) included in the imaging information supplied from the camera 111, and supplies the calculated amount of spatial blur (focus) to the layer division unit 293 and the background drawing unit 294.
[0114] The layer dividing unit 293 reads out the background 3D model stored in the storage unit 222, and divides the objects included in the background image into a plurality of layers 1 to n according to distance based on the 3D model. The division of layers will be described in detail later with reference to FIGS. 14 to 16.
[0115] The background drawing unit 294 divides the objects represented by the 3D model for each layer supplied by the layer dividing unit 293, taking into consideration the amount of motion blur supplied by the motion blur calculation unit 291 and the amount of spatial blur supplied by the spatial blur calculation unit 292, and draws a background image for each layer by rendering the objects represented by the 3D model for each divided layer by projective transformation.
[0116] The background rendering unit 294 may perform simplified rendering when the amount of movement of the camera 111 is large and the amount of blurring due to the movement is greater than a predetermined value.
[0117] 14 to 16, a description will be given of layer division by the layer division unit 293. The layer division unit 293 generates a histogram according to the distance from the camera 111 of an object generated based on a 3D model, and supplies each class as a layer to the background drawing unit 294 by clustering the histogram into a predetermined number of classes.
[0118] Here, as a 3D model, for example, consider the case where objects B11 to B17 exist within a range θ that is α (α > 1) times the angle of view of camera 111, centered on the optical axis of camera 111, and sandwiched between dotted lines starting from camera 111 in the figure, as shown in the left part of Figure 14.
[0119] 14, objects B12 and B13 are present, in order from left to right, at positions closest to and almost directly in front of camera 111. Also, object B11 is present further back than objects B12 and B13, on the left side of range θ in the figure, as viewed from camera 111. Furthermore, objects B13 to B16 are present, in order from left to right, further back than objects B11 to B13, as viewed from camera 111. Furthermore, object B17 is present further back than objects B11 to B13, to the right of object B16, and extends long in the depth direction, as viewed from camera 111.
[0120] The layer splitting unit 293 first quantizes the objects B11 to B17 according to distance by generating a histogram according to distance, as shown in the right part of Figure 14, based on the 3D model and according to the distance from the camera 111.
[0121] The bin width, which is the width of the distance unit for counting the histogram and is set according to the distance from the camera 111, expresses the depth resolution according to the distance from the camera 111. As shown in the center of FIG. 14, the bin width changes exponentially (y=a e x (y: bin width, x: distance) are set in advance.
[0122] As a result, as shown in the right part of FIG. 14, the bin widths W1 to W7, which are set based on the distance at which the object exists, are set so as to exponentially increase in accordance with the distance, the farther the object is.
[0123] For objects that exist across bins, which are the units for counting the histogram, the histogram is counted after dividing the objects into units of each bin. This allows all objects, including those that exist continuously in the distance direction as well as on the ground surface and in the horizontal direction, to be counted in the histogram using the set bin as a unit without omission.
[0124] Next, as shown in FIG. 15, the layer dividing unit 293 clusters the histogram based on the representative depth and depth width for each object, thereby setting each class to a layer.
[0125] In addition, when performing clustering, the number of classes is set in advance according to system resources. That is, the greater the number of classes, the greater the number of background images generated, which improves the accuracy of the position of each object in the depth direction from the camera 111. However, the higher the processing load, the longer the processing time and the more likely delays will occur. For this reason, it is desirable to set the number of classes in advance according to system resources.
[0126] In Figure 15, the left part shows a histogram (similar to the right part of Figure 14), and the right part shows an example of clustering the histogram on the left, showing that the histogram is classified into three classes, and each class is set to layers 1 to 3 (layer1 to 3).
[0127] Furthermore, layers 1 to 3 are set so that the boundaries of adjacent classes (adjacent layers) partially overlap by an overlap width dc, as shown in FIG.
[0128] More specifically, the layer dividing unit 293 calculates the overlap width dc (Figure 15) (= length_layer[i] = f(motion, focus, position_layer[i]) using a function f that multiplies the depth representative value position_layer[i] of the target layer (layer i) by a value obtained by multiplying the amount of motion blur and the amount of spatial blur by a predetermined coefficient, sets the layer by reflecting the calculated overlap width dc, and supplies it to the background drawing unit 294.
[0129] With this processing, the greater the amount of motion blur and spatial blur, and the greater the distance from camera 111, the larger the overlap width dc at the boundary between adjacent layers is set, making it possible to set the overlap width of layers corresponding to the distance in the depth direction and the amount of blur.
[0130] <Processing by Quality Verification Unit> Next, quality verification related to geometric correction by the quality verification unit 276 will be described with reference to FIG.
[0131] The background image generated for each layer is geometrically corrected based on the difference between the motion prediction result and the previous motion result, but the geometric correction may cause a malfunction, and the quality verification unit 276 verifies the quality according to the degree of malfunction.
[0132] For example, consider a case where a representative point Pd of a predetermined object in a 3D model, such as that indicated by a star in Figure 17, is defined by a center of gravity or the like, and a background image Rx of a predetermined layer is generated by rendering the optical axis center position Pc of the camera 111 through projective transformation based on the motion prediction results. Here, the coordinates of the 2D position of the representative point Pd of the predetermined object indicated by the star in the background image Rx are assumed to be the 2D position (u, v).
[0133] Here, when the position or imaging direction of the camera 111 changes based on the motion result after rendering, changing from the optical axis center position Pc to the optical axis center position P'c, the geometric correction unit 274 converts the background image Rx to a background image Rx' by geometrically correcting the background image Rx by the difference between the motion prediction result and the motion result. At this time, it is assumed that the 2D position coordinates of the representative point Pd indicated by a star are (u', v').
[0134] The quality verification unit 276 calculates the difference between the 2D position coordinates (u, v) of the representative point Pd indicated by a star in the background image Rx and the 2D position coordinates (u', v') of the representative point Pd indicated by a star in the background image Rx' as the deviation amount (error) before and after geometric correction, and judges the quality based on whether it is greater than a predetermined threshold.
[0135] For example, if a similar quality assessment is performed on all objects, and the difference (deviation) between the representative points before and after geometric correction is greater than a predetermined threshold for more than a predetermined number of objects (e.g., more than half) of all objects, and a decrease in the quality of the background image due to rendering is expected, the quality verification unit 276 controls the control monitor 223 to present information indicating that an error has occurred due to geometric correction.
[0136] In addition, if the quality verification unit 276 detects that the quality of the background image is expected to deteriorate due to rendering, it may notify the background synthesis unit 275, and the background synthesis unit 275 may output to the LED wall 114 the background image that was present immediately before the error was detected.
[0137] <VFX Display Processing> Next, VFX display processing by the imaging system 101 of FIG. 11 will be described with reference to the flowchart of FIG.
[0138] In step S31, the SLAM processing unit 152 of the camera track unit 121 estimates the position (self-position) of the camera 111 based on the image captured by the camera 151 and the angular velocity and acceleration information detected by the IMU 153, and outputs it to the motion prediction unit 271 of the rendering engine 221.
[0139] In step S32, the INS processing unit 154 outputs motion compensation information consisting of the movement and attitude of the camera 111 to the motion prediction unit 271 of the rendering engine 221 based on the acceleration and angular velocity information detected by the IMU 153.
[0140] In step S33, the motion prediction unit 271 of the rendering engine 221 predicts the position and orientation of the camera 111 at a timing after the rendering processing by the rendering processing unit 272 using linear prediction based on the position information and motion compensation information of the camera 111 supplied from the camera track unit 121, and outputs the motion prediction result to the rendering processing unit 272 and the motion prediction verification unit 273.
[0141] In step S34, the rendering processing unit 272 controls the layer dividing unit 293 to divide the object into a predetermined number of layers based on the predicted position and orientation of the camera 111 at the timing after the rendering processing by the rendering processing unit 272, and the 3D model for the background stored in the memory unit 222.
[0142] At this time, the motion blur calculation unit 291 of the rendering processing unit 272 calculates the amount of motion blur based on information on the amount of motion included in the motion compensation information. Also, the spatial blur calculation unit 292 calculates the amount of spatial blur based on information such as the F-number and focus position (focal length) supplied from the camera 111.
[0143] The layer dividing unit 293 determines the overlap width at the boundary between layers from the amount of motion blur, the amount of spatial blur, and the depth representative value of the target layer.
[0144] The layer dividing unit 293 then generates a histogram from the arrangement of the objects based on the 3D model, and divides the objects into multiple layers based on the generated histogram. At this time, the layer dividing unit 293 sets the layers taking into account the overlap width at the boundary between the layers determined by the above-mentioned processing.
[0145] In step S35, the background drawing unit 294 of the rendering processing unit 272 generates a plurality of background images for each layer using the objects divided into layers based on the motion compensation information, and outputs them to the geometric correction unit 274.
[0146] In step S36, the INS processing unit 154 supplies the motion compensation information of the camera 111 during rendering to the motion prediction verification unit 273 as the actual motion result during rendering.
[0147] In step S37 , the motion prediction verification unit 273 calculates the difference between the motion prediction result and the motion result, verifies the motion prediction result, and outputs the verification result to the geometric correction unit 274 .
[0148] In step S38, the geometric correction unit 274 performs geometric correction on the background image of each layer based on the difference between the motion prediction result, which is the verification result, and the motion result, and outputs the corrected image to the background synthesis unit 275 and the quality verification unit 276. At this time, the geometric correction unit 274 also outputs the background image before the geometric correction to the quality verification unit 276.
[0149] In step S39, the background synthesis unit 275 synthesizes the background images for each layer to generate one background image, and outputs it to the LED wall 114 to be displayed as a VFX image.
[0150] In step S40, the quality verification unit 276 calculates the distance between the coordinates of the representative positions of each object in the background image before and after the geometric correction in each layer as the amount of deviation.
[0151] In step S41, the quality verification unit 276 determines whether or not the number of objects in which the amount of deviation due to geometric correction is greater than a predetermined value, that is, the number of objects in which an error due to geometric correction has occurred, is equal to or greater than a predetermined number.
[0152] In step S41, if the number of objects in which the amount of deviation due to geometric correction is greater than a predetermined value, that is, the number of objects in which an error due to geometric correction has occurred, is equal to or greater than a predetermined number, the process proceeds to step S42.
[0153] In step S42, the quality verification unit 276 controls the control monitor 223 to display information indicating that an error has occurred due to geometric correction.
[0154] In step S41, if the number of objects whose deviation due to geometric correction is greater than a predetermined value, i.e., the number of objects in which an error due to geometric correction has occurred, is less than the predetermined number, no error due to geometric correction has occurred, and therefore the processing of step S42 is skipped.
[0155] In step S43, it is determined whether or not an instruction to end the process has been given. If an instruction to end the process has not been given, the process returns to step S31, and the subsequent steps are repeated.
[0156] Then, in step S43, if an instruction to end the process is given, the process ends.
[0157] Through the above processing, the background image, which is the rendering result based on the predicted camera movement results before rendering, is geometrically corrected based on the difference with the camera movement results during rendering, thereby making it possible to reduce the apparent delay.
[0158] Furthermore, the background image is generated by dividing it into layers according to the distance from the camera, and each layer is geometrically corrected before being combined, so there is no inconsistency in the anteroposterior relationship of objects, and it is possible to generate a natural, seamless background image that corresponds to changes in the camera position and imaging direction.
[0159] As a result, in an imaging system that images the performers and the background by displaying a background image on a large display device installed in a studio and having the performers perform in front of it, even if there are changes in the camera position or imaging direction relative to the display device, it is possible to display a natural, appropriate background image without delay and without any sense of incongruity.
[0160] <<3. Example of Execution by Software>> The above-described series of processes can be executed by hardware, but can also be executed by software. When the series of processes is executed by software, the program constituting the software is installed from a recording medium into a computer incorporated in dedicated hardware, or into, for example, a general-purpose computer that can execute various functions by installing various programs.
[0161] 19 shows an example of the configuration of a general-purpose computer. This computer has a built-in CPU (Central Processing Unit) 1001. An input / output interface 1005 is connected to the CPU 1001 via a bus 1004. A ROM (Read Only Memory) 1002 and a RAM (Random Access Memory) 1003 are connected to the bus 1004.
[0162] The input / output interface 1005 is connected to an input unit 1006 including input devices such as a keyboard and a mouse through which a user inputs operation commands, an output unit 1007 that outputs a processing operation screen and images of processing results to a display device, a storage unit 1008 including a hard disk drive or the like that stores programs and various data, and a communication unit 1009 including a LAN (Local Area Network) adapter or the like that executes communication processing via a network typified by the Internet. Also connected is a drive 1010 that reads and writes data from / to a removable storage medium 1011 such as a magnetic disk (including a flexible disk), an optical disk (including a CD-ROM (Compact Disc-Read Only Memory) and a DVD (Digital Versatile Disc)), a magneto-optical disk (including an MD (Mini Disc)), or a semiconductor memory.
[0163] The CPU 1001 executes various processes in accordance with a program stored in a ROM 1002 or a program read from a removable storage medium 1011 such as a magnetic disk, optical disk, magneto-optical disk, or semiconductor memory, installed in a storage unit 1008, and loaded from the storage unit 1008 into a RAM 1003. The RAM 1003 also stores data necessary for the CPU 1001 to execute various processes as appropriate.
[0164] In a computer configured as described above, the CPU 1001 performs the above-described series of processes by, for example, loading a program stored in the memory unit 1008 into the RAM 1003 via the input / output interface 1005 and the bus 1004 and executing it.
[0165] The program executed by the computer (CPU 1001) can be provided by being recorded on a removable storage medium 1011 such as a package medium, for example. The program can also be provided via a wired or wireless transmission medium such as a local area network, the Internet, or digital satellite broadcasting.
[0166] In a computer, a program can be installed in the storage unit 1008 via the input / output interface 1005 by inserting a removable storage medium 1011 into the drive 1010. The program can also be received by the communication unit 1009 via a wired or wireless transmission medium and installed in the storage unit 1008. Alternatively, the program can be installed in advance in the ROM 1002 or the storage unit 1008.
[0167] The program executed by the computer may be a program that processes in chronological order according to the order described in this specification, or may be a program that processes in parallel or at the required timing, such as when called.
[0168] 19 realizes the function of the rendering engine 221 in FIG. 11, and the storage unit 1008 realizes the function of the storage unit 222 that stores the background 3D model in FIG.
[0169] In this specification, a system refers to a collection of multiple components (devices, modules (components), etc.), regardless of whether all of the components are contained in the same housing. Therefore, multiple devices housed in separate housings and connected via a network, and a single device with multiple modules housed in a single housing, are both systems.
[0170] Furthermore, the embodiments of the present disclosure are not limited to the above-described embodiments, and various modifications are possible within the scope of the gist of the present disclosure.
[0171] For example, the present disclosure can be configured as a cloud computing system in which a single function is shared and processed collaboratively by multiple devices via a network.
[0172] Furthermore, each step described in the above flowchart can be executed by one device, or can be shared and executed by a plurality of devices.
[0173] Furthermore, when one step includes multiple processes, the multiple processes included in that one step can be executed by one device or can be shared and executed by multiple devices.
[0174] The present disclosure may also be configured as follows. <1> An information processing device comprising: a rendering processing unit that divides an object into a plurality of layers based on a position and movement of an imaging device that captures a scene including the background image displayed on a display device, the background image being made up of objects based on a 3D model, relative to the display device; and a background synthesis unit that synthesizes the background images for the plurality of layers rendered by the rendering processing unit into a single background image and displays the single background image on the display device. <2> The information processing device described in <1>, further including a motion prediction unit that predicts a position and movement of the imaging device at a timing when rendering by the rendering processing unit is completed, based on the position and movement of the imaging device relative to the display device, wherein the rendering processing unit divides the object into a plurality of layers and renders the background image for each layer based on a position and motion prediction result of the motion prediction unit. <3> The information processing device described in <1>, wherein the rendering processing unit divides the object into a plurality of layers based on a distance from the imaging device in the 3D model, and renders the background image for each layer. <4> The information processing device according to <2>, further comprising: a motion prediction result verification unit that verifies a motion prediction result based on a motion result based on a position and motion of the image capture device relative to the display device at a timing when rendering by the rendering processing unit is completed, and a difference between the position and motion that constitutes the motion prediction result; a geometric correction unit that geometrically corrects the background image rendered by the rendering processing unit for each of the layers based on the difference between the motion prediction result and the motion result, which constitutes the verification result of the motion prediction result verification unit; and the background synthesis unit synthesizes the background images geometrically corrected by the geometric correction unit into a single background image for each of the plurality of layers. <5> The information processing device according to <4>, further comprising: a quality verification unit that verifies quality of the geometrically corrected background image for each of the plurality of layers, based on a distance between the coordinates of the objects in the background image geometrically corrected by the geometric correction unit and the background image before the geometric correction.<6> The information processing device described in <5>, wherein the quality verification unit verifies the quality of the geometrically corrected background image for each of the plurality of layers based on the number of objects for which the coordinate distance between the background image geometrically corrected by the geometric correction unit and the background image before the geometric correction is greater than a predetermined distance. <7> The rendering processing unit quantizes the objects by generating a histogram with a unit width set according to the distance of the object from the imaging device, divides the histogram into a plurality of layers, and renders the background image for each layer. <8> The information processing device described in <7>, wherein adjacent layers are set to overlap by a predetermined width. <9> The information processing device described in <8>, wherein the predetermined overlap width set between adjacent layers is set based on the movement of the imaging device and the focal length of the imaging device. <10> The information processing device described in <9>, wherein the predetermined overlap width set between adjacent layers is set based on the amount of motion blur corresponding to the movement of the imaging device and the amount of spatial blur corresponding to the focal length of the imaging device. <11> The information processing device described in <9>, wherein the rendering processing unit divides the object into a plurality of layers based on a position and movement of the imaging device relative to the display device, and renders the background image for each layer based on an amount of motion blur corresponding to the movement of the imaging device and an amount of spatial blur corresponding to a focal length of the imaging device. <12> The information processing device described in <7>, wherein a width that is set according to the distance from the imaging device and that serves as a unit of the histogram is set to increase exponentially according to the distance to the imaging device. <13> The information processing device described in <1>, wherein the imaging device includes a position and motion detection unit that detects the position and movement of the imaging device relative to the display device, and the rendering processing unit divides the object into a plurality of layers based on the position and movement of the imaging device relative to the display device detected by the position and motion detection unit, and renders the background image for each layer.<14> The information processing device according to <13>, wherein the position and motion detection unit includes: an IMU (Inertial Measurement Unit) that detects acceleration and angular velocity of the imaging device; a motion detection unit that detects motion of the imaging device based on the acceleration and the angular velocity detected by the IMU; an imaging unit that captures an image of a periphery of the imaging device; and a SLAM processing unit that detects a position of the imaging device by SLAM (Simultaneous Localization and Mapping) based on the image captured by the imaging unit. <15> An information processing method, comprising: a rendering process that divides objects into a plurality of layers and renders the background image for each layer, based on a position and movement of an imaging device that captures a scene including the background image displayed on a display device that displays a background image made of objects based on a 3D model; and a background synthesis process that synthesizes the background images for the plurality of layers rendered by the rendering process into a single background image and displays the background image on the display device. <16> A program that causes a computer to function as: a rendering processing unit that divides objects into a plurality of layers based on the position and movement of an imaging device that captures a scene including a background image displayed on a display device, the background image being made up of objects based on a 3D model, and renders the background image for each of the layers; and a background synthesis unit that synthesizes the background images for each of the layers rendered by the rendering processing unit into a single background image and displays the single background image on the display device.
[0175] 101 Imaging system, 111 Camera, 112 Rendering server, 113 External monitor, 114 LED wall, 121 Camera track unit, 131 LED panel, 141 Performance area, 151, 151-1 to 151-5 Camera, 152 SLAM processing unit, 153 IMU, 154 INS processing unit, 201 Imaging unit, 221 Rendering engine, 222 Memory unit, 223 Control monitor, 241 Display control unit, 242 LED processor, 271 Motion prediction unit, 272 Rendering processing unit, 273 Motion prediction verification unit, 274 Geometric correction unit, 275 Background synthesis unit, 276 Quality verification unit, 291 Motion blur calculation unit, 292 Spatial blur calculation unit, 293 Layer division unit, 294 Background drawing section
Claims
1. An information processing device comprising: a rendering processing unit that divides objects into multiple layers based on the position and movement of an imaging device that captures a scene including a background image displayed on a display device, the background image being made up of objects based on a 3D model, and renders the background image for each layer; and a background synthesis unit that synthesizes the background images for each of the multiple layers rendered by the rendering processing unit into a single background image and displays it on the display device.
2. The information processing device of claim 1, further comprising a motion prediction unit that predicts the position and movement of the imaging device at the time rendering by the rendering processing unit is completed based on the position and movement of the imaging device relative to the display device, and the rendering processing unit divides the object into multiple layers based on the position and motion prediction results of the motion prediction unit, and renders the background image for each of the layers.
3. The information processing device according to claim 1, wherein the rendering processing unit divides the object into a plurality of layers based on a distance from the imaging device in the 3D model, and renders the background image for each layer.
4. The information processing device according to claim 2, further comprising: a motion prediction result verification unit that verifies the motion prediction result based on a motion result based on the position and movement of the imaging device relative to the display device at the time rendering by the rendering processing unit is completed, and a difference between the position and movement that constitutes the motion prediction result; and a geometric correction unit that geometrically corrects the background image rendered by the rendering processing unit for each layer based on the difference between the motion prediction result and the motion result that constitutes the verification result of the motion prediction result verification unit, and the background synthesis unit that synthesizes the background images geometrically corrected by the geometric correction unit for each of the multiple layers into a single background image.
5. The information processing device according to claim 4, further comprising a quality verification unit that verifies the quality of the geometrically corrected background image for each of the plurality of layers based on the distance between the coordinates of the objects in the background image geometrically corrected by the geometric correction unit and the background image before the geometric correction.
6. The information processing device described in claim 5, wherein the quality verification unit verifies the quality of the geometrically corrected background image for each of the plurality of layers according to the number of objects whose coordinate inter-distance between the background image geometrically corrected by the geometric correction unit and the background image before the geometric correction is greater than a predetermined distance.
7. The information processing device according to claim 3, wherein the rendering processing unit quantizes the object by generating a histogram with a unit width set according to the distance of the object from the imaging device, divides the histogram into a plurality of layers, divides the object into the plurality of layers, and renders the background image for each of the layers.
8. The information processing device according to claim 7, wherein adjacent layers are set to overlap by a predetermined width.
9. The information processing device according to claim 8, wherein the predetermined overlap width set between adjacent layers is set based on the movement of the imaging device and the focal length of the imaging device.
10. An information processing device according to claim 9, wherein the predetermined overlap width set between adjacent layers is set based on the amount of motion blur according to the movement of the imaging device and the amount of spatial blur according to the focal length of the imaging device.
11. The information processing device according to claim 9, wherein the rendering processing unit divides the object into multiple layers based on the position and movement of the imaging device relative to the display device, and renders the background image for each layer based on the amount of motion blur corresponding to the movement of the imaging device and the amount of spatial blur corresponding to the focal length of the imaging device.
12. The information processing device according to claim 7, wherein the width of the histogram unit, which is set according to the distance from the imaging device, is set to increase exponentially according to the distance from the imaging device.
13. The information processing device according to claim 1, wherein the imaging device includes a position and motion detection unit that detects the position and motion of the imaging device relative to the display device, and the rendering processing unit divides the object into multiple layers based on the position and motion of the imaging device relative to the display device detected by the position and motion detection unit, and renders the background image for each of the layers.
14. The information processing device of claim 13, wherein the position and motion detection unit comprises: an IMU (Inertial Measurement Unit) that detects the acceleration and angular velocity of the imaging device; a motion detection unit that detects the motion of the imaging device based on the acceleration and angular velocity detected by the IMU; an imaging unit that captures images of the surroundings of the imaging device; and a SLAM processing unit that detects the position of the imaging device by SLAM (Simultaneous Localization and Mapping) based on the image captured by the imaging unit.
15. An information processing method comprising: a rendering process for dividing an object into multiple layers and rendering the background image for each layer based on the position and movement of an imaging device that captures a scene including the background image displayed on a display device that displays the background image consisting of an object based on a 3D model; and a background compositing process for compositing the background images for each of the multiple layers rendered by the rendering process into a single background image and displaying the resulting image on the display device.
16. A program that causes a computer to function as a rendering processing unit that divides objects into multiple layers based on the position and movement of an imaging device that captures a scene including a background image displayed on a display device, the background image being composed of objects based on a 3D model, and renders the background image for each layer; and a background synthesis unit that synthesizes the background images for each of the multiple layers rendered by the rendering processing unit into a single background image and displays it on the display device.
Citation Information
Patent Citations
Electronic device and method for displaying electronic document
US20150067483A1
Immersive content production system with multiple targets
US20200145644A1
Information processing device, video processing method, and program
WO2023007817A1
Information processing apparatus, image processing method, and program
WO2023090038A1