Information processing device and information processing method
By generating an environment map from a hemispherical panoramic image, the challenges of capturing a spherical panoramic image are overcome, enabling efficient and cost-effective live-action CG synthesis through accurate room layout and light source estimation.
Patent Information
- Application Number
- PCT/JP2025/020794
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-06-27
- Filing Date
- 2025-06-09
- Publication Date
- 2026-01-02
AI Technical Summary
Capturing a spherical panoramic image requires two cameras, which is costly and inefficient due to the inclusion of unnecessary objects, making it difficult to create an environment map for live-action CG synthesis.
Creating an environment map from a hemispherical panoramic image captured by a wide-angle camera, which estimates room layout and light sources using upper and lower room layout information and camera position, allowing for easier and cost-effective live-action CG synthesis.
Enables the creation of a celestial sphere panoramic image from a hemispherical image, facilitating accurate room layout estimation and light source analysis, thereby enhancing the efficiency and reducing costs in live-action CG synthesis.
Smart Images

Figure JP2025020794_02012026_PF_FP_ABST
Abstract
Description
Information processing device and information processing method
[0001] The present disclosure relates to an information processing device and an information processing method, and more particularly to an information processing device and an information processing method that enable creation of an environment map from a hemispherical panoramic image captured by a wide-angle camera.
[0002] Non-Patent Documents 1 to 4 are examples of live-action CG compositing techniques for compositing CG objects with background images of live-action footage.
[0003] Non-Patent Document 1 discloses a technology that uses an image with a normal angle of view and room layout information corresponding to the image with the normal angle of view as input, and outputs a celestial sphere panoramic image, room layout information corresponding to the celestial sphere panoramic image, and information on one main light source by a machine learning-based method. With the method of Non-Patent Document 1, there is a risk that an image that does not reflect a real scene may be obtained as an image of a region outside the normal angle of view in the output celestial sphere panoramic image.
[0004] Non-Patent Document 2 describes a method in which a celestial sphere panoramic image is used as an input, corner points of a room, boundaries between a ceiling and a wall, and boundaries between a floor and a wall are estimated as room layout information, and a 3D mesh model of the room is constructed from the estimated room layout information.
[0005] Non-Patent Document 3 discloses a method to assist in using light source settings in application software that synthesizes live-action CG, in which an image with a normal angle of view is input and light source parameters of the photographed scene (such as the direction, distance, size, and color of the light source as seen from the camera) are estimated.
[0006] Non-patent document 4 describes a method for estimating the quality and materials of walls and floors that appear in an image that is used to synthesize a CG object, and describes a means for estimating the shape (normal, depth) and material (albedo, roughness) of the object from a single captured image.
[0007] Image-based lighting (IBL) is one of the technologies for compositing live-action CG. IBL uses a spherical panoramic image of the entire scene as an environment map, allowing the lighting conditions of the scene to be realistically reflected on CG objects.
[0008] Henrique Weber, Mathieu Garon, Jean-Francois Lalonde, “Editable Indoor Lighting Estimation”, European Conference on Computer Vision (ECCV), 2022, Internet <URL: https: / / arxiv.org / pdf / 2211.03928> Cheng Sun, Chi-Wei Hsiao, Min Sun, Hwann-Tzong Chen, “HorizonNet: Learning Room Layout With 1D Representation and Pano Stretch Data Augmentation”, The IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2019, Internet <URL: https: / / arxiv.org / pdf / 1901.03861> Gardner, M.A., Hold-Geoffroy, Y., Sunkavalli, K., Gagne, C., Lalonde, J.F., “Deep parametric indoor lighting estimation”, ICCV, 2019, Internet <URL: Shen Sang and M. Chandraker, “Single-Shot Neural Relighting and SVBRDF Estimation”, ECCV, 2020>Shen Sang and M. Chandraker, “Single-Shot Neural Relighting and SVBRDF Estimation”, ECCV, 2020, Internet<URL: https: / / cseweb.ucsd.edu / ~viscomp / projects / ECCV20NeuralRelighting / >
[0009] As described above, a spherical panoramic image is required as an environment map to combine a background image of a real-life video with a CG object, but capturing a spherical panoramic image in one shot usually requires two cameras, which is costly. Also, unnecessary objects such as the photographer and a tripod are captured in the spherical panoramic image, making part of the image unusable and inefficient.
[0010] Therefore, if an environment map could be created from a hemispherical panoramic image, which is an image of a hemisphere taken with a wide-angle camera that has a field of view of 360 degrees horizontally and 90 degrees vertically, rather than 360 degrees, it would make live-action CG synthesis easier and reduce costs.
[0011] The present disclosure has been made in light of these circumstances, and makes it possible to create an environment map from a hemispherical panoramic image captured by a wide-angle camera.
[0012] An information processing device according to one aspect of the present disclosure includes: an upper room layout estimation unit that estimates upper room layout information including coordinate information indicating boundary lines between the ceiling and walls of the room and coordinate information indicating corner points between the ceiling and walls based on an upper hemispherical panoramic image, which is an image of the upper part of the room; and a lower room layout estimation unit that estimates lower room layout information including coordinate information indicating boundary lines between the floor and walls of the room and coordinate information indicating corner points between the floor and walls using the upper room layout information and camera position information of a camera that captured the upper hemispherical panoramic image.
[0013] An information processing method according to one aspect of the present disclosure includes an information processing device estimating upper room layout information based on an upper hemispherical panoramic image, which is an image of the upper part of a room, including coordinate information indicating the boundary line between the ceiling and walls of the room and coordinate information indicating the corner points of the ceiling and walls; and estimating lower room layout information using the upper room layout information and camera position information of a camera that captured the upper hemispherical panoramic image, including coordinate information indicating the boundary line between the floor and walls of the room and coordinate information indicating the corner points of the floor and walls.
[0014] In one aspect of the present disclosure, upper room layout information is estimated based on an upper hemispherical panoramic image, which is an image of the upper part of a room, including coordinate information indicating the boundary line between the ceiling and walls of the room and coordinate information indicating the corner points of the ceiling and walls, and lower room layout information is estimated using the upper room layout information and camera position information of the camera that captured the upper hemispherical panoramic image, including coordinate information indicating the boundary line between the floor and walls of the room and coordinate information indicating the corner points of the floor and walls.
[0015] The information processing device according to one aspect of the present disclosure can be realized by causing a computer to execute a program. The program executed by the computer to realize the information processing device can be provided by transmitting it via a transmission medium or by recording it on a recording medium.
[0016] The information processing device may be an independent device or an internal block constituting a single device.
[0017] FIG. 1 is a diagram illustrating live-action CG synthesis. FIG. 2 is a diagram illustrating an example of the configuration of a shooting camera. FIG. 3 is a block diagram illustrating an example of the configuration of an information processing device according to a first embodiment. FIG. 4 is a block diagram illustrating an example of the detailed configuration of a room layout estimation unit. FIG. 5 is a diagram illustrating the room layout estimation unit. FIG. 6 is a diagram illustrating a method for estimating upper room layout information. FIG. 7 is a diagram illustrating a method for estimating lower room layout information. FIG. 8 is a block diagram illustrating an example of the detailed configuration of a texture estimation unit. FIG. 9 is a diagram illustrating estimation of a lower hemispherical panoramic image. FIG. 10 is a diagram illustrating correction of a light source region. FIG. 11 is a flowchart illustrating live-action CG synthesis image processing by the information processing device according to the first embodiment. FIG. 12 is a block diagram illustrating an example of the configuration of an information processing device according to a second embodiment. FIG. 13 is a block diagram illustrating an example of the configuration of an information processing device according to a third embodiment. FIG. 14 is a block diagram illustrating an example of the hardware configuration of a computer.
[0018] Hereinafter, with reference to the accompanying drawings, a description will be given of modes for carrying out the technology of the present disclosure (hereinafter referred to as embodiments). Note that in this specification and the drawings, components having substantially the same functional configuration are assigned the same reference numerals to avoid redundant description. The description will be given in the following order: 1. Overview of live-action CG synthesis 2. Configuration of the shooting camera 3. First embodiment of information processing device 4. Detailed description of the room layout estimation unit 5. Detailed description of the texture estimation unit 6. Processing flow of live-action CG synthesis image processing 7. Second embodiment of information processing device 8. Third embodiment of information processing device 9. Summary of information processing device 10. Example computer configuration
[0019] 1. Overview of Live-Action CG Composition First, with reference to FIG. 1, live-action CG composition, which combines a CG object with a background image of live-action video, will be described.
[0020] A 3DCG application is application software that performs live-action CG compositing. The 3DCG application uses a normal-angle image, a spherical panoramic image, and a CG object to generate a live-action CG composite image by using the normal-angle image as a background image and compositing the CG object on the background image. The normal-angle image is an image captured with a camera that does not have a wide-angle angle like a panoramic image. The live-action CG composite image is a rendered image in which a CG object is rendered from the viewpoint of the camera that captured the normal-angle image and composited with the background image. In the example shown in Figure 1, the CG object is a rabbit object, a 3DCG model created in advance. A spherical panoramic image is an image in which the entire 360-degree surrounding scene is expanded flat as an equirectangular image. The spherical panoramic image is used as an environment map that represents the lighting environment at the three-dimensional position where the CG object is inserted. Image-based lighting (IBL) is one technique for realistically reflecting the lighting conditions of a scene in a CG object by using a spherical panoramic image that captures the entire scene as an environment map. For environment maps, images with a wide dynamic range called HDRI (High Dynamic Range Image) are used, but images with a normal 256-level range are also acceptable. Images with a normal 256-level range are sometimes called LDRI (Low Dynamic Range Image) in contrast to HDRI.
[0021] As described above, to combine a background image of a real-life video with a CG object, a spherical panoramic image capturing the entire scene is required as an environment map. Acquiring a spherical panoramic image in one shot typically requires two cameras, which increases costs. Furthermore, unnecessary objects such as the photographer or a tripod appear in the spherical panoramic image, making part of the image unusable and inefficient.
[0022] Therefore, in the following, we propose a technology to create an environment map from an image of the upper part taken with a wide-angle camera with a field of view of 360 degrees horizontally and 90 degrees vertically above, rather than a 360-degree image of the surroundings, and then synthesize it with live-action CG. This makes live-action CG synthesis easier and reduces costs.
[0023] 2. Configuration of the Camera FIG. 2 is a diagram showing an example of the configuration of a camera that captures a normal angle of view image and an upper hemispherical panoramic image.
[0024] The photographing camera 10 includes a main camera 21 and a wide-angle camera 22. The main camera 21 is a photographing camera with a predetermined angle of view that is not wide-angle, and photographs a predetermined scene to generate a normal angle of view image as the photographed image. The wide-angle camera 22 is a photographing camera that photographs an area equivalent to a hemisphere of the entire scene, and generates an upper hemispherical panoramic image as the photographed image.
[0025] The photographic camera 10 also has a depth sensor 23, an IMU (Inertial Measurement Unit) 24, and an interface device 25. The depth sensor 23 is a distance sensor that acquires distance information to a subject using a ToF method. The IMU 24 is an inertial sensor that acquires inertial information such as the acceleration and angular velocity of the photographic camera 10. The depth sensor 23 and the IMU 24 form a tracking system that tracks the position and movement of the photographic camera 10. The interface device 25 may be a device dedicated to the photographic camera 10, or may be a device that can be linked to the photographic camera 10 using an information processing device such as a smartphone. The display of the interface device 25 displays, for example, a normal angle of view image (through image) captured by the main line camera 21. The display of the interface device 25 is a touch panel that can also accept user input.
[0026] The wide-angle camera 22, depth sensor 23, IMU 24, and interface device 25 are attached to the main line camera 21 by a fixing jig such as a mounting unit 26. In a tracking system that tracks the position of the camera, the offset positions of the main line camera 21 and the wide-angle camera 22 are known.
[0027] 3. First Embodiment of Information Processing Apparatus> FIG. 3 is a block diagram showing an example of the configuration of an information processing apparatus according to a first embodiment.
[0028] 3 is a device that performs processing to generate a rendering image by combining the normal angle of view image, which is a real-life image, with a CG object, using the normal angle of view image and upper hemispherical panoramic image obtained by the photographic camera 10. The information processing device 40 may be, for example, the interface device 25 of the photographic camera 10, or may be a personal computer or server device provided separately from the photographic camera 10.
[0029] The information processing device 40 receives input of the normal angle of view image, the upper hemispherical panoramic image, and camera position information obtained by the photographing camera 10. The camera position information includes position information for the main camera 21 and the wide-angle camera 22. In the present embodiment, the input normal angle of view image is the normal angle of view image shown in FIG. 1. The normal angle of view image shown in FIG. 1 is an image captured with the photographing camera 10 facing toward the floor of a room. The input upper hemispherical panoramic image is the upper half of the celestial sphere panoramic image shown in FIG. 1. The celestial sphere panoramic image shown in FIG. 1 is an image captured by capturing a 360-degree image of a specific room.
[0030] The information processing device 40 includes a room layout estimation unit 51 , a light source estimation unit 52 , a texture estimation unit 53 , a 3D mesh model generation unit 54 , a material estimation unit 55 , an image rendering unit 56 , and a storage unit 57 .
[0031] The room layout estimation unit 51 generates room layout information from the input upper hemispherical panoramic image and camera position information, and outputs the information to the texture estimation unit 53, the 3D mesh model generation unit 54, and the material estimation unit 55. The room layout information, which will be described later with reference to Fig. 5, is composed of (1) coordinate information indicating the boundary line between the ceiling and the walls of the room shown in the upper hemispherical panoramic image, (2) coordinate information indicating the boundary line between the floor and the walls, and (3) coordinate information indicating the corner points between the ceiling and the walls and the corner points between the floor and the walls. The camera position information input to the room layout estimation unit 51 needs to include at least the camera position information of the wide-angle camera 22.
[0032] The light source estimation unit 52 estimates light source information from the upper hemispherical panoramic image and outputs it to the texture estimation unit 53 and the 3D mesh model generation unit 54. The light source information is composed of parameters for one or more light sources in the room captured in the upper hemispherical panoramic image. Light source parameters include, for example, position, direction, color, size, etc. The light source position indicates the location where the light source is located. The light source direction indicates the direction of light as seen from the main camera 21. The light source color indicates the color of light emitted by the light source. The light source size indicates how the light spreads. As a method for estimating light source parameters by the light source estimation unit 52, for example, the method of estimating light source parameters from an input image disclosed in Non-Patent Document 3 can be adopted. Of course, methods other than Non-Patent Document 3 may also be adopted.
[0033] The texture estimation unit 53 receives the upper hemispherical panoramic image, room layout information, and light source information as input, and complements the missing portion of the lower hemispherical panoramic image to generate a celestial sphere panoramic image. The generated celestial sphere panoramic image is output to the 3D mesh model generation unit 54. Coordinate information indicating the boundary line between the floor and the wall and coordinate information indicating the corner points of the floor and the wall in the lower hemispherical panoramic image portion are estimated by the room layout estimation unit 51 and supplied as room layout information. The texture estimation unit 53 generates a celestial sphere panoramic image by estimating and complementing the texture of the wall and floor regions in the lower hemispherical panoramic image portion based on the room layout information. Note that, as will be described in detail later, in the process of the texture estimation unit 53 generating the celestial sphere panoramic image, the texture of light sources in the room that satisfy predetermined conditions is deleted.
[0034] The 3D mesh model generation unit 54 generates a textured 3D mesh model of the room using the celestial sphere panoramic image, room layout information, and camera position information. The textured 3D mesh model of the room is represented by a 3D mesh model of a three-dimensional shape consisting of a ceiling, a floor, and multiple walls. For simplicity's sake, this embodiment describes an example in which the textured 3D mesh model of the room is represented by a rectangular parallelepiped mesh model consisting of a ceiling, a floor, and four walls. However, the shape of the room in the technology disclosed herein is not limited to a rectangular parallelepiped. The celestial sphere panoramic image is input from the texture estimation unit 53, the room layout information is input from the room layout estimation unit 51, and the camera position information is input from outside the device. The camera position information may be the camera position information of the wide-angle camera 22 that captured the upper hemispherical panoramic image that serves as the base for the celestial sphere panoramic image. A method for generating the textured 3D mesh model of the room can be, for example, the method disclosed in Non-Patent Document 2. The generated textured 3D mesh model of the room is imported into a 3DCG application that constitutes a part of the 3D mesh model generation unit 54. In the 3DCG application, a light source can be set, and all light sources estimated by the light source estimation unit 52 can be additionally set.
[0035] The material estimation unit 55 estimates the material of the floor area on which the CG object is placed based on the normal angle of view image, and outputs the estimation result to the image rendering unit 56. Non-Patent Document 4 discloses a method for estimating the quality and materials of walls and floors that appear in an image on which a CG object is to be composited. For example, the method disclosed in Non-Patent Document 4 is used to estimate the material of the floor area on which the CG object is placed.
[0036] The image rendering unit 56 uses a normal angle of view image, which is a live-action image, as a background image, and generates a live-action CG composite image by compositing a CG object on the background image. More specifically, the image rendering unit 56 renders a predetermined CG object stored in the storage unit 57 from the same viewpoint as the main camera 21 that captured the normal angle of view image. The material estimated by the material estimation unit 55 is set in the floor area where the CG object is placed. The image rendering unit 56 generates a live-action CG composite image by compositing the rendered CG object and floor area with the normal angle of view image. The viewpoint of the main camera 21 is set using the camera position information of the main camera 21 among the camera position information input from outside. The textured 3D mesh model of the room input from the 3D mesh model generation unit 54 is used as an environment map. The 3D mesh model generation unit 54 and a portion of the image rendering unit 56 are configured as a 3DCG application.
[0037] The storage unit 57 stores (data of) CG objects created in advance. The storage unit 57 outputs (data of) predetermined CG objects designated by the image rendering unit 56 to the image rendering unit 56.
[0038] The information processing device 40 is configured as described above. The information processing device 40 is characterized by generating a celestial sphere panoramic image from an upper hemispherical panoramic image captured by the wide-angle camera 22 and using the generated celestial sphere panoramic image as an environment map. The part that generates the celestial sphere panoramic image based on the upper hemispherical panoramic image will be described in detail below.
[0039] 4. Detailed Description of Room Layout Estimation Unit The room layout estimation unit 51 will be described with reference to Fig. 4 and Fig. 5. Fig. 4 is a block diagram showing an example of the detailed configuration of the room layout estimation unit 51.
[0040] The room layout estimation unit 51 has an upper room layout estimation unit 71 and a lower room layout estimation unit 72. The input upper hemispherical panoramic image is supplied to the upper room layout estimation unit 71. Camera position information including at least the camera position information of the wide-angle camera 22 is supplied to the lower room layout estimation unit 72.
[0041] For example, an upper hemispherical panoramic image 81 shown in Fig. 5 is input to the upper room layout estimation unit 71 of the room layout estimation unit 51. The upper hemispherical panoramic image 81 is an image obtained by capturing the upper hemispherical portion of the celestial sphere panoramic image shown in Fig. 1, and the lower hemispherical portion of the celestial sphere panoramic image that is not captured is displayed as an image missing portion.
[0042] The upper room layout estimation unit 71 estimates upper room layout information from the upper hemispherical panoramic image and outputs it to the lower room layout estimation unit 72. The upper room layout information consists of (1) coordinate information indicating the boundary lines between the ceiling and the walls, and (3A) coordinate information indicating the corner points between the ceiling and the walls.
[0043] 5, (1) coordinate information indicating a boundary line 91 between the ceiling and the walls, and (3A) coordinate information indicating corner points 92A, 92B, 92C, and 92D between the ceiling and the walls are estimated as upper room layout information for the upper hemispherical panoramic image 81. The corner points 92A to 92D between the ceiling and the walls are corners where two adjacent walls intersect with the ceiling.
[0044] The lower room layout estimation unit 72 estimates lower room layout information using the upper room layout information estimated by the upper room layout estimation unit 71 and camera position information. The lower room layout information consists of (2) coordinate information indicating the boundary line between the floor and the wall, and (3B) coordinate information indicating the corner points between the floor and the wall. In the example of Figure 5, (2) coordinate information indicating the boundary line 101 between the floor and the wall, and (3B) coordinate information indicating the corner points 102A, 102B, 102C, and 102D between the floor and the wall are estimated as the lower room layout information. The corner points 102A to 102D between the ceiling and the wall are corners where two adjacent walls intersect with the floor.
[0045] The lower room layout estimation unit 72 outputs, as room layout information, the upper room layout information estimated by the upper room layout estimation unit 71 and the lower room layout information estimated by itself. In the example of Fig. 5, the lower room layout estimation unit 72 outputs, as room layout information, (1) coordinate information indicating boundary line 91 between the ceiling and the wall, (2) coordinate information indicating boundary line 101 between the floor and the wall, (3A) coordinate information indicating corner points 92A to 92D between the ceiling and the wall, and (3B) coordinate information indicating corner points 102A to 102D between the floor and the wall.
[0046] <Method of Estimating Upper Room Layout Information> A method of estimating upper room layout information by the upper room layout estimation unit 71 will be described with reference to FIG.
[0047] First, the upper room layout estimation unit 71 generates an inverted upper hemispherical panoramic image 82 by inverting the input upper hemispherical panoramic image 81 upside down using the inverted upper hemispherical panoramic image generation unit 121 .
[0048] Next, the upper room layout estimation unit 71 generates a pseudo omnidirectional panoramic image 83, which is a pseudo omnidirectional panoramic image, by combining the input upper hemispherical panoramic image 81 and the inverted upper hemispherical panoramic image 82 using the pseudo omnidirectional panoramic image generation unit 122.
[0049] The generated pseudo omnidirectional panoramic image 83 is input to a room layout estimator 123, which estimates upper room layout information and lower room layout information based on the pseudo omnidirectional panoramic image 83. The room layout estimator 123, which estimates the upper room layout information and lower room layout information using the omnidirectional panoramic image as input, can employ a method using a machine learning model, as disclosed in Non-Patent Document 2. In the example of FIG. 6 , the upper room layout information is estimated as (1) coordinate information indicating a boundary line 91 between the ceiling and the wall, and (3A) coordinate information indicating corner points 92A to 92D between the ceiling and the wall. Furthermore, the lower room layout information is estimated as (2) coordinate information indicating a boundary line 101′ between the floor and the wall, and (3B) coordinate information indicating corner points 102A′ to 102D′ between the floor and the wall. This lower room layout information is estimated based on an inverted upper hemispherical panoramic image 82, which is obtained by inverting the upper hemispherical panoramic image 81, and may differ from the actual room situation.
[0050] The upper room layout estimation unit 71 deletes the lower room layout information related to the inverted upper hemispherical panoramic image 82 using the pseudo lower room layout information deletion unit 124 , and outputs only the upper room layout information to the lower room layout estimation unit 72 .
[0051] As described above, in upper room layout estimation unit 71, by generating pseudo spherical panoramic image 83, it becomes possible to apply an existing method for estimating room layout information from a spherical panoramic image, such as the machine learning-based method disclosed in Non-Patent Document 2. The upper room layout information can be generated by applying an existing method without retraining a machine learning model.
[0052] If a celestial sphere panoramic image is captured by the wide-angle camera 22, the floor portion captured in the lower half of the lower hemispherical panoramic image will often have a desk, chair, or the like placed on it, and part of the boundary line between the floor and the wall will be hidden. In contrast, in the upper hemispherical panoramic image 81, the boundary line between the ceiling and the wall is often completely visible, making it easier to estimate the boundary line. By using an inverted upper hemispherical panoramic image 82, which is the upper hemispherical panoramic image 81 flipped upside down, it becomes easier to estimate the boundary line between the floor and the wall, and the accuracy of inferring the corner points between the floor and the wall improves.
[0053] As another method of the upper room layout estimation unit 71, an inference device of a machine learning model that estimates the upper room layout information from an upper hemispherical panoramic image may be generated, following the machine learning-based method of estimating room layout information from a celestial sphere panoramic image, and the upper room layout information may be generated using the inference device.
[0054] <Method of Estimating Lower Room Layout Information> Next, a method of estimating lower room layout information by the lower room layout estimation unit 72 will be described with reference to FIGS. 7 and 8. FIG.
[0055] As shown in FIG. 7, the coordinates of a point on the boundary between the ceiling and the wall detected in the upper hemispherical panoramic image are (x, y c ), then the point (x, y) on the boundary between the ceiling and the wall c ) the coordinates (x, y) of the point on the boundary between the floor and the wall f ) will be explained.
[0056] The lower room layout estimation unit 72 assumes that "the wide-angle camera 22 is placed horizontally with respect to the floor of the room" and that "the floor and walls of the room are perpendicular to each other and the floor and ceiling are parallel," and estimates a point (x, y) on the boundary line between the ceiling and the wall. c ) are the coordinates (x, y) of the point on the boundary between the floor and the wall, f ) is calculated.
[0057] The lower room layout estimation unit 72 receives the upper room layout information and the camera position information of the wide-angle camera 22 as input.
[0058] First, the lower room layout estimation unit 72 calculates the distance d from the position o of the wide-angle camera 22 to the ceiling based on the camera position information of the wide-angle camera 22. c and the distance d from the position o of the wide-angle camera 22 to the floor. f Calculate.
[0059] Next, the lower room layout estimation unit 72 calculates the distance between the ceiling and the wall by calculating the distance (x, y) from the point (x, y) on the boundary line between the ceiling and the wall. c ) y coordinate position y c Calculate.
[0060] FIG. 8 shows the coordinates of a point (x, y) on the boundary line between the ceiling and the wall on the upper hemispherical panoramic image. c 1 is a diagram illustrating the relationship between the celestial sphere and a celestial sphere panoramic image.
[0061] The upper hemispherical panoramic image captured by the wide-angle camera 22 corresponds to the upper half of the celestial sphere panoramic image. The celestial sphere panoramic image is an image in which the entire surrounding 360-degree scene is developed on a plane as an equirectangular cylinder, and is a two-dimensional image in which the horizontal direction is defined as longitude φ and the vertical direction is defined as latitude θ. The number of pixels H in the vertical direction and the number of pixels W in the horizontal direction of the celestial sphere panoramic image are known.
[0062] A point (x, y) on the boundary between the ceiling and the wall c ) represents the pixel position when the top left corner of the spherical panoramic image is the origin (0,0), and the y coordinate of the center of the pixel is offset by 0.5. c It is expressed as +0.5”.
[0063] The number of pixels H in the vertical direction of a spherical panoramic image corresponds to 180 degrees (π), so the number of pixels H and "y c +0.5”, the latitudinal angle θ c is required.
[0064] The π in equation (1) represents the ratio of the circumference of a circle to its circumference. The angle θ in the latitudinal direction in equation (1) c is in the range from 0 degrees to π (0≦θ c ≦π), as shown in Figure 8, (-π / 2≦θ c As shown in (2), subtracting π / 2 from the right side of equation (1) and inverting the sign (multiplying by -1) gives equation (2). The latitudinal angle θ in equation (2)c is calculated in the range from (-π / 2) to (+π / 2) (-π / 2≦θ c ≦+π / 2), but the point (x, y c ) is a position on the upper hemisphere panoramic image, so it is practically in the range from 0 to +π / 2 (0≦θ c ≤ +π / 2).
[0065] From the above, the point (x, y) on the boundary line between the ceiling and the wall is c ) latitudinal angle θ c is obtained, the lower room layout estimation unit 72 calculates the angle θ c , the distance d from the position o of the wide-angle camera 22 to the wall is calculated. w is calculated by the following equation (3).
[0066] Next, the lower room layout estimation unit 72 calculates a predetermined point (x, y c ) on the floor surface, the point (x, y) on the boundary line between the floor and the wall f ) and the horizontal plane of the wide-angle camera 22, the angle θ in the latitudinal direction f is calculated by the following equation (4).
[0067] Finally, the latitudinal angle θ f Using the above, a point (x, y) on the boundary between the floor and the wall is f ) row coordinate y f is calculated by the following equation (5).
[0068] As described above, the lower room layout estimation unit 72 assumes that "the wide-angle camera 22 is placed horizontally with respect to the floor of the room" and that "the floor and walls of the room are perpendicular to each other and the floor and ceiling are parallel," and estimates a predetermined point (x, y) on the boundary line between the ceiling and the wall. c ) on the boundary between the floor and the wall, f By performing a similar calculation for each point on the boundary line between the ceiling and the wall, it is possible to estimate (2) coordinate information indicating the boundary line 101 between the floor and the wall, and (3B) coordinate information indicating the corner points 102A, 102B, 102C, and 102D between the floor and the wall, as shown in the example of FIG. 5, and these are used as the lower room layout information.
[0069] The lower room layout estimation unit 72 outputs, as room layout information, (1) coordinate information indicating the boundary line 91 between the ceiling and the wall, (2) coordinate information indicating the boundary line 101 between the floor and the wall, (3A) coordinate information indicating corner points 92A to 92D between the ceiling and the wall, and (3B) coordinate information indicating corner points 102A to 102D between the floor and the wall.
[0070] 5. Detailed Description of Texture Estimation Unit Next, the texture estimation unit 53 will be described with reference to Fig. 9 to Fig. 11. Fig. 9 is a block diagram showing an example of the detailed configuration of the texture estimation unit 53.
[0071] The texture estimation unit 53 has a lower hemisphere panoramic image estimation unit 151 and a light source region correction unit 152. The upper hemisphere panoramic image, room layout information, and normal angle of view image are input to the lower hemisphere panoramic image estimation unit 151. The light source information from the light source estimation unit 52 is input to the light source region correction unit 152.
[0072] The lower hemisphere panoramic image estimation unit 151 estimates a lower hemisphere panoramic image based on the upper hemisphere panoramic image, room layout information, and the normal angle of view image. The lower hemisphere panoramic image estimation unit 151 combines the upper hemisphere panoramic image and the lower hemisphere panoramic image to generate a celestial sphere panoramic image, and supplies the generated image together with the room layout information to the light source region correction unit 152.
[0073] The light source region correction unit 152 corrects a light source region that satisfies a predetermined condition in the omnidirectional panoramic image. More specifically, the light source region correction unit 152 corrects the texture of a light source that is depicted as being stuck to a wall, among one or more light sources that appear in the omnidirectional panoramic image, and outputs the corrected omnidirectional panoramic image to the 3D mesh model generation unit 54 ( FIG. 3 ).
[0074] Referring to FIG. 10, the estimation of the lower hemispheric panoramic image by the lower hemispheric panoramic image estimation unit 151 will be described.
[0075] First, the lower hemisphere panoramic image estimation unit 151 combines the region of the lower hemisphere panoramic image, which is an image missing portion, with the input upper hemisphere panoramic image to generate a celestial sphere panoramic image. At this point, the pixel values (pixel color information and brightness information) of the region of the lower hemisphere panoramic image are set to predetermined initial values (e.g., zero). Then, based on the room layout information from the room layout estimation unit 51, the lower hemisphere panoramic image estimation unit 151 divides the celestial sphere panoramic image into six regions consisting of the ceiling, floor, and four walls. Here, if the six divided regions are referred to as the ceiling region (Ceil), floor region (Floor), wall 1 region (Wall 1), wall 2 region (Wall 2), wall 3 region (Wall 3), and wall 4 region (Wall 4), the lower hemisphere panoramic image includes the floor region, wall 1 region, wall 2 region, wall 3 region, and wall 4 region. As described above, this embodiment is an example in which the shape of the room is a rectangular parallelepiped, so there are four wall regions (four walls), but the number of wall regions can vary depending on the shape of the room.
[0076] The lower hemispherical panoramic image estimation unit 151 complements pixel values of image-missing portions within the wall regions for each of the four wall regions, Wall 1 region, Wall 2 region, Wall 3 region, and Wall 4 region, using, for example, one of the following first to third complementation methods. The first method calculates the average color within the wall region of the upper hemispherical panoramic image and complements the pixel values of the image-missing portions using that value. The second method complements pixel values of the image-missing portions by extracting patches within the wall region of the upper hemispherical panoramic image and tiling the patches over the image-missing portions. The third method inputs the upper hemispherical panoramic image and region information of the wall region to be complemented, generates a machine learning model that inpaints and outputs pixel values of the image-missing portions, and complements pixel values of the image-missing portions using the machine learning model.
[0077] Furthermore, the lower hemispherical panoramic image estimation unit 151 complements pixel values of image-missing portions in wall regions for the floor region, for example, using one of the following first to third complementation methods. The first method calculates the average color of the floor region from the normal angle of view image and complements the pixel values of the image-missing portion using that value. The second method complements pixel values of the image-missing portion by extracting a patch of the floor region from the normal angle of view image and tiling the patch over the image-missing portion. The third method inputs area information of the upper hemispherical panoramic image and the floor region to be complemented, generates a machine learning model that inpaints and outputs pixel values of the image-missing portion, and complements pixel values of the image-missing portion using the machine learning model.
[0078] The first to third interpolation methods for the wall region or floor region described above may be selected by a user input before starting the processing, or the interpolation method to be executed may be selected in advance using a setting value, etc. The celestial sphere panoramic image generated by interpolating the pixel values of the region of the lower hemisphere panoramic image in the above-described manner is output to the light source region correction unit 152.
[0079] Next, the correction of the light source area by the light source area correction unit 152 will be described with reference to FIG.
[0080] The spherical panoramic image estimated by the texture estimation unit 53 is used as a texture for a 3D mesh model in the subsequent 3D mesh model generation unit 54. The 3D mesh model represents a three-dimensional mesh model using the fact that it is a room as known information, and therefore, for example, a light source (lighting) hanging from the ceiling as shown in A of Fig. 11 will be represented as being stuck to the ceiling or wall as shown in B of Fig. 11 .
[0081] Therefore, based on the light source information from the light source estimation unit 52, the light source region correction unit 152 identifies a light source that would be depicted as sticking to a wall when a 3D mesh model of the room is generated. Then, the light source region correction unit 152 corrects the texture of the celestial sphere panoramic image in the region of the identified light source with the texture of the surrounding pixels. C in Fig. 11 shows an example of the celestial sphere panoramic image after the correction. In the celestial sphere panoramic image after the correction, the light source regions of cords hanging from the ceiling and lighting fixtures are corrected by inpainting with the texture of the ceiling and wall surrounding the light source region.
[0082] Among one or more light sources appearing in the celestial sphere panoramic image, light sources that appear to be stuck to the wall are identified using light source information supplied from the light source estimation unit 52. Specifically, the positions of six surfaces, namely, the ceiling area (Ceil), floor area (Floor), wall 1 area (Wall 1), wall 2 area (Wall 2), wall 3 area (Wall 3), and wall 4 area (Wall 4), are known based on the room layout information, and the position of each light source appearing in the celestial sphere panoramic image is also known based on the light source parameters supplied as light source information from the light source estimation unit 52. Therefore, the light source area correction unit 152 identifies light sources that are farther away from both the ceiling and the wall than a predetermined value as light sources that are hanging from the ceiling and appear to be stuck to the wall. The pixel value of the texture to be corrected may be an average color obtained by extracting a predetermined number of peripheral pixels of the light source area to be corrected, or may be an average color of the entire area of the ceiling area or wall area to which the light source is stuck.
[0083] 12, a description will be given of processing of a real-life CG composite image using an upper hemispherical panoramic image by the information processing device 40. This processing is started, for example, when a normal angle of view image, an upper hemispherical panoramic image, and camera position information are input from the photographing camera 10 to the information processing device 40.
[0084] First, in step S1, the information processing device 40 acquires the normal angle of view image, the upper hemispherical panoramic image, and the camera position information obtained by the photographing camera 10.
[0085] In step S2, the upper room layout estimation unit 71 of the room layout estimation unit 51 estimates upper room layout information based on the upper hemispherical panoramic image and outputs it to the lower room layout estimation unit 72. The upper room layout information consists of (1) coordinate information indicating the boundary line between the ceiling and the wall, and (3A) coordinate information indicating the corner points between the ceiling and the wall.
[0086] For example, the upper room layout estimation unit 71 generates a pseudo omnidirectional panoramic image by combining an inverted upper hemispherical panoramic image, which is obtained by flipping an upper hemispherical panoramic image upside down, with the upper hemispherical panoramic image, and estimates (1) coordinate information indicating the boundary line between the ceiling and the wall, (2) coordinate information indicating the boundary line between the floor and the wall, (3A) coordinate information indicating the corner points between the ceiling and the wall, and (3B) coordinate information indicating the corner points between the floor and the wall using the machine learning-based technique disclosed in Non-Patent Document 2. Then, the upper room layout estimation unit 71 deletes (2) the coordinate information indicating the boundary line between the floor and the wall and (3B) the coordinate information indicating the corner points between the floor and the wall, which are related to the inverted upper hemispherical panoramic image, from the estimation results, and outputs (1) the coordinate information indicating the boundary line between the ceiling and the wall and (3A) the coordinate information indicating the corner points between the ceiling and the wall to the lower room layout estimation unit 72 as upper room layout information.
[0087] In step S3, the lower room layout estimation unit 72 estimates lower room layout information based on the upper room layout information and camera position information of the wide-angle camera 22. The lower room layout estimation unit 72 outputs room layout information combining the upper room layout information and the lower room layout information to the texture estimation unit 53, the 3D mesh model generation unit 54, and the material estimation unit 55. The lower room layout information consists of (2) coordinate information indicating the boundary lines between the floor and the walls, and (3B) coordinate information indicating the corner points between the floor and the walls.
[0088] Specifically, based on the assumptions that "wide-angle camera 22 is placed horizontally relative to the floor of the room" and "the floor and walls of the room are perpendicular to each other, and the floor and ceiling are parallel," lower room layout estimation unit 72 calculates coordinates corresponding to (1) coordinate information indicating the boundary line between the ceiling and the walls and (3A) coordinate information indicating the corner points between the ceiling and the walls, and calculates (2) coordinate information indicating the boundary line between the floor and the walls and (3B) coordinate information indicating the corner points between the floor and the walls.
[0089] In step S4, the light source estimation unit 52 estimates light source information from the upper hemispherical panoramic image and outputs the information to the texture estimation unit 53 and the 3D mesh model generation unit 54. The light source information is composed of light source parameters including, for example, the position, direction, color, size, etc. of the light source. For example, the light source estimation unit 52 can estimate the parameters of each light source appearing in the upper hemispherical panoramic image using the method for estimating light source parameters from an input image disclosed in Non-Patent Document 3.
[0090] In step S5, the lower hemisphere panoramic image estimation unit 151 of the texture estimation unit 53 estimates a lower hemisphere panoramic image based on the upper hemisphere panoramic image, room layout information, and the normal angle of view image, and combines the upper hemisphere panoramic image and the lower hemisphere panoramic image to generate a celestial sphere panoramic image. Specifically, the lower hemisphere panoramic image estimation unit 151 generates a celestial sphere panoramic image by complementing the lower hemisphere panoramic image portion of the upper hemisphere panoramic image. Next, the lower hemisphere panoramic image estimation unit 151 divides the celestial sphere panoramic image into a ceiling region, a floor region, a wall 1 region, a wall 2 region, a wall 3 region, and a wall 4 region based on the room layout information. Then, the lower hemisphere panoramic image estimation unit 151 completes the lower hemisphere panoramic image portion using any of a first complementation method that complements pixel values of image-missing portions in each of the wall region and the floor region with an average color of a known region, a second complementation method that tiling patches, and a third complementation method that inpaints using a machine learning model.
[0091] In step S6, the light source area correction unit 152 of the texture estimation unit 53 corrects the texture of the light source area, which is depicted as being stuck to the wall when the 3D mesh model is generated, with the texture of the surrounding pixels. Specifically, the light source area correction unit 152 corrects the texture of the light source area that is a predetermined distance or more from both the ceiling and the wall, with the texture of the surrounding pixels.
[0092] In step S7, the 3D mesh model generation unit 54 generates a textured 3D mesh model of the room using the omnidirectional panoramic image, the room layout information, and the camera position information of the wide-angle camera 22. For example, the 3D mesh model generation unit 54 generates the textured 3D mesh model of the room using the technique disclosed in Non-Patent Document 2. The 3D mesh model generation unit 54 also imports the generated textured 3D mesh model of the room into a 3DCG application and sets all of the light sources estimated by the light source estimation unit 52.
[0093] In step S8, the material estimation unit 55 estimates the material of the floor area on which the CG object is placed, based on the normal angle of view image, and outputs the estimation result to the image rendering unit 56. For example, the material estimation unit 55 estimates the material of the floor area on which the CG object is placed, using the method disclosed in Non-Patent Document 4.
[0094] The processes of steps S7 and S8 may be executed in parallel, or the order of the processes may be reversed.
[0095] In step S9, the image rendering unit 56 uses the normal angle of view image, which is a live-action image, as a background image, and generates a live-action CG composite image by combining a CG object on the background image. More specifically, the image rendering unit 56 acquires a predetermined CG object from the storage unit 57 and renders it from the same viewpoint as the main camera 21 that captured the normal angle of view image. The material estimated by the material estimation unit 55 is set in the floor area in which the CG object is placed. The image rendering unit 56 generates a live-action CG composite image by combining the rendered CG object and floor area with the normal angle of view image. The viewpoint of the main camera 21 is set using the camera position information of the main camera 21. The textured 3D mesh model of the room input from the 3D mesh model generation unit 54 is used as an environment map.
[0096] In step S10, the image rendering unit 56 outputs the generated live-action CG composite image to an external display device for display, or may store data of the generated live-action CG composite image in an external device, the storage unit 57, or the like.
[0097] This completes the live-action CG composite image processing. The information processing device 40 can execute the live-action CG composite image processing shown in FIG. 12 every time a normal angle of view image, an upper hemispherical panoramic image, and camera position information are input from the photographing camera 10.
[0098] 7. Second Embodiment of Information Processing Apparatus> FIG. 13 is a block diagram showing an example of the configuration of an information processing apparatus according to a second embodiment.
[0099] In the second embodiment of FIG. 13, parts common to the first embodiment described above are given the same reference numerals, and the description of those parts will be omitted as appropriate.
[0100] 13, an information processing device 40 according to the second embodiment is newly provided with a mode switching unit 171 and an input unit 172. The image rendering unit 56 of the first embodiment is replaced with an image rendering unit 56'. The other configurations are the same as those of the first embodiment.
[0101] The mode switching unit 171 is provided between the 3D mesh model generation unit 54 and the image rendering unit 56′. The mode switching unit 171 is supplied with a textured 3D mesh model of the room from the 3D mesh model generation unit 54 and with a celestial sphere panoramic image from the texture estimation unit 53. Furthermore, light source information is supplied to the mode switching unit 171 from the light source estimation unit 52 as needed.
[0102] The mode switching unit 171 selects, automatically or based on a user instruction, which of the textured 3D mesh model of the room supplied from the 3D mesh model generation unit 54 and the celestial sphere panoramic image supplied from the texture estimation unit 53 to use as an environment map. The mode switching unit 171 outputs the selected textured 3D mesh model of the room or the celestial sphere panoramic image to the image rendering unit 56. When the mode switching unit 171 makes a selection based on a user instruction, the user issues an instruction to select either the textured 3D mesh model of the room or the celestial sphere panoramic image via the input unit 172. The input unit 172 supplies selection information based on the user instruction to the mode switching unit 171.
[0103] For example, when the estimation result of any one of the room layout estimation unit 51, the light source estimation unit 52, and the texture estimation unit 53 is significantly different from the actual scene and a decrease in quality of the rendered image of the textured 3D mesh model of the room is expected, the user gives an instruction to select a celestial sphere panoramic image via the input unit 172. The mode switching unit 171 selects the celestial sphere panoramic image based on the user instruction, and outputs it to the image rendering unit 56′.
[0104] For example, if the distance from the position of the main camera 21 to the insertion position of the CG object is sufficiently close to the distance from the position of the dominant light source of the scene to the insertion position of the CG object, the mode switching unit 171 selects the celestial sphere panoramic image and outputs it to the image rendering unit 56'. For example, based on the parameters of the light source, the light source closest to the position of the main camera 21 or the light source with the strongest light is determined to be the dominant light source of the scene. If the scene is captured outdoors, the distance from the position of the main camera 21 to the insertion position of the CG object is sufficiently close to the distance from the position of the dominant light source of the scene to the insertion position of the CG object. If the scene is assumed to be captured outdoors, the celestial sphere panoramic image is selected instead of the rendered image of the textured 3D mesh model.
[0105] The image rendering unit 56′ generates a live-action CG composite image by using, as an environment map, the textured 3D mesh model of the room or the omnidirectional panoramic image supplied from the mode switching unit 171. As a method of using the omnidirectional panoramic image as the environment map, for example, IBL can be used.
[0106] Next, a description will be given of live-action CG composite image processing by the information processing device 40 according to the second embodiment. In the live-action CG composite image processing according to the second embodiment, processing for selecting either a textured 3D mesh model of a room or a celestial sphere panoramic image by the mode switching unit 171 is added between steps S8 and S9 of the live-action CG composite image processing in Fig. 12. Then, in step S9, processing is executed for generating a live-action CG composite image using the textured 3D mesh model of a room or the celestial sphere panoramic image selected by the mode switching unit 171.
[0107] According to the second embodiment, when the quality of the textured 3D mesh model of the room is low, a live-action CG composite image can be generated by IBL using a celestial sphere panoramic image as an environment map. In other words, a live-action CG composite image can be generated by selecting either the rendering image of the textured 3D mesh model or the celestial sphere panoramic image.
[0108] 8. Third Embodiment of Information Processing Apparatus> FIG. 14 is a block diagram showing an example of the configuration of an information processing apparatus according to a third embodiment.
[0109] In the third embodiment of Fig. 14, the same components as those in the first embodiment are denoted by the same reference numerals, and the description of those components will be omitted where appropriate. In the third embodiment, the 3D mesh model generation unit 54 is changed to a 3D mesh model generation unit 54'. The other configurations are the same as those in the first embodiment.
[0110] In the first embodiment described above, the upper hemispherical panoramic image input to the information processing device 40 may be either an HDRI or an LDRI, which have a wide dynamic range. However, when an LDRI upper hemispherical panoramic image is input in the first embodiment, the textured 3D mesh model generated is a 3D mesh model with a normal 256-level range, and when an HDRI upper hemispherical panoramic image is input, the textured 3D mesh model generated is a 3D mesh model with a wide dynamic range.
[0111] On the other hand, instead of HDRI, multiple upper hemispherical panoramic images of LDRI captured with different exposure times are input to the information processing device 40 according to the third embodiment. Alternatively, multiple upper hemispherical panoramic images of LDRI captured at different locations are input. The multiple upper hemispherical panoramic images may be captured continuously and input sequentially at approximately the same timing, or may be captured and input at different timings separated by a predetermined time. In the following description, for example, it is assumed that an upper hemispherical panoramic image captured at time T1 and an upper hemispherical panoramic image captured at time T2, which is later than time T1, are input to the information processing device 40.
[0112] The 3D mesh model generation unit 54' generates a textured 3D mesh model of the room using the upper hemispherical panoramic image captured at time T1. The textured 3D mesh model generated at this time is referred to as the textured 3D mesh model at time T1. The 3D mesh model generation unit 54' stores the generated textured 3D mesh model at time T1 in the storage unit 57.
[0113] The 3D mesh model generation unit 54' also generates a textured 3D mesh model of the room using the upper hemispherical panoramic image captured at time T2. The textured 3D mesh model generated at this time is referred to as the textured 3D mesh model at time T2.
[0114] The 3D mesh model generation unit 54' retrieves the previously generated textured 3D mesh model at time T1 from the storage unit 57, and generates a composite 3D mesh model by combining the textured 3D mesh model at time T1 and the textured 3D mesh model at time T2. The composite 3D mesh model generated at this time is referred to as the composite 3D mesh model at time T2.
[0115] Examples of methods for combining a textured 3D mesh model at time T1 and a textured 3D mesh model at time T2 include (1) averaging the two textured 3D mesh models, (2) weighted averaging the two textured 3D mesh models, and (3) partial replacement, which replaces a portion of one of the two textured 3D mesh models with the other. With the weighted averaging method (2), for example, the weight of a textured 3D mesh model generated using a newer upper hemispherical panoramic image at the capture time can be increased, the weight of a textured 3D mesh model generated using an upper hemispherical panoramic image with a long (or short) exposure time can be increased, or the weight of a textured 3D mesh model generated using an upper hemispherical panoramic image captured at a predetermined location can be increased. With the partial replacement method (3), for example, a textured 3D mesh model generated previously at time T1 can be replaced with a textured 3D mesh model at time T2 for a predetermined area corresponding to the camera position of the main camera 21 or the wide-angle camera 22.
[0116] The 3D mesh model generation unit 54′ stores the generated composite 3D mesh model at time T2 in the storage unit 57. Thereafter, when an upper hemispherical panoramic image captured at time T3, which is after time T2, is input to the information processing device 40, the textured 3D mesh model at time T3 and the composite 3D mesh model at time T2 stored in the storage unit 57 are combined, and the composite 3D mesh model stored in the storage unit 57 is updated from the composite 3D mesh model at time T2 to the composite 3D mesh model at time T3.
[0117] The 3D mesh model generating unit 54 ′ outputs the latest synthesized 3D mesh model generated using the input upper hemispherical panoramic image to the image rendering unit 56 .
[0118] A description will be given of live-action CG composite image processing by information processing device 40 according to the third embodiment. In the live-action CG composite image processing according to the third embodiment, after step S7 of the live-action CG composite image processing in FIG. 12 , a process is added in which the textured 3D mesh model generated in step S7 is combined with a composite 3D mesh model acquired from storage unit 57 to update the composite 3D mesh model. The updated composite 3D mesh model is stored in storage unit 57. Then, in step S9, a process is executed in which a live-action CG composite image is generated using the updated composite 3D mesh model.
[0119] According to the third embodiment, when a plurality of upper hemispherical panoramic images input to the information processing device 40 are upper hemispherical panoramic images taken with different exposure times, it is possible to convert the 3D mesh models into HDR by combining a plurality of textured 3D mesh models. Furthermore, when a plurality of input upper hemispherical panoramic images are upper hemispherical panoramic images taken at different locations, it is possible to increase the resolution of the 3D mesh models by combining a plurality of textured 3D mesh models.
[0120] 9. Summary of Information Processing Device The information processing device 40 includes a room layout estimation unit 51 having: an upper room layout estimation unit 71 that estimates upper room layout information, including coordinate information indicating boundary lines between the ceiling and walls of the room and coordinate information indicating corner points of the ceiling and walls, based on an upper hemispherical panoramic image captured by a wide-angle camera 22 with a field of view of 360 degrees horizontally and 90 degrees vertically; and a lower room layout estimation unit 72 that estimates lower room layout information, including coordinate information indicating boundary lines between the floor and walls of the room and coordinate information indicating corner points of the floor and walls, using the upper room layout information and camera position information of the wide-angle camera 22 that captured the upper hemispherical panoramic image. The room layout information indicating the layout of the room can be estimated from the geometric relationship between the upper hemispherical panoramic image and the wide-angle camera 22.
[0121] The information processing device 40 further includes a texture estimation unit 53 having: a lower hemispherical panoramic image estimation unit 151 that estimates a lower hemispherical panoramic image based on the upper hemispherical panoramic image, room layout information including upper room layout information and lower room layout information, and a captured image (normal angle of view image) captured by the main camera 21 different from the wide-angle camera 22 that captured the upper hemispherical panoramic image, and generates a celestial sphere panoramic image by combining the upper hemispherical panoramic image and the lower hemispherical panoramic image; and a light source area correction unit 152 that corrects a light source area that satisfies a predetermined condition in the celestial sphere panoramic image. The lower hemispherical panoramic image can be estimated from the upper hemispherical panoramic image and the room layout information, and a celestial sphere panoramic image can be generated.
[0122] The information processing device 40 further includes a 3D mesh model generation unit 54 that generates a textured 3D mesh model of the room using the omnidirectional panoramic image after correcting the light source area that satisfies predetermined conditions, and an image rendering unit 56 that generates a composite image by combining a captured image and a CG object using the textured 3D mesh model of the room. Using the textured 3D mesh model of the room, the captured image, and camera position information of the main camera 21, a real-life CG composite image can be generated by combining a normal angle of view image, which is a real-life image, with a CG object on the background image. By using the textured 3D mesh model of the room as an environment map, ambient light from any position can be reproduced.
[0123] According to the information processing device 40, an environment map can be created from an upper hemispherical panoramic image captured by the wide-angle camera 22. When capturing an upper portion with the wide-angle camera 22, there is no need to worry about unwanted objects such as the photographer or a tripod appearing in the image, and the number of cameras can be reduced compared to a spherical camera, thereby reducing device costs. Because an environment map can be created from an upper hemispherical panoramic image, for example, it is possible to capture an upper hemispherical panoramic image in real time at a remote location using the wide-angle camera 22, and reflect the lighting conditions of the scene in the rendering of a CG object in real time. The position of the CG object can also be freely changed.
[0124] 10. Computer Configuration Example The series of processes executed by the information processing device 40 can be executed by hardware or software. When the series of processes are executed by software, the programs that make up the software are installed on a computer. Here, the computer includes a microcomputer built into dedicated hardware, and a general-purpose personal computer, for example, that can execute various functions by installing various programs.
[0125] FIG. 15 is a block diagram showing an example of the hardware configuration of a computer as an information processing device 40 when the above-described series of processes are executed by a program.
[0126] The computer 200 includes a central processing unit (CPU) 201, a read-only memory (ROM) 202, and a random access memory (RAM) 203. The CPU 201, the ROM 202, and the RAM 203 are connected to one another by a bus 204.
[0127] An input / output interface 205 is also connected to the bus 204. An input unit 206, an output unit 207, a storage unit 208, a communication unit 209, and a drive 210 are connected to the input / output interface 205.
[0128] The input unit 206 includes a keyboard, mouse, microphone, touch panel, input terminal, etc. The output unit 207 includes a display, speaker, output terminal, etc. The storage unit 208 includes a hard disk, SSD (Solid State Drive), RAM disk, non-volatile memory, etc. The communication unit 209 includes a network interface, etc. The drive 210 drives removable media 211 such as a magnetic disk, optical disk, magneto-optical disk, or semiconductor memory.
[0129] In the computer 200 configured as above, the CPU 201 performs the above-described series of processes by, for example, loading a program stored in the storage unit 208 into the RAM 203 via the input / output interface 205 and the bus 204 and executing the program. The RAM 203 also stores data and the like necessary for the CPU 201 to execute various processes.
[0130] The program executed by the CPU 201 of the computer 200 can be provided by being recorded on a removable medium 211 such as a package medium, for example. The program can also be provided via a wired or wireless transmission medium such as a local area network, the Internet, or digital satellite broadcasting.
[0131] In the computer 200, the program can be installed in the storage unit 208 via the input / output interface 205 by inserting the removable medium 211 into the drive 210. The program can also be received by the communication unit 209 via a wired or wireless transmission medium and installed in the storage unit 208. Alternatively, the program can be installed in the ROM 202 or the storage unit 208 in advance.
[0132] In this specification, the steps described in the flowcharts may be performed in chronological order in the order described, but they do not necessarily have to be processed in chronological order, and may be performed in parallel or at any necessary timing, such as when a call is made.
[0133] In this specification, a system refers to a collection of multiple components (devices, modules (components), etc.), regardless of whether all of the components are housed in the same housing. Therefore, multiple devices housed in separate housings and connected via a network, and a single device housed in a single housing with multiple modules, are both systems.
[0134] The embodiments of the present disclosure are not limited to the above-described embodiments, and various modifications are possible within the scope of the gist of the technology of the present disclosure.
[0135] For example, it is possible to adopt a form that combines all or part of the above-described multiple embodiments. The information processing device 40 may operate by switching between the configurations of the above-described first to third embodiments by setting an operation mode or the like.
[0136] For example, the technology of the present disclosure can be configured as a cloud computing system in which a single function is shared and processed collaboratively by multiple devices via a network.
[0137] Each step described in the above flowchart can be executed by one device or can be shared and executed by multiple devices. Furthermore, if one step includes multiple processes, the multiple processes included in that one step can be executed by one device or can be shared and executed by multiple devices.
[0138] The effects described in this specification are merely examples and are not limiting, and there may be effects other than those described in this specification.
[0139] The technology disclosed herein may employ the following configuration: (1) An information processing device having an upper room layout estimation unit that estimates upper room layout information including coordinate information indicating boundary lines between a ceiling and a wall of a room and coordinate information indicating corner points of the ceiling and the wall based on an upper hemispherical panoramic image that is an image of the upper part of the room, and a lower room layout estimation unit that estimates lower room layout information including coordinate information indicating boundary lines between a floor and a wall of the room and coordinate information indicating corner points of the floor and the wall, using the upper room layout information and camera position information of a camera that captured the upper hemispherical panoramic image. (2) The information processing device described in (1), further comprising: a lower hemispherical panoramic image estimation unit that estimates a lower hemispherical panoramic image based on the upper hemispherical panoramic image, room layout information including the upper room layout information and the lower room layout information, and a captured image captured by a camera different from that which captured the upper hemispherical panoramic image, and generates a celestial sphere panoramic image by combining the upper hemispherical panoramic image and the lower hemispherical panoramic image; and a light source region correction unit that corrects a region of a light source that satisfies a predetermined condition in the celestial sphere panoramic image. (3) The information processing device described in (1) or (2), wherein the upper room layout estimation unit generates a pseudo celestial sphere panoramic image using an inverted upper hemispherical panoramic image obtained by inverting the upper hemispherical panoramic image upside down, and estimates the upper room layout information based on the pseudo celestial sphere panoramic image. (4) The information processing device according to any one of (1) to (5), wherein the upper room layout estimation unit estimates room layout information based on the pseudo spherical panoramic image and estimates the upper room layout information by deleting lower room layout information from the estimated room layout information. (5) The information processing device according to (4), wherein the upper room layout estimation unit estimates room layout information based on the pseudo spherical panoramic image using a machine learning model. (6) The information processing device according to any one of (1) to (5), wherein the upper room layout estimation unit estimates the upper room layout information from the upper hemispherical panoramic image.(7) The information processing device according to (6), wherein the upper room layout estimation unit estimates the upper room layout information from the upper hemispherical panoramic image using a machine learning model. (8) The information processing device according to any of (1) to (8), wherein the lower room layout estimation unit estimates the lower room layout information by calculating points on a boundary line between the floor and walls that correspond to points on a boundary line between the ceiling and walls in the upper room layout information. (9) The information processing device according to (2), wherein the lower hemispherical panoramic image estimation unit estimates the lower hemispherical panoramic image by dividing an area of the lower hemispherical panoramic image into a floor area and a plurality of wall areas based on the room layout information and interpolating pixel values of the floor area and the plurality of wall areas. (10) The information processing device according to (9), wherein a method for interpolating the pixel values of the wall areas is a method of interpolating using an average color within the same wall area of the upper hemispherical panoramic image. (11) The information processing device according to any one of (9) to (10), wherein the method for complementing the pixel values of the wall region is a method for complementing by cutting out patches within the same wall region of the upper hemispherical panoramic image and tiling the patches. (12) The information processing device according to any one of (9) to (11), wherein the method for complementing the pixel values of the wall region is a method for inputting the upper hemispherical panoramic image to a machine learning model and complementing the values using the machine learning model. (13) The information processing device according to any one of (9) to (12), wherein the method for complementing the pixel values of the floor region is a method for complementing by an average color of the floor region calculated from the captured image. (14) The information processing device according to any one of (9) to (13), wherein the method for complementing the pixel values of the floor region is a method for complementing by cutting out patches of the floor region from the captured image and tiling the patches. (15) The information processing device according to any one of (9) to (14), wherein the method for complementing the pixel values of the floor region is a method for inputting the upper hemispherical panoramic image to a machine learning model and complementing the values using the machine learning model.(16) The information processing device according to any one of (2) to (18), further comprising: a light source estimation unit that estimates light source information from the upper hemispherical panoramic image, wherein the light source region correction unit corrects a light source region where the position of the light source indicated by the light source information satisfies the predetermined condition. (17) The information processing device according to (16), wherein the light source that satisfies the predetermined condition is a light source that is away from both a ceiling and a wall of the room by a distance equal to or greater than a predetermined value. (18) The information processing device according to (2), (16), or (17), wherein the light source region correction unit corrects the light source region that satisfies the predetermined condition with pixel values of pixels surrounding the region. (19) The information processing device according to any one of (2) to (18), further comprising: a 3D mesh model generation unit that generates a textured 3D mesh model of the room using the celestial sphere panoramic image after the light source region that satisfies the predetermined condition has been corrected; and an image rendering unit that generates a composite image by combining the captured image and a CG object using the textured 3D mesh model of the room. (20) An information processing method including: an information processing device estimating upper room layout information including coordinate information indicating the boundary line between the ceiling and walls of the room and coordinate information indicating the corner points of the ceiling and walls based on an upper hemispherical panoramic image that is an image of the upper part of the room; and estimating lower room layout information including coordinate information indicating the boundary line between the floor and walls of the room and coordinate information indicating the corner points of the floor and walls using the upper room layout information and camera position information of a camera that captured the upper hemispherical panoramic image.
[0140] 10 Shooting camera, 21 Main camera, 22 Wide-angle camera, 23 Depth sensor, 25 Interface device, 26 Mounting unit, 40 Information processing device, 51 Room layout estimation unit, 52 Light source estimation unit, 53 Texture estimation unit, 54, 54' 3D mesh model generation unit, 3D mesh model generation unit, 55 Material estimation unit, 56, 56' Image rendering unit, 57 Memory unit, 71 Upper room layout estimation unit, 72 Lower room layout estimation unit, 151 Lower hemisphere panoramic image estimation unit, 152 Light source area correction unit, 171 Mode switching unit, 172 Input unit, 201 CPU, 202 ROM, 203 RAM, 206 Input unit, 207 Output unit, 208 Memory unit, 209 Communication unit, 210 Drive, 211 Removable Media
Claims
1. An information processing device having: an upper room layout estimation unit that estimates upper room layout information including coordinate information indicating the boundary lines between the ceiling and walls of the room and coordinate information indicating the corner points of the ceiling and walls based on an upper hemispherical panoramic image, which is an image of the upper part of the room; and a lower room layout estimation unit that estimates lower room layout information including coordinate information indicating the boundary lines between the floor and walls of the room and coordinate information indicating the corner points of the floor and walls using the upper room layout information and camera position information of a camera that captured the upper hemispherical panoramic image.
2. The information processing device according to claim 1, further comprising: a lower hemispherical panoramic image estimation unit that estimates a lower hemispherical panoramic image based on the upper hemispherical panoramic image, room layout information consisting of the upper room layout information and the lower room layout information, and a captured image captured by a camera different from that which captured the upper hemispherical panoramic image, and generates a celestial sphere panoramic image by combining the upper hemispherical panoramic image and the lower hemispherical panoramic image; and a light source area correction unit that corrects an area of a light source that satisfies a predetermined condition in the celestial sphere panoramic image.
3. The information processing device according to claim 1, wherein the upper room layout estimation unit generates a pseudo omnidirectional panoramic image using an inverted upper hemispherical panoramic image obtained by inverting the upper hemispherical panoramic image upside down, and estimates the upper room layout information based on the pseudo omnidirectional panoramic image.
4. The information processing device according to claim 3, wherein the upper room layout estimation unit estimates room layout information based on the pseudo spherical panoramic image, and estimates the upper room layout information by deleting lower room layout information from the estimated room layout information.
5. The information processing device according to claim 4, wherein the upper room layout estimation unit estimates room layout information based on the pseudo spherical panoramic image using a machine learning model.
6. The information processing device according to claim 1, wherein the upper room layout estimation unit estimates the upper room layout information from the upper hemispherical panoramic image.
7. The information processing device according to claim 6, wherein the upper room layout estimation unit estimates the upper room layout information from the upper hemispherical panoramic image using a machine learning model.
8. The information processing device according to claim 1, wherein the lower room layout estimation unit estimates the lower room layout information by calculating points on the boundary line between the floor and walls that correspond to points on the boundary line between the ceiling and walls of the upper room layout information.
9. The information processing device according to claim 2, wherein the lower hemispherical panoramic image estimation unit divides the area of the lower hemispherical panoramic image into a floor area and multiple wall areas based on the room layout information, and estimates the lower hemispherical panoramic image by interpolating pixel values of the floor area and multiple wall areas.
10. The information processing device according to claim 9, wherein the pixel values of the wall region are complemented by an average color within the same wall region of the upper hemispherical panoramic image.
11. The information processing device according to claim 9, wherein the method for complementing the pixel values of the wall region is a method for complementing by extracting patches within the same wall region of the upper hemispherical panoramic image and tiling the extracted patches.
12. The information processing device according to claim 9, wherein the method for complementing the pixel values of the wall region is a method of inputting the upper hemispherical panoramic image into a machine learning model and performing completion using the machine learning model.
13. The information processing device according to claim 9, wherein the pixel values of the floor area are interpolated using an average color of the floor area calculated from the captured image.
14. The information processing device according to claim 9, wherein the method for interpolating the pixel values of the floor area is a method for interpolating by cutting out a patch of the floor area from the captured image and tiling it.
15. The information processing device according to claim 9, wherein the method for complementing the pixel values of the floor area is a method of inputting the upper hemispherical panoramic image into a machine learning model and performing completion using the machine learning model.
16. The information processing device according to claim 2, further comprising a light source estimation unit that estimates light source information from the upper hemispherical panoramic image, and wherein the light source area correction unit corrects the light source area where the position of the light source indicated by the light source information satisfies the predetermined condition.
17. The information processing device according to claim 16, wherein the light source that satisfies the predetermined condition is a light source that is at a distance equal to or greater than a predetermined value from both the ceiling and the walls of the room.
18. The information processing device according to claim 2, wherein the light source area correction unit corrects the light source area that satisfies the predetermined conditions using pixel values of pixels surrounding the area.
19. The information processing device according to claim 2, further comprising: a 3D mesh model generation unit that generates a textured 3D mesh model of the room using the spherical panoramic image after correcting the light source area that satisfies the predetermined condition; and an image rendering unit that generates a composite image by combining the captured image with a CG object using the textured 3D mesh model of the room.
20. An information processing method comprising: an information processing device estimating upper room layout information including coordinate information indicating the boundary line between the ceiling and walls of the room and coordinate information indicating the corner points of the ceiling and walls based on an upper hemispherical panoramic image, which is an image of the upper part of the room; and estimating lower room layout information including coordinate information indicating the boundary line between the floor and walls of the room and coordinate information indicating the corner points of the floor and walls using the upper room layout information and camera position information of a camera that captured the upper hemispherical panoramic image.
Citation Information
Patent Citations
Image processing device, image processing method, and program
JP2013127774A
Image processing device, program, and image processing method
JP2023110454A