Image processing method, program, and image processing system
Patent Information
- Application Number
- JP2025107970
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2025-06-26
- Publication Date
- 2025-12-02
- Estimated Expiration
- 2040-11-11
AI Technical Summary
Conventional methods for combining virtual objects with captured images often place them in unnatural positions, lacking accuracy in automatic placement.
An image processing method that includes structure estimation, area estimation, and image synthesis to automatically place virtual objects at appropriate positions within a space based on the estimated structure.
Enables accurate and natural placement of virtual objects within a structure, enhancing the realism of online property viewings.
Smart Images

Figure 00000000_0000_ABST
Abstract
Description
[Technical Field]
[0001] The present disclosure relates to an image processing method, a program, and an image processing system. [Background technology]
[0002] A system is known that distributes image data captured using an omnidirectional camera, allowing a user to view the situation at a remote location from another location. Omnidirectional images of a specific location can be viewed by the viewer in any direction, conveying realistic information. Such systems are used, for example, in the real estate industry, for online property viewings.
[0003] Such systems are used, for example, in the real estate industry, in areas such as online property viewing. There is also a service called "home staging," which arranges furniture and accessories in a property to create a spatial presentation, giving viewers an attractive image of the home and encouraging smooth transactions. In this service, in order to reduce costs, time, or the risk of damage to the property, a service is already known in which, rather than arranging actual furniture in the property, three-dimensional models (CG) of furniture are superimposed on photographed images of the property (e.g., Patent Documents 1 to 3). Summary of the Invention [Problem to be solved by the invention]
[0004] However, with conventional methods, when combining an image of a virtual object such as furniture with a captured image, the virtual object may be placed in an unnatural position for the viewer viewing the image, leaving room for improvement in terms of the accuracy of automatic placement of the virtual object. [Means for solving the problem]
[0005] In order to solve the above-mentioned problems, the invention of claim 1 is an image processing method executed by an image processing system, which executes a structure estimation step of estimating the structure of a space from a background image showing the internal space of a structure in all directions, an area estimation step of estimating an area in the space where a virtual object can be placed based on the estimated structure, and an image processing step of synthesizing the virtual object in the estimated area on the background image. [Effects of the Invention]
[0006] According to the present invention, it is possible to automatically place a virtual object at an appropriate position in the space inside a structure. [Brief explanation of the drawings]
[0007] [Figure 1] FIG. 1 is a diagram illustrating an example of the overall configuration of an image display system. [Figure 2] FIG. 10 is a diagram illustrating an example of a spherical image before a virtual object is placed. [Figure 3] FIG. 10 is a diagram showing an example of a processed image in which a virtual object is placed. [Figure 4] (A) is a hemispherical image (front) taken with a camera, (B) is a hemispherical image (back) taken with a camera, and (C) is an image represented by equirectangular projection. [Figure 5] (A) is a conceptual diagram showing how a sphere is covered with an equirectangular projection image, and (B) is a diagram showing a spherical image. [Figure 6] FIG. 10 is a diagram showing the positions of a virtual camera and a predetermined area when the celestial sphere image is a three-dimensional sphere. [Figure 7] 10 is a diagram showing the relationship between predetermined area information and an image of a predetermined area T. FIG. [Figure 8] FIG. 2 is a diagram illustrating an example of a state during shooting by the imaging device. [Figure 9] FIG. 1 is a diagram illustrating an example of a spherical image. [Figure 10] FIG. 10 is a diagram illustrating an example of a planar image converted from a spherical image. [Figure 11] FIG. 2 is a diagram illustrating an example of a hardware configuration of an image processing device, an image delivery device, and a display device. [Figure 12] FIG. 1 is a diagram illustrating an example of a functional configuration of an image display system. [Figure 13] FIG. 1 is a diagram illustrating an example of a functional configuration of an image display system. [Figure 14] FIG. 10 is a schematic diagram illustrating an example of an image data management table. [Figure 15] FIG. 10 is a conceptual diagram illustrating an example of a condition information management table. [Figure 16] 10 is a flowchart illustrating an example of processing in the image processing device. [Figure 17] 10 is a flowchart illustrating an example of a structure estimation process. [Figure 18] 10A and 10B are diagrams for explaining an example of a structure estimation result for a photographed image. [Figure 19] FIG. 10 is a diagram showing an example of the shape of a spatial structure estimated by the structure estimation process. [Figure 20] 10A and 10B are diagrams for explaining an example of a method for calculating the size of a space in the structure estimation process. [Figure 21] FIG. 10 is a diagram illustrating an example of an image when placement of a virtual object fails. [Figure 22] 10A and 10B are diagrams for explaining an example of a subject detection result for a captured image. [Figure 23] 10A and 10B are diagrams illustrating an example of a process for projecting a subject onto an estimated spatial structure. [Figure 24] 10 is a flowchart illustrating an example of layout processing of a virtual object. [Figure 25] 10A and 10B are diagrams illustrating an example of a layout algorithm for a 3D model of furniture. [Figure 26] 10A, 10B, and 10C are diagrams illustrating an example of a layout algorithm for a 3D model of furniture. [Figure 27] FIG. 1 is a diagram illustrating an example of a 3D model of furniture. [Figure 28]FIG. 10 is a diagram showing an example of a layout result of a 3D model of furniture. [Figure 29] FIG. 10 is a diagram showing an example of a layout result of 3D models of furniture on a 3D space model. [Figure 30] FIG. 10 is a diagram showing an example of a processed image using a shadow catcher. [Figure 31] (A)(B) is a diagram showing an example of image processing using image-based lighting. [Figure 32] 10A and 10B are diagrams showing an example of a spatial estimation result and a subject detection result in a processed image into which a virtual object is synthesized. [Figure 33] FIG. 10 is a diagram illustrating an example in which a virtual object is too close to the image capture device. [Figure 34] 10A and 10B are diagrams illustrating an example of a prohibited area for placement of a virtual object. [Figure 35] FIG. 10A is a diagram showing an example of a captured image in which a support member is captured, and FIG. 10B is a diagram showing an example of an image in which a virtual object is placed on the captured support member. [Figure 36] FIG. 10 is a diagram illustrating an example of the size of a support member. [Figure 37] FIG. 10 is a sequence diagram illustrating an example of an image display process. [Figure 38] (A) is an example of a screen of a captured image displayed on a display device, and (B) is an example of a screen of a processed image displayed on a display device. [Figure 39] FIG. 10 is a conceptual diagram illustrating an example of an additional information management table. [Figure 40] FIG. 10 is a diagram illustrating an example of the arrangement position of additional information. [Figure 41] (A) (B) are screen examples of processed images on which additional information is superimposed. [Figure 42] FIG. 10 is a diagram showing another example of a processed image displayed on the display device. [Figure 43] FIG. 10 is a diagram showing another example of a processed image displayed on the display device. DETAILED DESCRIPTION OF THE INVENTION
[0008] Hereinafter, an embodiment of the present invention will be described with reference to the drawings. In the description of the drawings, the same elements are given the same reference numerals, and duplicated explanations will be omitted.
[0009] ●Embodiment● ●Outline of the image display system First, an outline of the configuration of an image display system according to an embodiment will be described with reference to Fig. 1. Fig. 1 is a diagram showing an example of the overall configuration of an image display system. The image display system 1 shown in Fig. 1 is a system that allows viewers to view real estate properties online by displaying an image of the interior space of a structure such as a real estate property on a display device 90.
[0010] 1, the image display system 1 includes an image processing device 10, an image delivery device 30, a photographing device 70, a communication terminal 80, and a display device 90. The image processing device 10, the image delivery device 30, the photographing device 70, the communication terminal 80, and the display device 90 that constitute the image display system 1 can communicate with each other via a communication network 5. The communication network 5 is constructed using the Internet, a mobile communication network, a LAN (Local Area Network), or the like. Note that the communication network 5 may include not only wired communication networks but also wireless communication networks such as 3G (3rd Generation), 4G (4th Generation), 5G (5th Generation), Wi-Fi (Wireless Fidelity) (registered trademark), WiMAX (Worldwide Interoperability for Microwave Access), or LTE (Long Term Evolution).
[0011] The image processing device 10 is a server computer that performs image processing on captured images of the interior space of a structure such as a real estate property. The image processing device 10 synthesizes a virtual object onto the captured image based on, for example, captured image data transmitted from the image capture device 70, usage information indicating the usage of the space captured by the image capture device 70, and furniture information transmitted from the communication terminal 80. Here, the furniture information includes data indicating a 3D model of the furniture, furniture setting data indicating rules regarding the placement of the furniture, etc. The 3D model of the furniture is an example of a virtual object, and the furniture information is an example of object information. The virtual object may be a 3D model of a home appliance, or may be a 3D model of an electrical appliance, decoration, painting, lighting, fixture, or fixture, for example.
[0012] The image delivery device 30 is a server computer that delivers the processed image data that has been processed in the image processing device 10.
[0013] Here, the image processing device 10 and the image delivery device 30 are referred to as an image processing system 3. Note that the image processing system 3 may be, for example, a computer that integrates all or part of the functions of the image processing device 10 and the image delivery device 30. Also, each of the image processing device 10 and the image delivery device 30 may be configured to realize each function by distributing it among multiple computers. Furthermore, although the image processing device 10 and the image delivery device 30 will be described as server computers existing in a cloud environment, they may also be servers existing in an on-premise environment.
[0014] The image capturing device 70 is a special digital camera (spherical image capturing device) capable of capturing spherical (360°) images by capturing images of the interior space of a structure such as a real estate property in all directions. The image capturing device 70 is used, for example, by a real estate agent who manages or sells real estate properties. The image capturing device 70 may also be a wide-angle camera or a stereo camera capable of capturing wide-angle images having a predetermined angle of view or greater. A wide-angle image is generally an image captured using a wide-angle lens, which is an image captured with a lens that can capture a wider range than what the human eye can perceive. In other words, the image capturing device 70 is a capturing means capable of capturing images (spherical images, wide-angle images) captured using a lens with a focal length shorter than a predetermined value. A wide-angle image generally refers to an image captured with a lens with a focal length of 35 mm or less in 35 mm film equivalent. Furthermore, the captured images obtained by the image capturing device 70 may be video or still images, or may be both video and still images. The captured images may also include audio along with the images.
[0015] The communication terminal 80 is a computer such as a notebook PC that provides the image processing device 10 with information about virtual objects to be placed in the space shown in the captured image. The communication terminal 80 is used, for example, by a furniture manufacturer that manufactures or sells the furniture to be placed.
[0016] The display device 90 is a computer such as a smartphone used by an image viewer. The display device 90 displays the image distributed from the image distribution device 30. Note that the display device 90 is not limited to a smartphone, and may be, for example, a PC, a tablet terminal, a wearable terminal, an HMD (head mounted display), a PJ (Projector), or an IWB (Interactive White Board: an electronic white board with a blackboard function that allows mutual communication), etc.
[0017] Here, an image displayed on the display device 90 in the image display system 1 will be described with reference to FIGS. 2 and 3. FIG. 2 is a diagram showing an example of a celestial sphere image before a virtual object is placed. The image shown in FIG. 2 is a celestial sphere image obtained by capturing an image of a room of a real estate property, which is an example of the interior space of a structure, using the image capturing device 70. A celestial sphere image is suitable for viewing real estate properties because it can capture the interior of a room in all directions. There are various forms of celestial sphere images, but they are often generated using the equirectangular (equirectangular projection) projection method described below. An advantage of an image generated using this equirectangular projection is that the outer shape of the image is rectangular, making it efficient and easy to store image data, and that there is little distortion near the equator and no distortion of vertical straight lines, making it look relatively natural.
[0018] FIG. 3 is a diagram showing an example of a processed image in which virtual objects are arranged. The image shown in FIG. 3 shows a state in which furniture is arranged in the room shown in the image of FIG. 2. In the image of FIG. 3, a 3D model of furniture, which is an example of a virtual object, is synthesized against the background image, using the omnidirectional image shown in FIG. 2 as a background image. The image processing device 10 arranges the 3D model of furniture in a natural state based on the structure of the floor, walls, ceiling, etc. of the room photographed by the photographing device 70. As shown in FIG. 3, a desk, a bed, etc. are arranged along the walls in the room, and the passageway that is used daily is not blocked by furniture.
[0019] Conventionally, in order to place 3D models of furniture in a spherical image of a room in a real estate property, the furniture needs to be placed in a natural position as seen from the shooting position of the camera, so manual operation by the user is required to adjust the placement position and orientation. Also, although there are methods for automatically placing furniture models, these require input of a floor plan of the room and user input operations to understand the structure of the room where the furniture will be placed, so there is room for improvement in terms of increasing the accuracy of automatic placement of virtual objects without the hassle.
[0020] Therefore, the image processing system 3 detects the structure of the room and objects installed in the room using a spherical image of the interior of the room, and estimates an area where a virtual object can be placed. Then, the image processing system 3 places the virtual object in the estimated area where it can be placed, and generates a processed image as shown in Fig. 3 in which the placed virtual object and the spherical image are combined. This allows the image processing system 3 to naturally place furniture based on the general state of the room estimated based on the spherical image.
[0021] Here, a room in a real estate property is an example of an interior space of a structure. The structure is, for example, a building such as a house, an office, or a store. The omnidirectional image is an image captured by the image capturing device 70, and is an example of a background image showing the interior space of the structure in all directions.
[0022] ○How to generate spherical images○ Here, a method for generating a spherical image will be described with reference to Fig. 4 to Fig. 10. First, an outline of the process for generating a spherical image from an image captured by the image capturing device 70 will be described with reference to Fig. 4 and Fig. 5. Fig. 4(A) is a diagram showing a hemispherical image (front side) captured by the image capturing device, Fig. 4(B) is a hemispherical image (rear side) captured by the image capturing device, and Fig. 4(C) is a diagram showing an image expressed by equirectangular projection (hereinafter referred to as "equirectangular projection image"). Fig. 5(A) is a conceptual diagram showing a state in which a sphere is covered with an equirectangular projection image, and Fig. 5(B) is a diagram showing a spherical image.
[0023] The photographing device 70 has image sensors on both the front side (front side) and the back side (rear side). These image sensors (image sensors) are used in conjunction with optical components such as lenses that can capture hemispherical images (angle of view of 180° or more). The photographing device 70 can obtain two hemispherical images by capturing images of subjects around the user using the two image sensors.
[0024] 4(A) and (B), the images captured by the imaging element of the image capturing device 70 are curved hemispherical images (front and rear). The image capturing device 70 then combines the hemispherical image (front) with a hemispherical image (rear) flipped 180 degrees to create an equirectangular projection image EC as shown in FIG. 4(C).
[0025] Then, by using OpenGL ES (Open Graphics Library for Embedded Systems), the image capturing device 70 applies an equirectangular projection image EC to cover the spherical surface as shown in FIG. 5A, thereby creating a celestial sphere image (celestial sphere panoramic image) CE as shown in FIG. 5B. In this way, the celestial sphere image CE is expressed as an image in which the equirectangular projection image EC faces the center of the sphere. Note that OpenGL ES is a graphics library used to visualize 2D (2-Dimensions) and 3D (3-Dimensions) data. Furthermore, the celestial sphere image CE may be a still image or a video. Furthermore, the conversion method is not limited to OpenGL ES, and any method capable of converting a hemispherical image into an equirectangular projection may be used, for example, a CPU operation or OpenCL operation.
[0026] As described above, the spherical image CE is an image pasted to cover the spherical surface, which gives a sense of incongruity to people. Therefore, the image capturing device 70 can display a predetermined region T (hereinafter referred to as a "predetermined region image") that is a part of the spherical image CE as a planar image with little curvature, thereby enabling a display that does not give a sense of incongruity to people. This will be described with reference to FIGS. 6 and 7.
[0027] FIG. 6 is a diagram showing the positions of the virtual camera and the predetermined region when the celestial sphere image is a three-dimensional sphere. The virtual camera IC corresponds to the viewpoint of a user viewing the celestial sphere image CE displayed as a three-dimensional sphere. FIG. 6 represents the celestial sphere image CE as a three-dimensional sphere CS. If the celestial sphere image CE generated in this manner is a three-dimensional sphere CS, as shown in FIG. 6, the virtual camera IC is located inside the celestial sphere image CE. The predetermined region T in the celestial sphere image CE is the shooting region of the virtual camera IC and is specified by predetermined region information that indicates the shooting direction and angle of view of the virtual camera IC in a three-dimensional virtual space including the celestial sphere image CE. Zooming of the predetermined region T can also be expressed by moving the virtual camera IC closer to or farther away from the celestial sphere image CE. The predetermined region image Q is an image of the predetermined region T in the celestial sphere image CE. Therefore, the predetermined region T can be specified by the angle of view α and the distance f from the virtual camera IC to the celestial sphere image CE.
[0028] The predetermined area image Q is then displayed on a predetermined display as an image of the shooting area of the virtual camera IC. The following description will be given using the shooting direction (ea, aa) and angle of view (α) of the virtual camera IC. Note that the predetermined area T may be represented by the shooting area (X, Y, Z) of the virtual camera IC, which is the predetermined area T, instead of the angle of view α and the distance f.
[0029] Next, the relationship between the predetermined area information and the image of the predetermined area T will be described with reference to FIG. 7. FIG. 7 is a diagram showing the relationship between the predetermined area information and the image of the predetermined area T. As shown in FIG. 7, "ea" represents the elevation angle, "aa" represents the azimuth angle, and "α" represents the angle of view (Angle). That is, the attitude of the virtual camera IC is changed so that the gaze point of the virtual camera IC, indicated by the shooting direction (ea, aa), becomes the center point CP(x, y) of the predetermined area T, which is the shooting area of the virtual camera IC. As shown in FIG. 7, when the diagonal angle of view of the predetermined area T, represented by the angle of view α of the virtual camera IC, is α, the center point CP(x, y) becomes the parameter ((x, y)) of the predetermined area information. The predetermined area image Q is an image of the predetermined area T in the celestial sphere image CE. f represents the distance from the virtual camera IC to the center point CP(x, y). L represents the distance between any vertex of the predetermined area T and the center point CP(x, y) (2L is the diagonal). In FIG. 7, the trigonometric function shown in the following (Equation 1) generally holds.
[0030]
number
[0031] Next, a state in which an image is captured by the image capturing device 70 will be described with reference to FIG. 8. FIG. 8 is a diagram illustrating an example of a state in which an image is captured by the image capturing device. In order to capture an image that allows an entire room of a real estate property or the like to be viewed, it is preferable to install the image capturing device 70 at a position close to the height of the human eye. For this reason, as shown in FIG. 8, the image capturing device 70 is generally fixed to a support member 7 such as a monopod or a tripod when capturing an image. As described above, the image capturing device 70 is a celestial sphere image capturing device capable of capturing light rays in all directions around the entire periphery, and can also be said to capture an image on a unit sphere around the image capturing device 70 (a celestial sphere image CE). Once the image capturing direction of the image capturing device 70 is determined, the coordinates of the celestial sphere image are determined. For example, in FIG. 8, point A is located at a distance (d, -h) from the center point C of the image capturing device 70. If the angle between the line segment AC and the horizontal direction is θ, the angle θ can be expressed by the following (Equation 2).
[0032]
number
[0033] Assuming that point A is at a depression angle θ, the distance d between points A and B can be expressed by the following (Equation 3) using the installation height h of the image capturing device 70.
[0034]
number
[0035] Here, a process of converting position information on a celestial sphere image into coordinates on a planar image converted from the celestial sphere image will be outlined. Fig. 9 is a diagram for explaining an example of a celestial sphere image. Fig. 9(A) is a diagram showing the hemispherical image shown in Fig. 4(A) by connecting points where the horizontal and vertical incident angles with respect to the optical axis are equal. Hereinafter, the horizontal incident angle with respect to the optical axis will be referred to as "θ", and the vertical incident angle with respect to the optical axis will be referred to as "φ".
[0036] Moreover, Fig. 10(A) is a diagram illustrating an example of an image processed by equirectangular projection. Specifically, the images shown in Fig. 9 are associated with a pre-generated LUT (Look Up Table) or the like, processed by equirectangular projection, and the processed images shown in Fig. 9(A) and (B) are combined, whereby the planar image shown in Fig. 10(A) corresponding to the celestial sphere image is generated by the imaging device 70. The equirectangular projection image EC shown in Fig. 4(C) is an example of the planar image shown in Fig. 10(A).
[0037] As shown in FIG. 10(A), in an image processed by equirectangular projection, latitude (θ) and longitude (φ) are orthogonal to each other. In the example shown in FIG. 10(A), any position in the omnidirectional image can be indicated by defining the center of the image as (0,0), expressing the latitude direction as -90 to +90, and the longitude direction as -180 to +180. For example, the coordinates of the upper left corner of the image are (-180, -90). Note that the coordinates of the omnidirectional image may be expressed in a format using 360 degrees as shown in FIG. 10(A), or may be expressed in radians or in pixel numbers as in real images. Furthermore, the coordinates of the omnidirectional image may be converted into two-dimensional coordinates (x, y) as shown in FIG. 10(B) and expressed.
[0038] 10(A) or (B) is not limited to simply arranging the hemispherical images shown in Fig. 9(A) and (B) consecutively. For example, if the horizontal center of the celestial sphere image is not θ=180°, in the compositing process, the image capturing device 70 first preprocesses the hemispherical image shown in Fig. 4(C) and arranges it at the center of the celestial sphere image. Next, the image capturing device 70 may divide the hemispherical image shown in Fig. 4(B) into left and right portions of the image to be generated, each having a size that allows the preprocessed images to be arranged on the left and right portions, and then combine the hemispherical images to generate the equirectangular projection image EC shown in Fig. 4(C).
[0039] 10(A), the point corresponding to the pole (PL1 or PL2) of the hemispherical image (spherical image) shown in Figures 9(A) and 9(B) is the line segment CT1 or CT2. This is because, as shown in Figures 5(A) and 5(B), the celestial sphere image (for example, celestial sphere image CE) is created by pasting the planar image (equirectangular projection image EC) shown in Figure 10(A) onto a spherical surface by using OpenGL ES.
[0040] ●Hardware configuration Next, the hardware configuration of each device constituting the image display system according to the embodiment will be described with reference to Fig. 11. Note that components may be added or deleted from the hardware configuration shown in Fig. 11 as needed.
[0041] ○Hardware configuration of image processing device○ First, the hardware configuration of the image processing device 10 will be described with reference to Fig. 11. Fig. 11 is a diagram showing an example of the hardware configuration of the image processing device. Each piece of hardware configuration of the image processing device 10 is indicated by a reference number in the 100s. The image processing device 10 is constructed by a computer, and as shown in Fig. 11, includes a CPU (Central Processing Unit) 101, a ROM (Read Only Memory) 102, a RAM (Random Access Memory) 103, a HD (Hard Disk) 104, a HDD (Hard Disk Drive) controller 105, a display 106, an external device connection I / F (Interface) 108, a network I / F 109, a bus line 110, a keyboard 111, a pointing device 112, a DVD-RW (Digital Versatile Disk Rewritable) drive 114, and a media I / F 116.
[0042] Of these, the CPU 101 controls the overall operation of the image processing device 10. The ROM 102 stores programs used to drive the CPU 101, such as an IPL (Initial Program Loader). The RAM 103 is used as a work area for the CPU 101. The HD 104 stores various data, such as programs. The HDD controller 105 controls the reading and writing of various data from and to the HD 104 under the control of the CPU 101. The display 106 displays various information, such as a cursor, menus, windows, characters, or images. The display 106 may also be a touch panel display equipped with input means. The external device connection I / F 108 is an interface for connecting various external devices. In this case, the external device is, for example, a USB memory or a printer. The network I / F 109 is an interface for data communication using the communication network 5. The bus line 110 is an address bus, a data bus, or the like, for electrically connecting the components, such as the CPU 101, shown in FIG. 11 .
[0043] The keyboard 111 is a type of input means having multiple keys for inputting characters, numbers, various instructions, etc. The pointing device 112 is a type of input means for selecting or executing various instructions, selecting a processing target, moving a cursor, etc. The input means may be not only the keyboard 111 and the pointing device 112, but also a touch panel, a voice input device, etc. The DVD-RW drive 114 controls reading and writing of various data from a DVD-RW 113, which is an example of a removable recording medium. The removable recording medium is not limited to a DVD-RW, but may also be a DVD-R or a Blu-ray (registered trademark) Disc, etc. The media I / F 116 controls reading and writing (storing) of data from a recording medium 115, such as a flash memory.
[0044] ○Hardware configuration of image distribution device○ Fig. 11 is a diagram showing an example of the hardware configuration of an image delivery device. Each piece of hardware configuration of the image delivery device 30 is indicated by a reference number in the 300 range in parentheses. The image delivery device 30 is constructed using a computer, and as shown in Fig. 11, has the same configuration as the image processing device 10, so a description of each piece of hardware configuration will be omitted.
[0045] ○Display device hardware configuration○ Fig. 11 is a diagram showing an example of the hardware configuration of a display device. Each piece of hardware configuration of the display device 90 is indicated by a reference number in the 900s in parentheses. The display device 90 is constructed by a computer, and as shown in Fig. 11, has the same configuration as the image processing device 10, so a description of each piece of hardware configuration will be omitted.
[0046] Each of the above programs may be recorded on a computer-readable recording medium as an installable or executable file and distributed. Examples of recording media include CD-Rs (Compact Disc Recordables), DVDs (Digital Versatile Disks), Blu-ray Discs, SD cards, and USB memory. The recording media may also be provided domestically or internationally as a program product. For example, the image processing system 3 realizes the image processing method according to the present invention by executing the program according to the present invention.
[0047] ●Function configuration Next, the functional configuration of the image display system according to the embodiment will be described with reference to Fig. 12 to Fig. 15. Fig. 12 and Fig. 13 are diagrams showing an example of the functional configuration of the image display system. Fig. 12 and Fig. 13 show devices or terminals shown in Fig. 1 that are related to the processing or operation described below.
[0048] ○Functional configuration of image processing device○ First, the functional configuration of the image processing device 10 will be described with reference to Fig. 12. The image processing device 10 includes a transmitting / receiving unit 11, a receiving unit 12, a determining unit 13, a structure estimating unit 14, a detecting unit 15, a position estimating unit 16, a region estimating unit 17, a determining unit 18, a locating unit 19, an image processing unit 20, an input unit 21, and a storing / reading unit 29. Each of these units is a function or means realized when any of the components shown in Fig. 11 operates in response to an instruction from the CPU 101 in accordance with a program for the image processing device loaded from the HD 104 onto the RAM 103. The image processing device 10 also includes a storage unit 1000 configured by the ROM 102, RAM 103, and HD 104 shown in Fig. 11.
[0049] The transmitting / receiving unit 11 is mainly realized by the processing of the CPU 101 for the network I / F 109, and transmits and receives various data or information to and from other devices or terminals via the communication network 5.
[0050] The reception unit 12 is realized mainly by the processing of the CPU 101 on the keyboard 111 or the pointing device 112, and receives various selections or inputs from the user. The determination unit 13 is realized by the processing of the CPU 101, and makes various determinations.
[0051] The structure estimation unit 14 is realized by processing of the CPU 101, and estimates the structure of the space based on a background image showing the interior space of the structure in all directions.
[0052] The detection unit 15 is realized by the processing of the CPU 101, and detects the subject shown in the background image.
[0053] The position estimation unit 16 is realized by processing of the CPU 101, and estimates the spatial position of the subject detected by the detection unit 15.
[0054] The area estimation unit 17 is realized by processing of the CPU 101, and estimates an area in the space where a virtual object can be placed, based on the structure of the space estimated by the structure estimation unit .
[0055] The determination unit 18 is realized by processing of the CPU 101, and determines a virtual object to be placed in the space based on the use of the space shown in the background image.
[0056] The placement unit 19 is realized by processing of the CPU 101, and places a virtual object in the area estimated by the area estimation unit 17. The placement unit 19 performs a layout of the virtual object determined by the determination unit 18 in the area where placement is possible estimated by the area estimation unit 17, for example.
[0057] The image processing unit 20 is realized by processing of the CPU 101, and synthesizes a virtual object with respect to a background image in the area estimated by the area estimation unit 17. The image processing unit 20 performs rendering processing on the arranged virtual object, for example, based on the layout result of the virtual object by the arrangement unit 19.
[0058] The input unit 21 is mainly realized by the processing of the CPU 101 on the external device connection I / F 108, and receives input of various data or information from external devices.
[0059] The storage / readout unit 29 is mainly realized by the processing of the CPU 101 , and stores various data (or information) in the storage unit 1000 and reads out various data (or information) from the storage unit 1000 .
[0060] Image data management table Fig. 13 is a conceptual diagram showing an example of an image data management table. An image data management DB 1001 configured by the image data management table shown in Fig. 13 is constructed in the storage unit 1000. This image data management table manages an image ID that identifies image data, a condition ID that identifies the selection condition of a virtual object, captured image data, and processed image data in association with each other.
[0061] Condition information management table Fig. 14 is a conceptual diagram showing an example of a condition information management table. The condition information management table manages condition information that indicates placement conditions for virtual objects. A condition information management DB 1002 configured by the condition information management table shown in Fig. 14 is constructed in the storage unit 1000. This condition information management table associates and manages a condition ID that identifies selection conditions for a virtual object, the use and size of a room, and information on the style and set of furniture, which is an example of a virtual object to be selected.
[0062] ○Functional configuration of image distribution device○ Next, the functional configuration of image delivery device 30 will be described with reference to Fig. 13. Image delivery device 30 has a transmission / reception unit 31, a display control unit 32, a determination unit 33, a coordinate detection unit 34, a calculation unit 35, an image processing unit 36, and a storage / readout unit 39. Each of these units is a function or means realized when any of the components shown in Fig. 11 operates in response to an instruction from CPU 301 in accordance with an image delivery device program loaded from HD 304 onto RAM 303. Image delivery device 30 also has a storage unit 3000 constructed by ROM 302, RAM 303, and HD 304 shown in Fig. 11.
[0063] The transmitting / receiving unit 31 is mainly realized by the processing of the CPU 301 for the network I / F 309, and transmits and receives various data or information to and from other devices or terminals via the communication network 5.
[0064] The display control unit 32 is mainly realized by the processing of the CPU 301, and causes the display device 90 to display various images, characters, etc. The display control unit 32 uses, for example, a web browser or a dedicated application to distribute (transmit) image data to the display device 90, thereby causing the display device 90 to display various screens. The various screens displayed on the display device 90 are defined, for example, by HTML (HyperText Markup Language), XHTML (Extensible HyperText Markup Language), CSS (Cascading Style Sheets), JavaScript (registered trademark), or the like. The determination unit 33 is realized by the processing of the CPU 301, and makes various determinations.
[0065] The coordinate detection unit 34 is realized by processing by the CPU 101, and detects the coordinate position of a virtual object shown in a processed image generated by the image processing device 10. The calculation unit 35 is realized by processing by the CPU 301, and calculates the center position of the virtual object for superimposing additional information, which will be described later, on the processed image, based on the coordinate position detected by the coordinate detection unit 34. The image processing unit 36 is realized by processing by the CPU 301, and performs predetermined image processing on the processed image generated by the image processing device 10.
[0066] The storage / readout unit 39 is mainly realized by the processing of the CPU 301 , and stores various data (or information) in the storage unit 3000 and reads out various data (or information) from the storage unit 3000 .
[0067] ○Functional configuration of the display device○ Next, the functional configuration of the display device 90 will be described with reference to Fig. 13. The display device 90 has a transmitting / receiving unit 91, a receiving unit 92, and a display control unit 93. Each of these units is a function or means realized when any of the components shown in Fig. 11 operates in response to an instruction from the CPU 901 in accordance with a display device program loaded from the HD 904 onto the RAM 903.
[0068] The transmitting / receiving unit 91 is mainly realized by the processing of the CPU 901 for the network I / F 909, and transmits and receives various data or information to and from other devices or terminals via the communication network 5.
[0069] The reception unit 92 is mainly realized by the processing of the CPU 901 on the keyboard 911 or the pointing device 912, and receives various selections or inputs from the user.
[0070] The display control unit 93 is mainly realized by the processing of the CPU 901, and causes various images, characters, etc. to be displayed on the display 906. The display control unit 93 accesses the image delivery device 30 using, for example, a web browser or a dedicated application, and causes an image corresponding to data delivered from the image delivery device 30 to be displayed on the display 906, which is an example of a display means.
[0071] Processing or operation of the embodiment ○Image synthesis processing○ Next, the processing or operation of the image display system according to the embodiment will be described with reference to Figs. 16 to 43. First, the image synthesis processing in the image processing device 10 will be described with reference to Figs. 16 to 32. In the following description, an example of a room that is a real estate property is shown as an example of the interior space of a structure, and an example of furniture placed in the room is shown as an example of a virtual object. Fig. 16 is a flowchart showing an example of processing in the image processing device.
[0072] First, the image processing device 10 accepts input of a photographed image of a specific room, which is an example of the interior space of a structure (step S1). Specifically, the transmitter / receiver 11 of the image processing device 10 receives, for example, a photographed image of the interior space of a specific structure photographed by the photographing device 70 from the photographing device 70 via the communication network 5. The image processing device 10 may be configured to accept input of the photographed image to be processed from the photographing device 70 when performing image synthesis processing, or may be configured to store the photographed image received from the photographing device 70 in advance in the storage unit 1000 and read the stored photographed image when performing image synthesis processing. The image processing device 10 may also be configured to accept input of the photographed image via the input unit 21 by directly connecting to the photographing device 70 via the external device connection I / F 108. Furthermore, since the photographing device 70 may not have a communication function, for example, the image processing device 10 is not limited to accepting input of the photographed image directly from the photographing device 70, but may also be configured to accept input of the photographed image via a specific communication device owned by a real estate agent.
[0073] Next, the determination unit 13 uses the captured image input in step S1 to determine whether the furniture arrangement in the room shown in the captured image is appropriate (step S2). The room shown in the captured image is preferably, for example, an empty room with no furniture already placed therein, or a room with a predetermined space for arranging furniture. Therefore, if the determination unit 13 determines that the room shown in the captured image is an external space of a structure, such as outdoors, or is extremely narrow or has objects placed therein, and therefore no space for furniture arrangement can be secured, the determination unit 13 determines that the room shown in the captured image is not appropriate for furniture arrangement.
[0074] If the determination unit 13 determines that the image is suitable for furniture arrangement (YES in step S2), the process proceeds to step S3. On the other hand, if the determination unit 13 determines that the image is not suitable for furniture arrangement (NO in step S2), the process proceeds to step S9. In step S9, the image processing device 10 does not perform image synthesis processing and outputs an error message indicating that the image is not suitable for furniture arrangement. Specifically, the storage / readout unit 29 of the image processing device 10 associates the error message with the captured image input in step S1 and stores the error message in the storage unit 1000. This allows a viewer viewing the corresponding captured image to understand the error message along with the captured image. Note that the image processing device 10 may be configured to execute processing from step S3 onwards after outputting the error message. However, in this case, there is a high possibility that no furniture that can be arranged in the image synthesis processing in step S7 described below, and the processed image data stored in step S8 described below will be an image in which no furniture is arranged.
[0075] Next, the structure estimation unit 14 estimates the structure of the room shown in the captured image using the captured image input in step S1 (step S3). A known method for estimating the structure of a room involves, for example, detecting straight lines of a subject shown in the captured image using image processing, determining the vanishing points of the detected straight lines, and estimating the structure of the room from the boundaries of the floor, walls, or ceiling. When using a spherical image, all elements necessary for estimating the structure of the room, such as the ceiling, floor, and walls, are captured. This has the advantage of increasing the accuracy of structure estimation compared to when using a normal planar image, which captures only a portion of the room and makes estimation other than detecting the vanishing point difficult. Also known is a method using machine learning to detect vanishing points, boundaries between the floor and walls or the ceiling and walls, or to estimate the three-dimensional structure from the detection results. The structure estimation unit 14 may perform structure estimation using any known method.
[0076] ○Structure estimation processing An example of the structure estimation process in the image processing device 10 will now be described in detail with reference to Figures 17 to 20. Figure 17 is a flowchart showing an example of the structure estimation process.
[0077] First, the structure estimation unit 14 estimates the vertices of the space shown in the photographed image using the photographed image (step S31). Specifically, for example, as described above, the structure estimation unit 14 detects the lines of the subject shown in the photographed image by image processing of the photographed image, and estimates the vanishing points calculated from the detected lines as the vertices of the space.
[0078] FIG. 18 is a diagram illustrating an example of a structure estimation result for a captured image. FIG. 18 shows an example of a room structure represented by equirectangular projection. As described above, in equirectangular projection, vertical lines are projected as straight lines, and horizontal lines are projected as curved lines. When these are applied to the structure of a room, many rooms have shapes in which straight lines intersect vertically with each other. Therefore, the structure estimation unit 14 can estimate the general structure of a room by using an image represented by equirectangular projection. The structure estimation unit 14 estimates the general structure of the room by detecting the elements, lines, and surfaces that make up the room. The example in FIG. 18 is an example of an image of a rectangular parallelepiped room, and the structure estimation unit 14 estimates that there are four horizontal surfaces and two vertical surfaces.
[0079] Next, if the structure estimation unit 14 determines that the room shape can be classified based on the vertex estimation results obtained in step S31 (YES in step S32), it proceeds to step S33. On the other hand, if the structure estimation unit 14 determines that the room shape cannot be classified (NO in step S32), it continues the process of step S31. FIG. 19 illustrates an example of the shape of a spatial structure estimated by the structure estimation process. Real rooms vary in shape, and obtaining detailed three-dimensional information requires measurements using a laser scanner, total station, or the like, which is a time-consuming and expensive process. When virtually arranging furniture, detailed reconstruction is unnecessary; a simplified reconstruction based on narrowed conditions is sufficient. In other words, since it is sufficient to know the rough structure of the room, the structure estimation unit 14 narrows down the conditions by, for example, assuming that the room is composed of straight lines and planes, and that these lines essentially intersect at 90° angles (the Manhattan World Assumption). Furthermore, in order to aim for restoration to the extent that furniture can be placed, the structure estimation unit 14 classifies the room as, for example, a rectangular parallelepiped with eight vertices as shown in FIG. 19, or an L-shaped room with 12 vertices.
[0080] Next, the structure estimation unit 14 estimates the size (scale) of the space depicted in the captured image (step S33). Specifically, the structure estimation unit 14 acquires the equirectangular projection coordinates of each vertex of the room using the techniques of steps S31 and S32. The structure estimation unit 14 converts the acquired equirectangular projection coordinates into three-dimensional space coordinates.
[0081] The structure estimation unit 14 detects whether the image capture device 70 is installed vertically or detects the direction of gravitational acceleration and performs correction. Then, the structure estimation unit 14 can estimate the structure of the room by assuming that the south pole of the equirectangular projection (for example, PL1 shown in FIG. 9(A)) coincides with the direction of gravitational acceleration and that the Manhattan world hypothesis is followed.
[0082] The Manhattan World hypothesis is a hypothesis that many man-made artifacts are constructed parallel to a Cartesian coordinate system, which allows us to assume that walls, ceilings, etc. are constrained to be parallel in the x, y, and z directions. Based on this assumption, as shown in Figure 20, assuming that the height of the camera 70 is h from the floor, the distance between point A, where the floor and wall meet, and point B, where the ceiling and wall meet, can be expressed using the installation height h of the camera 70. However, this method only gives us a rough idea of the shape of the room, but does not provide an accurate indication of its size (scale). In an extreme example, we do not know whether the room is a miniature 20 cm high or a typical 2 m high room, so we need to have some idea of the scale of the room in order to arrange the furniture.
[0083] As a method for calculating the scale of the room, the structure estimation unit 14, for example, assumes that the installation height h of the image capture device 70 is known, and calculates point A located at the depression angle θ using the above (Equation 3) shown in FIG. 8. As a method for measuring the installation height of the image capture device 70 by physical means, the structure estimation unit 14 may measure the distance to the optical center of the image capture device 70 using a technique such as laser ranging. As a method for measuring the installation height of the image capture device 70 by image processing, the structure estimation unit 14 may measure the distance to the scale by placing a scale of known length on the floor and photographing it with the image capture device 70.
[0084] The structure estimation unit 14 may also estimate the scale of a room using the room height as a given value. In Japan, the Building Standards Act stipulates that ceiling height must be at least 210 cm, while the ceiling height of a typical apartment building is 240 cm to 250 cm. In the United States, the ceiling height is approximately 8 feet (243 cm), which is similar to that in Japan. Although room heights vary, a scale accuracy of ±10 cm will match within 5% or less, functioning as a rough scale. Furthermore, as a method for measuring the distance to an object using stereoscopic vision, the structure estimation unit 14 may utilize the parallax of the optical centers of the multiple lenses of the image capture device 70 to measure the distance to a specific object using the common area between the lenses. Furthermore, as a method for estimating the scale from so-called structure from motion, which estimates a three-dimensional structure from multiple images, and IMU (Inertial Measurement Unit) data, the structure estimation unit 14 may estimate the scale based on the travel distance, which can be roughly estimated from the IMU data of the image capture device 70.
[0085] The structure estimation unit 14 may use any of the above methods as a method for calculating the scale of the room in step S33.
[0086] Then, the structure estimation unit 14 acquires coordinate information of each vertex based on the room structure estimated in steps S31 to S33 (step S34). The structure estimation unit 14 acquires coordinate information of each vertex of n rooms (n=8 or 12) as a result of a series of processes. For example, the structure estimation unit 14 acquires coordinates Cn (Cn=((x0, y0, z0), (x1, y1, z1), ... (xn, yn, zn)) expressed in XYZ coordinates as shown in FIG. 10(B) using the optical center of the image capture device 70 as a limit. Note that the structure estimation unit 14 may also acquire coordinates displayed in polar coordinates as shown in FIG. 10(A).
[0087] In this way, the structure estimation unit 14 can use the captured image input to the image processing device 10 to estimate the general structure of the room shown in the captured image.
[0088] Returning to FIG. 16, the detection unit 15 of the image processing device 10 detects objects present in the room captured in the captured image input in step 1 (step S4). Here, the image processing device 10 may not be able to properly arrange furniture just by knowing the structure of the room. FIG. 21 is a diagram showing an example of an image when the arrangement of virtual objects has failed. As shown in FIG. 21, furniture may be placed in a position that is not suitable for actual furniture arrangement, such as a bed being placed in the hallway of the room. Therefore, in order to realize a natural furniture layout, the image processing device 10 uses the detection unit 15 to detect objects captured in the captured image and estimates locations where furniture can be placed naturally.
[0089] Here, the subject detected by the detection unit 15 is a structural object of the room that appears in the captured image, such as an object installed in the room, i.e., an object that is previously installed in the room and that is related to the layout of the room. The subject detected by the detection unit 15 is, for example, a door, a window, a frame, a sliding door, an electric switch, a closet, a storage room, a kitchen, a passageway, an air conditioner, an electrical outlet, a lighting outlet, a fireplace, a ladder, a staircase, or a fire alarm.
[0090] Since the development of machine learning, many object detection algorithms have been known as methods for detecting objects in images. A typical method is to represent the object detection result as a rectangle (bounding box). Furthermore, by using a method called semantic segmentation, which represents the object as a region, it is possible to detect the object with higher accuracy. The detection unit 15 may use any known method for detecting the object in step S4. Furthermore, the detection unit 15 detects the type of object appearing in the image using a known method. The type of object appearing in the image is, for example, information for identifying the object appearing in the image (e.g., whether it is a door or a window). FIG. 22 is a diagram illustrating an example of an object detection result for a captured image. FIG. 22 shows the results of the detection unit 15 detecting a kitchen, an air conditioner, a window, a door, and a hallway from the objects appearing in the captured image.
[0091] In this way, the image processing device 10 can estimate the state of the room shown in the captured image by using the input captured image to estimate the structure of the room shown in the captured image and detect the object. Furthermore, by performing object detection based on the room structure estimated in step S3, the image processing device 10 can estimate the location in the room structure where the object may be installed, thereby improving processing efficiency. Note that the image processing device 10 may perform the processes of steps S3 and S4 in parallel, or may perform steps S3 and S4 in reverse order.
[0092] Next, the position estimation unit 16 of the image processing device 10 estimates the position of the object detected in step S4 within the room (step S5). The object detection result is expressed as a rectangle when a bounding box is used, or as filled pixels of the corresponding range when semantic segmentation is used. These are representations on a unit sphere of the image capture device 70 as shown in FIG. 23(A), and can also be represented on an equirectangular projection as shown in FIG. 22. The position estimation unit 16 projects the object detection result on the unit sphere as shown in FIG. 23(A) onto the three-dimensional reconstructed shape of the room. For example, the position estimation unit 16 projects, among the detected objects, objects that are normally present on walls, such as doors, windows, and corridors as shown in FIG. 23(B), onto the room structure estimated by the structure estimation unit 14.
[0093] In this case, the position estimation unit 16 projects a virtual object according to the type of subject detected by the detection unit 15. For example, the position estimation unit 16 places a virtual object that serves as a light source at the position of the detected window and synthesizes an image of the virtual object placed by the image processing unit 20 (described later), thereby making it possible to more naturally represent external light entering the room. In this way, the position estimation unit 16 ultimately estimates and assigns the position of the subject within the structure of the room.
[0094] It should be noted that the estimated position of the subject in terms of the room structure by position estimation unit 16 does not necessarily have to be accurate. When detecting a subject using equirectangular projection, there will be a deviation from the actual position of the subject, but the result of the subject detection will be detected as being larger than the subject itself, so there will be a margin for estimating the area in which furniture can be placed, which will be described later, and this will not be a major problem in terms of layout.
[0095] Next, the image processing device 10 performs furniture layout processing (step S6). When a person actually lays out furniture, the layout is based on the structure of the room and the position of the subject within the room structure. There are rough rules for furniture layout done by humans based on customs, etc., and when furniture layout is done automatically, there are known methods for laying out the furniture according to the rules for human layout or methods for optimizing the layout using machine learning based on numerous past layout results. The image processing device 10 lays out the furniture based on simple rules for the room structure estimated by the structure estimation unit 14 and the subjects detected by the detection unit 15.
[0096] Layout processing An example of layout processing in the image processing device 10 will now be described in detail with reference to Fig. 24 to Fig. 29. Fig. 24 is a flowchart showing an example of layout processing of a virtual object. Fig. 24 shows processing for determining furniture to be placed in accordance with the purpose of a room and automatically placing the furniture in order in areas where it can be placed.
[0097] First, the determination unit 18 of the image processing device 10 determines the furniture to be arranged (step S61). The determination unit 18 determines the furniture to be arranged based on, for example, the purpose and size of the room. Specifically, the determination unit 18 determines the furniture to be arranged based on the condition information stored in the condition information management DB 1002 and purpose information indicating the purpose of the room. The purpose information is information specified by a real estate agent or the like who photographed the target room. The image processing device 10 receives the purpose information transmitted from an external device such as the photographing device 70 via the transceiver 11. Note that the purpose information may be input to the image processing device 10 together with the photographed image input in step S1, or may be information directly specified to the image processing device 10.
[0098] Here, the use information includes, for example, information about the purpose of the room and the size of the room. The purpose of the room is the purpose for which the room is used, such as a living room, bedroom, or children's room. Since it is generally difficult to determine the purpose of a room from the state of the room itself, it is preferable to configure the system so that the purpose is selected based on the intention of the user, such as a real estate agent who photographed the room. Note that the purpose of the room may be automatically estimated by the structure estimation unit 14 based on the structure of the room and the subject captured in the photographed image, for example, by estimating a large room with a kitchen as a living room and a room with few windows as a bedroom.
[0099] Furniture layouts vary widely, and the types of furniture to be arranged vary depending on personal preferences and cultural backgrounds. Furthermore, the goal of home staging is to make a room look beautiful, and aesthetic considerations are also important, so there are a variety of furniture layout patterns. The types of furniture are determined by various factors, such as the room's purpose, size, furniture style, season, and color coordination.
[0100] Therefore, the determination unit 18 retrieves condition information associated with the same use and size as the use information, for example, by searching the condition information management DB 1002 (see FIG. 15 ) using the use information as a search key. Then, the determination unit 18 selects furniture to be placed from among the furniture items stored in the storage unit 1000 or indicated in the furniture information transmitted from the communication terminal 80, based on the furniture style or furniture set indicated in the retrieved condition information.
[0101] In the example shown in FIG. 15 , the condition information defines different furniture sets for each room's purpose and size. For example, living rooms and bedrooms are each classified into three levels (L, M, S) according to the room size. For example, a furniture set for a large room is defined as including a dining table and a large sofa, while a furniture set for a small room is defined as including a single sofa and a table. The condition information also defines furniture styles instead of furniture sets. In this case, the determination unit 18 selects a furniture set according to the defined furniture style from among the furniture items stored in the storage unit 1000 or included in the furniture information transmitted from the communication terminal 80. Furniture styles include, for example, natural, pop, modern, Japanese, Scandinavian, or Asian. The condition information may also include information such as color coordination or season in addition to the furniture set or furniture style.
[0102] Next, the image processing device 10 acquires furniture information, which is information about the furniture to be arranged determined in step S61 (step S62). The furniture information includes data indicating a 3D model of the furniture and furniture setting data indicating rules regarding the arrangement of the furniture. Specifically, the storage / reading unit 29 of the image processing device 10 acquires the furniture information of the determined furniture by reading the furniture information stored in the storage unit 1000. The furniture information is transmitted to the image processing device 10 from a communication terminal 80 owned by a furniture manufacturer or the like and stored in advance in the storage unit 1000. Note that the transmitting / receiving unit 11 of the image processing device 10 may be configured to acquire the furniture information of the determined furniture by receiving the furniture information transmitted from the communication terminal 80 in response to a request from the image processing device 10 in step S62.
[0103] Next, the area estimation unit 17 estimates an area where furniture can be placed based on the room structure estimated in step S3 and the position of the subject estimated in step S5 (step S63). The area where furniture can be placed will be described in detail using FIGS. 25 and 26. FIG. 25 is a diagram showing, as a specific example, a layout algorithm for a 3D model of a rug, which is an example of furniture. FIG. 25(A) shows the position of the image capture device 70 and the room structure estimated by the structure estimation unit 14, and FIG. 25(B) shows a state in which a rug is placed in the center of the room, which is the area where furniture can be placed estimated by the area estimation unit 17. Rugs and carpets are laid on the floor, so they are furniture that can be placed regardless of the positions of subjects such as doors and windows, as long as the room structure is known.
[0104] Figure 26 shows a layout algorithm for a 3D model of a bed, a piece of furniture that requires understanding of its surroundings, as a more complex case. Figure 26(A) shows the position of the camera 70 and the room structure estimated by the structure estimation unit 14, Figure 26(B) shows the possible placement area estimated by the area estimation unit 17, and Figure 26(C) shows the state in which the bed has been placed in the possible placement area. A rug is also shown placed in the center of the room.
[0105] The area estimation unit 17 estimates an area in which the target furniture can be placed based on basic rules for placing furniture indicated in the furniture setting data acquired in step S62. The rules for placing a bed include, for example, placing the bed on the floor (not in the air), placing it along a wall, not placing it in a hallway or in front of a door (it may be placed in front of a window), etc. Based on these rules, the area estimation unit 17 estimates an area in which the bed can be placed, such as that shown in FIG. 26(B), based on the room structure estimated by the structure estimation unit 14 and the position of the subject detected by the detection unit 15.
[0106] The rules for placing beds also include sub-rules such as placing the bed randomly, placing the bed in a corner of the room, or placing the bed in the middle of a side of the room. The area estimation unit 17 determines the position for placing the bed based on the area where the bed can be placed and the sub-rules. If the area estimation unit 17 cannot place a piece of furniture according to the rules indicated in the furniture setting data, it cancels the placement of that piece of furniture and places another piece of furniture instead.
[0107] Next, the placement unit 19 of the image processing device 10 determines the placement of the 3D furniture models based on the furniture information acquired in step S63 (step S64). 3D furniture models are available in various file formats, such as 3ds.max, .blend, .stl, or .fbx, and any of these formats may be used. On the other hand, since the installation direction and center position of typical 3D furniture models are not specified, it is necessary to set rules for the initial installation direction and center point, and then edit the 3D model according to the set rules, or prepare separate data for conversion. The rules for the furniture installation direction and center point may be included in the furniture setting data, or may be set as a separate database when selecting a 3D model.
[0108] FIG. 27 is a diagram showing an example of a 3D model of furniture. The 3D model of the table shown in FIG. 27 has coordinates (0,-1,0) in the virtual space as the front, and the direction a person faces is aligned with the front. Furthermore, the 3D model of the table shown in FIG. 27 defines the center of the surface facing the floor as the center of the furniture. The center of the furniture is determined based on the surface of the furniture that is in contact with the floor. Furthermore, for example, the center of a light hanging from the ceiling is the point where the light touches the ceiling.
[0109] 28 is a diagram showing an example of the layout result of the 3D model of furniture. The placement unit 19 calculates the coordinates and orientation as the placement position of the furniture based on the placement area estimated in step S62 and the placement rules indicated in the furniture setting data. This allows the placement unit 19 to determine the furniture placement determined in step S61. Here, the layout information indicating the furniture placement determined by the placement unit 19 includes information on the type of furniture, as well as the orientation, position, and size of the furniture.
[0110] Then, when the arrangement of all the furniture determined in step S71 is completed (YES in step S75), the arrangement unit 19 ends the process. On the other hand, when the arrangement of all the furniture determined in step S71 is not completed (NO in step S75), the arrangement unit 19 repeats the process of step S74 until the arrangement of all the furniture is completed. Note that the furniture can be arranged in order from the largest to the smallest, allowing more furniture to be arranged.
[0111] In this way, the image processing device 10 can automatically arrange 3D models of furniture suitable for the purpose of the room shown in the captured image. Furthermore, the image processing device 10 can realize a more natural furniture arrangement for the viewer by arranging the determined 3D models of furniture in an arrangement possible area estimated based on the estimated room structure and the detected position of the subject.
[0112] Returning to FIG. 16, the image processing unit 20 of the image processing device 10 performs image synthesis processing of the captured image input in step S1 with the 3D model of the furniture arranged in step S6 (step S7). Specifically, the image processing unit 20 performs rendering using the captured image input in step S1, the room structure estimated in step S3, and the furniture layout information in step S6. For rendering, a CG tool such as 3dsMax, Blender, Maya, various CAD tools, Unity, or a web browser is used, preferably a tool with a function that allows the operation of the CG tool to be controlled by script. Furthermore, rendering is preferably performed in a format based on equirectangular projection, but rendering may also be performed using a projection method in which part of the image is converted using perspective projection or other projection methods. Furthermore, rendering may be performed using either a rasterization method or a ray tracing method, but ray tracing is preferred from the perspective of improving quality.
[0113] First, the placement unit 19 places 3D furniture models on the CG tool based on the layout results from step S6. FIG. 29 is a diagram showing an example of the layout results of 3D furniture models on a 3D space model. As described above, furniture layout information includes the type of furniture, as well as the orientation, position, and size of the furniture. The image processing unit 20 places the 3D furniture models decoded and specified by the script on the CG tool in the 3D space. Note that the furniture layout information may also include correction information for the 3D furniture models, such as the color or texture of the furniture. The example in FIG. 29 shows a bed, rug, desk, and potted plants placed in the 3D space.
[0114] Furthermore, the placement unit 19 may represent all or part of the room structure estimated in step S3 on a CG tool. The room structure, i.e., the floor, ceiling, and walls, may be represented using texture or may be represented transparently. The example in FIG. 29 represents everything except the foreground wall and ceiling. In the case of a transparent representation, the image processing unit 20 sets the transparent surface to function as a shadow catcher and renders a shadow, thereby enhancing the texture of the CG. Then, the image processing unit 20 performs a process of compositing the 3D model placed by the placement unit 19 with the captured image input in step S1.
[0115] Fig. 30 is a diagram showing an example of a processed image using the shadow catcher. As shown in Fig. 30, by using the shadow catcher function, the image processing unit 20 can cast a shadow where there is nothing or cast a shadow on a subject appearing in a captured image. In this way, by performing rendering processing, the image processing unit 20 can maintain consistent shadow quality as a single rendered image.
[0116] Fig. 31 is a diagram showing an example of image processing using Image Based Lighting. As shown in Fig. 31(B), the image processing unit 20 can express more natural light and improve texture by, for example, setting a captured image as the data background of Image Based Lighting, compared to the normal lighting processing shown in Fig. 30(A).
[0117] Then, the storage / readout unit 29 stores the processed image data combined in step S7 in the image data management DB 1001 (see FIG. 14) (step S8). In this case, the storage / readout unit 29 stores the processed image data combined in step S7 in the image data management DB 1001 in association with the photographed image data before the combining process and a condition ID that identifies the furniture selection condition.
[0118] In this way, the image processing device 10 estimates the structure of the room and the position of the subject using the input captured image, and places 3D models of furniture in the possible placement area according to the estimation results, thereby achieving a more natural placement of virtual objects.
[0119] Fig. 32 is a diagram showing an example of a space estimation result and an object detection result in a processed image into which a virtual object has been synthesized. In the example of Fig. 32, the dotted line indicates the space estimation result, which is the structure of a room, and the thick frame indicates the object detection result. As shown in Fig. 32, the image processing device 10 can arrange furniture in an arrangement possible area based on the space estimation result and the object detection result.
[0120] *Examples of image synthesis processing applications 33 to 36, an application example of the image synthesis process in the image processing device 10 will be described. First, a process of estimating a prohibited area for placement of a virtual object such as furniture will be described with reference to Figs. 33 and 34.
[0121] Fig. 33 is a diagram showing an example in which a virtual object is too close to the image capture device. As shown in Fig. 33, when a virtual object is placed in a placement possible area based on the room structure and subject detection results as described above, the virtual object may be too close to the image capture device 70, resulting in a poor appearance. Therefore, as shown in Figs. 34(A) and (B), the image processing device 10 estimates the area around the image capture device 70 as a placement prohibited area in which placement of a virtual object is prohibited, thereby improving the appearance of the processed image after image synthesis.
[0122] In this case, in step S4 described above, the detection unit 15 detects the position of the image capture device 70. In step S63, the area estimation unit 17 estimates the area around the image capture device 70 detected by the detection unit 15 as a placement prohibited area. Then, the area estimation unit 17 estimates an area where the virtual object can be placed, taking into account the estimated placement prohibited area as well as the room structure estimated by the structure estimation unit 14 and the subject position estimated by the position estimation unit 16. Note that the placement prohibited area may be defined in two dimensions or three dimensions.
[0123] Next, using FIGS. 35 and 36, a process for hiding a subject in a captured image will be described. FIGS. 35 and 36 illustrate an example in which a camera device 70 or a support member 7 is captured in the captured image. As shown in FIG. 35(A), because the camera device 70 captures the surroundings in all directions, the camera's hand holding the camera device 70 or a support member 7, such as a tripod or monopod, may appear in the captured image. This is undesirable from the perspective of making the room look beautiful. Therefore, the image processing device 10, for example, detects the support member 7 using the detection unit 15 and places a virtual object of a size that can hide the detected support member 7. The image processing unit 20 of the image processing device 10 then synthesizes an image of the placed virtual object with the captured image as a background, thereby eliminating the inclusion of the support member 7, as shown in FIG. 35(B).
[0124] FIG. 36 is a diagram illustrating an example of the size of a support member. The detection unit 15 detects the support member 7 in step S4. The support member 7 may be detected directly in equirectangular projection, or may be detected by converting the vertically downward direction using perspective projection. As a result, the detection unit 15 obtains the field of view angle φ of the support member 7. As shown in FIG. 36, if the installation height of the image capture device 70 from the floor is h and the width of the support member 7 is w, the width w of the support member 7 can be expressed as w = 2 tan(φ / 2). The placement unit 19 places any virtual object with a size equal to or greater than the width w on the floor, and the image processing unit 20 combines an image of the placed virtual object with the captured image to prevent the support member 7 from appearing in the image. Note that the object detected by the detection unit 15 is not limited to the support member 7. The detection unit 15 may detect the image capture device 70 or the photographer using the image capture device 70 and place any virtual object that hides the detected object.
[0125] ○Image display processing○ Next, the image display processing in the image processing system 3 will be described with reference to Fig. 37 to Fig. 43. Fig. 37 is a sequence diagram showing an example of the image display processing. Fig. 37 shows the processing when the processed image data stored in the image processing device 10 by the above-mentioned processing is distributed to a viewer using the image distribution device 30.
[0126] First, the transmitter / receiver 91 of the display device 90 transmits an image display request indicating a request for display of an image to the image delivery device 30 based on an input operation by the viewer on an input device or the like (step S51). This image display request includes an image ID that identifies an image in which the requested structure was captured. As a result, the transmitter / receiver 31 of the image delivery device 30 receives the image display request transmitted from the display device 90.
[0127] Next, the transmitting / receiving unit 31 of the image delivery device 30 transmits an image acquisition request to the image processing device 10, requesting acquisition of image data to be delivered to the display device 90 (step S52). This image acquisition request includes the image ID received in step S51. As a result, the transmitting / receiving unit 11 of the image processing device 10 receives the image acquisition request transmitted from the image delivery device 30.
[0128] Next, the storage / readout unit 29 of the image processing device 10 searches the image data management DB 1001 (see FIG. 14) using the image ID received in step S52 as a search key, and reads out the captured image data and processed image data associated with the same image ID as the received image ID (step S54). The transmission / reception unit 11 transmits the captured image data and processed image data read out in step S54 to the image delivery device 30. As a result, the transmission / reception unit 31 of the image delivery device 30 receives the captured image data and processed image data transmitted from the image processing device 10.
[0129] The display control unit 32 then transmits (distributes) the received captured image data or processed image data to the display device 90 via the transmitting / receiving unit 31, thereby causing the display device 90 to display the captured image or processed image (step S55). The display control unit 93 of the display device 90 then causes the display 906 to display the captured image or processed image corresponding to the data transmitted (distributed) from the image delivery device 30 (step S56).
[0130] Fig. 38(A) is an example of a screen of a captured image displayed on a display device, and Fig. 38(B) is an example of a screen of a processed image displayed on a display device. Captured image 400 shown in Fig. 38(A) is an image showing the state of a room before furniture has been arranged. On the other hand, processed image 600 shown in Fig. 38(B) is an image after furniture has been arranged in captured image 400 shown in Fig. 38(A).
[0131] Furthermore, the reception unit 92 of the display device 90 receives a selection of whether or not furniture is arranged in response to a predetermined input operation using the input means of the display device 90 (step S57). This allows the viewer to select whether or not furniture is arranged in the image displayed on the display device 90, and to view the state of the room before and after the furniture is arranged by switching between images.
[0132] In this way, the image processing system 3 can convey a more specific image of the room to the viewer by displaying on the display device 90 a processed image that has been combined with a 3D model of the furniture.
[0133] ○ Display of additional information Next, a process for superimposing and displaying additional information corresponding to the placed furniture on the processed image 600 as shown in Fig. 38(B) will be described using Fig. 39 to Fig. 41. When displaying the processed image 600 displayed on the display device 90, the image delivery device 30 can display additional information such as an icon to call attention, a link to a website selling the furniture, or a description of the furniture on the image rendered with the furniture model placed.
[0134] FIG. 39 is a conceptual diagram showing an example of an additional information management table. As shown in FIG. 13, an additional information management DB 3001 configured by the additional information management table shown in FIG. 39 is constructed in the storage unit 3000. This additional information management table manages, for each image ID that identifies image data, an additional ID that identifies additional information, a furniture type, coordinate information indicating the placement position of the additional information, and a website link in association with each other. Of these, the coordinate position is calculated by the calculation unit 35 based on the coordinate position of a 3D model of the furniture corresponding to the additional information. Furthermore, the website link is included, for example, in the furniture information transmitted from the above-mentioned communication terminal 80.
[0135] FIG. 40 is a diagram illustrating an example of the placement position of additional information. When additional information such as a description, an icon, or a link is superimposed on a captured image, the additional information needs to be accurately superimposed on the position of the furniture. Therefore, as shown in FIG. 40, the coordinate detection unit 34 detects the coordinates on the processed image of the furniture shown in the processed image. The calculation unit 35 calculates the coordinates of the center position D of the furniture and calculates the direction pointing to the center position D calculated from the photographing device 70. Then, the image processing unit 36 superimposes the additional information corresponding to the furniture on the coordinate position on the processed image that indicates the direction calculated by the calculation unit 35.
[0136] The additional information to be superimposed is, for example, an icon that alerts the viewer, a description of the furniture, or an image for accepting access to a website link. For example, when the image delivery device 30 receives the processed image data via the transmitter / receiver 31 in step S54 shown in Fig. 36, the image delivery device 30 executes the above-described process of superimposing the additional information, and in step S55 causes the display device 90 to display the processed image on which the additional information has been superimposed.
[0137] 41A and 41B show examples of a screen of a processed image on which additional information is superimposed. A processed image 600a shown in FIG. 41A displays, as additional information, an image 710 for accepting access to a website corresponding to the furniture shown in the processed image 600a. The image 710 includes, for example, a link to a website such as an EC (Electronic Commerce) site where the furniture shown in the processed image 600a can be purchased. A viewer viewing the processed image 600a displayed on the display device 90 can access the page of the corresponding EC site by, for example, pressing the image 710.
[0138] In the processed image 600a shown in FIG. 41(B), an icon 730 is displayed as additional information indicating that the furniture shown in the processed image 600a is a synthesized image. When an image of a virtual object is synthesized using ray tracing or image-based lighting technology, it can be difficult for the viewer to tell which parts of the processed image are CG. Therefore, in the processed image 600b, an icon 730 is displayed on the synthesized furniture image to draw attention, allowing the viewer to clearly distinguish between objects installed in the room and synthesized objects.
[0139] The icon 730 may not be displayed all the time, but may be hidden after a certain period of time has elapsed, or may be switched between display and non-display in response to an input operation by the viewer. The icon 730 may also be given an effect such as a blinking display to further draw the viewer's attention. Furthermore, the processed image 600b may be configured to display an explanation of the furniture when the viewer selects the icon 730.
[0140] ○Example of image display application Next, an application example of the processed image displayed on the display device 90 will be described with reference to Figs. 42 and 43. The processed image 600c shown in Fig. 42 shows a state in which an image of placed furniture is outlined. For example, when the image processing unit 36 of the image delivery device 30 receives the processed image data by the transmitter / receiver 31 in step S54 shown in Fig. 36, the image processing unit 36 generates a composite image in which the image of the furniture is outlined, and displays the generated processed image 600c on the display device 90. This allows a viewer of the processed image 600c to clearly grasp the location of the composited virtual object (furniture).
[0141] Furthermore, processed image 600d shown in Fig. 43 shows a state in which the color of the image of the placed furniture has been changed. For example, when the image processing unit 36 of the image delivery device 30 receives the processed image data via the transmitter / receiver 31 in step S54 shown in Fig. 36, the image processing unit 36 performs processing to change the color of the image of the furniture, and displays the processed processed image 600d on the display device 90. This allows a viewer of processed image 600c to clearly identify the location of the synthesized virtual object (furniture) by viewing an image with an intentionally unnatural appearance due to the changed color.
[0142] Effect of the embodiment As described above, the image display system 1 estimates the structure of the room and the position of the subject using the captured image taken by the photographing device 70, and places 3D models of furniture in the possible placement area according to the estimated results, thereby realizing a more natural placement of virtual objects.
[0143] Furthermore, the image display system 1 displays on the display device 90 a processed image in which 3D models of furniture have been synthesized by the image processing system 3, allowing the viewer of the image to view not only the empty room but also the room with furniture arranged, thereby conveying a more concrete image of the room to the viewer.
[0144] ●Summary● As described above, an image processing method according to one embodiment of the present invention is an image processing method executed by the image processing system 3, and executes the following steps: a structure estimation step of estimating the structure of a space from a background image (e.g., a celestial sphere image) in which the internal space of a structure (e.g., a room) is shown in all directions; an area estimation step of estimating an area in the space where a virtual object (e.g., a 3D model of furniture) can be placed based on the estimated structure; and an image processing step of overlaying the virtual object on the background image with the estimated area. This allows the image processing method to automatically place the virtual object at an appropriate position in the internal space of the structure.
[0145] Moreover, an image processing method according to an embodiment of the present invention further includes a detection step of detecting a subject shown in a background image (e.g., a spherical image) and a position estimation step of estimating the spatial position of the detected subject, and the region estimation step estimates a region based on the estimated spatial structure and the estimated position of the subject. Thus, the image processing method can estimate the state of the space by estimating the spatial structure shown in the background image and detecting the subject. Furthermore, the image processing method can improve processing efficiency by detecting the subject based on the estimated spatial structure, thereby estimating a location in the spatial structure where the subject may be located.
[0146] Furthermore, in an image processing method according to one embodiment of the present invention, the image processing system 3 includes a condition information management DB 1002 (an example of a storage means) that stores condition information indicating placement conditions for virtual objects (e.g., 3D models of furniture). The image processing method executed by the image processing system 3 then executes a determination step of selecting a virtual object according to the purpose of the space (e.g., a room) inside the structure from the stored condition information. This allows the image processing method to automatically place a virtual object suitable for the purpose according to the purpose of the space shown in the background image.
[0147] Moreover, an image processing system according to one embodiment of the present invention includes a structure estimation unit 14 (an example of a structure estimation means) that estimates the structure of the space from a background image (e.g., a celestial sphere image) in which the internal space of a structure (e.g., a room) is shown in all directions, an area estimation unit 17 (an example of an area estimation means) that estimates an area in the space where a virtual object (e.g., a 3D model of furniture) can be placed based on the estimated structure, and an image processing unit 20 (an example of an image processing means) that combines the virtual object with the estimated area in the background image. This enables the image processing system 3 to automatically place the virtual object in an appropriate position in the internal space of the structure.
[0148] Furthermore, the image processing system according to one embodiment of the present invention includes a display control unit 32 (an example of a display control means) that displays the processed image synthesized by the image processing unit 20 (an example of an image processing means) on the display device 90. This allows the image processing system 3 to switch between images to allow the viewer to view the state of the space before and after the virtual object is placed.
[0149] ●Additional Information● Each function of the above-described embodiments can be realized by one or more processing circuits. Here, the term "processing circuit" in the present embodiment includes a processor programmed to perform each function by software, such as a processor implemented by an electronic circuit, as well as devices designed to perform each function described above, such as an application specific integrated circuit (ASIC), a digital signal processor (DSP), a field programmable gate array (FPGA), a system on a chip (SOC), a graphics processing unit (GPU), and a conventional circuit module.
[0150] Furthermore, the various tables in the above-described embodiments may be generated by the learning effects of machine learning, and tables may not be used by classifying data for each associated item using machine learning. Here, machine learning refers to a technology that allows a computer to acquire human-like learning capabilities, in which the computer autonomously generates algorithms necessary for judgments such as data classification from previously acquired learning data and applies these algorithms to new data to make predictions. The learning method for machine learning may be any of supervised learning, unsupervised learning, semi-supervised learning, reinforcement learning, and deep learning, or may be a combination of these learning methods. Any learning method for machine learning is acceptable.
[0151] So far, we have described an image processing method, program, and image processing system according to one embodiment of the present invention, but the present invention is not limited to the above-described embodiment, and other modifications, such as additions, changes, or deletions, can be made within the scope of what a person skilled in the art can conceive, and any aspect is included in the scope of the present invention as long as it achieves the functions and effects of the present invention. [Explanation of symbols]
[0152] 1 Image display system 3. Image Processing System 5. Communication Network 7 Support member 10 Image processing device 11 Transmitter / Receiver 14 Structure estimation unit (an example of a structure estimation means) 15. Detection unit 16 Position estimation part 17 Area estimation unit (an example of area estimation means) 18 Decision Section 19 Placement section 20 Image processing unit (an example of image processing means) 30 Image distribution device 31 Transmitter / Receiver 32 Display control unit (an example of display control means) 35 Calculation unit (an example of calculation means) 70 Imaging equipment 80 Communication terminal 90 Display device 1002 Condition information management DB (an example of storage means) [Prior art documents] [Patent documents]
[0153] [Patent Document 1] Patent No. 6570161 [Patent Document 2] Patent No. 6116746 [Patent Document 3] Patent No. 3720587
Claims
1. An image processing method executed by an image processing system, comprising: a structure estimation step of estimating a structure and a size of a room of real estate from a celestial sphere image showing the room in all directions; a detection step of detecting structural objects in the room that appear in the omnidirectional image and further detecting the type of the structural objects in the room; a position estimation step of estimating a position in the room of the detected structural object of the room; Run Furthermore, the structure estimation step estimates a use of the room shown in all directions in the celestial sphere image based on the estimated structure and size of the room and the detected type of structural object in the room; a determining step of determining furniture to be synthesized into the celestial sphere image based on the estimated use; an area estimation step of estimating an area in the room in which the determined furniture can be placed, based on the estimated structure and size of the room, the estimated positions of structural objects in the room, and the determined furniture placement rule for the structure of the room and the type of structural object detected in the room; an image processing step of synthesizing the determined 3D model of the furniture with the estimated area for the spherical image; An image processing method that performs
2. the image processing system includes a storage means for storing condition information in which the purpose and size of a room are associated with furniture; 2. The image processing method according to claim 1, wherein said determining step selects furniture according to the estimated use of the room from the stored condition information.
3. the detecting step detects a support member that supports an image capturing device that captures an image of the room; The image processing method according to claim 1 or 2, wherein the image processing step comprises synthesizing a predetermined image onto the spherical image so as to hide the detected support member.
4. On the computer, a structure estimation step of estimating a structure and a size of a room of real estate from a celestial sphere image showing the room in all directions; a detection step of detecting structural objects in the room that appear in the omnidirectional image and further detecting the type of the structural objects in the room; a position estimation step of estimating a position in the room of the detected structural object of the room; Run Furthermore, the structure estimation step estimates a use of the room shown in all directions in the celestial sphere image based on the estimated structure and size of the room and the detected type of structural object in the room; a determining step of determining furniture to be synthesized into the celestial sphere image based on the estimated use; an area estimation step of estimating an area in the room in which the determined furniture can be placed, based on the estimated structure and size of the room, the estimated positions of structural objects in the room, and the determined furniture placement rule for the structure of the room and the type of structural object detected in the room; an image processing step of synthesizing the determined 3D model of the furniture with the estimated area for the spherical image; A program that executes the following.
5. A structure estimation means for estimating the structure and size of a room in real estate from a celestial sphere image showing the room in all directions; a detection means for detecting structural objects of the room that appear in the omnidirectional image and for detecting the type of the structural objects of the room; a position estimation means for estimating a position in the room of a detected structural object in the room; Equipped with the structure estimation means estimates a use of the room shown in all directions in the celestial sphere image based on the estimated structure and size of the room and the detected types of structural objects in the room; Furthermore, a determination means for determining furniture to be synthesized with the omnidirectional image based on the estimated purpose; an area estimation means for estimating an area in the room in which the determined furniture can be placed, based on the estimated structure and size of the room, the estimated positions of structural objects in the room, and the determined furniture placement rule for the room structure and the detected types of structural objects in the room; an image processing means for synthesizing the determined 3D model of the furniture with the estimated area in the celestial sphere image; An image processing system comprising:
6. Furthermore, 6. The image processing system according to claim 5, further comprising display control means for displaying the processed image synthesized by said image processing means on a display device.
7. The image processing system according to claim 6 , wherein the display control means switches between displaying the spherical image and displaying the processed image.
8. a calculation means for calculating a center position of the 3D model of the furniture on the displayed processed image and a direction pointing to the calculated center position; The display control means displays the additional information superimposed at the calculated center position of the coordinate position of the processed image indicating the calculated direction.
8. The image processing system according to claim 6 or 7.
9. 9. The image processing system according to claim 8, wherein the additional information is an icon corresponding to the furniture or a link to a website.
10. The image processing system according to any one of claims 6 to 9, wherein the display control means changes the color of the 3D model of the furniture shown in the processed image or displays an image in which the 3D model of the furniture is outlined.