Image processing device, program, and image processing method
The image processing device generates a 3D model of a space by geometrically processing panoramic images, addressing the limitations of deep learning-based methods and providing accurate VR stereoscopic representations.
Patent Information
- Application Number
- JP2022011923
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2022-01-28
- Publication Date
- 2026-01-06
- Estimated Expiration
- 2042-01-28
AI Technical Summary
Existing technologies for generating VR stereoscopic images from panoramic images rely on deep learning algorithms, lacking geometric calculations for accurate 3D model generation.
An image processing device that acquires a panoramic image, estimates the shape of a polygonal prism within a real space, and generates a 3D model by performing geometric operations on the image.
Enables the appropriate generation of a 3D model of the photographed space, allowing users to intuitively grasp the room's layout and furniture arrangement.
Smart Images

Figure 0007793999000009 
Figure 0007793999000010 
Figure 0007793999000011
Abstract
Description
[Technical Field]
[0001] The present invention relates to an image processing device, a program, and an image processing method. [Background technology]
[0002] Patent Document 1 discloses a technology for generating a VR stereoscopic image showing the shape of a room from a panoramic image of the room, and arranging 3D images of furniture in the VR stereoscopic image under conditions similar to the actual scale of the furniture and the room. The technology disclosed in Patent Document 1 adjusts the size of the furniture modeling data based on the length of a reference line in the panoramic image and the ratio of the actual measurement value of the reference line and the dimensions of the furniture, and can render the furniture arranged in the VR stereoscopic image based on the modeling data. [Prior art documents] [Patent documents]
[0003] [Patent Document 1] Patent No. 6872660 Summary of the Invention [Problem to be solved by the invention]
[0004] Patent Document 1 discloses the automatic generation of a VR stereoscopic image from a panoramic image using a deep learning algorithm. In other words, the technology disclosed in Patent Document 1 is not a configuration for generating a VR stereoscopic image by geometric calculations on a panoramic image.
[0005] The present invention has been made in consideration of the above circumstances, and its purpose is to provide an image processing device, etc., that can appropriately generate a 3D model of the space being photographed by performing geometric operations on a panoramic image. [Means for solving the problem]
[0006] An image processing device according to one aspect of the present disclosure includes an image acquisition unit that acquires a panoramic image taken within a closed real space, a spatial information acquisition unit that acquires the height of the real space and the number of corners of a polygon in a planar view, an estimation unit that identifies a position on the panoramic image corresponding to a vertical side of the real space and estimates the shape of a polygonal prism in the real space using the acquired height of the real space, and a generation unit that generates a 3D model constituting the estimated polygonal prism. [Effects of the Invention]
[0007] According to one aspect of the present disclosure, a 3D model of the space of the subject being photographed can be appropriately generated by performing geometric operations on the photographed image, and thus the appropriately generated 3D model can be provided to the user. [Brief explanation of the drawings]
[0008] [Figure 1] FIG. 1 is a schematic diagram illustrating an example of the configuration of an image processing system. [Figure 2] FIG. 2 is a block diagram showing an example of the configuration of a camera and a registration terminal. [Figure 3] FIG. 2 is a block diagram showing an example of the configuration of a server and a viewing terminal. [Figure 4] FIG. 2 is a schematic diagram illustrating an example of the configuration of a DB stored in a server. [Figure 5] FIG. 1 is a schematic diagram illustrating an example of the appearance of a camera. [Figure 6] FIG. 2 is an explanatory diagram of an example of the configuration of a captured image. [Figure 7] 10 is a flowchart illustrating an example of a procedure for registering a captured image. [Figure 8] FIG. 10 is a schematic diagram showing an example of a registration screen. [Figure 9] 10 is a flowchart illustrating an example of a procedure for generating an image for browsing. [Figure 10] FIG. 10 is an explanatory diagram of a process for generating an image for browsing. [Figure 11] FIG. 10 is an explanatory diagram of a process for generating an image for browsing. [Figure 12] 10 is a flowchart illustrating an example of a processing procedure for zenith correction. [Figure 13] 10 is a flowchart illustrating an example of a procedure for specifying a rotation axis included in zenith correction. [Figure 14] 10 is a flowchart illustrating an example of a procedure for specifying a rotation axis included in zenith correction. [Figure 15] 10 is a flowchart illustrating an example of a procedure for specifying a rotation amount included in zenith correction. [Figure 16] FIG. 10 is an explanatory diagram of zenith correction. [Figure 17] FIG. 2 is a schematic diagram showing the relationship between a polar coordinate system and a world Cartesian coordinate system. [Figure 18] 10 is a flowchart illustrating an example of a procedure for estimating a shooting space. [Figure 19] FIG. 10 is an explanatory diagram of a process for estimating a shooting space. [Figure 20] FIG. 10 is an explanatory diagram of a process for estimating a shooting space. [Figure 21] FIG. 2 is an explanatory diagram of polygon data. [Figure 22] FIG. 2 is an explanatory diagram of a texture image. [Figure 23] FIG. 2 is an explanatory diagram of a texture image. [Figure 24] 10 is a flowchart illustrating an example of a panoramic image playback processing procedure. [Figure 25] FIG. 10 is a schematic diagram showing an example of a display screen for a panoramic image. [Figure 26] 10 is a flowchart showing an example of a procedure for generating an image for browsing according to the second embodiment. [Figure 27] FIG. 10 is an explanatory diagram of a process for generating an image for browsing according to the second embodiment. [Figure 28] 10 is a flowchart showing an example of a procedure for estimating a shooting space according to the second embodiment. [Figure 29] FIG. 10 is a schematic diagram showing an example of the configuration of an image processing system according to a third embodiment. DETAILED DESCRIPTION OF THE INVENTION
[0009] An image processing device, a program, and an image processing method according to the present disclosure will be described in detail below with reference to the drawings illustrating embodiments thereof.
[0010] (Embodiment 1) 1 is a schematic diagram showing an example of the configuration of an image processing system. The image processing system 100 of this embodiment includes a server 10 (image processing device), a registration terminal 20, a camera 30, and a viewing terminal 40. The server 10, the registration terminal 20, and the viewing terminal 40 are connectable to a network N such as the Internet, a public communication line, or a LAN (Local Area Network), and transmit and receive information via the network N. The registration terminal 20 and the camera 30 are capable of, for example, wireless communication, and transmit and receive information directly via wireless communication. Note that if the camera 30 is connectable to the network N, the camera 30 and the registration terminal 20 may be configured to transmit and receive information via the network N.
[0011] The server 10 is an information processing device such as a server computer or a personal computer, and may be provided in multiple units, may be realized by multiple virtual machines provided in a single device such as a mainframe computer, or may be realized using a cloud server. The registration terminal 20 is a personal computer or a tablet terminal, and may be provided in multiple units. The camera 30 is an example of an imaging device. In this embodiment, the camera 30 is used to capture images at any position within a space (real space), such as a room in a real estate property such as a rental house or a house for sale, or a room in an accommodation facility such as a hotel or a guesthouse. The space captured by the camera 30 may be a museum, art gallery, theme park, amusement facility, tourist spot, commercial facility, etc., but in this embodiment, the imaging space is a closed, approximately polygonal prism-shaped room. The camera 30 is a spherical camera that captures images in all directions with a single shutter.
[0012] In this embodiment, for example, an employee of a real estate property management company (a person who registers the panoramic image) uses a camera 30 to take a photograph of the interior (space) of a real estate property. The photographed image (panoramic image) taken by the camera 30 is transmitted to the server 10 via the registration terminal 20 and recorded in the photographed image DB 12a (see FIG. 3) of the server 10. Note that although the camera 30 acquires image data (photographed image data) by photographing, the image data may be simply referred to as an image hereinafter. Therefore, a photographed image may refer to photographed image data, and a panoramic image may refer to panoramic image data. A photographed image includes multiple pixels, and the photographed image data has position information indicating the position (coordinates) of each pixel in the photographed image associated with the brightness (pixel value) of each pixel.
[0013] Server 10 makes the panoramic images registered in photographed image DB 12a publicly available via network N. Users (viewers) who wish to view rooms in a real estate property can access server 10 using viewing terminal 40 to view the panoramic image of the desired room. Server 10 generates a viewing image to be played on viewing terminal 40 from the panoramic image registered in photographed image DB 12a and provides the viewing image to viewing terminal 40. Viewing terminal 40 is a general-purpose information processing device such as a smartphone, tablet terminal, or personal computer, and multiple viewing terminals may be provided. Viewing terminal 40 receives input of an arbitrary position (virtual position) in the photographed space and the direction of the room from the virtual position via input unit 44 (see FIG. 3 ), for example, and generates and displays a 3D model (virtual image) of the room as viewed from the received virtual position in the received direction (the line of sight from the virtual position). Viewing terminal 40 may also be an HMD (Head Mounted Display) type information processing device worn by the viewer. In this case, viewing terminal 40 has various built-in sensors that detect the head position and orientation (gaze direction) of the viewer wearing it, and generates a 3D model based on the detected head position and orientation of the viewer. Viewing terminal 40 may also be a combination of a general-purpose information processing device and a display device such as an HMD, or a combination of a lightweight portable information processing device such as a smartphone and a fixing member that fixes the portable information processing device in front of the viewer's eyes.
[0014] With this configuration, the viewer can virtually walk around the room that has been photographed, face in any direction, and intuitively grasp the size of the room, the arrangement of furniture, the view outside the window, etc. The viewer may also move the virtual viewpoint (virtual position) and the direction from the virtual viewpoint by operating a joystick or the like while seated in a swivel chair. In this case, the viewer can virtually walk around or fly around the room, face in any direction, and grasp the indoor space without the risk of falling or colliding.
[0015] FIG. 2 is a block diagram showing an example configuration of the camera 30 and the registration terminal 20. The camera 30 includes a control unit 31, a memory unit 32, a communication unit 33, a shutter button 34, a photographing unit 35, and other components, all of which are interconnected via a bus. The control unit 31 includes one or more processors, such as a central processing unit (CPU), a microprocessing unit (MPU), or a graphics processing unit (GPU). The control unit 31 executes a control program stored in the memory unit 32 and controls the operation of each hardware component constituting the camera 30 via the bus. This allows the control unit 31 to execute various control processes and information processing operations to be performed by the camera 30. The memory unit 32 includes a random access memory (RAM), a flash memory, a hard disk, and other components. The memory unit 32 pre-stores the control program executed by the control unit 31 and various data required for executing the control program. The memory unit 32 also temporarily stores data generated when the control unit 31 executes the control program.
[0016] The communication unit 33 is an interface for wireless communication such as Bluetooth (registered trademark) or Wi-Fi, and directly transmits and receives information to and from an external device (e.g., the registration terminal 20) via wireless communication. The communication unit 33 may also be an interface for transmitting and receiving information to and from an external device via wired communication via a cable. The communication unit 33 may also be an interface for connecting to a network N via wireless or wired communication, in which case it transmits and receives information to and from an external device via the network N.
[0017] The shutter button 34 is a button for receiving an instruction to capture a still image. The control unit 31 may receive the operation of the shutter button 34 via wireless communication or through the network N. The photographing unit 35 is an imaging device having an optical system for capturing images, an image sensor, etc. The control unit 31 causes the photographing unit 35 to capture an image based on the instruction received from the shutter button 34, and performs various image processing on data obtained by photoelectrically converting light incident via the optical system in the image sensor, thereby acquiring the captured image. In addition to the above-described configuration, the camera 30 may also have a display unit such as a liquid crystal display, an input unit for receiving various operational inputs from the user, etc.
[0018] The registration terminal 20 includes a control unit 21, a memory unit 22, a communication unit 23, an input unit 24, a display unit 25, a reading unit 26, and other components, all of which are interconnected via a bus. The control unit 21 includes one or more processors, such as a CPU, an MPU, or a GPU. The control unit 21 executes a control program 22P stored in the memory unit 22 and controls the operation of each hardware component constituting the registration terminal 20 via the bus. This allows the control unit 21 to execute various control processes and information processing operations to be performed by the registration terminal 20. The memory unit 22 includes RAM, flash memory, a hard disk, an SSD (Solid State Drive), and other components. The memory unit 22 pre-stores the control program 22P (program product) executed by the control unit 21 and various data required for executing the control program 22P. The memory unit 22 also temporarily stores data generated when the control unit 21 executes the control program 22P. The memory unit 22 also stores a web browser 22a for browsing websites published via the network N.
[0019] The communication unit 23 is an interface for connecting to the network N by wired or wireless communication, and transmits and receives information to and from an external device via the network N. The communication unit 23 may also be configured to have an interface for directly transmitting and receiving information to and from an external device (e.g., the camera 30) by wireless or wired communication. The input unit 24 includes a mouse, a keyboard, etc., and receives operation input by a user (registered person), and sends a control signal corresponding to the operation content to the control unit 21. The display unit 25 is a liquid crystal display, an organic EL display, etc., and displays various information in accordance with instructions from the control unit 21. The input unit 24 and the display unit 25 may be a touch panel configured as an integrated unit.
[0020] The reading unit 26 reads information stored in a portable storage medium 2a, which may include a CD (Compact Disc)-ROM, a DVD (Digital Versatile Disc)-ROM, a USB (Universal Serial Bus) memory, an SD (Secure Digital) card, etc. The control program 22P and data stored in the storage unit 22 may be read by the control unit 21 from the portable storage medium 2a via the reading unit 26 and stored in the storage unit 22. The control program 22P and data stored in the storage unit 22 may also be downloaded by the control unit 21 from an external device via the communication unit 23 and the network N and stored in the storage unit 22. Furthermore, the control unit 21 may read the control program 22P and data from the semiconductor memory 2b.
[0021] FIG. 3 is a block diagram showing an example configuration of the server 10 and the viewing terminal 40. The server 10 includes a control unit 11, a storage unit 12, a communication unit 13, an input unit 14, a display unit 15, a reading unit 16, and the like, which are interconnected via a bus. These units 11 to 16 have the same configuration as the units 21 to 26 of the registration terminal 20, and therefore detailed description of their configurations will be omitted. The storage unit 12 of the server 10 stores a control program 12P (program product) as well as a captured image DB 12a and a browsing image DB 12b, which will be described later. The captured image DB 12a and the browsing image DB 12b may be stored in an external storage device connected to the server 10, or may be stored in an external storage device with which the server 10 can communicate via the network N. The input unit 14 and the display unit 15 of the server 10 are not essential, and the server 10 may be configured to accept operations via a connected computer or to output information to be displayed to an external display device.
[0022] Browsing terminal 40 includes control unit 41, memory unit 42, communication unit 43, input unit 44, display unit 45, etc., which are interconnected via a bus. These units 41 to 45 have the same configuration as units 21 to 25 of registration terminal 20, and therefore detailed description of their configurations will be omitted. If viewing terminal 40 is an HMD-type information processing device, it has a sensor unit (not shown) that detects the position and orientation of the head of a user wearing viewing terminal 40. The sensor unit is a combination of multiple sensors, such as a geomagnetic sensor, a tilt sensor, and a GPS (Global Positioning System) sensor, and detects the position and orientation of viewing terminal 40. For example, a GPS sensor detects the current position of viewing terminal 40, a geomagnetic sensor detects the orientation of viewing terminal 40, and a tilt sensor detects the tilt of viewing terminal 40 with respect to the direction of gravity (i.e., vertically downward). By detecting the position and orientation of viewing terminal 40, the sensor unit detects the position and orientation of the head of a user wearing viewing terminal 40.
[0023] FIG. 4 is a schematic diagram showing an example of the configuration of DBs 12a and 12b stored in the server 10, with FIG. 4A showing the captured image DB 12a and FIG. 4B showing the viewable image DB 12b. The captured image DB 12a stores panoramic images (photographed images) captured using the camera 30. The captured image DB 12a shown in FIG. 4A includes an image ID column, a shooting location information column, a captured image column, etc., and stores information on the shooting location and the captured image (panoramic image) in association with the image ID. The image ID column stores identification information assigned to the captured image. The shooting location information column stores information on the location, room, and space where the captured image was captured, such as the address of the room where the image was captured, the building name and room number, the landmark name, the latitude and longitude, etc. The captured image column stores data (photographed image data) of the panoramic image captured using the camera 30. Note that the data of the panoramic image (photographed image data) may be stored in a predetermined area of the storage unit 12 or in an external storage device connected to the server 10, rather than being stored in the photographed image DB 12a. In this case, the photographed image sequence stores information for reading out the data of the photographed image (panoramic image) (for example, a file name indicating the storage location of the data). The image ID stored in the photographed image DB 12a is issued and stored by the control unit 11 when the server 10 acquires a new panoramic image via the communication unit 13. Other information stored in the photographing information DB 12a is stored by the control unit 11 when the control unit 11 acquires the other information from the registration terminal 20 via the communication unit 13. Note that photographing position information is transmitted from the registration terminal 20 to the server 10 together with the panoramic image. The stored contents of the photographed image DB 12a are not limited to the example shown in FIG. 4A , and for example, the photographing date and time, information about the photographer, information about the space (room) of the photographed subject, etc. may be stored.
[0024] The image for browsing DB 12b stores panoramic images for browsing (images for browsing) provided by the server 10 via the network N. The images for browsing are data obtained by the server 10 performing predetermined image processing on the photographed images registered in the photographed image DB 12a, and specifically include polygon data and texture image data. The image for browsing DB 12b shown in FIG. 4B has a configuration similar to that of the photographed image DB 12a, and stores images for browsing generated from the photographed images instead of the photographed images in the photographed image DB 12a. The data of the images for browsing may also be stored in a predetermined area of the storage unit 12 or in an external storage device connected to the server 10, rather than in the image for browsing DB 12b. In this case, the image for browsing sequence stores information for reading the data of the images for browsing (polygon data and texture image data) (e.g., a file name indicating the storage location of the data). When the server 10 generates an image for browsing from a photographed image, the image ID and the photographing location information corresponding to the photographed image are read and stored by the control unit 11 from the photographed image DB 12a by the control unit 11. The images for viewing stored in the image for viewing DB 12b are stored by the control unit 11 when the control unit 11 generates the images for viewing from the captured images. The stored contents of the image for viewing DB 12b are not limited to the example shown in Fig. 4B, and may store, for example, the date and time of the captured image before processing, information about the photographer, information about the space (room) of the subject, etc. Furthermore, the captured image DB 12a and the image for viewing DB 12b may be configured as a single DB, or the images for viewing may be stored in the captured image DB 12a.
[0025] The following describes the camera 30 in the image processing system 100 of this embodiment and the panoramic image captured by the camera 30. FIG. 5 is a schematic diagram showing an example of the appearance of the camera 30. The camera 30 of this embodiment has a plate-shaped portion 38 that is a substantially rectangular plate. As shown in FIG. 5, the camera 30 of this embodiment captures images with the long side of the plate-shaped portion 38 facing up and down, and therefore, in the following description, the long side of the plate-shaped portion 38 will be referred to as the up and down direction of the camera 30. Furthermore, in the camera 30 with the long side of the plate-shaped portion 38 facing up and down, the first wide surface 381 of the plate-shaped portion 38 will be referred to as the front side (front face side) of the camera 30, the second wide surface 382 will be referred to as the rear side (rear face side) of the camera 30, and the short side of the plate-shaped portion 38 will be referred to as the left and right direction of the camera 30.
[0026] The camera 30 has a dome-shaped first lens 371 provided near the top of the first wide surface 381. A dome-shaped second lens 372 provided near the top of the second wide surface 382. A plurality of optical components, such as lenses and prisms (not shown), are arranged inside the first lens 371 and the second lens 372 to form an optical system with a field of view of 180 degrees or more. In the following description, the optical axis of the optical system on the first lens 371 side will be referred to as the first optical axis, and the optical axis of the optical system on the second lens 372 side will be referred to as the second optical axis. The first optical axis and the second optical axis are arranged on the same straight line. A shutter button 34 is arranged on the first wide surface 381.
[0027] The following describes processing performed by camera 30 configured as described above when capturing an image. Fig. 6 is an explanatory diagram of a configuration example of a captured image. Fig. 6A is a schematic diagram showing a celestial sphere image. Fig. 6B is a schematic diagram showing an image obtained by developing the celestial sphere image of Fig. 6A using equirectangular projection. When capturing an image using capturing unit 35 based on an instruction received from shutter button 34, control unit 31 of camera 30 generates a celestial sphere image with the capturing position as the capturing center C as shown in Fig. 6A by combining an image captured on the first optical axis side and an image captured on the second optical axis side.
[0028] The coordinate system that determines the position of each pixel on the celestial sphere image will be described. The position of a pixel on the celestial sphere image can be expressed using a Cartesian coordinate system and a polar coordinate system. The Cartesian coordinate system uses a right-handed Cartesian coordinate system in which the right direction shown in FIG. 5 is the positive side of the X axis, the upward direction is the positive side of the Y axis, and the backward direction is the positive side of the Z axis. The polar coordinate system uses a coordinate system in which coordinates are represented by an azimuth angle φ, a zenith angle θ, and a distance r from the imaging center C. The azimuth angle in the polar coordinate system indicates the counterclockwise angle φ in the XZ plane (a plane including the front, rear, left, and right directions of the camera 30) with the second optical axis side (positive side of the Z axis) as the reference. The first optical axis (negative side of the Z axis) is located at a position where the azimuth angle φ is π radians. Here, π means the ratio of the circumference of a circle to its circumference. The second optical axis is located at a position where the azimuth angle φ is 0 radians and 2π radians. The zenith angle in the polar coordinate system indicates the angle θ pointing downward from the center of the image capture C as the center, with the upward direction (positive side of the Y axis) from the center of the image capture C to the zenith as the reference. The upward direction is at a position where the zenith angle θ is 0 radians. The downward direction (negative side of the Y axis) is at a position where the zenith angle θ is π radians. The position of any pixel a on the omnidirectional image can be expressed as coordinates (x a ,y a ,z a ) and the coordinates (r,θ a ,φ a The coordinates (x, y, z) of the Cartesian coordinate system and the coordinates (r, θ, φ) of the polar coordinate system can be converted to each other using the following equations (1) and (2).
[0029]
number
[0030] The control unit 31 of the camera 30 cuts the celestial sphere image along the line where the azimuth angle φ=0, and generates a rectangular captured image (hereinafter referred to as a panoramic image) that is developed on a plane with the azimuth angle φ on the horizontal axis and the zenith angle θ on the vertical axis using equirectangular projection, as shown in FIG. 6B. The luminance (pixel value) of each pixel in the celestial sphere image is assigned to each pixel that constitutes this panoramic image. As shown in FIG. 6B, the rectangular (planar) panoramic image has the origin (0,0) at the bottom left, and the origin (0,0) indicates a pixel at a position where the azimuth angle φ is 2π radians and the zenith angle θ is π radians. In addition, in the panoramic image, the azimuth angle φ decreases to the right, so that the right edge of the panoramic image is a pixel at a position where the azimuth angle φ is 0 radians, and the zenith angle θ decreases to the upward direction, so that the top edge of the panoramic image is a pixel at a position where the zenith angle θ is 0 radians.
[0031] For example, a registrant uses camera 30 to take an image within the space of the target to be photographed, and imports the acquired panoramic image into registration terminal 20. At this time, control unit 31 of camera 30 transmits a panoramic image as shown in FIG. 6B from communication unit 33 to registration terminal 20. Note that control unit 31 of camera 30 may transmit a captured image of a celestial sphere image as shown in FIG. 6A to registration terminal 20, instead of the panoramic image shown in FIG. 6B. In this case, registration terminal 20 may perform a process of generating the panoramic image shown in FIG. 6B from the celestial sphere image. Furthermore, the process of generating the panoramic image from the celestial sphere image may be performed by server 10 that stores the panoramic image. In this case, the celestial sphere image photographed by camera 30 is transmitted by registration terminal 20 to server 10, and server 10 performs a process of generating the panoramic image from the celestial sphere image. The camera 30 may transmit position information of the photographing center C (photographing position information) together with the panoramic image to the registration terminal 20. For example, by using the Exif (Exchangeable image file format), it is possible to transmit the photographing position information and the photographed image in a single file. The photographing position information is acquired, for example, using a GPS sensor built into the camera 30.
[0032] The following describes the processing performed by each device in the image processing system 100 of this embodiment. FIG. 7 is a flowchart illustrating an example of the registration processing procedure for a captured image, and FIG. 8 is a schematic diagram illustrating an example of a registration screen. In FIG. 7, the left side shows the processing performed by the registration terminal 20, and the right side shows the processing performed by the server 10. The registrant first takes a picture of the interior of the room to be photographed using the camera 30 to acquire a panoramic image (photographed image). At this time, the camera 30 may acquire information about the shooting position along with the photographed image. The control unit 31 of the camera 30 performs shooting using the photographing unit 35 based on an instruction from the shutter button 34, generates a celestial sphere image as shown in FIG. 6A, and further generates a planar panoramic image as shown in FIG. 6B from the celestial sphere image. The control unit 31 transmits the acquired photographed image to the registration terminal 20 in response to a request from the registration terminal 20, for example. Note that if information about the shooting position is measured by a sensor built into the camera 30, the control unit 31 also transmits information about the measured shooting position together with the photographed image to the registration terminal 20.
[0033] The control unit 21 of the registration terminal 20 acquires the photographed image captured by the camera 30 (S11). If the camera 30 also acquires information about the photographing location, the control unit 21 acquires the information about the photographing location along with the photographed image. When registering a photographed image, the registrant uses the registration terminal 20 to access a website provided by the server 10 and requests the server 10 to register the photographed image via the website. The server 10 has a function as a web server, and makes available a website via the network N for accepting instructions to register photographed images. The control unit 21 of the registration terminal 20 accesses the server 10 and requests the server 10 to register the photographed image (S12).
[0034] When a request to register a photographed image is received, the control unit 11 of the server 10 transmits a registration screen as shown in FIG. 8 to the registration terminal 20 (S13). The control unit 21 of the registration terminal 20 receives the registration screen transmitted by the server 10 and displays the received registration screen on the display unit 25 (S14). The registration screen shown in FIG. 8 has an input field 25a for specifying a photographed image (panoramic image) to be registered, and input fields 25b and 25c for inputting information about the photographing location (photography location). Note that input field 25b includes an input field for inputting the address, building name, and room number of the room where the photograph was taken, and input field 25c includes an input field for inputting the latitude and longitude of the photographing location. When a file name is input in input field 25a to specify a photographed image to be registered, the specified photographed image is displayed on the registration screen. When the specified photographed image includes information about the photographing location (latitude and longitude), the latitude and longitude of the photographing location are displayed in input field 25c. Note that if information about the shooting location is not added to the captured image received by the registration terminal 20 from the camera 30, the information about the shooting location (latitude and longitude) is not displayed in the input field 25c. In this case, the input field 25c may be left blank, or, for example, the registrant may input the latitude and longitude of the shooting location via the input unit 24. On the registration screen, the registrant inputs the address, building name, and room number of the shooting location into the input field 25b via the input unit 24. The registration screen has a registration button for instructing the registration of each piece of information input via the screen. The configuration of the registration screen is not limited to the example shown in FIG. 8, and the information about the shooting location may be configured so that landmark names, store names, etc. are input in addition to the address, etc.
[0035] The control unit 21 accepts the designation of a captured image in the input field 25a on the registration screen, and accepts input of information about the shooting location (such as an address) in the input field 25b (S15). The control unit 21 determines whether an instruction to register the information input via the registration screen has been accepted, depending on whether a registration button on the registration screen has been operated (S16). If the control unit 21 determines that a registration instruction has not been accepted (S16: NO), it returns to the processing of step S15 and continues accepting each piece of information via the registration screen. If the control unit 21 determines that a registration instruction has been accepted (S16: YES), that is, if the registration button on the registration screen has been operated, the control unit 21 associates the captured image and the information about the shooting location accepted via the registration screen and transmits them to the server 10 (S17).
[0036] The control unit 11 (image acquisition unit) of the server 10 receives (acquires) the photographed image and the information on the photographing position transmitted by the registration terminal 20, and registers the received information in association with it (S18). Specifically, the control unit 11 issues an image ID to the received photographed image, and stores the received information on the photographing position and the photographed image in the photographed image DB 12a in association with the image ID. Through the above-described processing, a panoramic image (photographed image) taken in the room to be photographed is registered in the server 10 together with the information on the photographing position. Note that photographed images may be taken from multiple different photographing positions in one room, and each photographed image is registered in the server 10 by performing the above-described processing for each photographed image.
[0037] Next, a process for generating an image for viewing (image for viewing) from a photographed image registered in the photographed image DB 12a of the server 10 by the above-described process will be described. FIG. 9 is a flowchart showing an example of the process procedure for generating an image for viewing, and FIGS. 10 and 11 are explanatory diagrams of the process for generating an image for viewing. The server 10 executes the following process using the control unit 11 in accordance with a control program 12P stored in the storage unit 12, but part of the process may be implemented by a dedicated hardware circuit. For example, when receiving a photographed image from the registration terminal 20, the control unit 11 of the server 10 registers the received photographed image in the photographed image DB 12a and performs the following process. The control unit 11 may also perform the following process periodically or after a certain number of photographed images have been registered in the photographed image DB 12a.
[0038] When the control unit 11 of the server 10 receives an instruction to execute a process for generating an image for browsing, for example, through a user operation via the input unit 14, the control unit 11 displays a screen for generating an image for browsing on the display unit 15 (S21). The screen shown in FIG. 10A has an input field 15a for specifying a panoramic image for which an image for browsing is to be generated, an input field 15b for inputting the height from the floor to the ceiling of the room to be photographed for the panoramic image, and an input field 15c for inputting the number of corners (number of sides) of a polygon in a planar view of the room to be photographed. Each of the input fields 15a-15c may be configured to allow input of any letters and numbers, or may have a pull-down menu from which an arbitrary one of multiple options can be selected. The screen shown in FIG. 10A also has an OK button for instructing execution of a process for generating an image for browsing based on the input information.
[0039] When a file name is entered in the input field 15a on the generation screen, the control unit 11 accepts the designation of the captured image to be processed (S22). When the designation of the captured image to be processed is accepted, the control unit 11 reads the designated captured image from the captured image DB 12a and displays the read captured image on the currently displayed generation screen as shown in FIG. 10A (S23). In addition, the control unit 11 (spatial information acquisition unit) accepts input of the height of the room to be photographed and the number of corners of the polygon in plan view via input fields 15b and 15c on the screen, and displays the input height and number of corners in the respective input fields 15b and 15c (S24). Below, a process for generating an image for viewing will be described, using a panoramic image taken in a room that is a roughly octagonal prism, which is a convex polygon in plan view, as the processing target. Therefore, "8" is entered in the input field 15c on the screen shown in FIG. 10A.
[0040] Note that the execution of the process of generating an image for browsing by the server 10 may be instructed, for example, via the registration terminal 20. In this case, the server 10 may transmit a screen for generating an image for browsing to the registration terminal 20, display it on the display unit 25 of the registration terminal 20, and accept input of information necessary for the process of generating an image for browsing via the input unit 24 of the registration terminal 20.
[0041] The control unit 11 determines whether the OK button on the screen shown in FIG. 10A has been operated (S25). If it determines that the OK button has not been operated (S25: NO), the process returns to step S22 and continues accepting information via the generation screen. If it determines that the OK button has been operated (S25: YES), the control unit 11 executes a process for generating an image for viewing based on the information accepted via the generation screen. First, the control unit 11 (line segment display unit) displays the captured image (panoramic image) read from the captured image DB 12a in step S23, and displays, superimposed on the captured image, line segments extending in the vertical direction of the captured image, the number of which is the number of angles (eight in this case) input via the generation screen (S26). FIG. 10B shows an example of a captured image in which eight line segments extending in the vertical direction of the captured image are displayed. In the example shown in FIG. 10B, the eight line segments are displayed at equal intervals in the horizontal direction of the captured image, and the upper and lower ends of each line segment are connected by line segments extending in the horizontal direction of the captured image. It is sufficient that only the line segments extending in the vertical direction of the captured image are displayed, and each line segment may be displayed anywhere in the captured image, and may even be displayed in an area on the screen other than the display area of the captured image.
[0042] The control unit 11 (receiving unit) receives an instruction to move the eight line segments displayed on the captured image to eight line segments (vertical lines) that exist in the vertical direction in the real space at the eight corners of the captured space (S27). The control unit 11 then moves the eight line segments displayed on the captured image in accordance with the received movement instruction (S28). FIG. 11A shows a state after the eight line segments in the vertical direction shown in FIG. 10B have been moved to eight vertical lines L0 to L8 in the captured image. The user (registrant) of the registration terminal 20 or the user of the server 10 issues an instruction to move the end points (upper and lower ends) of the eight line segments displayed on the screen shown in FIG. 10B to the end points of the eight vertical lines L0 to L7 in the captured image displayed on the screen shown in FIG. 11A by a predetermined operation using the input units 24 and 14 (e.g., a drag operation with a mouse). The upper end point of each of the vertical lines L0 to L7 is the position where the vertical line L0 to L7 touches the ceiling, and the lower end point of each of the vertical lines L0 to L7 is the position where the vertical line L0 to L7 touches the floor. As a result, eight vertical lines L0 to L7 corresponding to the eight corners of the shooting space are set in the shot image, as shown in FIG. 11A.
[0043] The control unit 11 determines whether the OK button on the screen shown in Fig. 11A has been operated (S29), and if it determines that the OK button has not been operated (S29: NO), the process returns to step S27 and continues to accept instructions to move the line segments displayed on the captured image. If it determines that the OK button has been operated (S29: YES), the control unit 11 acquires the eight vertical lines L0 to L7 after movement (S30). Specifically, the control unit 11 acquires the coordinate values of both end points (upper and lower end points) of each of the eight vertical lines L0 to L7, i.e., the coordinate values of a total of 16 vertices.
[0044] As shown in FIG. 11B, the position of each pixel in the captured image is expressed as coordinates (x I ,y I ) where the height (length in the vertical direction) of the captured image is H and the width (length in the horizontal direction) is W. Therefore, the coordinate values of both end points (vertices) of the vertical lines L0 to L7 are expressed as coordinate values (x I,y I ) As a result, the control unit 11 can identify eight line segments (vertical lines) that exist in the vertical direction at the eight corners of the photographed space (real space) on the photographed image. The control unit 11 stores the coordinate values of the acquired 16 vertices in the storage unit 12.
[0045] The control unit 11 may automatically identify eight vertical lines (eight vertical lines) that exist at the corners of the photographed space in the vertical direction by image recognition processing based on the photographed image read out from the photographed image DB 12a. For example, feature amounts of line segments (vertical lines) in the photographed image that correspond to vertical lines that extend in the vertical direction at the corners of the photographed space (real space) may be extracted in advance and stored in the storage unit 12, and the control unit 11 may identify the vertical lines in the photographed image by automatically extracting locations having the stored feature amounts from the photographed image.
[0046] Furthermore, when a captured image (panoramic image) is input, the control unit 11 may identify a vertical line in the captured image using a trained model that has been trained to classify each pixel in the captured image into a vertical line region (a line segment corresponding to a vertical line in the captured space) or another region. The trained model here can be configured using a segmentation DNN (Deep Neural Network) such as a SegNet model, a Fully Convolutional Network (FCN) model, or a U-Net model. The segmentation DNN trains using training data that includes an input image (captured image) and a labeled image in which information (class label) indicating whether each pixel in the input image is a vertical line is associated with the pixel. When an input image included in the training data is input, the segmentation DNN trains to output the labeled image included in the training data. In the training process, the segmentation DNN optimizes data such as coefficients and thresholds of various functions that define predetermined operations to be performed on input values using steepest descent, backpropagation, etc. This results in a segmentation DNN that has been trained to classify each pixel in a captured image (panoramic image) into a vertical line region or another region when the image is input.
[0047] When the control unit 11 automatically identifies the vertical lines in the photographed image based on the photographed image, the control unit 11 does not need to display the photographed image read out from the photographed image DB 12a on the display units 15 and 25, and does not need to display the eight line segments as shown in Fig. 10B. Therefore, in this case, the control unit 11 skips the processes of steps S26 to S29, and identifies and acquires the eight vertical lines in the photographed image based on the photographed image in step S30.
[0048] Next, the control unit 11 (correction unit) performs zenith correction on the captured image read out in step S23 based on the eight vertical lines acquired in step S30 (S31). Fig. 12 is a flowchart showing an example of the processing procedure for zenith correction, Figs. 13 and 14 are flowcharts showing an example of the processing procedure for specifying the rotation axis included in the zenith correction, Fig. 15 is a flowchart showing an example of the processing procedure for specifying the rotation amount included in the zenith correction, and Fig. 16 is an explanatory diagram of the zenith correction.
[0049] Here, the zenith correction process will be described. FIG. 16A shows an example of a captured image before zenith correction, and FIG. 16B shows an example of a captured image after zenith correction. When the up-down direction of the camera 30 (the long side direction of the plate-like portion 38) is inclined with respect to the vertical direction in real space, the first optical axis and the second optical axis of the camera 30 are inclined with respect to the horizontal direction in real space. When capturing an image in this state, a plane including the front, back, left, and right directions of the camera 30 (hereinafter referred to as the capture plane) is inclined with respect to the horizontal plane in real space. Therefore, an object on the horizontal plane including the capturing position in real space is not captured in the vertical center of the captured image, resulting in a distorted panoramic image. Therefore, zenith correction is performed to correct the distorted panoramic image shown in FIG. 16A to a panoramic image shown in FIG. 16B, in which an object on the horizontal plane including the capturing position in real space is captured in the vertical center of the captured image. Note that, as can be seen from FIG. 16B, in the panoramic image after zenith correction, eight vertical lines L0 to L7 are arranged in the vertical direction of the image.
[0050] In the zenith correction process, the control unit 11 first acquires the coordinates of the 16 vertices, which are the endpoints of the eight vertical lines acquired in step S30 (S41). Specifically, the control unit 11 reads out the coordinate values of the 16 vertices stored in the storage unit 12. The control unit 11 then calculates parameters representing each of the eight vertical lines based on the coordinate values of the 16 vertices that have been read out (S42). Here, the control unit 11 calculates parameters representing each vertical line by the distance ρ from the origin of the image (for example, the lower left corner) and the slope θ. Therefore, the control unit 11 calculates the parameters (θ, ρ) representing each of the eight vertical lines. Note that the slope θ is set to a value in the range of 0 to π radians, for example, and the distance ρ is set to a value in the range of 0 to L (L is the length of the diagonal of the image).
[0051] Next, the control unit 11 specifies the amount of correction to correct the panoramic image shown in FIG. 16A to the panoramic image shown in FIG. 16B based on the eight vertical lines. Here, the control unit 11 specifies the amount of correction in zenith correction to be performed so that each of the vertical lines in the captured image is positioned in the up-down direction in the corrected captured image (panoramic image) as shown in FIG. 16B. Specifically, the control unit 11 specifies the central axis (rotation axis) and rotation angle (rotation amount) of the shooting plane, which includes the front-back and left-right directions of the camera 30 at the time of shooting, inclined with respect to the horizontal plane in real space. That is, the control unit 11 specifies the central axis and rotation angle about which the shooting plane rotates with respect to the horizontal plane in real space as the amount of correction in zenith correction.
[0052] As can be seen from FIG. 11B, the slope of each vertical line can be expressed as an increasing / decreasing function with one period equal to the horizontal image width of the panoramic image. The position where the slope is greatest on the center line passing through the vertical center of the image is the center of rotation (rotation axis), and this slope can be used to represent the amount of rotation (rotation angle). Therefore, we consider approximating the slope of each vertical line with a sine wave. Because the period of this sine wave matches the image width of the panoramic image, the sine wave can be specified using two parameters: phase φ and amplitude A. Furthermore, the slope of each vertical line was calculated in step S42 using a parameter θ (0≦θ≦π) that represents each vertical line. Therefore, the sine wave to be estimated can be expressed as y = A·sin(x+φ). Based on each vertical line, a sine wave is estimated using the least-squares method with a voting method for each combination of phase φ and amplitude A (first estimation process). Since θ=0 and θ=π both indicate the vertical direction, for subsequent calculations, if θ>π / 2, then θ:=θ-π, and the range is -π / 2≦θ≦π / 2.
[0053] In specifying the correction amount, the control unit 11 first identifies the position of the central axis (rotation axis) along which the photographing plane is tilted with respect to the horizontal plane in real space (S43). Specifically, the control unit 11 identifies the position of the rotation axis along which the photographing plane should be rotated so that each vertical line is aligned in the vertical direction of the image in the corrected panoramic image, based on eight vertical lines, so that the plane coincides with the horizontal plane in real space. In the process of specifying the rotation axis shown in FIGS. 13 and 14, the control unit 11 first sets a two-dimensional array (voting space) with two parameters, amplitude A and phase φ, representing sine waves approximating the slope of each vertical line as each axis, and initializes all elements in the array to 0 (S51). Specifically, the control unit 11 defines the two-dimensional array as the voting space, for example, in the storage unit 12, assigns amplitude A and phase φ to each axis of the array, and initializes all elements in the array (A, φ) represented by the amplitude A and phase φ to 0. The number of divisions S of the array is, for example, 1000 divisions for the range of 0 to π / 2 for the amplitude A, and W divisions (in units of 1 pixel) for the range of the image width of the panoramic image from 0 to W (W is the horizontal length of the image) for the phase φ. Each element of the array (A, φ) accumulates and stores the square of the error (difference) of the slope θ of each vertical line with respect to y = A sin(x + φ).
[0054] Next, the control unit 11 sets a candidate value for the amplitude A of the sine wave to be estimated to 0 (S52), and sets a candidate value for the phase φ to 0 (S53). The control unit 11 acquires one of the vertical lines whose parameters were calculated in step S42 (S54), and calculates the error between the value based on the sine wave represented by the set amplitude A and phase φ for the acquired vertical line and the slope θ (parameter) calculated in step S42 (S55). C (X coordinate value of the image coordinate system) is the X coordinate value of the position where each vertical line intersects with the center line of the panoramic image, and the horizontal position x of each vertical line is C Since is known, to normalize it to 0 to 2π, we use x=2π×x C In addition, the i-th phase in the φ direction (φ axis) in the voting space is expressed as 2π×i / S (where S is the number of divisions, W in this case). From these, the horizontal position x of the vertical line corresponding to the element of the array (A, φ) is C The value of the sine wave at is y=A·sin(2π×x C / W-2π×i / S) Therefore, the control unit 11 calculates the difference (error) between the value y of the sine wave calculated in this way and the gradient θ already calculated for this vertical line.
[0055] The control unit 11 calculates the square of the calculated difference (error) (y-θ) 2 The control unit 11 calculates the error, and adds the square of the calculated error to the element of the array (A, φ) (S56). The control unit 11 determines whether the processing of steps S55 to S56 has been completed for all vertical lines for which parameters were calculated in step S42 (S57). If it determines that the processing has not been completed (S57: NO), the control unit 11 returns to the processing of step S54, acquires one unprocessed vertical line (S54), and repeats the processing of steps S55 to S56 for the acquired vertical line. As a result, for one element of the array (A, φ), the square of the error between the function value of each vertical line based on a sine wave based on the phase φ and amplitude A and the calculated slope of each vertical line is accumulated.
[0056] If control unit 11 determines that the processing of steps S55 to S56 has been completed for all vertical lines calculated in step S42 (S57: YES), it adds W / S (W is the horizontal width of the image, and S is the number of divisions into the array) to the phase φ at this point in time (S58), and determines whether the phase φ after the addition is equal to or greater than W (image width) (S59). If control unit 11 determines that the phase φ after the addition is less than W (S59: NO), it returns to the processing of step S54. Then, control unit 11 repeats the processing of steps S54 to S57 for all vertical lines whose parameters were calculated in step S42 based on the phase φ and amplitude A at this point in time.
[0057] If it is determined that the phase φ after the addition is equal to or greater than W (S59: YES), control unit 11 adds π / 2S (S is the number of divisions of the array) to the amplitude A at this point (S60), determines whether the amplitude A after the addition is equal to or greater than π / 2 (S61), and returns to the processing of step S53 if it is determined that the amplitude A after the addition is less than π / 2 (S61: NO). Then, control unit 11 repeats the processing of steps S53 to S59 based on the amplitude A at this point. In this way, by performing the processing of steps S54 to S57 while increasing the candidate value of phase φ by W / S and the candidate value of amplitude A by π / 2S, control unit 11 accumulates, for each element of array (A, φ), the square of the error between the function value of each vertical line due to a sine wave based on the candidate value of phase φ and the candidate value of amplitude A and the calculated slope of each vertical line.
[0058] If it is determined in step S60 that the amplitude A after the addition is π / 2 or more (S61: YES), that is, if the accumulation of the squares of the errors for all elements of the array (A, φ) for all vertical lines is completed, the control unit 11 identifies the minimum value from the elements of the array (A, φ) and calculates the amplitude A and phase φ corresponding to the element with the minimum value as the formal parameters (A t ,φ t ) (S62). t ,φ t) is the first estimation result. The control unit 11 uses the first estimation result to examine the parameters representing the sine wave by a method called robust estimation (second estimation process). In robust estimation, the influence of each vertical line on the estimation of the sine wave is examined by using the estimated results of the tentative parameters (A t ,φ t ) is expressed using weights based on the difference with the formal parameter (A t ,φ t ) the horizontal position x of each vertical line C The sine wave function value y=A t sin(2π×x C / W-φ t A weight ω for the difference d=y-θ between the vertical line θ (parameter) calculated in step S42 and the difference d=y-θ is set using a predetermined tolerance Ω as shown in the following equation (3). Note that the tolerance Ω can be set to, for example, 1.25 times the median of the differences d among all vertical lines. Then, the control unit 11 performs a second sine wave estimation process using the weight ω determined using the following equation (3).
[0059]
number
[0060] Therefore, the control unit 11 performs the same processes as in steps S51 to S54 in the first estimation process (S63 to S66), and determines a weight for the acquired vertical line according to the difference between the function value based on the sine wave indicated by the formal parameters and the gradient θ calculated in step S42 (S67). The control unit 11 also performs the same process as in step S55 (S68), calculates a value by multiplying the square of the calculated error by the weight determined in step S67, and adds this to the corresponding element of the array (A, φ) (S69). Therefore, in the second estimation process, for each element of the array (A, φ), the square of the error (difference) between the sine wave function y=A·sin(x+φ) and the gradient θ calculated for each vertical line is accumulated, but at this time, it is multiplied by the weight ω expressed by the above equation (3) before being added. As a result, the first estimation result, the formal parameters (A t ,φ tThis reduces the influence of a vertical line that is farther away from the sine wave represented by Ω than the tolerance value Ω on the estimation of the sine wave, making it possible to estimate the sine wave with higher accuracy.
[0061] The control unit 11 determines whether the processing of steps S67 to S69 has been completed for all vertical lines for which parameters were calculated in step S42 (S70). If it is determined that the processing has not been completed (S70: NO), the control unit 11 returns to the processing of step S66, acquires one unprocessed vertical line (S66), and repeats the processing of steps S67 to S69 for the acquired vertical line. If the control unit 11 determines that the processing of steps S67 to S69 has been completed for all vertical lines for which parameters were calculated in step S42 (S70: YES), the control unit 11 performs the same processing as steps S58 to S61 (S71 to S74), increasing the candidate value for phase φ by W / S and the candidate value for amplitude A by π / 2S, while performing the processing of steps S66 to S70. This allows the accumulation of a value obtained by multiplying the square of the error between the function value of each vertical line due to a sine wave based on the candidate values for phase φ and amplitude A and the calculated slope of each vertical line by the weight set for this vertical line. If it is determined in step S73 that the amplitude A after addition is π / 2 or greater (S74: YES), that is, when the accumulation of the values obtained by multiplying the square of the error for all vertical lines by the weight ω for all elements of array (A, φ) is completed, control unit 11 identifies the smallest value among the elements of array (A, φ) and specifies the amplitude A and phase φ corresponding to the element with the smallest value as parameters (A, φ) that represent a sine wave that approximates the slope of each vertical line (S75).
[0062] FIG. 16A shows an example of a sine wave estimated by the above-described process. In FIG. 16A, the dashed line indicates the sine wave estimated by the first estimation process, the solid line indicates the sine wave estimated by the second estimation process, and the dashed-dotted line indicates the center line of the image. The above-described process allows for estimation of a sine wave that approximates the gradient θ of each vertical line. Note that in the second estimation process, highly reliable estimation results can be obtained by reducing the influence of vertical lines that are far from the sine wave represented by the tentative parameters, which are the first estimation results, on the estimation of the sine wave. Note that in the process of identifying the rotation axis, the control unit 11 may use the sine wave obtained by the first estimation process as a sine wave that approximates the gradient θ of each vertical line, without performing robust estimation (second estimation process). In this case, the phase φ, which is a parameter obtained by the first estimation process, may be used as the correction amount (position of the rotation axis) used for zenith correction.
[0063] Here, the phase φ of the sine wave calculated (estimated) as described above indicates the position on the center line of the image where the function value y of the sine wave is 0 (the position where the gradient of the vertical line is 0), so the rotation axis is a position shifted by π / 2 on the center line from the position of the phase φ. This is because the position on the center line where the gradient of the vertical line is maximum becomes the rotation axis. Therefore, the control unit 11 specifies the position of the rotation axis (phase φ) by adding π / 2 to the phase φ of the sine wave calculated (estimated) as described above (φ:=φ+π / 2). Note that up to this point, the left end of the panoramic image has been treated as φ=0 and the right end as φ=2π. However, in the polar coordinate system used in the zenith correction performed in step S45, the left end of the panoramic image becomes φ=2π and the right end becomes φ=0. Therefore, the control unit 11 subtracts the identified phase (φ:=φ+π / 2) from 2π to calculate the phase φ (φ:=3π / 2-φ), and identifies the calculated phase φ as the correction amount (phase φ indicating the position of the rotation axis) to be used for the zenith correction performed in step S45 (S76). The control unit 11 ends the process of identifying the rotation axis, and returns to the process of FIG.
[0064] Next, the control unit 11 specifies the rotation amount (rotation angle) α to be used for the zenith correction performed in step S45 based on the position (phase φ) of the rotation axis specified in step S43 (S44). Specifically, the control unit 11 specifies the rotation amount by which the plane at the time of shooting should be rotated around the specified rotation axis so that each vertical line is positioned in the up-down direction in the corrected panoramic image. Note that the least squares method is also used in estimating the rotation amount α by using a voting method. That is, the control unit 11 calculates the error of each vertical line based on the sine wave specified in step S43, and sets a weight ω for each vertical line according to the above equation (3) depending on the calculated error. Specifically, the horizontal position x of each vertical line in the parameters (A, φ) of the sine wave specified in step S43 is C The function value of a sine wave at y=A·sin(2π×x C A weight ω for each vertical line is set according to the difference d=y-θ between the rotation amount α (α / W-φ) and the slope θ of the vertical line calculated in step S42. The voting process using the least squares method to estimate (identify) the rotation amount α is the same as the voting process used to estimate a sine wave, except that a one-dimensional array is used.
[0065] In the process of specifying the amount of rotation shown in FIG. 15, the control unit 11 sets a one-dimensional array with the amount of rotation α as a parameter and initializes all elements in the array α to 0 (S81). Specifically, the control unit 11 defines a one-dimensional array as a voting space, for example, in the storage unit 12, and initializes all elements in the array α to 0. Note that the number of elements S in the array is set to, for example, 1000, and the amount of rotation within a range of ±30° (±π / 6 radians) is divided into 1000 equal parts. The ±30° range used in the array is set based on the assumption that the tilt occurring during shooting using the camera 30 is at most 30°, but is not limited to this range; for example, a range of ±45° may be used. In the array α in the voting space, the i-th amount of rotation α is expressed as α=(i / 1000)×(π / 3)−(π / 6). Each element of the array stores an accumulated value obtained by multiplying the square of the tilt of each vertical line in the panoramic image after rotation (after zenith correction) when each vertical line is rotated by each rotation amount α around the phase φ specified in step S43 as the central axis by the weight ω previously set for each vertical line. When rotated by an appropriate rotation amount around the axis indicated by the specified phase φ as the central axis, each vertical line should be positioned in the up-down direction on the panoramic image after zenith correction, so the rotation amount α corresponding to the smallest value among the elements of the array can be used as the rotation amount used for zenith correction performed in step S45.
[0066] The control unit 11 sets the candidate value for the rotation amount α to −π / 6 (S82). The control unit 11 acquires one of the vertical lines whose parameters were calculated in step S42 (S83), and for the acquired vertical line, identifies a weight according to the difference between the function value based on the sine wave indicated by the parameters (amplitude A and phase φ) identified by the rotation axis identification process and the slope calculated in step S42 (S84).
[0067] The control unit 11 performs zenith correction (rotation) on the acquired vertical line based on the rotation axis φ identified in step S43 and the rotation amount α set in step S82 (S85), and calculates the tilt of the corrected vertical line (S86). Specifically, the control unit 11 identifies pixels P1 and P2 corresponding to both end points of the vertical line to be processed, and calculates the positions of pixels P1' and P2' to which the end points P1 and P2 have moved in the panoramic image after rotation by the rotation amount α around the phase φ as the rotation axis (after zenith correction). When performing zenith correction (rotation) on the panoramic image based on the phase φ and the rotation amount α, the coordinate values of the end points P1 and P2 in the panoramic image are converted from the image coordinate system to a world Cartesian coordinate system with the shooting position as the origin, and the zenith correction is performed in the world Cartesian coordinate system. When converting from the image coordinate system to the world Cartesian coordinate system, the coordinates are first converted to a polar coordinate system and then converted to the world Cartesian coordinate system.
[0068] Therefore, the control unit 11 calculates the coordinates (x I ,y I ) into coordinates (r, θ, φ) in the polar coordinate system using the following equation (4). Since coordinate values in the polar coordinate system correspond one-to-one to coordinate values in the image coordinate system, they can be converted using the following equation (4), and the radius r in the polar coordinate system is set to 1 for convenience. Then, the control unit 11 converts the coordinates (r, θ, φ) in the polar coordinate system of each of the end points P1 and P2 into coordinates (x W ,y W ,z W )
[0069]
number
[0070] Figure 17 is a schematic diagram showing the relationship between the polar coordinate system and the world orthogonal coordinate system. In the world orthogonal coordinate system, the vertical direction in the real space is the Y W The positive side of the axis, the horizontal plane in real space is X W axis and Z WAs shown in Figure 17, the polar coordinate system is a three-dimensional right-handed Cartesian coordinate system represented by the vertical direction (Y W The zenith angle θ is based on the positive side of the axis, and the Z angle is W The horizontal plane is based on the positive side of the X axis. W Z W It is expressed by the counterclockwise azimuth angle φ in the plane and the distance r from the origin 0. Therefore, any position P in the three-dimensional space can be expressed by the coordinates (x W ,y W ,z W ) and the coordinates (r,θ,φ) of the polar coordinate system. W ,y W ,z W ) can be converted to polar coordinate system coordinates (r, θ, φ) using the following equation (6).
[0071]
number
[0072] Then, the control unit 11 calculates the coordinates (x W ,y W ,z W ) based on the phase φ and the rotation amount α, the control unit 11 first performs zenith correction (rotation) on the panoramic image (end points P1, P2). W The rotation axis based on the phase φ is rotated clockwise around the axis by φ, and the rotation axis is temporarily set to Z in the world Cartesian coordinate system. W axis, then the Z axis of the world Cartesian coordinate system W Then, the control unit 11 rotates the first Y W To return the rotation around the axis, use the Y coordinate system in the world Cartesian coordinate system. WThe image is rotated counterclockwise around the axis by φ, and the coordinate values of the end points (pixels) P1' and P2' in the rotated panoramic image are converted back to the coordinate values of the image coordinate system via the polar coordinate system by coordinate transformation. This allows the positions of the pixels P1' and P2' in the corrected panoramic image to be calculated by zenith correction based on the phase φ and the rotation amount α, to which the end points P1 and P2 in the panoramic image before correction have moved. Note that the Y W Rotation around the axis by an angle θ, and Z W The rotation of the angle θ around the axis is expressed as the following equation (7) using a 3 × 3 matrix. Therefore, the coordinates (x W ,y W ,z W ) and calculate the coordinates after rotation (x W ´,y W ´,z W The coordinates (x W ,y W ,z W ) are converted to the coordinates (r,θ,φ) of the polar coordinate system by the above equation (6), and the coordinates (r,θ,φ) of the polar coordinate system are converted to the coordinates (x I ,y I ) As a result, positions P1' and P2' in the corrected panoramic image to which the end points P1 and P2 in the pre-correction panoramic image move due to zenith correction (rotation) based on the rotation axis φ and the rotation amount α are calculated.
[0073]
number
[0074] The control unit 11 calculates the slope of the line segment P1'P2' based on the coordinate values of pixels P1' and P2' in the panoramic image after zenith correction calculated as described above. The control unit 11 then adds a value obtained by multiplying the square of the calculated slope of the line segment P1'P2' by the weight specified in step S84 to the corresponding element of array α (S87). The control unit 11 determines whether the processing of steps S84 to S87 has been completed for all vertical lines whose parameters were calculated in step S42 (S88). If it is determined that the processing has not been completed (S88: NO), the control unit 11 returns to the processing of step S83, acquires one unprocessed vertical line (S83), and repeats the processing of steps S84 to S87 for the acquired vertical line. As a result, for one element of array α, a value obtained by multiplying the square of the slope of each vertical line after zenith correction based on the rotation axis φ specified in step S43 and the rotation amount α here by the weight set for each vertical line is accumulated. If it is determined that the processing of steps S84 to S87 has been completed for all vertical lines whose parameters were calculated in step S42 (S88: YES), the control unit 11 adds π / 3S (S is the number of divisions of the array) to the rotation amount α at this point (S89), and determines whether the rotation amount α after the addition is π / 6 or more (S90).
[0075] If it is determined that the rotation amount α is less than π / 6 (S90: NO), the control unit 11 returns to the processing of step S83. Then, the control unit 11 repeats the processing of steps S83 to S88 for all vertical lines for which parameters were calculated in step S42 based on the rotation amount α at this time. By performing the processing of steps S83 to S88 while increasing the candidate value of the rotation amount α by π / 3S, the control unit 11 can accumulate, for each element of the array α, a value obtained by multiplying the square of the slope of each vertical line after zenith correction based on each candidate value of the rotation axis φ and rotation amount α identified in step S43 by the weight set for each vertical line. If it is determined in step S89 that the rotation amount α after addition is equal to or greater than π / 6 (S90: YES), that is, if the accumulation of the values obtained by multiplying the square of the corrected slope by the weight ω for all vertical lines for each element of array α is completed, the control unit 11 identifies the minimum value from the elements of array α, and identifies the rotation amount α corresponding to the element with the minimum value as the correction amount (rotation amount α) to be used for the zenith correction performed in step S45 (S91). The control unit 11 ends the process of identifying the rotation amount, and returns to the process of FIG.
[0076] The control unit 11 performs zenith correction on the panoramic image read out in step S23 based on the position (phase φ) of the rotation axis identified in step S43 and the rotation amount α identified in step S44 (S45). Here, the control unit 11 performs the same process as the zenith correction in step S85 in FIG. 15 on the panoramic image. Specifically, the control unit 11 calculates the coordinates (x I ,y I ) is converted to the coordinates (r,θ,φ) of the polar coordinate system by the above equation (4), and then converted to the coordinates (x W ,y W ,z W Then, the control unit 11 converts the coordinates (x W ,y W ,z W ) in the world Cartesian coordinate system W Rotate clockwise around the axis φ, and set the rotation axis based on the phase φ to Z WAfter aligning each pixel with the Z axis of the world Cartesian coordinate system, W The control unit 11 then rotates the camera by a rotation angle α around the axis, and performs zenith correction (rotation) based on the rotation amount α. W To restore the rotation around the axis, we convert each pixel after zenith correction into the Y coordinate of the world Cartesian coordinate system. W The coordinate values of each pixel after zenith correction are then converted back to coordinate values in the image coordinate system via the polar coordinate system by coordinate transformation. This allows the pixel positions in the corrected panoramic image to be calculated, where each pixel in the pre-correction panoramic image has moved due to zenith correction based on the phase φ and the amount of rotation α.
[0077] The control unit 11 then copies the luminance of each pixel in the pre-correction panoramic image to each position in the calculated post-correction panoramic image, thereby generating a post-correction panoramic image. Because a panoramic image is a digital image, the position of a pixel in the post-correction panoramic image may not be identified based on the calculated post-correction coordinate values. Therefore, the luminance of each pixel in the post-correction panoramic image may be calculated based on the calculated luminance of surrounding pixels, for example, by performing an interpolation operation such as a bilinear method. Alternatively, the position of each pixel in the post-correction panoramic image may be calculated from the position of each pixel in the corresponding pre-correction panoramic image, and the positions (coordinate values) of each pixel before and after correction may be associated. In this case, the post-correction panoramic image is obtained by copying the calculated luminance of the pre-correction pixel position to the post-correction pixel position.
[0078] The control unit 11 performs the same zenith correction as in step S45 on the eight vertical lines acquired in step S30 (S46), and calculates the pixel positions in the panoramic image after zenith correction where both end points of each vertical line in the panoramic image before correction have moved due to the zenith correction. The control unit 11 then terminates the zenith correction process and returns to the process of FIG. 9. The above-described zenith correction corrects the distorted panoramic image as shown in FIG. 16A to a panoramic image as shown in FIG. 16B. Furthermore, the eight vertical lines L0 to L7 set as shown in FIG. 11B each extend in the up and down directions in the panoramic image shown in FIG. 16B.
[0079] When zenith correction is performed on the captured image and the vertical line, the control unit 11 (estimation unit) estimates the shape of the captured space, with the height of the vertical line after zenith correction, based on the captured image after zenith correction (S32). Here, the control unit 11 estimates the shape of an octagonal prism. FIG. 18 is a flowchart showing an example of the procedure for estimating the captured space, and FIGS. 19 and 20 are explanatory diagrams of the estimation process of the captured space. In the estimation process of the captured space shown in FIG. 18, the control unit 11 acquires the captured image after zenith correction and the vertical line after zenith correction, which are obtained by the zenith correction process of step S31 (S101). Specifically, the control unit 11 acquires the captured image after zenith correction and the positions in the captured image of both end points of each vertical line after zenith correction (coordinate values in the image coordinate system). Note that this information is stored in the memory unit 12 after the zenith correction is performed.
[0080] FIG. 19A shows a captured image after zenith correction and eight vertical lines L0 to L7. The dashed-dotted line in FIG. 19A indicates the center line of the captured image, and a subject at the same height as the shooting position is captured on the center line. That is, for each of the vertical lines L0 to L7 in the captured image, the length from the center line to the top end (length on the image) corresponds to the distance from the shooting position to the ceiling (distance in real space), and the length from the center line to the bottom end corresponds to the distance from the shooting position to the floor. Therefore, the control unit 11 calculates the height of the shooting position based on the ratio of the length from the center line to the top end of one vertical line to the length from the center line to the bottom end (S102).
[0081] For example, the control unit 11 calculates the height of the photographing position based on the vertical line L0 shown in Fig. 19A. Specifically, the control unit 11 calculates the length L of the vertical line L0 and the length L from the center line to the top end of the vertical line L0 based on the coordinate values of the image coordinate system of the top and bottom ends of the vertical line L0 acquired in step S101. C and the length L from the center line to the bottom of the vertical line L0 F FIG. 19B shows the relationship between the corner of the room (a corner extending in the vertical direction) corresponding to the vertical line L0 in the real space and the shooting position, and the length of the corner is the height H of the room from the floor to the ceiling, and the distance H from the height of the shooting position to the ceiling at the corner is C and the distance H from the shooting position to the floor F The ratio of the length L shown in FIG. C and length L F Therefore, the control unit 11 (photographing position calculation unit) calculates the length L C and length L F Based on the ratio of the height H of the shooting space to the distance from the shooting position to the ceiling and the distance from the shooting position to the floor, the control unit 11 calculates the height of the shooting position. Specifically, the control unit 11 calculates the height H of the shooting space by dividing the height H of the shooting space into the distance from the shooting position to the ceiling and the distance from the shooting position to the floor using the following equation (10): C and the distance H from the shooting position to the floor F Here, the height H is the height of the room of the subject of photography input in step S24 in FIG. 9. The control unit 11 calculates the distance H from the photography position to the ceiling. C and the distance H from the shooting position to the floor F It is sufficient to calculate either one of the following.
[0082]
number
[0083] Next, the control unit 11 calculates the distance in real space from the shooting position to each of the vertical lines L0 to L7. As shown in FIG. 19B, the distance d from the shooting position in real space to the vertical line (the corner of the room) is calculated by multiplying the distance H from the shooting position to the ceiling. Cand the elevation angle θ of the line segment connecting the shooting position and the top of the vertical line based on the horizontal direction including the shooting position. C Based on this, it can be calculated using, for example, the following equation (11): C is calculated from the zenith angle θ indicated on the vertical axis in the panoramic image expressed in equirectangular projection as shown in FIG. 6B. Therefore, the control unit 11 calculates the coordinates (x I ,y I ) is extracted and converted to the coordinates (r, θ, φ) of the polar coordinate system using the above equation (4), to obtain the zenith angle θ of the pixel corresponding to the top end of the vertical line, and θ C Elevation angle θ by calculating =π / 2-θ C The distance d from the shooting position to the vertical line is calculated by multiplying the distance H from the shooting position to the floor. F and the depression angle θ of the line segment connecting the shooting position and the bottom of the vertical line based on the horizontal direction including the shooting position. F It can be calculated similarly based on the above.
[0084]
number
[0085] The control unit 11 acquires the zenith angle θ of the pixel at the top end position or the pixel at the bottom end position as information about one vertical line in the captured image (S103). Specifically, the control unit 11 acquires the coordinates (x I ,y I ) is extracted and converted into coordinates (r, θ, φ) in the polar coordinate system using the above equation (4), and the zenith angle θ of the pixel at the top end position or the pixel at the bottom end position is obtained. Then, the control unit 11 (distance calculation unit) calculates the distance from the shooting position to the vertical line using the above equation (11) based on the calculated zenith angle θ and the height of the shooting position (S104).
[0086] The control unit 11 determines whether or not there are any vertical lines in the captured image for which the distance calculation process described above has not been performed (unprocessed vertical lines) (S105). If it is determined that there are unprocessed vertical lines (S105: YES), the control unit 11 returns to the process of step S103 and repeats the processes of steps S103 to S104 for the unprocessed vertical lines. This makes it possible to calculate the distance d from the shooting position for each vertical line in the captured image. If it is determined that there are no unprocessed vertical lines (S105: NO), the control unit 11 (plane estimation unit) estimates the shape of the shooting space in a planar view (here, an octagonal shape) (S106). The shape of the shooting space in a planar view is determined by the Y x -axis direction, which indicates the vertical direction in real space in a world Cartesian coordinate system with the shooting position as the origin, as shown in FIG. 20A. W X, looking at the shooting space from the positive side of the axis W Z W It can be expressed on a plane. W When each vertical line is viewed from the positive side of the axis, each vertical line can be regarded as a point, and as shown in Fig. 20A, each of the vertical lines L0 to L7 can be represented by vertices P0 to P7. Therefore, the control unit 11 calculates each of the vertices P0 to P7 as X based on the azimuth angles φ0 to φ7 of each vertical line L0 to L7 and the distances d0 to d7 from the shooting position to each vertical line L0 to L7. W Z W The planar shape of the shooting space is estimated by plotting it on a plane. Note that, for each of the vertical lines L0 to L7 in the shot image as shown in Fig. 20B, the azimuth angles φ0 to φ7 with respect to the right edge of the image (the direction behind the camera 30) as the reference point are known.
[0087] The control unit 11 is X W Z W By calculating the coordinate values (coordinate values in the world Cartesian coordinate system) of each vertex P0 to P7 plotted on the plane, the plane shape of the shooting space (X W Z W Specifically, the control unit 11 calculates the coordinate values of each vertex P0 to P7 in the world Cartesian coordinate system based on the azimuth angles φ0 to φ7 of each vertical line L0 to L7 and the distances d0 to d7 from the shooting position to each vertical line L0 to L7, in accordance with the following equation (12). Note that k in equation (12) ranges from 0 to 7.
[0088]
number
[0089] Since the vertical lines L0 to L7 may not be arranged in the vertical direction of the image strictly, the azimuth angles φ0 to φ7 of the vertical lines L0 to L7 may be specified (calculated) from the average value of the X coordinate values (X coordinate values in the image coordinate system) of both end points of the vertical lines L0 to L7. Specifically, the control unit 11 determines the X coordinate values (X coordinate values in the image coordinate system x) of both end points of the vertical lines after the zenith correction. I ) and converts the calculated average value of the X coordinate values into a coordinate value φ in the polar coordinate system using the above equation (4). The control unit 11 calculates the coordinate values φ0 to φ7 in the polar coordinate system for each of the vertical lines L0 to L7, thereby converting the X W Z W In the plane, for vertices P0 to P7 corresponding to each vertical line L0 to L7, Z W Azimuth angles φ0 to φ7 based on the positive direction of the axis can be identified. After estimating the shape of the imaging space, the control unit 11 returns to the processing of FIG.
[0090] Next, the control unit 11 calculates the coordinates (X W ,Y W ,Z W ), polygon data is generated for the ten faces that make up the octagonal prism, specifying the coordinate values of the vertices of each face and the connectivity of the vertices (S33). Note that, since eight of the ten sides that make up the octagonal prism are rectangular, control unit 11 generates polygon data that specifies the coordinate values of the four vertices of each side face and the connectivity of the vertices. Since the top and floor faces are convex octagons, control unit 11 generates polygon data for the top and floor faces that specifies the coordinate values of the four vertices of a circumscribed rectangle that encompasses the octagon and the connectivity of the vertices, as shown in FIG. 21B.
[0091] Fig. 21 is an explanatory diagram of polygon data. Fig. 21A shows an octagonal prism in the world orthogonal coordinate system, and Fig. 21B shows the shooting space in the Y coordinate system of the world orthogonal coordinate system. W The control unit 11 generates the vertex data and polygon data shown in FIG. 21C. First, as shown in FIG. 21A, the control unit 11 (coordinate value acquisition unit) assigns vertex IDs (V0 to V15) to each vertex of the estimated octagonal prism, and calculates the coordinates (X W ,Y W ,Z W) and stores them in association with the vertex IDs. Furthermore, the control unit 11 specifies a circumscribing rectangle that includes the ceiling surface as shown in FIG. 21B, assigns vertex IDs (V16 to V19) to each of the four vertices of the specified circumscribing rectangle, and stores the coordinates of each vertex in the world Cartesian coordinate system in association with the vertex ID. Similarly, the control unit 11 specifies a circumscribing rectangle that includes the floor surface, assigns vertex IDs to each of the four vertices of the specified circumscribing rectangle, and stores the coordinates of each vertex in the world Cartesian coordinate system in association with the vertex ID. Note that in this embodiment, since the ceiling surface and the floor surface are convex-shaped, as shown in FIG. 21B, two of the four vertices of the circumscribing rectangle of the ceiling surface and the floor surface (the two vertices with vertex IDs V18 and V19 for the ceiling surface shown in FIG. 21B) are common to the vertices of the octagonal prism. Therefore, in this embodiment, in addition to the 16 vertices of the octagonal prism, the control unit 11 assigns vertex IDs to two of the four vertices of the circumscribing rectangle of the ceiling surface that are not shared with any vertices of the octagonal prism (the two vertices with vertex IDs V16 and V17 in FIG. 21B) and two of the four vertices of the circumscribing rectangle of the floor surface that are not shared with any vertices of the octagonal prism, and stores the coordinates of each vertex in the world Cartesian coordinate system in association with the vertex ID. Note that the coordinate values of the two vertices in the circumscribing rectangle of the ceiling surface that are not shared with any vertices of the octagonal prism can be calculated from the coordinates of other vertices on the ceiling surface, and the coordinate values of the two vertices in the circumscribing rectangle of the floor surface that are not shared with any vertices of the octagonal prism can be calculated from the coordinates of other vertices on the floor surface. This generates vertex data such as that shown in FIG. 21C. Next, the control unit 11 assigns polygon IDs (Po0 to Po9) to each face of the octagonal prism and stores the vertex IDs (first vertex ID to fourth vertex ID) of the four vertices of each face in association with the polygon ID. In this embodiment, the polygons for the ceiling and floor surfaces are circumscribed rectangles that encompass the octagonal ceiling and floor surfaces, but polygon data may also be generated using octagonal polygons. The control unit 11 stores a texture ID assigned to a texture image (described later) in association with the polygon ID. This generates polygon data such as that shown in FIG. 21C.
[0092] Next, the control unit 11 (texture image generation unit) assigns each pixel of the panoramic image to each face (each polygon) of the octagonal prism to generate a texture image (texture image data) (S34). FIGS. 22 and 23 are explanatory diagrams of texture images. FIG. 22A shows the captured image that has been subjected to zenith correction in step S31. The control unit 11 generates texture images for 10 faces by mapping the captured image (panoramic image) shown in FIG. 22A onto each polygon based on the coordinate values of the vertices of each polygon. Note that in this embodiment, each texture image is square, but a rectangular texture image may also be used. Also, in this embodiment, the eight texture images on the sides are the same size, but the size of the texture image may be different depending on the size of each side of the shooting space.
[0093] First, the control unit 11 acquires the coordinate values of the four vertices of the polygon to be processed in the world Cartesian coordinate system from the vertex data and polygon data. Next, based on the coordinate values of the four vertices, the control unit 11 divides this polygon equally in both the vertical and horizontal directions, setting, for example, 1024 pixels x 1024 pixels, and calculates the coordinate values (world Cartesian coordinate system) of each pixel. Next, the control unit 11 calculates the coordinate values (X W ,Y W ,Z W ) are converted into coordinate values (θ, φ) in the polar coordinate system using the above equation (6), and the coordinate values (θ, φ) in the polar coordinate system are converted into coordinate values (X I ,Y I ) and each coordinate value (X I ,Y I ) to each pixel of the polygon. As a result, a texture image for one polygon is generated, and the generated texture image is stored in storage unit 12 in association with the texture ID stored in the polygon data. For example, if each pixel in the front side surface (polygon) indicated by the thick rectangle in FIG. 22B corresponds to each pixel in the area in the captured image indicated by the thick closed area in FIG. 22A, a texture image such as that shown in FIG. 22C can be generated by copying each pixel in the thick closed area in FIG. 22A.
[0094] Because panoramic images are digital images, it may be impossible to identify the position of a pixel in the panoramic image based on the coordinate values of the image coordinate system calculated as described above. In this case, the brightness of each pixel in the polygon may be calculated based on the pixel values of pixels surrounding the calculated coordinate values, using interpolation calculations such as the bilinear method or nearest neighbor method. Alternatively, the position of each pixel in the polygon may be calculated from the position of each pixel in the captured image, and the pixel value of each pixel in the captured image may be copied to each pixel in the calculated polygon. The control unit 11 performs the above-described processing for all polygons to generate texture images for the 10 faces of the octagonal prism. Figure 23 shows the texture images for the 10 faces. The eight texture images arranged in two rows on the top correspond to the eight side faces of the captured space, the texture image on the lower left corresponds to the bottom surface (floor) of the captured space, and the texture image on the lower right corresponds to the top surface (ceiling) of the captured space.
[0095] For the polygons on the ceiling and floor surfaces, for example, the area outside the octagons on the ceiling and floor surfaces, which is the hatched area in Fig. 22D, is assigned the pixels of the image of the side surface connected to this area. Any pixel value, such as white pixels or black pixels, may be assigned to this area.
[0096] After generating the 10 texture images, the control unit 11 stores the vertex data and polygon data generated in step S33 and the 10 texture images generated in step S34 as images for viewing in the image for viewing DB 12b of the storage unit 12 (S35). Note that the control unit 11 may display the generated 10 texture images on the display unit 15 or the display unit 25 of the registration terminal 20, and store them in the image for viewing DB 12b after receiving an instruction to register the images for viewing from the user of the server 10 or the registration terminal 20 via the input unit 14 or 24. The images for viewing can be stored as file data in an existing format for 3D models, such as the OBJ format.
[0097] Through the above processing, a viewable image to be provided to the user (viewer) of viewing terminal 40 can be generated from the captured image taken by camera 30 and registered in storage unit 12 (viewable image DB 12b). In this embodiment, the eight vertical lines L0 to L7 specified for the captured image are used not only in the process of estimating the polygonal shape of the captured space but also in the zenith correction performed on the captured image, enabling efficient processing. In this embodiment, server 10 stores the viewable image and provides the viewable image to viewing terminal 40 in response to a request from viewing terminal 40. However, an image providing server separate from server 10 may be provided, and the image providing server may acquire the viewable image from server 10 and provide the viewable image to viewing terminal 40 in response to a request from viewing terminal 40. After generating the viewable image, server 10 may delete the captured image before generating the viewable image from captured image DB 12a.
[0098] Next, the processing performed by server 10 and viewing terminal 40 when a panoramic image captured by camera 30 is played back (displayed) on viewing terminal 40 based on the images for viewing registered in viewing image DB 12b of server 10 will be described. Fig. 24 is a flowchart showing an example of the playback processing procedure for a panoramic image, and Fig. 25 is a schematic diagram showing an example of a display screen for the panoramic image. In Fig. 24, the processing performed by viewing terminal 40 is shown on the left, and the processing performed by server 10 is shown on the right.
[0099] When a viewer wishes to view one of the panoramic images registered on server 10, the viewer accesses server 10 using, for example, a web browser (not shown) on viewing terminal 40 and requests transmission of the desired panoramic image (image for viewing) via a website provided by server 10. Server 10 functions as a web server and publishes a website via network N for accepting requests for panoramic images (images for viewing). Therefore, control unit 41 of viewing terminal 40 requests the panoramic image (image for viewing) via the website provided by server 10 in accordance with instructions from the viewer via input unit 44 (S111). In addition to accessing server 10 using the web browser, control unit 41 may also use an application program for accessing the website provided by server 10.
[0100] When a panoramic image is requested from the viewing terminal 40, the control unit 11 of the server 10 selects and reads out from the viewing image DB 12b an image for viewing that has been generated from the requested panoramic image (photographed image) (S112). The control unit 41 of the viewing terminal 40 may accept a specification of the room in which the panoramic image that the viewer wishes to view was taken, and request the panoramic image from the server 10 based on information about the specified room (photographing location information). In this case, the control unit 11 of the server 10 may read out the image for viewing that is stored in the viewing image DB 12b in association with information about the specified room (photographing location information). The image for viewing includes the vertex data and polygon data shown in FIG. 21C and the texture image shown in FIG. 23.
[0101] The control unit 11 transmits the image for viewing read from the image for viewing DB 12b and a viewing program for playing (displaying) the image for viewing to the viewing terminal 40 via the communication unit 13 (S113). This allows the control unit 11 (transmission unit) to transmit information indicating the shape of the shooting space (vertex data and polygon data) and a texture image to an external device. The viewing program is written in, for example, JavaScript and WebGL, and is stored in advance in the storage unit 12. The control unit 41 of the viewing terminal 40 receives the image for viewing and the viewing program transmitted by the server 10 via the communication unit 43 and stores them in the storage unit 42 (S114). By executing the viewing program, the control unit 41 can generate a 3D model of the state as viewed from a virtual position in the shooting space based on the image for viewing and display the 3D model on the display unit 45.
[0102] For example, the control unit 41 reads out a preset default virtual position and line of sight direction (orientation) (S115). The default virtual position and line of sight direction are set, for example, in a viewing program. Then, the control unit 41 uses the rendering function (texture mapping) of the viewing program to generate a 3D model viewed from the default virtual position in the default line of sight direction based on the viewing image (vertex data, polygon data, texture image) (S116). Specifically, the control unit 41 (generation unit) generates an octagonal prism shape of the shooting space in a three-dimensional (3D) virtual space based on the vertex data and polygon data, and maps each pixel of the texture image onto each face of the generated octagonal prism to generate a three-dimensional virtual image (3D model). Note that the control unit 41 generates a 3D model viewed from the default virtual position in the default line of sight direction by associating each pixel of the texture image with each pixel position in the virtual position. When the control unit 41 generates the 3D model, the control unit 41 displays the generated 3D model on the display unit 45 (S117).
[0103] FIG. 25 shows an example screen of a 3D model, and the control unit 41 displays the 3D model, a line of sight change button 441, and a movement button 442. The line of sight change button 441 has four buttons: an up-line sight button, a down-line sight button, a left-line sight button, and a right-line sight button. When an operation on each button is accepted, the control unit 41 accepts an instruction to change the line of sight direction upward, downward, leftward, or rightward. The movement button 442 has two buttons: a forward button and a backward button. When an operation on each button is accepted, the control unit 41 accepts an instruction to change the virtual position forward or backward. Note that the movement button 442 may have a button for instructing movement leftward or rightward. The viewer can instruct a change in the line of sight direction and a change in the virtual position by operating either the line of sight change button 441 or the movement button 442 via the input unit 44.
[0104] The control unit 41 receives an instruction to change the virtual position or the line of sight direction by receiving an operation on either the movement button 442 or the line of sight change button 441 via the input unit 44. Therefore, the control unit 41 determines whether an instruction to change the virtual position or the line of sight direction has been received by operating the movement button 442 or the line of sight change button 441 (S118). If the control unit 41 determines that an instruction to change the virtual viewpoint or line of sight direction has been received (S118: YES), the control unit 41 changes the virtual position or the line of sight direction in accordance with the received instruction (S119). For example, when the forward button or the backward button of the movement buttons 442 is operated, the control unit 41 sets the virtual position to a position obtained by moving the current virtual position forward or backward by a predetermined distance from the current line of sight direction. Furthermore, when the up line of sight button, down line of sight button, left line of sight button, or right line of sight button of the line of sight change buttons 441 is operated, the control unit 41 sets the line of sight direction to a direction obtained by changing the current line of sight direction by a predetermined angle upward, downward, leftward, or rightward. Thereafter, control unit 41 returns to the process of step S116 and performs the processes of steps S116 to S117, thereby generating a 3D model of the state seen from the changed virtual position in the changed line of sight direction, and displaying it on display unit 45.
[0105] When the control unit 41 determines that an instruction to change the virtual viewpoint or the line of sight direction has not been received (S118: NO), it determines whether an instruction to end the panoramic image viewing process has been received via the input unit 44 (S120). For example, when an instruction to end access to the server 10 (website) is received via the input unit 44, the control unit 41 determines that an instruction to end the panoramic image viewing process has been received. When it determines that an instruction to end the panoramic image viewing process has not been received (S120: NO), the control unit 41 returns to the processing of step S118. Each time the control unit 41 receives an operation from the viewer on the line of sight change button 441 or the movement button 442 via the input unit 44, the control unit 41 performs the processing of steps S119 and S116 to S117, thereby updating and displaying the 3D model. When it determines that an instruction to end the panoramic image viewing process has been received (S120: YES), the control unit 41 ends the processing. Through the above-described processing, it is possible to provide the viewer with an image (3D model) viewed from any viewpoint within the shooting space based on the actual data captured by the camera 30. Therefore, the viewer can virtually move around the room in the shooting space, change their line of sight in any direction, and intuitively grasp the size of the room, the arrangement of furniture, the view outside the window, etc.
[0106] In this embodiment, zenith correction can be performed on a panoramic image captured by camera 30 by inputting vertical lines indicating corners of the captured space extending in the vertical direction, and a viewing image that can be played back on viewing terminal 40 can be generated. Therefore, zenith correction can be performed appropriately regardless of the shape of the captured space in a planar view. Furthermore, the shape of the captured space in a planar view is estimated by calculating the distance and direction from the capturing position to each corner, so the planar shape of the captured space can be appropriately estimated regardless of its shape. Therefore, an appropriate texture image can be generated based on the appropriately estimated shape of the captured space. Furthermore, a 3D model viewed from any position can be appropriately generated based on the appropriately generated texture image. In this embodiment, viewing terminal 40 downloads a viewing image (vertex data, polygon data, and texture image) generated by server 10, and playback processing is performed on viewing terminal 40. Therefore, server 10 generates a viewing image from a panoramic image (captured image) registered via registration terminal 20 and provides the viewing image to viewing terminal 40. Note that server 10 may generate a 3D model from the image for viewing by texture mapping and transmit the generated 3D model to viewing terminal 40. In this case, viewing terminal 40 may display the 3D model received from server 10 on display unit 45.
[0107] In this embodiment, an HMD-type information processing device may be used as the viewing terminal 40. In this case, an image processing system 100 can be provided that allows even a user who is unfamiliar with operating a computer or the like to easily check images of a room viewed from various viewpoints. According to this embodiment, an image processing system 100 can be provided that can provide images of a room viewed from various viewpoints using images in a general-purpose format such as JPEG (Joint Photographic Experts Group) format captured using a commercially available omnidirectional camera or the like. Note that the camera 30 is not limited to a omnidirectional camera. For example, it may be a camera that can capture a hemispherical range or a camera for capturing ordinary planar images.
[0108] (Embodiment 2) An image processing system 100 will be described that generates an image for viewing based on a captured image in which the upper or lower end of a line segment (vertical line) that indicates a corner of the captured space that exists in the vertical direction in real space is not captured due to furniture placed in the captured space. The image processing system 100 of this embodiment is realized using devices similar to the image processing system 100 of embodiment 1 shown in Figure 1, so a description of the configuration will be omitted. In this embodiment, only the content that is different from the processing performed by each device in the image processing system 100 of embodiment 1 described above will be described.
[0109] FIG. 26 is a flowchart showing an example of a processing procedure for generating an image for viewing according to the second embodiment, FIG. 27 is an explanatory diagram of the processing for generating an image for viewing according to the second embodiment, and FIG. 28 is a flowchart showing an example of a processing procedure for estimating a shooting space according to the second embodiment. The processing shown in FIG. 26 is obtained by adding steps S131 to S132 between steps S28 and S29 and adding step S133 instead of step S30 in the processing shown in FIG. 9. Descriptions of the same steps as in FIG. 9 will be omitted. Also, steps S33 to S35 in FIG. 9 are not shown in FIG. 26. The processing for estimating a shooting space shown in FIG. 28 is the processing of step S32 in FIG. 26. The processing shown in FIG. 28 is obtained by adding steps S141 to S142 after step S106 in the processing shown in FIG. 18.
[0110] In this embodiment, the control unit 11 of the server 10 performs the same processing as steps S21 to S28 in FIG. 9. As a result, eight vertical lines L0 to L7 corresponding to the eight corners of the captured space are set in the captured image. In this embodiment, if the upper or lower end of any vertical line is not visible in the captured image due to furniture placed in the room, the missing upper or lower end is set as an impossible viewpoint. FIG. 27 shows a state in which the eight vertical line segments shown in FIG. 10B have been moved to the eight vertical lines L0 to L7 in the captured image, and the lower end of the vertical line L5 has been set as an impossible viewpoint. The screen shown in FIG. 27 has an impossible viewpoint setting button for instructing the setting of an impossible viewpoint. For example, by operating the impossible viewpoint setting button and then selecting and operating the endpoint (upper or lower end) of any vertical line in the captured image, an instruction can be given to set the selected endpoint as an impossible viewpoint. In FIG. 27, the endpoint of the vertical line set as an impossible viewpoint has been moved to a location close to the exact position of the endpoint. However, the position of the endpoint set as an impossible viewpoint may be any position.
[0111] After processing step S28, the control unit 11 receives an instruction to set the upper or lower end of any vertical line on the captured image as an impossible viewpoint (S131). Then, the control unit 11 displays a mark indicating an impossible viewpoint at the position of the upper or lower end of the vertical line instructed to be set as an impossible viewpoint in the captured image on the screen (S132). In the example shown in FIG. 27, the position of the impossible viewpoint is indicated by a white square mark. Thereafter, if the control unit 11 determines that the OK button on the screen shown in FIG. 27 has been operated (S29: YES), it acquires the vertical lines after movement (S133). Here, the control unit 11 acquires the coordinate values in the image coordinate system of the end points (upper and lower end points) of the eight vertical lines L0 to L7 whose positions after movement have been specified, and acquires a label indicating that the end point set as an impossible viewpoint is an impossible viewpoint (information indicating that the point cannot be instructed to move and its position cannot be specified).
[0112] Thereafter, the control unit 11 executes the processes from step S31 onward. In this embodiment, the upper or lower ends of the vertical lines L0 to L7 that are not visible in the captured image are set as invisible viewpoints, so in step S31, the control unit 11 performs zenith correction using the vertical lines whose upper and lower end positions have been specified in step S133. In a room, furniture is often installed on the floor, so the upper end of a vertical line (the end point on the ceiling side) is likely to be a visible point that appears in the captured image. Furthermore, among the multiple vertical lines (eight in this case) in the room, it is thought that there are multiple vertical lines whose both upper and lower end positions are specified. Therefore, the control unit 11 can perform zenith correction of the captured image based on the multiple vertical lines whose both upper and lower end positions have been specified.
[0113] 12, in step S41, control unit 11 acquires the coordinate values of both ends of a vertical line whose upper and lower end positions are specified, and in step S42, calculates parameters representing each vertical line based on the acquired coordinate values of both ends of the vertical line. Then, control unit 11 executes the processes of steps S43 to S44 based on the calculated parameters (θ, ρ) of the multiple vertical lines, and performs zenith correction on the captured image and the vertical lines based on the obtained position of the rotation axis (phase φ) and rotation amount α (S45, S46).
[0114] 26, the control unit 11 executes the process shown in FIG. 28. In the process shown in FIG. 28, the control unit 11 performs the same process as steps S101 to S106 in FIG. 18. In step S101, the control unit 11 acquires the captured image after zenith correction and the coordinate values of the image coordinate system of the upper and lower ends of each vertical line after zenith correction. Here, the control unit 11 associates a label indicating an impossible viewpoint with the end point set as an impossible viewpoint. In step S102, the control unit 11 calculates the height of the capturing position based on the ratio between the length from the center line to the upper end and the length from the center line to the lower end of one vertical line whose coordinate values of the upper and lower ends were acquired in step S101. In step S103, the control unit 11 acquires the zenith angle θ of the pixel of the end point (upper end or lower end) that is not set as an impossible viewpoint, as information of one vertical line. Specifically, the control unit 11 calculates the coordinates (x I ,y I ) and calculates the zenith angle θ of the pixel at the endpoint. Furthermore, in step S104, the control unit 11 calculates the distance from the shooting position to the vertical line based on the calculated zenith angle θ and the height of the shooting position calculated in step S102.
[0115] After estimating the planar shape of the shooting space in step S106, the control unit 11 determines whether or not there is an endpoint set at an impossible viewpoint based on the information (coordinate values or labels indicating impossible viewpoints) of the upper and lower ends of each vertical line acquired in step S101 (S141). If it is determined that there is an endpoint set at an impossible viewpoint (S141: YES), the control unit 11 calculates the coordinate values of the endpoint set at an impossible viewpoint (S142). The Xw and Zw coordinate values in the world Cartesian coordinate system of the lower end of the vertical line L5 in FIG. 27 are the same as the Xw and Zw coordinate values in the world Cartesian coordinate system of the upper end of the vertical line L5 as shown in FIG. 20A. Furthermore, the Yw coordinate value in the world Cartesian coordinate system of the lower end of the vertical line L5 is the same as the Yw coordinate value in the world Cartesian coordinate system of the lower ends of the other vertical lines L0 to L4 and L6 to L7. Therefore, the control unit 11 (coordinate value calculation unit) can calculate the coordinate value in the world Cartesian coordinate system of an endpoint set at an impossible viewpoint from the coordinate values of other endpoints in the vicinity of the endpoint, and stores the calculated coordinate values in the storage unit 12. Specifically, the control unit 11 can calculate the coordinate value of the endpoint set at an impossible viewpoint based on the coordinate value of other endpoints located vertically relative to the endpoint set at an impossible viewpoint and the coordinate value of other endpoints located horizontally.
[0116] If it is determined that there is no end point set at an impossible viewpoint (S141: NO), the control unit 11 skips the process of step S142 and returns to the process of Fig. 26. Thereafter, the control unit 11 executes the processes of steps S33 to S35 based on the shape of the shooting space estimated by the process shown in Fig. 28. As a result, it is possible to generate and store the estimated shape of the shooting space, specifically, vertex data and polygon data indicating an octagonal prism, and texture images for 10 faces.
[0117] By the above-described processing, in this embodiment as well, a viewing image can be generated from a captured image taken by camera 30. Furthermore, in this embodiment, even if the upper or lower end of a corner (vertical line) of the captured space is not visible in the captured image due to furniture placed in the room, the coordinate value of this upper or lower end can be calculated based on the coordinate values of other end points (upper or lower ends of other vertical lines) visible in the captured image, thereby enabling accurate estimation of the polygonal shape of the captured space. Also, in this embodiment as well, server 10 can provide a viewing image registered in viewing image DB 12b in response to a request from viewing terminal 40. Furthermore, viewing terminal 40 can generate a 3D model of the captured space as viewed from a virtual position based on the viewing image and display it on display unit 45.
[0118] (Embodiment 3) An image processing system 100 will be described in which an image for viewing generated by the server 10 is provided to a viewing terminal 40 by a server (image for viewing providing server) separate from the server 10. The image processing system 100 of this embodiment is similar to that of the first embodiment, except that, of the processes performed by the server 10 of the first embodiment, the process of providing the image for viewing to the viewing terminal 40 is performed by a separate server. FIG. 29 is a schematic diagram showing an example of the configuration of the image processing system 100 of a third embodiment. The image processing system 100 of this embodiment includes, in addition to devices similar to those of the image processing system 100 of the first embodiment, a server for providing images for viewing 50 that provides images for viewing via a network N. The server for providing images for viewing 50 has a configuration similar to that of the server 10, and transmits and receives information to and from other devices via the network N.
[0119] In the image processing system 100 of this embodiment, the server 10 performs a process of registering a photographed image shown in FIG. 7 and registers the photographed image (panoramic image) acquired via the registration terminal 20 in the storage unit 12 (photographed image DB 12a). The server 10 also performs a process of generating images for browsing shown in FIGS. 9, 12-15, and 18 and generates a view image from each of the registered photographed images. Here, the server 10 of this embodiment associates the view images generated from the photographed images with information about the shooting location (shooting position information) and transmits them to the view image providing server 50. Specifically, the server 10 transmits a view image DB 12b as shown in FIG. 4B to the view image providing server 50, and the view image providing server 50 stores the received view image DB 12b in the storage unit. Then, in this embodiment, the view image providing server 50 performs a process of providing images for browsing shown in FIG. 24 and provides the view images to the viewing terminal 40 in response to a request from the viewing terminal 40.
[0120] In this embodiment, the processing load on the server 10 can be reduced by distributing the processing performed by the server 10 in the first embodiment. Furthermore, for example, when panoramic images of real estate properties are to be provided (viewed), a view image providing server 50 may be provided for each region, and each view image providing server 50 may be configured to manage view images of properties in each region. In this case, for example, an employee of a real estate company takes photos and registers the obtained panoramic images in the server 10. The server 10 generates view images from the registered panoramic images and registers the generated view images in the view image providing server 50 corresponding to the region of the property to be photographed. With this configuration, view images of each property can be managed by a view image providing server 50 for each region.
[0121] The embodiments disclosed herein are illustrative in all respects and should not be considered limiting. The scope of the present invention is defined by the claims, not by the above meaning, and is intended to include all modifications within the meaning and scope of the claims. [Explanation of symbols]
[0122] 10 Servers 11 Control section 12 Storage section 13 Communications Department 20 Registration terminal 21 Control Unit 22 Memory section 23 Communications Department 30 Camera 40 Viewing terminals 41 Control Unit 42 Storage section 43 Communications Department 100 Image Processing System 12a Photographic image database 12b Image DB for viewing
Claims
1. an image acquisition unit that acquires a panoramic image taken within a closed real space; a space information acquisition unit that acquires the height of the real space and the number of corners of a polygon in a planar view; a photographing position calculation unit that specifies a position on the panoramic image corresponding to a vertical side of the real space, and calculates the distance of the photographing position from the ceiling or the floor by dividing the height of the real space corresponding to one line segment in the panoramic image corresponding to the vertical side of the real space into a distance from a photographing position to a ceiling and a distance from the photographing position to a floor based on a ratio between a length from a center to an upper end of the panoramic image in the vertical direction and a length from the center to a lower end of the line segment; a distance calculation unit that calculates the distance from the imaging position to the line segment based on the distance from the ceiling of the imaging position and an elevation angle from the imaging position to an upper end of the line segment on the ceiling side, or calculates the distance from the imaging position to the line segment based on the distance from the floor of the imaging position and an depression angle from the imaging position to a lower end of the line segment on the floor side; a plane estimation unit that estimates the shape of the polygon based on the distances calculated for each line segment; an estimation unit that estimates a shape of a polygonal prism in the real space using the estimated shape of the polygon and the acquired height in the real space; a generation unit that generates a 3D model that constitutes the estimated polygonal prism; An image processing device comprising:
2. The photographing position calculation unit identifies a position corresponding to a vertical side of the real space by image recognition. The image processing device according to claim 1 .
3. The photographing position calculation unit receives and identifies an input of a position of an end of a vertical side of the real space on the panoramic image. The image processing device according to claim 1 .
4. a texture image generating unit that generates a texture image by associating each pixel of the panoramic image with each face that constitutes the estimated polygonal prism, The generation unit generates the estimated shape of the polygonal prism in a 3D virtual space, and generates the 3D model by mapping each pixel of the texture image onto each surface of the generated shape.
4. The image processing device according to claim 1.
5. a transmitting unit that transmits information indicating the estimated shape of the polygonal prism and the generated texture image to an external device; The image processing device according to claim 4 , comprising:
6. a line segment display unit that displays, on the panoramic image, line segments in the vertical direction of the panoramic image, the number of line segments corresponding to the number of corners of the acquired polygon; a reception unit that receives a movement instruction for moving upper and lower ends of the line segment displayed on the panoramic image, The photographing position calculation unit receives the positions of the upper and lower ends of the line segment after the movement in accordance with the movement instruction.
4. The image processing device according to claim 1.
7. the receiving unit receives information indicating that the upper and lower ends of the line segment on the panoramic image are points that can be instructed to move or that cannot be instructed to move; The imaging position calculation unit receives the positions of the upper and lower ends of the line segment after the movement of the line segment for which the movement instruction has been received. The image processing device according to claim 6 .
8. a correction unit that corrects the panoramic image so that each of the line segments corresponding to the vertical sides of the acquired real space exists in the up and down directions of the panoramic image. The image processing device according to claim 1 , further comprising:
9. the imaging position calculation unit calculates a distance from a ceiling or a floor surface to the imaging position based on the ratio of a line segment that can identify the positions of both the upper and lower ends; When the position of the lower end of the line segment on the floor side cannot be specified, the distance calculation unit calculates the distance from the imaging position to the line segment based on the distance from the imaging position to the ceiling and the angle of elevation from the imaging position to the upper end of the line segment on the ceiling side, and when the position of the upper end of the line segment on the ceiling side cannot be specified, the distance calculation unit calculates the distance from the imaging position to the line segment based on the distance from the imaging position to the floor and the angle of depression from the imaging position to the lower end of the line segment on the floor side.
9. An image processing device according to claim 1.
10. a coordinate value acquisition unit that acquires coordinate values on the panoramic image of vertices of each of a plurality of side surfaces, a ceiling surface, and a floor surface that constitute the estimated polygonal prism; The texture image generation unit generates a texture image of each surface based on coordinate values of vertices of each of the side surfaces, the ceiling surface, and the floor surface.
6. The image processing device according to claim 4 or 5.
11. a coordinate value calculation unit that, when any of the vertices of the estimated side surfaces, ceiling surface, and floor surface constituting the polygonal prism is a point whose position cannot be specified, calculates a coordinate value of the vertex whose position cannot be specified based on the coordinate values of the other vertices. The image processing device according to claim 10 , comprising:
12. Acquire panoramic images taken in a closed real space, Acquire the height of the real space and the number of corners of the polygon in a planar view; a position on the panoramic image corresponding to a vertical edge of the real space is identified, and based on a ratio of a length from a center to an upper end of a line segment in the panoramic image corresponding to the vertical edge of the real space to a length from the center to a lower end, the height of the real space corresponding to the line segment is proportionally divided into a distance from a shooting position to a ceiling and a distance from the shooting position to a floor, thereby calculating a distance from the shooting position to the ceiling or the floor; Calculating the distance from the photographing position to the line segment based on the distance from the ceiling of the photographing position and the angle of elevation from the photographing position to the upper end of the line segment on the ceiling side, or calculating the distance from the photographing position to the line segment based on the distance from the floor of the photographing position and the angle of depression from the photographing position to the lower end of the line segment on the floor side, estimating the shape of the polygon based on the distances calculated for each line segment; estimating a shape of a polygonal prism in the real space using the estimated shape of the polygon and the acquired height in the real space; A 3D model of the estimated polygonal prism is generated. A program that causes a computer to perform a process.
13. Acquire panoramic images taken in a closed real space, Acquire the height of the real space and the number of corners of the polygon in a planar view; a position on the panoramic image corresponding to a vertical edge of the real space is identified, and based on a ratio of a length from a center to an upper end of a line segment in the panoramic image corresponding to the vertical edge of the real space to a length from the center to a lower end, the height of the real space corresponding to the line segment is proportionally divided into a distance from a shooting position to a ceiling and a distance from the shooting position to a floor, thereby calculating a distance from the shooting position to the ceiling or the floor; Calculating the distance from the photographing position to the line segment based on the distance from the ceiling of the photographing position and the angle of elevation from the photographing position to the upper end of the line segment on the ceiling side, or calculating the distance from the photographing position to the line segment based on the distance from the floor of the photographing position and the angle of depression from the photographing position to the lower end of the line segment on the floor side, estimating the shape of the polygon based on the distances calculated for each line segment; estimating a shape of a polygonal prism in the real space using the estimated shape of the polygon and the acquired height in the real space; A 3D model of the estimated polygonal prism is generated. An image processing method in which processing is performed by a computer.
Citation Information
Patent Citations
Image processing apparatus, image processing method, and program
JP2020071854A
Information processing device, program, and information processing system
JP2020204973A
Image processing device, program and image processing system
JP2021089487A
Information processing device and program
JP6872660B1
Three-dimensional model generation system, three-dimensional model generation method, and program
WO2017203710A1