Viewpoint image generation device, viewpoint image generation method, and program
Patent Information
- Application Number
- JP2024544073
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Priority Date
- 2023-08-07
- Filing Date
- 2023-08-07
- Publication Date
- 2025-06-09
Abstract
Description
Viewpoint image generating device, viewpoint image generating method, and program
[0001] The present invention relates to a viewpoint image generating device, a viewpoint image generating method, and a program.
[0002] In recent years, cameras mounted on drones (unmanned aerial vehicles) have been used to photograph buildings and other objects on the ground from the sky. The cameras mounted on the drones photograph buildings on the ground while the drones are flying. Drone photography generally results in the acquisition of a large number of images. From these large numbers of images, 3D point clouds and orthoimages of the target object (e.g., a building) are generated using 3D restoration techniques such as SfM (Structure from Motion) or MVS (Multi View Stereo).
[0003] For example, Patent Document 1 describes a technique for generating a three-dimensional model by SfM based on images captured by a drone.
[0004] Japanese Patent Application Laid-Open No. 2020-5186
[0005] The point cloud obtained by such 3D reconstruction technology has the advantage of being able to observe objects such as buildings from any viewpoint and to measure distances, etc. However, the point cloud obtained by 3D reconstruction technology also has the disadvantage that the resolution is too coarse for detecting abnormalities (e.g., damage) in buildings, etc.
[0006] In addition, orthoimages (rectilinear projection images) can be generated after 3D reconstruction using 3D reconstruction technologies such as SfM and MVS. These images have the advantage of being able to be overlaid on a map, making it easy to identify objects, but they also have the disadvantage of not being able to observe the sides of the object.
[0007] The present invention has been made in consideration of these circumstances, and its purpose is to provide a viewpoint image generation device, a viewpoint image generation method, and a program that can generate viewpoint images that allow the side of a target object to be properly observed from images taken by a drone or the like.
[0008] In order to achieve the above-mentioned object, a first aspect of the present invention is a viewpoint image generation device that includes a processor, which acquires a plurality of images of a target object taken from different viewpoints, acquires the position and attitude of the camera that captured the plurality of images, acquires target object coordinates that are the coordinates of the target object, calculates a first camera field of view coordinate of the target object in the camera based on the camera position, attitude, and target object coordinates, generates observation viewpoint coordinates that are the coordinates of a plurality of virtual observation viewpoints based on the target object coordinates, determines observation directions of the target object according to the plurality of virtual observation viewpoints, calculates a second camera field of view coordinate of the target object when a camera is placed at the virtual observation viewpoint based on the observation viewpoint coordinates, observation directions, and target object coordinates, selects a transformed image to be subjected to projective transformation from the plurality of images based on the virtual observation viewpoint and observation direction, and performs projective transformation on the transformed image based on the first camera field of view coordinates and the second camera field of view coordinates to generate a viewpoint image observed from the virtual observation viewpoint.
[0009] According to this aspect, the extracted transformed image is projectively transformed based on the coordinates in the field of view of the first camera and the coordinates in the field of view of the second camera, and viewpoint images are generated. As a result, this aspect can generate viewpoint images that allow the target object to be appropriately observed from multiple viewpoints.
[0010] In a second aspect, preferably in the first aspect, the processor extracts the target object coordinates from map information of the target object.
[0011] A third aspect is preferably the first or second aspect, wherein the map information includes polygon data of the target object.
[0012] A fourth aspect is preferably any one of the first to third aspects, wherein the projective transformation includes an expansion transformation or a contraction transformation.
[0013] A fifth aspect is preferably any one of the first to fourth aspects, wherein the plurality of virtual observation viewpoints are set on the circumference of a circle that contains the target object.
[0014] A sixth aspect is preferably any one of the first to fifth aspects, wherein the processor identifies a plurality of virtual observation viewpoints from a group of polygon coordinates that define the outer shape of the target object.
[0015] In a seventh aspect, preferably, in the fifth aspect, the height of the circle relative to the target object, or the diameter or radius of the circle, is changeable.
[0016] The eighth aspect is preferably any one of the first to seventh aspects, in which the selection of the transformed image to be subjected to projective transformation is performed based on a virtual observation viewpoint, a camera position, an observation direction, and a camera attitude.
[0017] In a ninth aspect, preferably, in the eighth aspect, the processor calculates a projection transformation matrix from at least four coordinates in the field of view of the first camera and four corresponding coordinates in the field of view of the second camera, and performs projection transformation on the transformed image.
[0018] A tenth aspect is preferably any one of the first to ninth aspects, wherein the processor generates a new viewpoint image between viewpoint images by interpolation based on adjacent viewpoint images.
[0019] An eleventh aspect is preferably any one of the first to tenth aspects, wherein the processor displays the viewpoint images in a revolving order around the target object.
[0020] In a twelfth aspect, preferably in any one of the first to eleventh aspects, the processor displays a viewpoint image in response to a change in the virtual observation viewpoint performed in a sequence of revolving around the target object.
[0021] In a thirteenth aspect, preferably in any one of the first to twelfth aspects, a processor acquires an image group including a plurality of images in which the target object is visible, and extracts a plurality of images in which the target object is visible from the image group.
[0022] A viewpoint image generating method according to a fourteenth aspect of the present invention includes the steps of acquiring a plurality of images of a target object taken from different viewpoints, acquiring the position and orientation of a camera that captured the plurality of images, acquiring target object coordinates that are the coordinates of the target object, calculating first camera field of view coordinates of the target object in the camera based on the camera position, orientation, and target object coordinates, generating observation viewpoint coordinates that are the coordinates of a plurality of virtual observation viewpoints based on the target object coordinates, determining observation directions of the target object according to the plurality of virtual observation viewpoints, calculating second camera field of view coordinates of the target object when a camera is placed at the virtual observation viewpoint based on the observation viewpoint coordinates, observation direction, and target object coordinates, and selecting a transformed image to be subjected to projective transformation from the plurality of images based on the virtual observation viewpoint and observation direction, and performing projective transformation on the transformed image based on the first camera field of view coordinates and the second camera field of view coordinates to generate a viewpoint image observed from the virtual observation viewpoint.
[0023] A fifteenth aspect of the present invention is a program that executes the following steps: acquiring multiple images of a target object taken from different viewpoints; acquiring the position and orientation of the camera that captured the multiple images; acquiring target object coordinates, which are the coordinates of the target object; calculating first camera field of view coordinates of the target object in the camera based on the camera position, orientation, and target object coordinates; generating observation viewpoint coordinates, which are the coordinates of multiple virtual observation viewpoints, based on the target object coordinates; determining observation directions of the target object according to the multiple virtual observation viewpoints; calculating second camera field of view coordinates of the target object when the camera is placed at the virtual observation viewpoint, based on the observation viewpoint coordinates, observation direction, and target object coordinates; and selecting a transformed image to be subjected to projective transformation from the multiple images based on the virtual observation viewpoint and observation direction, and performing projective transformation on the transformed image based on the first camera field of view coordinates and the second camera field of view coordinates to generate a viewpoint image observed from the virtual observation viewpoint.
[0024] According to the present invention, viewpoint images can be generated from images captured by a drone or the like, allowing a target object to be properly observed from multiple viewpoints.
[0025] FIG. 1 is a conceptual diagram showing a drone for aerial photography and a viewpoint image generation device. FIG. 2 is a block diagram showing an embodiment of a hardware configuration. FIG. 3 is a functional block diagram showing the functions of a processor. FIG. 4 is a diagram explaining an example of a method for aerial photography by a drone. FIG. 5 is a diagram showing an example of a plurality of captured images. FIG. 6 is a diagram explaining target object coordinates. FIG. 7 is a diagram explaining target object coordinates. FIG. 8 is a diagram explaining determination of an observation direction by an observation direction determination unit. FIG. 9 is a diagram explaining an example of calculation of a projective transformation matrix. FIG. 10 is a diagram explaining an example of calculation of a projective transformation matrix. FIG. 11 is a diagram explaining an example of display of a viewpoint image. FIG. 12 is a diagram explaining an example of display of a viewpoint image. FIG. 13 is a flow diagram showing a viewpoint image generation method. FIG. 14 is a diagram explaining captured images. FIG. 15 is a diagram explaining a viewpoint image generation unit.
[0026] Hereinafter, preferred embodiments of a viewpoint image generating device, a viewpoint image generating method, and a program according to the present invention will be described with reference to the accompanying drawings.
[0027] 1 is a conceptual diagram showing a drone for aerial photography that photographs a target object of the present invention and a viewpoint image generation device of the present invention. In FIG. 1, the drone 12 and viewpoint image generation device 20 are connected via a network 22.
[0028] The drone 12 is an unmanned aerial vehicle that is remotely controlled using a remote controller 16 .
[0029] For example, when a disaster occurs, an investigator investigating the damage situation or a drone operator commissioned by a local government operates the drone 12 using a remote controller 16 and takes aerial photographs of a target area (disaster area) where multiple houses exist using the camera 14 mounted on the drone 12. Note that photography using the drone 12 is not limited to when the above-mentioned disaster occurs. Aerial photography using the drone 12 is also widely performed in other cases, and the images used in the present invention also include images taken by the drone 12 in various cases.
[0030] The camera 14 is mounted on the drone 12 via a gimbal head 13 .
[0031] An image captured using the camera 14 (hereinafter referred to as a "captured image IM") can be stored in an internal storage built into the camera 14 and / or a storage device such as a memory card detachably attached to the camera 14. The captured image IM can be transferred to the viewpoint image generation device 20 using wireless communication.
[0032] The viewpoint image generating device 20 is configured using a computer. The computer applied to the viewpoint image generating device 20 may be a server, a personal computer, or a workstation.
[0033] The viewpoint image generating device 20 can perform data communication with the drone 12 and the remote controller 16 via a network 22. The network 22 may be a local area network or a wide area network.
[0034] The viewpoint image generating device 20 can acquire the captured image IM from the drone 12 or the camera 14 via the network 22, or via the remote controller 16. The viewpoint image generating device 20 can also acquire the captured image IM from the memory card of the camera 14 or the like without going via the network 22. This is because it is conceivable that the network in some areas may be down due to a disaster.
[0035] Furthermore, the viewpoint image generating device 20 acquires the necessary map from an internal memory (database 220 (see FIG. 2)) or an external database that provides maps. In this case, the map is a map that corresponds to the captured image IM, and is preferably a map that includes the area captured by the captured image IM. Details of the viewpoint image generating device 20 and the map will be described later.
[0036] The display device 230 displays the captured image IM, and also displays a generated viewpoint image as will be described later.
[0037] FIG. 2 is a block diagram showing an embodiment of the hardware configuration of the viewpoint image generating device 20 shown in FIG.
[0038] The viewpoint image generating device 20 shown in FIG. 2 includes a processor 200 , a memory 210 , a database 220 , a display device 230 , an input / output interface 240 , and an operation unit 250 .
[0039] The processor 200 is composed of a CPU (Central Processing Unit) and the like, and controls each part of the viewpoint image generating device 20, and also performs processing to generate viewpoint images, which will be described later, and processing to display the generated viewpoint images on the display device 230. Details of the various processes of this processor 200 will be described later.
[0040] The memory 210 includes a flash memory, a read-only memory (ROM), a random access memory (RAM), a hard disk drive, etc. The flash memory, ROM, or hard disk drive is a non-volatile memory that stores various programs, etc., including an operating system. Note that these programs also include a program that executes a viewpoint image generation method, which will be described later. The RAM functions as a working area for processing by the processor 200, and also temporarily stores programs, etc., stored in the flash memory, etc. Note that the processor 200 may have a portion of the memory 210 (RAM) built in.
[0041] The memory 210 also functions as an image storage unit that stores the captured images IM, and can store and manage the captured images IM.
[0042] The database 220 is a part that manages map information MD of the aerial photographed area (area L), and in this example, manages the map GO and polygon data CO related to the houses on the map contained in the map information MD.
[0043] Details of the information managed by the database 220 will be described later. The database 220 may be configured inside the viewpoint image generating device 20 or may be an external database connected via communication, such as a database of the Geospatial Information Authority of Japan that manages basic map information or an OpenStreetMap database.
[0044] The display device 230 displays the captured image IM and also displays the viewpoint image in response to an instruction from the processor 200. The display device 230 is also used as part of a GUI (Graphical User Interface) when various types of information are received from the operation unit 250.
[0045] The display device 230 may be included in the viewpoint image generating device 20, or may be provided separately from the viewpoint image generating device 20 as shown in FIG.
[0046] The input / output interface 240 includes a connection unit connectable to an external device, a communication unit connectable to a network, etc. As the connection unit connectable to an external device, a Universal Serial Bus (USB), a High-Definition Multimedia Interface (HDMI) (HDMI is a registered trademark), etc. can be applied. The processor 200 can acquire a captured image IM via the input / output interface 240, or output necessary information in response to an external request.
[0047] The operation unit 250 includes a pointing device such as a mouse, a keyboard, and the like, and functions as part of a GUI that accepts input of various information and instructions by user operation.
[0048] FIG. 3 is a functional block diagram showing the functions of the processor 200 of the viewpoint image generating device 20 shown in FIG.
[0049] The processor 200 functions as a captured image acquisition unit 200A, a camera position and orientation acquisition unit 200B, a target object coordinate acquisition unit 200C, a first camera field of view coordinate calculation unit 200D, an observation viewpoint coordinate generation unit 200E, an observation direction determination unit 200F, a second camera field of view coordinate calculation unit 200G, a converted image selection unit 200H, a viewpoint image generation unit 200I, and a display control unit 200J. The first camera field of view coordinates are calculated mainly by the functional block group I, and the second camera field of view coordinates are calculated mainly by the functional block group II.
[0050] <Acquisition of Photographed Images> The photographed image acquisition unit 200A acquires a plurality of photographed images IM (see FIG. 1 ) captured by the camera 14 of the drone 12 from the drone 12 via the network 22. The photographed image acquisition unit 200A may also acquire a plurality of photographed images IM from the memory card of the camera 14.
[0051] FIG. 4 is a diagram illustrating an example of a method for the drone 12 to take aerial photographs.
[0052] The drone 12 flies along a flight path Z and takes aerial photographs of the area L during flight. The drone 12 then acquires multiple captured images IM. The drone 12 captures the area L using so-called horizontal grid photography. Here, for example, when using a 3D reconstruction technique such as SfM (Structure from Motion) or MVS (Multi View Stereo), the camera 14 of the drone 12 captures images so that the overlap of adjacent captured images IM (the overlap of captured images IM in the M direction in the figure) is approximately 75%, and the side overlap of adjacent captured images IM (the overlap of captured images IM in the N direction in the figure) is approximately 60%. Therefore, assuming that 3D reconstruction will be performed using SfM or the like, the drone 12 acquires a large number of captured images IM of the area L.
[0053] 5 is a diagram showing an example of a plurality of captured images IM captured by the drone 12 and acquired by the captured image acquisition unit 200A. In the example shown in FIG. 5, the captured images IM are cut out from a portion in which the target object, a house T, is captured. In this example, the captured image acquisition unit 200A acquires only the captured images IM in which the house T is captured.
[0054] The captured image acquisition unit 200A acquires multiple captured images IM of a house T, which is a target object, captured from different viewpoints. The captured images IM show the house T captured from various capturing positions while the drone 12 is moving along the flight path Z. Although the captured images IM show the house T, the house T may appear upside down or the walls of the house T may appear distorted depending on the capturing position of the drone 12. Therefore, as will be described later, the viewpoint image generation unit 200I performs projective transformation on the captured images IM to generate viewpoint images that are easy for the user to observe.
[0055] <Acquisition of coordinates within the first camera's field of view> First, acquisition of coordinates within the first camera's field of view will be described. Here, the coordinates within the first camera's field of view are coordinates within the field of view of the camera 14 of the house T, which is the target object, when photographed by the camera 14 of the drone 12.
[0056] The camera position and orientation acquisition unit 200B acquires the position and orientation of the camera 14 from the captured images IM acquired by the captured image acquisition unit 200A. For example, the camera position and orientation acquisition unit 200B acquires the position and orientation of the camera 14 when each captured image IM is captured using SfM. Here, SfM can acquire the three-dimensional coordinates of the object and information on the camera direction and camera position (camera parameters) using multiple images of the same object captured from multiple viewpoints (camera positions). SfM uses parallax based on camera movement to estimate and calculate three-dimensional coordinates based on the amount of displacement of the same feature point of the target object, and accordingly, the orientation and position of the camera 14 when each captured image IM is captured are estimated. Therefore, the camera position and orientation acquisition unit 200B uses SfM to acquire the position and orientation of the camera when each captured image IM was captured from the multiple captured images IM. The camera position and orientation acquisition unit 200B may calculate the position and orientation of the camera 14 by the processor 200 or may acquire them from the outside via the input / output interface 240.
[0057] The target object coordinate acquisition unit 200C acquires the target object coordinates, which are the coordinates of the house T. The target object coordinate acquisition unit 200C can acquire the target object coordinates by various methods. For example, the target object coordinate acquisition unit 200C extracts the target object coordinates from the map information MD of the target object stored in the database 220. The target object coordinate acquisition unit 200C may also acquire the target object coordinates input from outside via the input / output interface 240.
[0058] 6 and 7 are diagrams for explaining the target object coordinates.
[0059] FIG. 6 is a diagram conceptually showing a house T, which is a target object, and FIG. 7 is a diagram showing an example of polygon data CO stored in the database 220. As shown in FIG.
[0060] The world coordinates (latitude, longitude, altitude) of the vertices P1 to P8 that define the outline of the house T are stored as polygon data CO in the database 220.
[0061] The database 220 stores polygon data CO of the house T (vertex data indicating the vertices of the house T by latitude, longitude, and altitude: polygon coordinate group) in association with the house ID (Identification) of the house T. The database 220 also has a map GO of the area around the house T, and the map GO and the polygon data CO are associated with each other.
[0062] The polygon data CO has vertex data P1 to P8 that indicate the outer periphery of the house T, and the vertex data indicates the positions of the vertices and is expressed by coordinates of longitude, latitude, and altitude.
[0063] In the example shown in FIG. 7, polygon data CO of a house ID "OOXXΔΔ" corresponding to a house T is shown, and polygon data CO for other houses is not shown.
[0064] For example, the vertex data, which is polygon data CO for a house ID "XX△△", is composed of a first vertex (P1) to an eighth vertex (P8), which are represented by three-dimensional coordinates of longitude, latitude, and elevation. Regarding elevation, elevation values corresponding to latitude and longitude may be obtained using a DEM (digital elevation model). In the above example, the polygon data CO represents the perimeter of the house T using eight vertices, but the polygon data CO is not limited to eight vertices and may be composed of multiple vertices that can represent the perimeter of the house.
[0065] The first camera field of view coordinate calculation unit 200D calculates first camera field of view coordinates, which are field of view coordinates of the camera 14 of the house T, based on the position and orientation of the camera 14 acquired by the camera position and orientation acquisition unit 200B and the target object coordinates acquired by the target object coordinate acquisition unit 200C. The first camera field of view coordinates are calculated in numbers corresponding to the target object coordinates. For example, the first camera field of view coordinates are calculated corresponding to the above-mentioned P1 to P8 and are represented by pixel positions of the image. Note that the object coordinates are converted from longitude and latitude into UTM (Universal Transverse Mercator) coordinates (unit: m), and the camera position and virtual camera are also expressed in this unified coordinate system.
[0066] The first camera field of view coordinate calculation unit 200D calculates the first camera field of view coordinates x', y' using the following (Equation 1).
[0067] (Formula 1)
[0068] Note that matrix (A1) indicates the calculated coordinates within the first camera's field of view of the house T, matrix (B1) indicates a matrix for conversion into coordinates within the camera's field of view of the camera 14, matrix (C1) indicates a matrix indicating the focal length of the camera 14, matrix (D1) indicates a matrix indicating perspective projection, matrix (E1) indicates a matrix for conversion from world coordinates to camera coordinates of the camera 14, and matrix (F1) is a matrix indicating the three-dimensional coordinates (target object coordinates) of the house T. As described above, the first camera's field of view coordinate calculation unit 200D acquires the field of view coordinates of the house T in each captured image IM based on the above-mentioned (Equation 1).
[0069] <Coordinates in the Second Camera's Field of View> Next, acquisition of coordinates in the second camera's field of view will be described. The coordinates in the second camera's field of view are coordinates in the camera's field of view of the house T, which is the target object, when the camera 14 is placed at a virtual observation viewpoint (virtual camera position) and the attitude of the camera 14 (virtual camera attitude) is directed in a virtual observation direction.
[0070] The observation viewpoint coordinate generating unit 200E generates observation viewpoint coordinates, which are coordinates corresponding to a plurality of virtual observation viewpoints, based on the target object coordinates of the house T.
[0071] For example, the observation viewpoint coordinate generation unit 200E sets the center O of the polygon data CO (P1 to P8) of the house T, or the center O of the circumscribing rectangle of the polygon data CO, as (OX, OY, OZ), and sets the observation circle Q to have the center O, radius R, and height AA from the ground (see FIG. 8). Given the shooting step angle (SA), the position (observation viewpoint coordinates) (CX, CY, CZ) of the virtual camera can be calculated using the following (Equation 2). The height AA and radius R of the observation circle Q can be set arbitrarily by the user. The shooting step angle is determined according to the observation viewpoint, which is set by the user or determined automatically.
[0072] (Formula 2) CX=R×cos(SA×i)+OX CY=R×sin(SA×i)+OY CZ=A+OZ
[0073] In addition, i in (Equation 2) represents the shooting number, which is the same as the number of observation viewpoints.
[0074] The observation viewpoint coordinate generating unit 200E may also generate observation viewpoint coordinates based on values input from the operation unit 250. For example, the user may input information about the four directions of north, south, east, and west, or information about the direction facing the road from the operation unit 250, and the observation viewpoint coordinates may be generated based on the input values.
[0075] The observation direction determination unit 200F determines the observation directions of the target object according to a plurality of virtual observation viewpoints, based on the observation viewpoint coordinates.
[0076] FIG. 8 is a diagram for explaining the determination of the observation direction by the observation direction determination unit 200F.
[0077] 8, an observation circle Q having a height AA and a radius R is set to observe a house T, and virtual cameras 30A to 30D are placed on the circumference of the observation circle Q. Note that the observation circle Q is a circle that includes the house T.
[0078] The orientation VZ of each of the virtual cameras 30A to 30D shown in Fig. 8 is a vector from the center O of the house T to the position of the virtual camera. The orientation VX of each virtual camera shown in Fig. 8 is a tangent vector to the observation circle Q and is calculated from the relationship VX·VZ=0 (sinner product). The orientation VY of each camera shown in Fig. 8 can be calculated from the cross product of VZ and VX, and an upward vector (opposite to the ground) is used.
[0079] From the above, the virtual observation viewpoint, which is the position of the virtual camera, is calculated as (CX, CY, CZ), and the observation direction, which is the attitude of the virtual camera, is calculated as (VX (vector), VY (vector), VZ (vector). Note that in the above example, the virtual cameras 30A to 30D are arranged on the circumference of the observation circle Q, but this is not limiting. For example, instead of the observation circle Q, the virtual cameras 30A to 30D may be arranged on the circumference of a rectangle that contains the target object.
[0080] The parameters of the virtual camera are the focal length, sensor size, and pixel size of the sensor, and there are no physical constraints. For example, the sensor size of the virtual camera is the parameter of a standard 35 mm full-size sensor. The pixel size is sufficient as long as it is large enough to extract a photographed house, so if the ground resolution of the photo is 10 cm / pixel and the physical size of the house to be extracted is 15 m, the design value is approximately 150 pixels.
[0081] The focal length f can be calculated using the above parameters and the radius of the observation circle, and if the radius of the observation circle is 60 m, it can be calculated using the following (Equation 3): (Equation 3) Focal length f (mm) = radius R of the observation circle (m) × sensor size (mm) / cutout size of house (m) = 60 (m) × 35 (mm) / 15 (m) = 140 (mm).
[0082] From (Equation 3), it can be calculated that the virtual camera has a lens with a focal length of 140 mm.
[0083] The second camera field of view coordinate calculation unit 200G calculates the second camera field of view coordinates based on the observation viewpoint coordinates, the observation direction, and the target object coordinates.
[0084] Specifically, the second camera field of view coordinate calculation unit 200G calculates the coordinates x'', y'' within the first camera field of view using the following (Equation 4).
[0085] (Formula 4)
[0086] Note that matrix (A2) indicates the calculated coordinates within the field of view of the second camera of the house T, matrix (B2) indicates a matrix for conversion into coordinates within the field of view of the virtual camera, matrix (C2) indicates a matrix indicating the focal length of the virtual camera, matrix (D2) indicates a matrix indicating perspective projection, matrix (E2) indicates a matrix for conversion from world coordinates to camera coordinates of the virtual camera, and matrix (F2) is a matrix indicating the three-dimensional coordinates (target object coordinates) of the house T. As described above, the second camera field of view coordinate calculation unit 200G can acquire the field of view coordinates of the house T at the position of each virtual camera based on the above-mentioned (Equation 4).
[0087] <Generation of Viewpoint Image> Next, generation of the viewpoint image will be described. The viewpoint image is generated by projectively transforming the transformed image based on the coordinates in the field of view of the first camera and the coordinates in the field of view of the second camera.
[0088] The converted image selection unit 200H selects a converted image to be subjected to projective transformation from among the plurality of captured images IM based on the observation viewpoint and observation direction. The converted image selection unit 200H selects a converted image from the captured images IM using various methods. Note that the converted image is selected in accordance with the number of observation viewpoints described above.
[0089] The transformed image selection unit 200H selects a transformed image to be subjected to projective transformation from the multiple captured images IM acquired by the captured image acquisition unit 200A, based on the position and orientation of the virtual camera. That is, the transformed image selection unit 200H selects a captured image IM suitable for projective transformation. For example, the transformed image selection unit 200H selects, as the transformed image, a captured image IM whose position and orientation are close to those of the virtual camera 14. For example, the transformed image selection unit 200H selects a transformed image whose dot product between the vector indicating the orientation of the camera 14 and the vector indicating the orientation of the virtual camera is large.
[0090] The viewpoint image generation unit 200I performs projective transformation on the transformed image based on the coordinates in the field of view of the first camera and the coordinates in the field of view of the second camera, to generate a viewpoint image observed from a virtual observation viewpoint. For example, the viewpoint image generation unit 200I calculates a projective transformation matrix from the coordinates of at least four points in the field of view of the first camera and the coordinates of four points in the field of view of the second camera, and performs projective transformation on the transformed image to generate a viewpoint image. The method used by the viewpoint image generation unit 200I to calculate the projective transformation matrix from four corresponding points is a well-known technique known as the DLT (Direct Linear Transformation) algorithm (see Multiple View Geometry in Computer Vision, Cambridge University Press).
[0091] 9 is a diagram illustrating an example of calculation of a projective transformation matrix. In this example, the projective transformation matrix is calculated based on the vertices of a rectangle formed by the parts of the house T that contact the ground.
[0092] The converted image 42 selected by the converted image selection unit 200H from the multiple captured images IM described in Fig. 5 is displayed. The converted image 42 shows a house T, which is a target object, and indicates coordinates J1 to J4 within the first camera's field of view of the house T. Furthermore, the image 44 is an image that is assumed to have been captured by the virtual camera 30D, and coordinates K1 to K4 within the second camera's field of view of the virtual camera 30D are displayed on the image 44.
[0093] The viewpoint image generation unit 200I calculates a 3×3 projective transformation matrix H such that Ki=H×Ji (i is an integer from 1 to 4). Then, the viewpoint image generation unit 200I applies the projective transformation matrix H to the transformed image 42 to generate a viewpoint image corresponding to the transformed image 42.
[0094] 10 is a diagram illustrating another example of calculation of a projective transformation matrix. In this example, the projective transformation matrix is calculated using the vertices of a rectangle that forms the wall surface of the house T as a reference.
[0095] 5 by the converted image selection unit 200H. The converted image 46 shows a house T, which is the target object, and shows coordinates V1 to V4 within the first camera's field of view for the wall surface of the house T. Furthermore, the image 48 is an image that is assumed to have been captured by the virtual camera 30D, and the coordinates W1 to W4 within the second camera's field of view of the virtual camera 30D are shown on the image 48.
[0096] The viewpoint image generation unit 200I calculates a 3×3 projective transformation matrix H such that Wi=H×Vi (i is an integer from 1 to 4). Then, the viewpoint image generation unit 200I applies the projective transformation matrix H to the transformed image 42 to generate a viewpoint image corresponding to the transformed image 42.
[0097] As described above, the viewpoint image generating unit 200I calculates a projective transformation matrix from the coordinates in the field of view of the first camera and the coordinates in the field of view of the second camera of the transformed image, and generates a viewpoint image by transforming the transformed image using the projective transformation matrix. Note that the above-mentioned projective transformation matrix includes an enlargement or reduction transformation and a transformation that inverts the image.
[0098] <Display of viewpoint images> The viewpoint images generated by the viewpoint image generation unit 200I are displayed on the display device 230 by the display control unit 200J. The viewpoint images are displayed on the display device 230 in the order of circling around the house T. By displaying them in this manner, the viewpoint images of the house T can be displayed to the user from a natural viewpoint.
[0099] FIG. 11 is a diagram illustrating a display example of a viewpoint image.
[0100] Reference numeral 52 indicates a display example 1 of viewpoint images. Display example 1 shows an example in which viewpoint images are displayed in a list. In display example 1, from the left end of the drawing, a viewpoint image 52A of the road surface of house T, a viewpoint image 52B of the left side of the road of house T, a viewpoint image 52C of the back side of the road of house T, and a viewpoint image 52D of the right side of the road of house T are displayed in a list. In this case, the direction perpendicular to the road surface may be detected from polygon data of a map showing the road, and may be set as the front (viewpoint image 52A of the road surface of house T).
[0101] Reference numeral 54 indicates a display example 2 of viewpoint images. Display example 2 shows an example in which viewpoint images are displayed in a list. In display example 2, from the left end of the drawing, a viewpoint image 54A of the north face of the house T, a viewpoint image 54B of the east face of the house T, a viewpoint image 54C of the south face of the house T, and a viewpoint image 54D of the west face of the house T are displayed in a list.
[0102] In this way, a more natural display can be achieved by displaying the viewpoint images in the order of rotation around the house T (target object). Also, by displaying the viewpoint images in a list, the house T can be comprehensively observed.
[0103] FIG. 12 is a diagram illustrating a display example of a viewpoint image.
[0104] 12 shows a display example 3 of viewpoint images. In display example 3, viewpoint images corresponding to the respective observation viewpoints are displayed on the display device 230 in accordance with changes in the observation viewpoints that are made in the order of circling around the house T.
[0105] Specifically, the user changes the observation viewpoint from the north side to the east side, the south side, and the west side by operating the operation unit 250. Note that this change of observation viewpoint is performed in the same order as when circling the target object. Then, in response to this change, the display control unit 200J causes the display device 230 to display a viewpoint image 56A of the north side of the house T, a viewpoint image 56B of the east side of the house T, a viewpoint image 56C of the south side of the house T, and a viewpoint image 56D of the west side of the house T.
[0106] In this way, by displaying the corresponding viewpoint images in accordance with the viewpoints that are changed in the order of turning around the house T (target object), a more natural display can be achieved.
[0107] <Viewpoint image generation method> Next, a description will be given of a viewpoint image generation method using the viewpoint image generation device 20. Note that the viewpoint image generation method is performed, for example, by causing the processor 200 of the viewpoint image generation device 20 to execute a program that executes the viewpoint image generation method.
[0108] FIG. 13 is a flow diagram showing a viewpoint image generating method using the viewpoint image generating device 20.
[0109] First, the captured image acquisition unit 200A acquires multiple captured images IM captured by the drone 12 (step S01). Then, the camera position and orientation acquisition unit 200B acquires the position and orientation of the camera 14 for each captured image IM using SfM (step S02). Then, the target object coordinate acquisition unit 200C acquires the target object coordinates (polygon data CO) of the target object, a house T, from the database 220 (step S03). Then, the first camera field of view coordinate calculation unit 200D calculates the first camera field of view coordinates of the house T based on the position and orientation of the camera 14 and the target object coordinates of the house T (step S04).
[0110] Next, the observation viewpoint coordinate generation unit 200E sets an observation circle Q and calculates the position (observation viewpoint) of the virtual camera based on the observation circle Q and the number of observation viewpoints (step S05). The observation direction determination unit 200F then calculates the orientation (observation direction) of the virtual camera based on the observation circle Q and the position of the virtual camera (step S06). The second camera field of view coordinate calculation unit 200G then calculates second camera field of view coordinates based on the position (observation viewpoint) and the orientation (observation direction) of the virtual camera (step S07). The converted image selection unit 200H then selects a converted image from the multiple captured images IM acquired by the captured image acquisition unit 200A in accordance with the number of observation viewpoints (step S08). The viewpoint image generation unit 200I then calculates a projective transformation matrix based on the first camera field of view coordinates and the second camera field of view coordinates, and transforms the converted image using the projective transformation matrix to generate a viewpoint image (step S09).
[0111] As described above, the viewpoint image generating device 20 of the present disclosure can generate viewpoint images that allow a target object to be observed from multiple observation viewpoints from images captured by a drone or the like. Because these viewpoint images are images obtained by projectively transforming the captured image IM, degradation of resolution from the captured image IM is suppressed, making it possible to detect abnormalities (e.g., damage) in buildings and the like. Furthermore, because the observation viewpoint and observation direction are appropriately set for these viewpoint images, the walls of the house T can be appropriately observed.
[0112] In the above embodiment, the hardware structure of the processing units (captured image acquisition unit 200A, camera position and orientation acquisition unit 200B, target object coordinate acquisition unit 200C, first camera field of view coordinate calculation unit 200D, observation viewpoint coordinate generation unit 200E, observation direction determination unit 200F, second camera field of view coordinate calculation unit 200G, converted image selection unit 200H, viewpoint image generation unit 200I, and display control unit 200J (see FIG. 3 )) that perform various processes is the following various processors. The various processors include a CPU (Central Processing Unit), which is a general-purpose processor that executes software (programs) and functions as various processing units, a programmable logic device (PLD), such as an FPGA (Field Programmable Gate Array), whose circuit configuration can be changed after manufacture, and a dedicated electrical circuit, such as an ASIC (Application Specific Integrated Circuit), which is a processor having a circuit configuration designed specifically for executing specific processes.
[0113] A single processing unit may be configured with one of these various processors, or may be configured with two or more processors of the same or different types (e.g., multiple FPGAs, or a combination of a CPU and an FPGA). Multiple processing units may also be configured with a single processor. Examples of multiple processing units configured with a single processor include: a first configuration, as typified by client or server computers, in which a single processor is configured with a combination of one or more CPUs and software, and this processor functions as multiple processing units; and a second configuration, as typified by system-on-chip (SoC), in which a processor is used to realize the functions of an entire system including multiple processing units on a single IC (Integrated Circuit) chip. In this way, the various processing units are configured with one or more of the above-mentioned various processors as a hardware structure.
[0114] Furthermore, the hardware structure of these various processors is, more specifically, an electric circuit made up of a combination of circuit elements such as semiconductor elements.
[0115] The above-described configurations and functions can be realized by any hardware, software, or a combination of both. For example, the present invention can be applied to a program that causes a computer to execute the above-described processing steps (processing procedures), a computer-readable recording medium (non-transitory recording medium) on which such a program is recorded, or a computer on which such a program can be installed.
[0116] <Modification 1> Next, a description will be given of Modification 1. In this example, the captured image acquisition unit 200A acquires a group of images acquired by the drone 12. Specifically, the captured image acquisition unit 200A acquires a group of images made up of captured images IM in which the target object is captured, and a group of images made up of captured images IN in which the target object is not captured.
[0117] FIG. 14 is a diagram illustrating the captured image acquired by the captured image acquisition unit of this example.
[0118] The captured image acquisition unit 200A acquires a plurality of captured images IM that show the house T and a plurality of captured images IN that do not show the house T. The drone 12 performs aerial photography along the flight path Z as described in Fig. 4, for example, and therefore does not necessarily capture the house T. Therefore, the captured image acquisition unit 200A first acquires all captured images taken by the drone 12, and then extracts the captured images IM that show the house T.
[0119] The captured image acquisition unit 200A can use various methods to extract the captured image IM that shows the house T. For example, the captured image acquisition unit 200A may determine whether or not the target object is present in the field of view of the camera 14 based on the coordinates within the first camera field of view, and if the target object is present in the field of view of the camera 14, extract the captured image IM as containing the house T.
[0120] In this way, the photographed image acquisition unit 200A first acquires all images acquired by the drone 12. After that, the photographed image acquisition unit 200A extracts the photographed images IM that show the house T. This eliminates the need for the user to select the photographed images and input them into the photographed image acquisition unit 200A, allowing for efficient work.
[0121] <Modification 2> Next, a description will be given of Modification 2. In this example, the viewpoint image generating section 200I generates a new viewpoint image between viewpoint images by interpolating the viewpoint images based on adjacent viewpoint images.
[0122] FIG. 15 is a diagram illustrating the viewpoint image generating unit 200I of this example.
[0123] 11, reference numeral 54 indicates a viewpoint image 54A of the north face of the house T, a viewpoint image 54B of the east face of the house T, a viewpoint image 54C of the south face of the house T, and a viewpoint image 54D of the west face of the house T. The viewpoint image generation unit 200I, for example, interpolates between the viewpoint image 54A of the north face and the viewpoint image 54B of the east face to create a viewpoint image 55 of the northeast face of the house T. Note that a known image processing technique is used for the interpolation by the viewpoint image generation unit 200I.
[0124] In this way, the viewpoint image generating unit 200I generates new viewpoint images by interpolating based on the generated viewpoint images, thereby allowing the user to observe the house T from more viewpoints.
[0125] <Others> In the above explanation, a house T has been described as an example of a target object, but the target object of the present invention is not limited to the house T. For example, the present invention can be applied to artificial objects such as buildings and structures other than houses, and natural objects such as mountains and rivers as target objects.
[0126] An example of an application of the present invention is when an image of a building is acquired after a disaster occurs to create a disaster damage certificate. The present invention is not limited to when a disaster occurs, but can also be used to generate an image for observing a target object in other cases.
[0127] Although examples of the present invention have been described above, it goes without saying that the present invention is not limited to the above-described embodiments, and various modifications are possible within the scope of the present invention.
[0128] 12: Drone 13: Gimbal head 14: Camera 16: Remote controller 20: Viewpoint image generation device 22: Network 30A: Virtual camera 30B: Virtual camera 30C: Virtual camera 30D: Virtual camera 42: Image to be transformed 46: Image to be transformed 200: Processor 200A: Photographed image acquisition unit 200B: Attitude acquisition unit 200C: Target object coordinate acquisition unit 200D: First camera field of view coordinate calculation unit 200E: Observation viewpoint coordinate generation unit 200F: Observation direction determination unit 200G: Second camera field of view coordinate calculation unit 200H: Image to be transformed selection unit 200I: Viewpoint image generation unit 200J: Display control unit 210: Memory 220: Database 230: Display device 240 : Input / output interface 250 : Operation unit
Claims
1. A viewpoint image generation device including a processor, wherein the processor acquires a group of images composed of images having an overlap with each other by continuously photographing an area including a target object; extracts a plurality of images obtained by photographing the target object from different viewpoints from the group of images; acquires the positions and postures of the cameras that photographed the plurality of images; acquires target object coordinates that are the coordinates of the target object; calculates first in-camera field-of-view coordinates of the target object in the camera based on the position, the posture, and the target object coordinates of the camera; generates observation viewpoint coordinates that are the coordinates of a plurality of virtual observation viewpoints based on the target object coordinates; determines the observation directions of the target object corresponding to the plurality of virtual observation viewpoints; calculates second in-camera field-of-view coordinates of the target object when the camera is arranged at the virtual observation viewpoint based on the observation viewpoint coordinates, the observation directions, and the target object coordinates; selects a transformed image to be transformed that undergoes projective transformation from the plurality of images based on the virtual observation viewpoint and the observation directions; projects and transforms one of the transformed images based on the first in-camera field-of-view coordinates and the second in-camera field-of-view coordinates to generate one viewpoint image observed from the virtual observation viewpoint. Viewpoint image generation device.
2. The processor extracts the target object coordinates from map information of the target object, according to the viewpoint image generation device of Claim 1.
3. The map information has polygon data of the target object, according to the viewpoint image generation device of Claim 2.
4. The projective transformation includes magnification transformation or reduction transformation, according to the viewpoint image generation device of Claim 1.
5. The plurality of virtual observation viewpoints are set on the circumference of a circle enclosing the target object, according to the viewpoint image generation device of Claim 1.
6. The processor identifies the plurality of virtual observation viewpoints from a group of polygon coordinates defining the outer shape of the target object, according to the viewpoint image generation device of Claim 5.
7. The height of the circle with respect to the target object, or the diameter or radius of the circle is changeable, according to the viewpoint image generation device of Claim 5.
8. The viewpoint image generation apparatus according to claim 1, wherein the selection of the image to be transformed for performing the projective transformation is performed based on the virtual observation viewpoint, the position of the camera, the observation direction, and the orientation of the camera.
9. The viewpoint image generation apparatus according to claim 8, wherein the processor calculates a projective transformation matrix from at least four points of the first in-camera field coordinates and four corresponding points of the second in-camera field coordinates, and performs projective transformation on the image to be transformed.
10. The viewpoint image generation apparatus according to claim 1, wherein the processor generates an interpolated new viewpoint image between the viewpoint images based on adjacent viewpoint images.
11. The viewpoint image generation apparatus according to claim 1, wherein the processor displays the viewpoint images in the order of revolving around the target object.
12. The viewpoint image generation apparatus according to claim 1, wherein the processor displays the viewpoint images in response to a change in the virtual observation viewpoint that is performed in the order of revolving around the target object.
13. A step of acquiring an image group composed of images having an overlap with each other by continuously photographing an area including a target object; A step of extracting a plurality of images of the target object photographed from different viewpoints from the image group; A step of acquiring the positions and orientations of the cameras that photographed the plurality of images; A step of acquiring target object coordinates that are the coordinates of the target object; A step of calculating first in-camera field coordinates of the target object in the camera based on the position, the orientation, and the target object coordinates of the camera; A step of generating observation viewpoint coordinates that are coordinates of a plurality of virtual observation viewpoints based on the target object coordinates; A step of determining an observation direction of the target object corresponding to the plurality of virtual observation viewpoints; A step of calculating second in-camera field coordinates of the target object when the camera is arranged at the virtual observation viewpoint based on the observation viewpoint coordinates, the observation direction, and the target object coordinates; Selecting an image to be transformed for performing projective transformation from the plurality of images based on the virtual observation viewpoint and the observation direction; A step of performing projective transformation on one of the images to be transformed based on the first in-camera field coordinates and the second in-camera field coordinates to generate one viewpoint image observed from the virtual observation viewpoint; A viewpoint image generation method including the above steps. [
14. ] A step of obtaining an image group composed of images having an overlap with each other by continuously photographing an area including a target object; A step of extracting a plurality of images in which the target object is photographed from different viewpoints from the image group; A step of obtaining the positions and postures of the cameras that photographed the plurality of images; A step of obtaining target object coordinates that are the coordinates of the target object; A step of calculating the in-camera first camera field-of-view coordinates of the target object based on the position, the posture, and the target object coordinates of the camera; A step of generating observation viewpoint coordinates that are the coordinates of a plurality of virtual observation viewpoints based on the target object coordinates; A step of determining the observation direction of the target object corresponding to the plurality of virtual observation viewpoints; A step of calculating the in-camera second camera field-of-view coordinates of the target object when the camera is arranged at the virtual observation viewpoint based on the observation viewpoint coordinates, the observation direction, and the target object coordinates; Selecting a transformed image to be subjected to projective transformation from the plurality of images based on the virtual observation viewpoint and the observation direction; A step of generating one viewpoint image observed from the virtual observation viewpoint by performing projective transformation on one of the transformed images based on the first camera field-of-view coordinates and the second camera field-of-view coordinates; A program for causing the above to be executed. [
15. ] A non-transitory and computer-readable recording medium on which the program according to claim 14 is recorded.