Image processing device, image processing method and program
The image processing device addresses alignment inaccuracies by optimizing perspective projection transformation parameters for precise spatial alignment of aerial images with maps, enhancing mapping accuracy.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2025-12-19
- Publication Date
- 2026-03-11
AI Technical Summary
Existing image processing methods for aligning aerial images with maps suffer from inaccuracies due to discrepancies between calculated and actual camera positions and attitudes, leading to poor alignment between spatial positions and captured images.
An image processing device that uses processors to acquire images and three-dimensional position information, sets perspective projection transformation parameters, evaluates the degree of correspondence between line segments, and adjusts parameters for accurate alignment by searching for optimal values.
Achieves high-accuracy alignment of spatial positions with captured images by automatically determining camera matrix parameters, ensuring precise mapping of features onto images.
Smart Images

Figure 2026042796000001_ABST
Abstract
Description
[Technical Field]
[0001] The present disclosure relates to an image processing device, an image processing method, and a program, and more particularly to an image processing technique including processing for associating an image captured by a camera with the spatial position of a subject range. [Background technology]
[0002] Patent Document 1 describes a captured image processing method for capturing images of the ground surface from a camera mounted on an airborne aircraft and identifying the conditions present on the ground surface. The method described in Patent Document 1 identifies the aerial capture position in three dimensions, calculates and obtains the capture range of the captured ground surface, transforms the captured image to match the capture range, and then overlays it on a map in a map information system for display. [Prior art documents] [Patent documents]
[0003] [Patent Document 1] Japanese Patent Application Laid-Open No. 2003-316259 Summary of the Invention [Problem to be solved by the invention]
[0004] The technology described in Patent Document 1 calculates the shooting range by identifying the camera position and camera attitude from output signals obtained from detection units such as an aircraft position detection unit, an aircraft attitude detection unit, and a camera attitude detection unit provided on the aircraft, and aligns the image with a map. However, in an actual system, when calculations are performed based on the camera position and attitude determined from output signals obtained from detection units such as an aircraft position detection unit, an aircraft attitude detection unit, and a camera attitude detection unit, there is a large discrepancy between the actual shooting range and the calculation results, resulting in poor accuracy in aligning the image with the map.
[0005] The present disclosure has been made in consideration of these circumstances, and aims to provide an image processing device, an image processing method, and a program that are capable of highly accurate alignment between the spatial position of the range to be photographed and the captured image. [Means for solving the problem]
[0006] An image processing device according to one embodiment of the present disclosure includes one or more processors and one or more memories storing a program to be executed by the one or more processors, and the one or more processors execute instructions of the program to acquire a captured image captured using a camera, acquire three-dimensional position information indicating the positions of multiple specific points in the space of the captured range, set parameter values for a perspective projection transformation that converts the three-dimensional position information into two-dimensional image coordinates based on the shooting conditions of the captured image, convert the position information of the multiple specific points into image coordinate data using the perspective projection transformation, evaluate the degree of correspondence between a first line segment extracted based on the image coordinate data obtained by the transformation and a second line segment extracted from the captured image, change the parameter values of the perspective projection transformation and evaluate the degree of correspondence multiple times, and associate the captured image with the positions of the multiple specific points based on the results of the multiple evaluations.
[0007] According to the image processing device of this aspect, one or more processors set and change the values of parameters of the perspective projection transformation based on the shooting conditions, and for each transformation result, evaluate the degree of correspondence between a first line segment extracted from the data of the transformation result and a second line segment extracted from the captured image, and search for parameter values. This makes it possible to determine parameter values with a good evaluation of the degree of correspondence, and to align the position of a specific point in the space of the shooting range with the captured image captured by the camera with high accuracy.
[0008] The "photography conditions" include, for example, at least one condition related to the position and attitude of the camera at the time of photography. The photographed image may be an image photographed from the air. The term "aerial" includes the concept of "above the sky." An image photographed using a camera mounted on an aircraft is an example of an "image photographed from the air."
[0009] The specific points may be points in the geographical space of the imaging target area, points that specify the geographical position of a feature such as a building or a road, or virtual points that specify the position of a roof estimated from the height of a building.
[0010] In an image processing device according to another aspect of the present disclosure, one or more processors can be configured to acquire map data corresponding to the range to be photographed, and acquire position information of multiple specific points from the map data.
[0011] In an image processing device according to another aspect of the present disclosure, the map data may include latitude, longitude, and altitude data, and the one or more processors may convert the map data into Cartesian coordinate data. "Altitude" includes the concept of elevation. When the location information included in the map data is coordinate data in a geographic coordinate system, it is preferable that the one or more processors convert the geographic coordinate data into Cartesian coordinate data.
[0012] In an image processing device according to another aspect of the present disclosure, the plurality of specific points may include points that specify the shape of the house, including points that define the perimeter of the house and points that specify the height of the house.
[0013] In an image processing device according to another aspect of the present disclosure, the plurality of specific points may include points that specify the position of a road.
[0014] In an image processing device according to another aspect of the present disclosure, the transformation matrix used for the perspective projection transformation may include multiple parameters, and one or more processors may be configured to perform multiple evaluations of the degree of similarity by changing the combination of values of the multiple parameters.
[0015] The plurality of parameters may be parameters relating to the position and orientation of the camera that captured the captured image.
[0016] In an image processing device according to another aspect of the present disclosure, the captured image is an image captured using a camera mounted on an aircraft, and one or more processors can be configured to acquire camera position information indicating the position of the camera at the time the captured image was captured and attitude information indicating the attitude of the camera at the time of capture, and determine a search range in which to search for parameter values based on the camera position information and attitude information.
[0017] In an image processing device according to another aspect of the present disclosure, the camera position information can include latitude, longitude, and altitude data, and the attitude information can include azimuth angle, tilt angle, and roll angle data indicating the inclination from the horizontal.
[0018] In an image processing device according to another aspect of the present disclosure, the camera position information and attitude information can be configured to be obtained from sensor data obtained by a sensor disposed on at least one of the camera and the aircraft.
[0019] In an image processing device according to another aspect of the present disclosure, one or more processors may be configured to weight the evaluation of the degree of match differently between the central part and the peripheral part of the captured image. For example, when the accuracy of alignment in the central part of the captured image is emphasized, it is preferable to weight the evaluation of the central part relatively more heavily than the evaluation of the peripheral part.
[0020] In an image processing device according to another aspect of the present disclosure, the one or more processors may be configured to select parameter values that provide the highest degree of match based on the results of multiple evaluations. According to this aspect, parameter values for perspective projection transformation that provide good registration accuracy can be automatically selected.
[0021] In an image processing device according to another aspect of the present disclosure, one or more processors may be configured to generate a composite image by superimposing a first line segment generated using a perspective projection transformation defined by the values of selected parameters on the captured image.
[0022] In an image processing device according to another aspect of the present disclosure, one or more processors may be configured to perform a process of displaying the top evaluation results from multiple evaluations, and to accept an instruction to select one of the top evaluation results.
[0023] According to this aspect, a plurality of results with the highest evaluation scores are presented to the user, and the user can select one result from among these that he or she deems appropriate.
[0024] In an image processing device according to another aspect of the present disclosure, one or more processors may be configured to generate a composite image by superimposing a first line segment generated using a perspective projection transformation defined by a parameter value corresponding to the selected result on the captured image in accordance with the received instruction.
[0025] In an image processing device according to another aspect of the present disclosure, the plurality of specific points may include points that specify the shape of a house, and the composite image may be an image in which a figure indicating the area of the house is superimposed on the captured image using a first line segment.
[0026] In an image processing device according to another aspect of the present disclosure, one or more processors can be configured to accept input of an instruction to move a figure indicating the area of a house displayed superimposed on a captured image, and to move the figure on the captured image in accordance with the input instruction.
[0027] In an image processing device according to another aspect of the present disclosure, the one or more processors may be configured to cut out an image portion of a house surrounded by a graphic from the captured image.
[0028] According to this aspect, it is possible to accurately extract image portions of individual houses from the captured image.
[0029] An image processing device according to another aspect of the present disclosure can be configured to include a display unit that displays the results of associating a captured image with the positions of multiple specific points, and an input unit that inputs instructions from a user.
[0030] An image processing method according to another aspect of the present disclosure is an image processing method executed by one or more processors, the one or more processors including: acquiring a captured image captured using a camera; acquiring three-dimensional position information indicating the positions of multiple specific points in the space of the captured range; setting parameter values of a perspective projection transformation that converts the three-dimensional position information into two-dimensional image coordinates based on the shooting conditions of the captured image; converting the position information of the multiple specific points into image coordinate data using the perspective projection transformation; evaluating the degree of correspondence between a first line segment extracted based on the image coordinate data obtained by the transformation and a second line segment extracted from the captured image; changing the parameter values of the perspective projection transformation and evaluating the degree of correspondence multiple times; and associating the captured image with the positions of the multiple specific points based on the results of the multiple evaluations.
[0031] A program according to another aspect of the present disclosure enables a computer to implement the following functions: acquire an image captured using a camera; acquire three-dimensional positional information indicating the positions of multiple specific points in the space of the range to be captured; set parameter values for perspective projection transformation that converts the three-dimensional positional information into two-dimensional image coordinates based on the shooting conditions of the captured image; convert the positional information of the multiple specific points into image coordinate data using the perspective projection transformation; evaluate the degree of correspondence between a first line segment extracted based on the image coordinate data obtained by the transformation and a second line segment extracted from the captured image; change parameter values for the perspective projection transformation to evaluate the degree of correspondence multiple times; and associate the captured image with the positions of the multiple specific points based on the results of the evaluations performed multiple times. [Effects of the Invention]
[0032] According to the present disclosure, it is possible to align the spatial position of the imaging target range with the captured image with high accuracy. [Brief explanation of the drawings]
[0033] [Figure 1]FIG. 1 is a schematic diagram showing an example of the configuration of a captured image processing system according to an embodiment. [Figure 2] FIG. 2 is a block diagram illustrating an example of the electrical configuration of a drone equipped with a camera. [Figure 3] FIG. 3 shows an example of a captured image corresponding to map data including location data indicating the location of a house. [Figure 4] FIG. 4 shows an example of a composite image in which the positions of houses and roads, which are the result of converting map data into image coordinates by applying sensor data to the camera matrix parameters, are superimposed on the captured image. [Figure 5] FIG. 5 is a block diagram showing an example of the hardware configuration of the image processing apparatus according to the embodiment. [Figure 6] FIG. 6 is a functional block diagram showing the functional configuration of the image processing device. [Figure 7] FIG. 7 is an explanatory diagram of the definition of six parameters that indicate the camera position and orientation. [Figure 8] FIG. 8 is an explanatory diagram illustrating an example of the relationship between a three-dimensional space coordinate system converted into coordinates with the projection center as the origin and an image coordinate system. [Figure 9] FIG. 9 is an explanatory diagram showing an example of automatic alignment by line segment matching. [Figure 10] FIG. 10 is an explanatory diagram showing an example of line segment extraction when the azimuth angle value is changed, and an example of the number of matching line segments. [Figure 11] FIG. 11 shows an example of a composite image obtained by aligning a photographed image with map information as a result of automatic parameter value search using line segment matching. [Figure 12] FIG. 12 is a flowchart showing an example of the flow of processing in the image processing device. [Figure 13] FIG. 13 is a flowchart showing an example of the flow of processing in the image processing device. DETAILED DESCRIPTION OF THE INVENTION
[0034] DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS Preferred embodiments of the present invention will now be described in detail with reference to the accompanying drawings. In this specification, the same components are designated by the same reference numerals, and redundant explanations will be omitted where appropriate.
[0035] FIG. 1 is a schematic diagram showing an example configuration of a photographic image processing system 10 according to an embodiment. The photographic image processing system 10 includes a drone 12 for aerial photography, a camera 14 mounted on the drone 12, a remote controller 16, and an image processing device 20. The drone 12 is an unmanned aerial vehicle remotely controlled using the remote controller 16. The drone 12 may have an autopilot function for flying according to a program. The drone 12 is an example of an "air vehicle" in this disclosure.
[0036] The camera 14 is mounted on the drone 12 via a gimbal head 13. The camera 14 includes an optical system, an image sensor, and a signal processing circuit (not shown). The optical system includes one or more lenses such as a focus lens. The image sensor may be, for example, a CCD (Charge Coupled Device) image sensor or a CMOS (Complementary Metal-Oxide Semiconductor) image sensor.
[0037] The camera 14 generates digital image data of the photographed subject by processing signals obtained from the image sensor using a signal processing circuit. The digital image data generated by the camera 14 may be referred to as a "captured image." Images captured using the camera 14 may be stored in an internal storage device built into the drone 12 and / or a storage device such as a memory card removably attached to the drone 12. Images captured using the camera 14 may also be transferred to the remote controller 16, the image processing device 20, and other terminal devices 24 using wireless communication.
[0038] The remote controller 16 is a transmitter that controls the operation of the camera 14 and the drone 12 via wireless communication. The wireless communication format may be a wireless local area network (LAN) format, a communication format using radio waves in the 2.4 GHz or 5.7 GHz band, or a format using a mobile communication network. The communication format for the control signals for operating the drone 12 and the communication format for transferring images captured by the camera 14 may be different or may be the same.
[0039] The remote controller 16 includes left and right sticks for controlling the flight operation of the drone 12, a lever for operating the gimbal head 13, a shooting button for instructing the camera 14 to take a picture, and a shooting mode button for switching between video shooting and still image shooting. By employing a touch panel display for the display 16A, the shooting button and other operation buttons can be realized by the touch panel display.
[0040] The live video captured by the camera 14 can be displayed on the display 16A of the remote controller 16. The remote controller 16 can also grasp the status of the drone 12, such as its flight position and flight speed, in real time based on data from various sensors provided on the drone 12. Flight information indicating the status of the drone can be displayed on the display 16A.
[0041] 1 is an example of an image captured using the camera 14. In this embodiment, at least one still image is captured from the air, and the captured image IM is processed by the image processing device 20.
[0042] The image processing device 20 is configured using a computer. The computer applied to the image processing device 20 may be a server, a personal computer, or a workstation.
[0043] The image processing device 20 can perform data communication with the remote controller 16 and the terminal device 18 via a network 22. The network 22 may be a local area network or a wide area network. The image processing device 20 acquires various information from the drone 12 and the camera 14. The image processing device 20 can also acquire map data of the imaging target range from a geographic information system (not shown) via the network 22. The map data may be acquired in advance before imaging, or may be acquired after imaging.
[0044] The terminal device 24 may be a mobile information terminal such as a smartphone or a tablet terminal. The terminal device 24 includes a display 24A. The terminal device 24 may have the functions of the remote controller 16. The terminal device 24 may also have the processing functions of the image processing device 20.
[0045] [Example of camera drone configuration] 2 is a block diagram that schematically illustrates an example of the electrical configuration of a drone 12 equipped with a camera 14. The drone 12 includes a GPS (Global Positioning System) receiver 30, a barometric pressure sensor 32, a direction sensor 34, a gyro sensor 36, and a motor 38. The motor 38 is a power source that rotates a rotor (not shown), and the drone 12 includes multiple motors 38 that drive multiple rotors.
[0046] The GPS receiver 30 acquires location information including the latitude and longitude of the drone 12. The barometric pressure sensor 32 detects the barometric pressure in the drone 12. The drone 12 may acquire its altitude based on the barometric pressure detected using the barometric pressure sensor 32. The term "acquire" includes the concept of generating information by data processing such as calculation. The latitude, longitude, and altitude of the drone 12 constitute the location information of the drone 12 and the camera 14.
[0047] The orientation sensor 34 may be, for example, a geomagnetic sensor, and may detect the azimuth angle at which the lens of the camera 14 is facing.
[0048] The gyro sensor 36 detects a roll angle representing a rotation angle about a roll axis, a pitch angle representing a rotation angle about a pitch axis, and a yaw angle representing a rotation angle about a yaw axis. The drone 12 acquires attitude information of the drone 12 based on the rotation angles acquired using the gyro sensor 36. Note that some or all of the sensors, such as the GPS receiver 30, the barometric pressure sensor 32, the orientation sensor 34, and the gyro sensor 36, may be located on the camera 14 side.
[0049] The drone 12 includes a processor 40, a storage device 42, and a communication interface 44. The storage device 42 may be a memory, an internal storage device, an external storage device, or a combination thereof. The processor 40 serves as a flight controller and performs various calculations necessary for flight control of the drone 12 based on sensor data obtained from various sensors.
[0050] The communication interface 44 is a communication unit that performs wireless communication with the remote controller 16, etc. The communication interface 44 may also include a communication terminal that supports wired communication. Furthermore, the drone 12 includes a battery and a battery charging terminal (not shown).
[0051] <<Explanation of technical issues in processing captured images IM>> Here, we will explain an example of processing to identify the positions of houses in a captured image IM obtained by photographing the ground from the air. In this case, as shown in Figure 3, based on map data MP containing position data indicating the positions of the houses and the captured image IM, the positions on the captured image IM corresponding to each of multiple specific points indicated by black dots in the map data MP are identified.
[0052] In the map data MP, each house is assigned a house ID (Identification) as an identification code to identify the house, and location data indicating the respective positions of multiple specific points that make up the perimeter of the house is recorded, linked to the house ID. The location data for each specific point is three-dimensional data of latitude, longitude, and altitude. In Japan, map data MP containing such geographic coordinate data can be obtained, for example, from the basic map information provided by the Geospatial Information Authority of Japan. Alternatively, such map data MP can be obtained from the OpenStreetMap database.
[0053] The problem of identifying the position on the captured image IM that corresponds to a specific point on the map data MP can be understood as the problem of finding the correspondence between three-dimensional spatial coordinates and two-dimensional image coordinates.
[0054] About the camera matrix The problem of finding the correspondence between three-dimensional spatial coordinates and two-dimensional image coordinates can be solved by finding the camera matrix as the transformation matrix for perspective projection transformation using the following formula based on the camera model.
[0055] Image coordinates (u,v) = camera matrix * 3D coordinates (x,y,z) The camera matrix can be expressed as the product of an intrinsic parameter matrix and an extrinsic parameter matrix. The extrinsic parameter matrix is a matrix that converts three-dimensional coordinates (world coordinates) into camera coordinates. The extrinsic parameter matrix is a matrix determined by the camera position and orientation (photography angle) at the time of shooting, and includes translation parameters and rotation parameters.
[0056] The internal parameter matrix is a matrix that converts camera coordinates into image coordinates, and is determined by the specifications of the camera 14, such as the focal length of the camera, the sensor size and aberration (distortion) of the image sensor, etc.
[0057] By converting from 3D coordinates (x,y,z) to camera coordinates using an extrinsic parameter matrix, and then converting from camera coordinates to image coordinates (u,v) using an intrinsic parameter matrix, 3D coordinates (x,y,z) can be mapped (converted) to image coordinates (u,v).
[0058] The intrinsic parameter matrix can be specified in advance, whereas the extrinsic parameter matrix depends on the camera position and orientation at the time of shooting, and therefore must be set for each captured image.
[0059] The camera matrix can be calculated if there are six or more corresponding points between the 3D coordinates in the real 3D space and the image coordinates in the captured image. However, manually specifying these multiple corresponding points is a time-consuming process.
[0060] In this regard, the image processing device 20 of this embodiment can automatically determine the camera matrix (transformation matrix) based on the shooting conditions when the captured image IM was captured, without the need for a human to specify corresponding points. A specific processing method will be described in detail later.
[0061] <Issues when using sensor data for the extrinsic parameter matrix> It is conceivable to calculate the extrinsic parameter matrix using sensor data (sensor values) obtained from various sensors such as the GPS receiver 30, a direction sensor, and a gyro sensor mounted on the drone 12 as data indicating the position and attitude of the camera 14. However, there is a problem in that the camera matrix calculated using actual sensor data cannot correctly map the position on the map onto the captured image.
[0062] FIG. 4 shows an example of a composite image in which map data is converted into image coordinates using a camera matrix that uses sensor data as parameter values, and the positions of houses and roads are superimposed on the captured image. In FIG. 4, each of the multiple polygons PG superimposed on the captured image IMs represents the perimeter of a house on the map converted using a camera matrix that uses sensor data as parameter values. Furthermore, the line RL superimposed on the captured image IMs represents a road on the map converted using the same camera matrix. As shown in FIG. 4, the polygons PG and the line RL are significantly offset from the positions of the houses and roads in the captured image IMs. A camera matrix that uses sensor data (sensor values) as parameters for the position and orientation of the camera 14 contains errors in the sensor data, making it impossible to correctly map houses and other structures on the map onto the captured image IMs.
[0063] Overview of the image processing device 20 according to this embodiment The image processing device 20 automatically searches for the parameter values of the camera matrix based on the sensor data at the time of shooting, and finds the optimal parameter values, i.e., the camera matrix that can accurately match (align) positions on the map with positions on the captured image.
[0064] In the process of searching for the parameter values of the camera matrix, the image processing device 20 assigns parameter values based on the sensor data values, converts the map data into image coordinates using the camera matrix of those parameter values, evaluates the degree of correspondence between the conversion result and the position on the captured image, selects the parameter values that give the highest evaluation score, and determines the camera matrix.
[0065] In the process of evaluating the degree of match, the image processing device 20 extracts line segments such as the perimeter of houses and roads from both the result of converting map data into image coordinates and the captured image, and calculates an evaluation value that quantitatively evaluates the degree of match between the line segments. A line segment is specified by the coordinates of two points (start and end points). The "degree of match" here may be the degree of match, including an acceptable range, for at least one, and preferably multiple, of the distance between the line segments, the difference in length of the line segments, and the difference in inclination angle of the line segments.
[0066] 5 is a block diagram showing an example of the hardware configuration of the image processing device 20. The image processing device 20 includes a processor 202, a computer-readable medium 204 which is a non-transitory tangible entity, a communication interface 206, and an input / output interface 208.
[0067] The processor 202 includes a central processing unit (CPU). The processor 202 may also include a graphics processing unit (GPU). The processor 202 is connected to a computer-readable medium 204, a communication interface 206, and an input / output interface 208 via a bus 210.
[0068] The image processing device 20 may include an input device 214 and a display device 216. The input device 214 and the display device 216 are connected to the bus 210 via an input / output interface 208. The input device 214 is configured by, for example, a keyboard, a mouse, a multi-touch panel, or other pointing device, or a voice input device, or an appropriate combination of these. The input device 214 is an example of an "input unit" in the present disclosure.
[0069] The display device 216 is configured by, for example, a liquid crystal display, an organic electro-luminescence (OEL) display, a projector, or an appropriate combination of these. The display device 216 is an example of a "display unit" in the present disclosure.
[0070] The computer-readable medium 204 includes a memory serving as a primary storage device and a storage serving as an auxiliary storage device. The computer-readable medium 204 may be, for example, a semiconductor memory, a hard disk drive (HDD), a solid state drive (SSD), or a combination of these. The computer-readable medium 204 stores various programs, data, and the like, including an image processing program 220 and a display control program 250.
[0071] By executing instructions of the image processing program 220, the processor 202 functions as processing units such as an information acquisition unit 222, a coordinate conversion unit 224, a camera matrix parameter setting unit 226, a perspective projection conversion unit 228, a line segment extraction unit 230, a coincidence evaluation unit 234, an optimal parameter value selection unit 236, an image synthesis unit 238, a position adjustment unit 240, and a cropping unit 242. The computer-readable medium 204 includes a map information storage unit 260, a captured image storage unit 262, and a sensor data storage unit 264 that store map information, captured images, and sensor data acquired via the information acquisition unit 222.
[0072] 6 is a functional block diagram showing the functional configuration of the image processing device 20. The information acquisition unit 222 includes a map information acquisition unit 222A, a shooting condition acquisition unit 222B, and a captured image acquisition unit 222C. The map information acquisition unit 222A acquires map information 100. The map information 100 may be, for example, basic map information from the Geospatial Information Authority of Japan, map information from OpenStreetMap, or a combination of these.
[0073] The photographing condition acquisition unit 222B acquires camera position information 112 and attitude information 113 of the camera 14 as photographing conditions when the photographed image 110 was photographed. The photographed image 110 is linked to the camera position information 112 and attitude information 113 at the time of photographing. The camera position information 112 may be position information obtained from the GPS receiver 30 of the drone 12, and includes latitude, longitude, and altitude data. The altitude data in the camera position information 112 may be calculated based on data obtained from the barometric pressure sensor 32. The attitude information 113 includes data on the azimuth angle, tilt angle, and roll angle obtained from the orientation sensor 34 and gyro sensor 36. The tilt angle is the angle of the camera toward the ground and is synonymous with the "depression angle."
[0074] The coordinate conversion unit 224 converts position data including latitude and longitude data into Cartesian coordinate data. The Cartesian coordinate system may be, for example, the Universal Transverse Mercator (UTM) coordinate system. The coordinate conversion unit 224 converts three-dimensional map data including latitude, longitude, and altitude data into UTM coordinates. The coordinate conversion unit 224 also converts latitude and longitude data included in the camera position information 112 at the time of shooting into Cartesian coordinate data (xc, yc), and passes the data to the camera matrix parameter setting unit 226.
[0075] The camera matrix parameter setting unit 226 determines a search range for the parameter values of the camera matrix Mc based on the camera position information 112 and attitude information 113 acquired via the shooting condition acquisition unit 222B, and sets and changes the parameter values within the search range. The parameters of the camera matrix Mc include the camera position (xc, yc, zc) at the time of shooting, and the azimuth angle θh, tilt angle θt, and roll angle θr at the time of shooting. For the camera position (xc, yc, zc) at the time of shooting, the camera matrix parameter setting unit 226 sets the values of these six parameters. Furthermore, the camera matrix parameter setting unit 226 changes the values of each of these six parameters by a change amount (step size) predetermined for each parameter, and changes the combination of parameter values.
[0076] The perspective projection transformation unit 228 performs perspective projection transformation using the camera matrix Mc of the parameter values set by the camera matrix parameter setting unit 226, and transforms the three-dimensional Cartesian coordinate data (x, y, z) into two-dimensional image coordinates (u, v). The perspective projection transformation unit 228 transforms the three-dimensional Cartesian coordinate data (x, y, z) of each of a plurality of specific points included in the map information 100 into image coordinates (u, v). A transformed map image is obtained by mapping each point represented by the image coordinate data 104 resulting from the transformation by the perspective projection transformation unit 228 onto the image coordinate system.
[0077] The line segment extraction unit 230 includes a first line segment extraction unit 231 and a second line segment extraction unit 232. The first line segment extraction unit 231 performs processing to extract line segments such as the perimeter of a house from map information after perspective projection conversion (hereinafter referred to as a converted map) represented by image coordinate data 104 resulting from conversion by the perspective projection conversion unit 228.
[0078] The second line segment extraction unit 232 performs processing to extract line segments from the photographed image 110. For the processing to extract line segments from the photographed image 110, an existing method such as an LSD (Line Segment Detector) can be applied.
[0079] The matching degree evaluation unit 234 evaluates the degree of matching between the line segment (first line segment) extracted by the first line segment extraction unit 231 and the line segment (second line segment) extracted by the second line segment extraction unit 232. The matching degree evaluation unit 234 includes an evaluation value calculation unit 235 that calculates an evaluation value that quantifies the degree of matching between the first line segment and the second line segment.
[0080] The degree of match refers to the degree (level) of match, and is not limited to a perfect match, but may be a degree of match that is determined to be roughly the same with allowable differences. Various methods can be applied to quantify the degree of match between two line segments being compared. For example, the evaluation value calculation unit 235 may calculate the evaluation value by quantifying at least one characteristic item among the position, length, and inclination of the line segments.
[0081] The matching evaluation unit 234 comprehensively evaluates the matching of the multiple line segments extracted by the first line segment extraction unit 231 and the second line segment extraction unit 232, and calculates an evaluation value for each combination of parameter values (i.e., for each camera matrix Mc).
[0082] The optimal parameter value selection unit 236 selects the combination of parameter values that will give the highest evaluation score based on the evaluation results of the degree of match for each transformation result of multiple camera matrices Mc in which the parameter values are changed within the search range of parameter values.
[0083] The combination of optimal parameter values selected by the optimal parameter value selection unit 236 determines the camera matrix Mc that can align the captured image 110 and the map information 100 with high accuracy.
[0084] In this way, the three-dimensional coordinate data of the map information 100 is perspectively projected into image coordinates using the camera matrix Mc determined by the automatic parameter value search using line segment matching, and a converted map image 106 aligned with the captured image 110 is generated from the conversion result. The converted map image 106 may include at least one of polygons PG representing the shapes of houses and lines RL representing roads.
[0085] The image synthesis unit 238 performs processing to generate a synthetic image by superimposing the photographed image 110 and the converted map image 106 .
[0086] The display control unit 251 generates data for display on the display device 216. The composite image generated by the image synthesis unit 238 is displayed on the display device 216 via the display control unit 251.
[0087] The position adjustment unit 240 receives an instruction to move, on the captured image 110, each of the polygons PG representing the shapes of individual houses in the converted map image 106 that is displayed superimposed on the captured image 110, and performs processing to move the position of the polygons PG in accordance with the received instruction. "Movement" includes the concepts of translation and rotation. The user can select the polygon to be moved from the input device 214 and input an instruction to move the polygon.
[0088] 《Explanation of perspective projection transformation using camera matrix》 Here, we will describe in detail a calculation method for converting the three-dimensional Cartesian coordinate data (x, y, z) of points constituting the perimeter of the house contained in the map information into coordinates when projected onto the image sensor of the camera 14, i.e., image coordinates (u, v). In the points (x, y, z) constituting the perimeter of the house, x and y are latitude and longitude converted into UTM coordinates, which are Cartesian coordinate systems, and z is altitude. If height information is available for a building such as a house, it is desirable to use that height information to calculate the position of the roof on the image. Furthermore, for houses without height information, the roof position may be calculated assuming a height of, for example, 6 m.
[0089] The camera position at the time of shooting is assumed to be (xc, yc, zc), where xc and yc are the latitude and longitude of the camera position information 112 converted into UTM coordinates, and zc is the altitude.
[0090] The camera attitude during shooting is specified by the azimuth angle θh, tilt angle θt, and roll angle θr. The azimuth angle θh is the angle from north, with north as the reference point. The tilt angle θt is the camera angle (depression angle) toward the ground. The roll angle θr is the inclination from the horizontal.
[0091] FIG. 7 shows an explanatory diagram of the definition of six parameters that indicate the camera position and attitude. In the UTM coordinate system, the x-axis is defined as east and the y-axis is defined as north. In FIG. 7, the position of the camera 14 is defined as Pc(xc, yc, zc). Arrow A indicates the shooting direction of the camera 14.
[0092] The equation for converting the coordinates of the points (x, y, z) that make up the exterior perimeter of the house into the origin of the projection center (i.e., the camera position at the time of shooting) is expressed by the following equation (1).
[0093]
number
[0094] Furthermore, rotation matrices Mh, Mt, and Mr are defined as follows:
[0095]
number
[0096]
number
[0097]
number
[0098] The coordinates of the points constituting the exterior perimeter of the house, with the projection center as the origin, are converted into camera coordinates using the following equation (5).
[0099]
number
[0100] The origin of the camera coordinate system is the projection center, the X axis is the horizontal direction of the image sensor, the Y axis is the vertical direction of the image sensor, and the Z axis is the depth direction. Figure 8 shows an example of the relationship between a three-dimensional spatial coordinate system having three axes corresponding to the three-dimensional coordinates (x', y', z') obtained by the coordinate transformation of equation (1) and the image coordinate system of image sensor 140 of camera 14.
[0101] The camera coordinate points (unit: meters) obtained by the above formula (5) are converted to coordinates on the image (unit: pixels) by the following formula (6).
[0102]
number
[0103] In equation (6), f is the focal length, and p is the pixel pitch. The pixel pitch is the distance between pixels of the image sensor 140, and is usually the same in both the vertical and horizontal directions. Uc and Vc are the image center coordinates (in pixel units).
[0104] Example of searching for optimal parameter values An example of the specific procedure for calculating the camera matrix Mc in the image processing apparatus 20 will be described. [Step 1] The processor 202 obtains the camera position and orientation during shooting from the sensor data. The camera position (xc_0, yc_0, zc_0) and orientation (θh_0, θt_0, θr_0) obtained from the sensor data are used as reference values in the search for parameter values.
[0105] [Step 2] The processor 202 sets the search range and the step width during the search for each of the six parameter values of the camera position and orientation. For example, the processor 202 determines that the search range for the x coordinate of the camera position is in the range of ±10 m from the reference value, and the step width is 1 m. That is, the search range for the x coordinate of the camera position is set to "xc_0 - 10 < xc < xc_0 + 10", and the step width during the search is set to 1 (the unit is meter). Xc - 10 indicating the lower limit of the search range is an example of the search lower limit value, and xc + 10 indicating the upper limit of the search range is an example of the search upper limit value.
[0106] For each of the y and z coordinates of the camera position and the parameters of the orientation (θh, θt, θr), the search range and the step width are also set respectively. For example, for the azimuth angle θh, the search range is set to ±45° with respect to the reference value indicated by the sensor data, and the parameter value is changed with a step width of 1°. Different search ranges and step widths can be set for each parameter.
[0107] [Step 3] The processor 202 moves the step width within the respective search ranges for the six parameters of the camera position and orientation, and determines the combination of parameter values. Then, using the determined combination of parameter values (xc, yc, zc), (θh, θt, θr), the three-dimensional position data (latitude, longitude, altitude) of the houses and roads included in the map data is converted into coordinates on a two-dimensional image.
[0108] [Step 4] The processor 202 evaluates the converted map image (converted map image) obtained by mapping the positions of the houses and roads thus converted onto the image and the captured image by line segment matching.
[0109] [Step 5] Processor 202 performs steps 3 and 4 above, varying the camera position and orientation parameter values using all step sizes within the search range of each parameter, and adopts the camera position and orientation parameter values that yield the best line-segment matching evaluation score as the correct camera position and orientation. In this way, the optimal camera matrix is automatically calculated for each captured image, and a converted map image that is accurately aligned with each captured image is obtained.
[0110] It is not necessary to evaluate all possible combinations of parameter values for the search range of each parameter; a local search algorithm such as hill-climbing may be used to find the optimal combination of parameter values.
[0111] Overview of automatic alignment using line segment matching Image IMa shown on the left of Figure 9 is an image obtained by superimposing a transformed map image TMa, which is composed of line segments LS1a indicating the positions of houses and roads that have been perspectively projected using a camera matrix with a certain parameter value, on a photographed line segment image IML, which is composed of line segments LS2 extracted from a photographed image. There is a positional discrepancy between the two images, the transformed map image TMa and the photographed line segment image IML, and it can be seen that the image registration is insufficient.
[0112] On the other hand, image IMb shown on the right side of Figure 9 is an image obtained by superimposing a transformed map image TMb composed of line segment LS1b indicating the positions of houses and roads that have been perspectively projected using a camera matrix in which some of the parameter values of the camera matrix applied to generate image IMa have been changed, and a photographed line segment image IML composed of line segment LS2 extracted from the photographed image. In image IMb, the positions of the two images, the transformed map image TMb and the photographed line segment image IML, generally match, indicating that the images have been properly aligned. Line segment LS1a and line segment LS1b are each an example of a "first line segment" in the present disclosure, and line segment LS2 is an example of a "second line segment" in the present disclosure.
[0113] Here, for simplicity of explanation, the azimuth angle θh is taken as an example of a changed parameter, but in reality, not only one type of parameter but also a combination of values of multiple parameters is changed.
[0114] Assume that the azimuth angle θh of the camera matrix applied to generating image IMa is 122°, and the azimuth angle θh of the camera matrix applied to generating image IMb is 124°.
[0115] When evaluating whether the image registration between the captured image and the converted map image is appropriate (whether the positions of the two images match), the processor 202 quantifies the degree to which the image positions match between the images.
[0116] For example, the processor 202 compares line segments extracted from houses, roads, etc., resulting from the conversion with line segments extracted from a photographed image of the geographic space including the houses and roads, and calculates the number of matching line segments. To evaluate two compared line segments as "matching line segments," it is preferable to define "matching" not only when the two line segments perfectly match, but also to include an acceptable range of difference, and to treat line segments that satisfy the acceptable range as "matching line segments." The acceptable range for determining a match may be defined, for example, with respect to the position of the line segments (the distance between the line segments), the length of the line segments, or the slope of the line segments, or a combination of these.
[0117] The processor 202 performs the calculation for all houses, roads, etc., and adds up the number of matching line segments. This number of matching line segments is an example of an evaluation value.
[0118] The processor 202 repeats the same calculation by changing the parameter values of the camera position and orientation, and selects the parameter values that result in the largest number of matching line segments as the optimal parameter values. This makes it possible to determine a camera matrix that provides high accuracy in aligning the converted map image with the captured image.
[0119] Fig. 10 shows an example of extracted line segments when the value of the azimuth angle θh is changed, and an example of the number of matching line segments. Of the azimuth angles of 120°, 124°, and 128° shown in Fig. 10, the degree of line matching is highest when the azimuth angle is 124°. Line segments enclosed by dashed ellipses in Fig. 10 are evaluated as matching line segments. Using the line segment matching technique shown in Fig. 10, processor 202 calculates an evaluation value of the degree of line segment matching for each combination of six types of parameter values and determines the optimal combination of parameter values.
[0120] 11 shows an example of a composite image obtained by aligning a photographed image with map information as a result of automatic parameter value search using line segment matching. As is clear from a comparison with FIG. 4, the image processing device 20 of this embodiment can align a photographed image with map information with high precision.
[0121] Other Functions of Image Processing Device 20 In addition to the above-described processes, the image processing device 20 may also perform the following processes.
[0122] [1] Weighting function in line segment matching evaluation The processor 202 may place emphasis on alignment in the central part (near the center) of the screen of the captured image, and weight the evaluation of the degree of coincidence of line segments in the central part of the screen with the degree of coincidence of line segments in the peripheral part of the screen, thereby determining an overall evaluation value with emphasis on the degree of coincidence in the central part.
[0123] [2] Fine-tuning function for alignment The processor 202 may be equipped with a manual position adjustment function that automatically matches map data including the locations of houses to the captured image, and then accepts operations to move polygons PG indicating the locations of individual houses on the image, and fine-tunes the location to an even more optimal position according to the user's operations.
[0124] [3] Combining automatic matching with a user interface (UI) Instead of automatically determining the optimal parameter value with the best evaluation score as a result of the parameter value search, the configuration may be such that multiple results with the highest evaluation scores for line segment matching are presented to the user, and the user is allowed to select the one they deem most optimal from among the multiple candidates.
[0125] [4] Measures to speed up automatic matching processing Since evaluating the degree of line segment coincidence for all houses included in the photographed area takes a long processing time, it is possible to limit the houses to be subjected to the line segment matching process. For example, when photographing for the purpose of investigating damage caused by a disaster such as an earthquake or flood, it is possible to use only sturdy buildings for alignment based on attribute information of buildings such as houses. Furthermore, since it is expected that houses will be lost due to fire or the like, it is also possible to use position information of only elements other than houses, such as roads or rivers, for line segment matching.
[0126] [5] Collaboration with house segmentation processing Once the captured image IMs and the map data MP are correctly aligned, the area (partial image) of each house shown in the captured image IMs can be cut out by comparing it with the map data MP. The area of each house may be cut out, for example, by a circumscribing rectangle that contains the area of the house. When cutting out a house, it is desirable to use the position data of the points that make up the exterior perimeter of the house as well as the height data of the house to determine the image coordinates of the roof shape and to determine the entire area of the house including the roof. The cut-out image of the house is linked to the house ID and saved.
[0127] [6] Collaboration with the automatic house damage assessment function The images of houses extracted from the captured images IMs are input to a processing unit of an automatic residential damage assessment AI (Artificial Intelligence) that automatically determines the extent of damage to affected houses, thereby making it possible to streamline damage survey work.
[0128] <<Example of image processing method executed by image processing device 20>> 12 is a flowchart showing an example of the processing flow in the image processing device 20. In step S12, the processor 202 acquires map data of the imaging target range. The processor 202 may acquire map data including the geographic space to be imaged in advance before imaging, or may acquire map data including the imaged geographic space after imaging.
[0129] In step S14, the processor 202 converts the map data of the three-dimensional map, including geographic coordinate data of latitude and longitude, into Cartesian coordinate data (x, y, z) such as UTM coordinates.
[0130] Also, in step S16, the processor 202 acquires the captured image taken by the camera 14. Furthermore, in step S18, the processor 202 acquires sensor data indicating the position and orientation of the camera at the time of capturing the image.
[0131] The processing order of steps S12, S16, and S18 is not particularly limited, and they may be executed in parallel or in a parallel manner.
[0132] After step S18, in step S20, the processor 202 determines a search range for the camera matrix parameters based on the acquired sensor data. For each of the six parameters, the processor 202 determines a lower limit value and an upper limit value of the search range from the reference value indicated by the sensor data. The step size of the parameter value for each parameter may be determined in advance.
[0133] In step S22, the processor 202 sets the value of each parameter within the determined search range. The initial setting value of the parameter may be a reference value indicated by the sensor data, or may be a search lower limit value or a search upper limit value.
[0134] Next, in step S24, the processor 202 converts the Cartesian coordinate data (x, y, z) of multiple specific points contained in the map data into two-dimensional image coordinate data (u, v) by perspective projection transformation using the camera matrix of the set parameter values.
[0135] In step S26, processor 202 extracts line segments from the transformation result. Each point of the image coordinate data of the transformation result is mapped onto coordinates, and by connecting the points with straight lines (line segments) for each house, a polygon including line segments showing the shape of the house can be generated. Furthermore, by connecting multiple points indicating the positions of roads, rivers, etc. with straight lines, it is possible to generate line segments indicating the shape of roads, rivers, etc. In this way, generating line segments based on the image coordinate data of the conversion results of multiple specific points of the conversion results is included in the concept of "extracting" line segments. The line segments extracted from the conversion results are an example of a "first line segment" in the present disclosure.
[0136] Meanwhile, in step S28, the processor 202 extracts line segments from the acquired photographed image. The line segments extracted from the photographed image are an example of the "second line segments" in the present disclosure.
[0137] Next, in step S30, processor 202 evaluates the degree of match between the line segments extracted from the conversion result and the line segments extracted from the captured image. Processor 202 calculates an evaluation value that quantifies the degree of match between the line segments. In calculating the matching between the line segments, processor 202 uses not only position data of points that constitute the ground perimeter of the house, but also building height data and calculates the matching evaluation value using the line segments of the house roof shape.
[0138] In step S32, processor 202 determines whether to end the search for parameter values. If there is a combination of parameter values for each step size within the search range of multiple parameters for which an evaluation value has not been calculated, the determination result in step S32 may be No. If the determination result in step S32 is No, processor 202 proceeds to step S34.
[0139] In step S34, the processor 202 changes the parameter value within the search range, and returns to step S24. The processor 202 executes steps S24 to S34 multiple times until the determination in step S32 is Yes.
[0140] When steps S24 to S34 are repeatedly executed multiple times and evaluation values are calculated for all combinations of parameter values with each step size within the search range of each parameter, the determination result in step S32 may be Yes.
[0141] If the determination result of step S32 is Yes, the processor 202 proceeds to step S36.
[0142] In step S36, processor 202 changes the parameter values and selects the optimal parameter values that provide the highest degree of match based on the multiple evaluation values that are repeatedly calculated. When selecting the optimal parameter values, the parameter values that were actually used in the search may be used, or the maximum value may be estimated by interpolation or the like based on the parameter values that have been discretely changed in increments.
[0143] After step S36, the processor 202 proceeds to step S38 in FIG.
[0144] In step S38, the processor 202 superimposes a transformed map image, generated using the image coordinate data resulting from the perspective projection transformation defined by the optimal parameter values determined by automatic matching, onto the captured image. The transformed map image is precisely aligned with the captured image, and a composite image is obtained in which the positions of houses and other structures included in the map data are properly associated with the captured image.
[0145] In step S40, the processor 202 causes the generated composite image to be displayed on the display device 216. The processor 202 may also cause the generated composite image to be displayed on the display 16A of the remote controller 16 and / or the terminal device 24.
[0146] In step S42, processor 202 accepts an instruction to adjust the position of the figures constituting the converted map image. The figures here include line drawings of polygons PG representing the shapes of individual houses. The user can use a user interface such as input device 214 to select the figures to move and specify the destination position of the figures. If the user determines that position adjustment is not necessary, the user can input an instruction to save the alignment results.
[0147] In step S44, the processor 202 determines whether or not to adjust the position of the graphic. If a graphic to be moved is selected and a destination position is specified, the determination result in step S44 is Yes, and the process proceeds to step S46.
[0148] In step S46, the processor 202 moves the position of the figure based on the received instruction. After step S46, the processor 202 returns to step S44.
[0149] If the determination result in step S44 is No, that is, if further position adjustment is not required, the processor 202 proceeds to step S48.
[0150] In step S48, the processor 202 receives a designation of a partial region to be cut out from the captured image. The partial region to be cut out may be the region of an individual house.
[0151] The user can specify houses to be subject to the cutout process using a UI such as the input device 214. An operation to individually specify a house to be the subject may be accepted, or an area including multiple houses may be specified, and each of the multiple houses included in the specified area may be specified as a house to be subject to the cutout process. In addition to the operation to specify an individual house or the operation to comprehensively specify multiple houses in the specified area, an operation menu such as "select all houses at once" may be provided to specify all houses in the captured image.
[0152] In step S50, the processor 202 determines whether or not to perform cutting. If the determination result in step S50 is Yes, the processor 202 proceeds to step S52.
[0153] In step S52, the processor 202 performs a process of cutting out a partial area corresponding to the image portion of the house from the captured image as specified. The cut-out image of the house is associated with a house ID and stored in the computer-readable medium 204 of the image processing device 20 and / or a storage device (not shown).
[0154] The extracted image of the house is input to, for example, an image recognition device (not shown), and the damage status of the house is automatically determined by image recognition. The image recognition device may be configured to use a learned model trained by machine learning. The processing function of the image recognition device may be incorporated into the image processing device 20, or may be implemented in an image processing server or cloud server (not shown) connected via the network 22.
[0155] If the determination result in step S50 is No, the processor 202 ends the flowcharts of FIGS.
[0156] About the programs that run computers A program that causes a computer to realize the processing functions of the image processing device 20 can be recorded on a computer-readable medium such as an optical disk, a magnetic disk, a semiconductor memory, or other tangible non-transitory information storage medium, and the program can be provided through this information storage medium.
[0157] In addition, instead of providing the program by storing it on such a tangible, non-transitory computer-readable medium, it is also possible to provide the program signal as a download service using a telecommunications line such as the Internet.
[0158] Furthermore, some or all of the processing functions of the image processing device 20 may be realized by cloud computing, and may also be provided as a SaaS (Software as a Service) service.
[0159] <<Hardware configuration of each processing unit>> The hardware structure of the processing units in the image processing device 20 that execute various processes, such as the information acquisition unit 222, coordinate conversion unit 224, camera matrix parameter setting unit 226, perspective projection conversion unit 228, line segment extraction unit 230, matching evaluation unit 234, optimal parameter value selection unit 236, image synthesis unit 238, position adjustment unit 240, and display control unit 251, is, for example, various processors as shown below.
[0160] Various types of processors include CPUs, which are general-purpose processors that execute programs and function as various processing units, GPUs, which are processors specialized for image processing, programmable logic devices (PLDs), such as FPGAs (Field Programmable Gate Arrays), which are processors whose circuit configuration can be changed after manufacturing, and dedicated electrical circuits, such as ASICs (Application Specific Integrated Circuits), which are processors with circuit configurations designed specifically to execute specific processes.
[0161] A single processing unit may be configured with one of these various processors, or may be configured with two or more processors of the same or different types. For example, a single processing unit may be configured with multiple FPGAs, or a combination of a CPU and an FPGA, or a combination of a CPU and a GPU. Alternatively, multiple processing units may be configured with a single processor. A first example of multiple processing units configured with a single processor is a configuration in which one or more CPUs and software are combined to form a single processor, as typified by client or server computers, and this processor functions as multiple processing units. A second example is a configuration in which a processor is used to realize the functions of an entire system including multiple processing units on a single IC (Integrated Circuit) chip, as typified by a system-on-chip (SoC). In this way, the various processing units are configured with one or more of the above-mentioned various processors as a hardware structure.
[0162] Furthermore, the hardware structure of these various processors is, more specifically, an electric circuit made up of a combination of circuit elements such as semiconductor elements.
[0163] Advantages of this embodiment The image processing device 20 according to the embodiment has the following advantages.
[0164] [1] According to the image processing device 20, the parameter values of the camera matrix are automatically searched for and the optimal parameter values are selected based on the sensor data obtained from the drone 12, so that highly accurate alignment between the map data of the shooting range and the captured image is possible without the need for a human to specify corresponding points.
[0165] [2] The image processing device 20 displays a composite image obtained by automatic alignment, and the position of the figures showing the area of each house can be moved on the image according to instructions from the user to adjust it to an optimal position. This allows the result of automatic alignment to be further improved by manual operation by the user, and the accuracy of alignment for each house can be increased.
[0166] Variation 1 The processing functions of the image processing device 20 may be realized by multiple computers or by cloud computing. The processing functions of the image processing device 20 may be implemented in the remote controller 16 and / or the terminal device 24.
[0167] Variation 2 In the above embodiment, an example of processing a still image as a captured image is described, but the camera 14 may also capture a video, and the image processing device 20 may extract some frames from the captured video and perform similar processing.
[0168] Variation 3 The method of calculating the degree of coincidence described using FIG. 10 and the method of calculating the degree of coincidence described as a function of the coincidence evaluation unit 234 are merely examples, and other methods may be applied as a method of evaluating the degree of coincidence, without being limited to the above examples.
[0169] Other application examples In the above embodiment, an example is given of processing an image captured by the camera 14 mounted on the drone 12, but the scope of application of the present disclosure is not limited to this example. For example, an image captured using a camera installed at a high location overlooking the ground, such as on the roof of a building or on a steel tower, is included in the concept of "images captured from the air." Even in the case of a fixed camera, the camera's attitude can be changed by panning and tilting operations, etc. If the camera position is fixed, the values of the camera position parameters in the camera matrix may be fixed, and a configuration can be adopted in which only the values of the attitude-related parameters are searched for.
[0170] Furthermore, the technology disclosed herein is not limited to the association of geospatial location information (geographic coordinates) with the image coordinates of a captured image, but can be widely applied to the association of three-dimensional spatial coordinates with the image coordinates of a captured image. For example, the technology disclosed herein can be applied to the case where a three-dimensional coordinate system is defined in a specific space such as an indoor ball game stadium, indoor sports stadium, amusement facility, photography studio, or factory, and the coordinate data of multiple specific points in that space is associated with the image coordinates of a captured image. Images captured using a camera installed on the ceiling of an indoor ball game stadium or a camera suspended from a wire, etc., are included in the concept of "images captured from the air."
[0171] "others" The present disclosure is not limited to the above-described embodiments, and various modifications are possible within the scope of the gist of the technical idea of the present disclosure. [Explanation of symbols]
[0172] 10. Image Processing System 12. Drone 13 Gimbal Head 14 Camera 16 Remote Controller 16A Display 20 Image processing device 22 Network 24 Terminal Equipment 24A Display 30 GPS receivers 32 Barometric pressure sensor 34 Orientation sensor 36 Gyro sensor 38 Motor 40 processors 42 Storage device 44 Communication Interface 100 Map Information 104 Image coordinate data 106 Converted Map Images 110 captured images 112 Camera location information 113 Posture information 140 Image Sensor 202 processors 204 Computer-readable medium 206 Communication Interface 208 Input / Output Interface 210 Bus 214 Input Device 216 Display device 220 Image Processing Program 222 Information Acquisition Department 222A Map information acquisition unit 222B Shooting condition acquisition unit 222C Image acquisition unit 224 Coordinate conversion unit 226 Camera matrix parameter setting section 228 Perspective projection transformation unit 230 Line segment extraction section 231 First line segment extraction unit 232 Second line segment extraction unit 234 Matching Evaluation Unit 235 Evaluation value calculation unit 236 Optimal parameter value selection section 238 Image Synthesis Unit 240 Position adjustment section 242 Cutout 250 Display Control Program 251 Display control unit 260 Map information storage unit 262 Photographed image storage unit 264 Sensor data storage unit IM, IMs images IMa Images IMb Images IML Line Image TMa converted map image TMb converted map image LS1a Line Segment LS1b line segment LS2 line segment MP map data PG Polygon RL Line S12~S52 Image processing steps
Claims
1. one or more processors; one or more memories in which programs to be executed by the one or more processors are stored; Equipped with The one or more processors execute the instructions of the program to: Acquire a photographed image taken using a camera; Acquire three-dimensional position information indicating the positions of a plurality of specific points in the space of the photographed range; setting parameter values of a perspective projection transformation that converts the three-dimensional position information into two-dimensional image coordinates based on the photographing conditions of the photographed image; converting the position information of the plurality of specific points into image coordinate data using the perspective projection transformation; evaluating a degree of coincidence between a first line segment extracted based on the image coordinate data obtained by the transformation and a second line segment extracted from the captured image; Image processing device.
2. The one or more processors: evaluating the degree of match to evaluate whether the registration using the perspective projection transformation is appropriate; The image processing device according to claim 1 .
3. The one or more processors: Correlating the captured image with the positions of the plurality of specific points based on the result of the evaluation of the degree of coincidence. The image processing device according to claim 1 .
4. The one or more processors: determining values of the parameters that result in a good evaluation of the degree of match; The image processing device according to claim 1 .
5. The one or more processors: calculating an evaluation value quantifying the degree of match; Finding the parameter value that provides the best evaluation result for the evaluation value. The image processing device according to claim 1 .
6. The one or more processors: different weights for evaluating the degree of coincidence between a central portion and a peripheral portion of the photographed image; The image processing device according to claim 1 .
7. 1. An image processing method executed by one or more processors, comprising: the one or more processors: Acquiring a photographed image taken using a camera; Acquiring three-dimensional position information indicating the positions of a plurality of specific points in a space within a range to be photographed; setting parameter values of a perspective projection transformation that converts the three-dimensional position information into two-dimensional image coordinates based on the photographing conditions of the photographed image; converting the position information of the plurality of specific points into image coordinate data using the perspective projection transformation; evaluating a degree of coincidence between a first line segment extracted based on the image coordinate data obtained by the transformation and a second line segment extracted from the captured image; An image processing method comprising:
8. On the computer, A function of acquiring images taken using a camera; A function for acquiring three-dimensional position information indicating the positions of multiple specific points in the space of the shooting range; a function of setting parameter values of a perspective projection transformation that converts the three-dimensional position information into two-dimensional image coordinates based on the shooting conditions of the captured image; a function of converting the position information of the plurality of specific points into image coordinate data using the perspective projection transformation; a function of evaluating a degree of coincidence between a first line segment extracted based on the image coordinate data obtained by the transformation and a second line segment extracted from the captured image; A program to make this happen.
Citation Information
Patent Citations
Photography image processing method and system thereof
JP2003316259A