Image processing device, image processing method, and program

The image processing device addresses alignment inaccuracies by detecting objects and estimating transformation parameters, ensuring precise geospatial coordination of aerial images for effective damage assessment.

WO2025177900A1PCT designated stage Publication Date: 2025-08-28FUJIFILM CORP
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
PCT/JP2025/004491
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-02-20
Filing Date
2025-02-12
Publication Date
2025-08-28

AI Technical Summary

Technical Problem

Existing image processing systems face inaccuracies in aligning aerial images with geospatial coordinates due to positioning errors from GNSS, camera pan and tilt, and attitude discrepancies, making it difficult to accurately link captured images with geospatial information.

Method used

An image processing device that detects objects in images, matches them with map objects using camera information, and estimates transformation parameters to minimize positional errors between geospatial and image coordinates, without requiring ground control points.

Benefits of technology

Accurately links captured images with geospatial coordinates, enabling precise identification of objects like houses and facilitating efficient damage assessment after disasters.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure JP2025004491_28082025_PF_FP_ABST
    Figure JP2025004491_28082025_PF_FP_ABST
Patent Text Reader

Abstract

Provided are: an image processing device capable of precisely connecting a photographed image and geospatial coordinates; an image processing method; and a program. Specifically provided is an image processing device comprising one or more processors, wherein the one or more processors execute: a process for acquiring an image photographed by a camera; a process for acquiring camera information including information regarding the position and orientation of the camera when the image was photographed; a process for detecting an object from the image; a process for using the camera information to associate the object in the image with an object on a map included in map information; and a process for estimating a conversion parameter for a coordinate conversion between geospatial coordinates and in-image coordinates so as to reduce the positional error of the object in the coordinate conversion in light of the positional relationship between the object in the image and the object on the map, which have been associated.
Need to check novelty before this filing date? Find Prior Art

Description

Image processing device, image processing method and program

[0001] The present disclosure relates to an image processing device, an image processing method, and a program, and in particular to a georeferencing technology that links images taken from the air by a camera mounted on a drone or the like to geospatial coordinates, and a geocoding technology that links geospatial coordinates to information such as addresses.

[0002] Patent Literature 1 describes a captured image processing method for capturing an image of the ground surface from an imaging device mounted on an airborne aircraft and identifying the conditions present on the ground surface. The method described in Patent Literature 1 identifies the aerial imaging position three-dimensionally, calculates and obtains the imaging range of the captured ground surface, transforms the captured image to match the imaging range, and then overlays it on a map in a map information system for display.

[0003] Patent document 2 describes an image alignment device that extracts reference point data on an image for alignment purposes from an image taken by a camera mounted on a helicopter, and calculates conversion coefficients for a coordinate transformation formula from the reference point data on the image and reference point data on a map that has been registered in advance in map data.

[0004] Patent document 3 describes a program for causing a computer to execute the following steps on a map corresponding to a digital image: determining third and fourth reference points on the image that correspond to a plurality of first reference points on the map, each of which has known map coordinate information, and a plurality of second reference points selected for each of the first reference points; determining a fifth reference point with high local accuracy for the third reference point based on the fourth reference point; and correcting geometric distortion in the digital image based on the plurality of fifth reference points and the plurality of first reference points corresponding thereto.

[0005] Patent Document 4 describes a geocoding technology that detects the edges of roads or houses from images taken by a camera mounted on a drone and associates them with the edges of roads or houses in map data.

[0006] Japanese Patent Application Laid-Open No. 2003-316259 Japanese Patent Application Laid-Open No. 3-196372 Japanese Patent Application Laid-Open No. 2004-171413 International Publication No. 2023 / 047799

[0007] The technology described in Patent Document 1 calculates the shooting range by identifying the camera position and camera attitude from output signals obtained from detection units such as an aircraft position detection unit, an aircraft attitude detection unit, and a camera attitude detection unit provided on the aircraft, and aligns the shooting range with a map. However, in an actual system, when calculations are performed based on the camera position and attitude determined from output signals obtained from detection units such as an aircraft position detection unit, an aircraft attitude detection unit, and a camera attitude detection unit, there is a large discrepancy between the actual shooting range and the calculation result, resulting in poor accuracy in aligning the map and the captured image.

[0008] For example, the positioning provided by a Global Navigation Satellite System (GNSS) receiver mounted on a standard drone has an error of approximately 2 to 5 meters. To improve this, high-precision GNSS positioning, such as real-time kinematic (RTK) or post-processed kinematics (PPK), can be used.

[0009] However, RTK or PPK functions are costly because they require dual-frequency antennas, and are generally only installed in industrial drones used for industrial purposes such as construction or surveying, which require high-precision positioning.

[0010] Furthermore, the correspondence between the image coordinates and geospatial coordinates of a captured image is significantly affected by camera pan and tilt positioning errors and the aircraft's attitude. Furthermore, errors are caused by magnetic sensors when measuring azimuth angles. Therefore, when calculations are performed using transformation parameters that use external camera parameters that include these errors, a discrepancy occurs between the captured image and geospatial information. As a result, it becomes difficult to determine the correspondence between, for example, a damaged house shown in a captured image and its address, making it difficult to quickly grasp the situation at the scene.

[0011] The present disclosure has been made in consideration of the above circumstances, and aims to provide an image processing device, an image processing method, and a program that can accurately link a captured image with geospatial coordinates.

[0012] An image processing device according to a first aspect of the present disclosure is an image processing device having one or more processors, which perform the following processes: acquiring an image captured by a camera; acquiring camera information including information about the position and attitude of the camera when the image was captured; detecting an object from the image; using the camera information to match the object on the image with an object on a map included in the map information; and estimating transformation parameters for the coordinate transformation based on the positional relationship between the matched object on the image and the object on the map so as to reduce the positional error of the object due to the coordinate transformation between geospatial coordinates and coordinates in the image.

[0013] According to the first aspect, even if the information about the position and attitude of the camera obtained from the camera information is information about the approximate position and attitude including errors, it is possible to obtain transformation parameters that realize coordinate transformation that accurately corresponds image coordinates with geospatial coordinates based on the relationship between an object detected in an image and an object in the corresponding map information. Furthermore, according to the first aspect, it is not necessary to install ground control points (GCPs) or the like in advance, and it is possible to accurately correspond positions in the image with positions in geospatial space by using objects in the captured image.

[0014] The object may be, for example, a building such as a house. It is sufficient to utilize map information containing information about the position of the object in geographic space, and the type of object may vary. In-image coordinates are coordinates indicating a position in the image plane of a captured image, and are synonymous with coordinates within the camera's field of view. Geospatial coordinates are coordinates indicating a position in geographic space, and are synonymous with geographic coordinates. The position information included in the map information can be expressed in geospatial coordinates. Position error is synonymous with positional deviation. It is desirable for one or more processors to estimate transformation parameters that convert between geospatial coordinates and in-image coordinates so as to minimize the position error of the corresponding object due to the coordinate transformation between the geospatial coordinates and the in-image coordinates. "So as to minimize the position error" is not limited to strictly matching the minimum state, but also includes approaching the minimum state.

[0015] The image processing device of the second aspect may be configured such that, in the image processing device of the first aspect, the process of detecting an object from an image includes a process of obtaining information indicating the area of ​​the object on the image, and the matching process includes a first matching process that matches the area of ​​the object on the image with the object on the map by conversion using camera information.

[0016] An image processing device according to a third aspect may be configured such that, in the image processing device according to the second aspect, one or more processors acquire a representative point of an object on the map from map information, and the first matching process includes matching an area of ​​the object on the image with the representative point of the object on the map.

[0017] The representative point may be, for example, the center of gravity of the object, or a point that constitutes the periphery of the object. The center of gravity includes the concept of a center. The center of gravity of the object is not limited to a strict one, and may be a point within a range that can be substantially regarded as the center of gravity. The representative point may be a point calculated from position information of points that constitute the periphery of the object, which is included in the map information. Multiple representative points may be obtained for one object.

[0018] An image processing device according to a fourth aspect may be configured such that, in the image processing device according to the third aspect, the first matching process includes matching the area of ​​the object on the image with the representative point of the object on the map based on the positional relationship between the point, onto the image, of the representative point of the object on the map projected by conversion using camera information, and the area of ​​the object on the image.

[0019] An image processing device according to a fifth aspect may be the image processing device according to the third or fourth aspect, in which the representative point is a center of gravity.

[0020] An image processing device according to a sixth aspect may be configured such that, in an image processing device according to any one of the second to fifth aspects, the information indicating the area of ​​an object on an image includes information on an outer frame surrounding the area of ​​the object.

[0021] An image processing device according to a seventh aspect may be configured such that, in an image processing device according to any one of the second to fifth aspects, the information indicating the area of ​​an object on an image includes segmentation information in which the area of ​​the object is identified on a pixel-by-pixel basis.

[0022] An image processing device according to an eighth aspect may be configured such that, in an image processing device according to any one of the second to seventh aspects, the first matching process includes matching each of a plurality of objects shown in the image with an object on a map included in the map information.

[0023] The image processing device of the ninth aspect may be configured such that, in the image processing device of any one of the second to eighth aspects, the first matching process includes matching so as to increase the weight of objects in the center of the image.

[0024] An image processing device according to a tenth aspect may be configured such that, in the image processing device according to any one of the second to eighth aspects, the first matching process includes matching objects in the peripheral parts of the image so as to give them a higher weight.

[0025] An image processing device according to an eleventh aspect may be configured such that, in the image processing device according to any one of the second to tenth aspects, the process of estimating the transformation parameters includes a second matching process that corrects the positional relationship between the object on the image that has been matched by the first matching process and the object on the map.

[0026] An image processing device according to a twelfth aspect may be configured such that, in the image processing device according to the eleventh aspect, one or more processors further execute a process of calculating a representative point of an object on the image detected from the image, and the second matching process includes correcting the positional relationship between the representative point of the object on the image matched by the first matching process and a point obtained by projecting the representative point of the object on the corresponding map onto the image, thereby matching coordinates in the image with geospatial coordinates.

[0027] The image processing device according to the thirteenth aspect may be configured such that, in the image processing device according to the eleventh or twelfth aspect, the second matching process includes estimating transformation parameters using objects in each of the four corners and the central area of ​​the image so as to minimize positional error.

[0028] An image processing device according to a fourteenth aspect may be configured such that, in the image processing device according to the eleventh to thirteenth aspects, the second matching process includes using a plurality of objects on the image to remove outlying points and estimate transformation parameters so as to minimize positional errors.

[0029] The image processing device of the 15th aspect may be configured such that, in the image processing device of any one of the 1st to 14th aspects, one or more processors further execute a process of converting objects on a map into in-image coordinates of the image, and displaying the result of the conversion on the image.

[0030] The image processing device according to a sixteenth aspect may be the image processing device according to the fifteenth aspect, further comprising a display device that displays the image and the conversion result in a superimposed manner.

[0031] An image processing device according to a seventeenth aspect may be configured such that, in the image processing device according to the fifteenth or sixteenth aspect, the one or more processors further accept an operation to manually match an object on the image with a corresponding object on the map based on the transformation result overlaid on the image, and execute a process to match the object on the image with the object on the map in accordance with the instructions of the accepted operation.

[0032] The image processing device according to an eighteenth aspect may be configured such that the image processing device according to the seventeenth aspect further comprises an input device for manually inputting instructions.

[0033] The image processing device of the 19th aspect may be configured such that, in the image processing device of any one of the 1st to 18th aspects, the image is an aerial image taken by a camera mounted on an aircraft, and the camera information is obtained from sensor data obtained by a sensor placed on at least one of the camera and the aircraft.

[0034] An image processing device according to a twentieth aspect may be configured such that, in an image processing device according to any one of the first to nineteenth aspects, the object includes at least one of a house, a manhole, and a road intersection.

[0035] The image processing device of the 21st aspect may be an image processing device of any one of the first to 20th aspects, in which the object includes a house, and the one or more processors may further be configured to extract the object, that is, the house, from the image and execute a process to determine the degree of damage to the extracted house.

[0036] An image processing method according to a 22nd aspect of the present disclosure includes the steps of acquiring an image captured by a camera, acquiring camera information including information about the position and attitude of the camera when the image was captured, detecting an object from the image, using the camera information to match the object on the image with an object on a map included in the map information, and estimating transformation parameters for the coordinate transformation based on the positional relationship between the matched object on the image and the object on the map so as to reduce positional errors of the object due to coordinate transformation between geospatial coordinates and coordinates in the image.

[0037] Each step of the image processing method according to the twenty-second aspect may be executed by one or more processors. The one or more processors may automatically execute the processing of each step according to a program, or may accept instructions input from a user and execute processing for some or all of the steps in accordance with the accepted instructions.

[0038] The image processing method according to the twenty-second aspect may have a configuration including the same specific aspects as the image processing device according to any one of the second to twenty-first aspects.

[0039] A program according to a 23rd aspect of the present disclosure is a program that enables a computer to perform the following functions: acquire an image captured by a camera; acquire camera information including information about the position and attitude of the camera when the image was captured; detect an object from the image; use the camera information to match the object in the image with an object on a map included in the map information; and estimate transformation parameters for the coordinate transformation based on the positional relationship between the matched object in the image and the object on the map so as to reduce positional errors of the object due to coordinate transformation between geospatial coordinates and coordinates within the image.

[0040] The program according to the twenty-third aspect may be configured to include the same specific aspects as the image processing device according to any one of the second to twenty-first aspects.

[0041] According to the present disclosure, it is possible to link captured images with geospatial coordinates with high accuracy.

[0042] FIG. 1 is a schematic diagram showing an example of the configuration of a captured image processing system according to an embodiment. FIG. 2 is a block diagram showing an example of the electrical configuration of a drone equipped with a camera. FIG. 3 is a block diagram showing an example of the hardware configuration of an image processing device according to an embodiment. FIG. 4 is a flowchart showing an example of an image processing method executed by the image processing device. FIG. 5 is an example image showing a detection result of detecting a house as an object from a captured image. FIG. 6 is an explanatory diagram showing an example of a first association process for associating the outer frame of an object on an image with the object on a map. FIG. 7 is an explanatory diagram showing an example of a process for calculating the center of gravity of an object on an image. FIG. 8 is an explanatory diagram showing an example of a second association process for associating the center of gravity of an object on an image with the center of gravity of an object on a map. FIG. 9 is a functional block diagram showing the functional configuration of an image processing device. FIG. 10 is a flowchart showing an example of processing in a damage assessment investigation support system to which an image processing device according to an embodiment is applied. FIG. 11 is a functional block diagram showing the functional configuration of an image processing device applied to the damage assessment investigation support system.

[0043] DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS Preferred embodiments of the present invention will now be described in detail with reference to the accompanying drawings. In this specification, the same components are designated by the same reference numerals, and redundant explanations will be omitted where appropriate.

[0044] [Overview of the Embodiments] As an application example of the technology disclosed herein, a technology for linking objects in images captured from the air using a camera mounted on a drone with data from a geographic information system (GIS) is described. The objects are, for example, houses. By using an image processing device according to an embodiment of the present disclosure, it is possible to realize a system that grasps the external appearance of individual houses from captured images, identifies the addresses of each house in the images, and easily links them to various data held by municipal government agencies. Such a system is useful, for example, for grasping the damage status of buildings such as houses after a disaster, ultimately enabling the formulation of an efficient residential damage assessment survey plan.

[0045] [Configuration Example of a Photographic Image Processing System] Fig. 1 is a schematic diagram showing a configuration example of a photographic image processing system 10 according to an embodiment of the present disclosure. The photographic image processing system 10 includes a drone 12 for aerial photography, a camera 14 mounted on the drone 12, a remote controller 16, and an image processing device 20. The drone 12 is an unmanned aerial vehicle that is remotely controlled using the remote controller 16. The drone 12 may have an autopilot function that flies according to a program. The drone 12 is an example of an air vehicle.

[0046] The camera 14 is mounted on the drone 12 via a gimbal head 13. The camera 14 includes an optical system, an image sensor, and a signal processing circuit (not shown). The optical system includes one or more lenses such as a focus lens. The image sensor may be, for example, a charge-coupled device (CCD) image sensor or a complementary metal-oxide semiconductor (CMOS) image sensor.

[0047] The camera 14 generates digital image data of the photographed subject by processing signals obtained from the image sensor using a signal processing circuit. The digital image data generated by the camera 14 can be used as a photographed image. Images photographed using the camera 14 (photographed images) can be stored in an internal storage device built into the drone 12 and / or a storage device such as a memory card removably attached to the drone 12. Images photographed using the camera 14 can also be transferred to the remote controller 16, the image processing device 20, and other terminal devices 24 using wireless communication.

[0048] The remote controller 16 is a transmitter that controls the operation of the camera 14 and the drone 12 via wireless communication. The wireless communication may be in the form of a wireless local area network (LAN), a communication format using radio waves in the 2.4 GHz or 5.7 GHz band, or a format using a mobile communication network. The communication format for the control signals for operating the drone 12 and the communication format for transferring images captured by the camera 14 may be different or may be the same.

[0049] The remote controller 16 includes left and right sticks for controlling the flight operation of the drone 12, a lever for operating the gimbal head 13, a shooting button for instructing the camera 14 to take a picture, and a shooting mode button for switching between video shooting and still image shooting. By employing a touch panel display for the display 16A, the shooting button and other operation buttons can be realized by the touch panel display.

[0050] The live video captured by the camera 14 can be displayed on the display 16A of the remote controller 16. The remote controller 16 can also grasp the status of the drone 12, such as its flight position and speed, in real time based on data from various sensors provided on the drone 12. Flight information indicating the status of the drone can be displayed on the display 16A.

[0051] 1 is an example of an image captured using the camera 14. In this embodiment, at least one still image is captured from the air, and the captured image is processed by the image processing device 20.

[0052] The image processing device 20 is configured using a computer. The computer applied to the image processing device 20 may be a server, a personal computer, or a workstation.

[0053] The image processing device 20 can perform data communication with the remote controller 16 and the terminal device 24 via a network 22. The network 22 may be a local area network or a wide area network. The image processing device 20 acquires various information from the drone 12 and the camera 14. The image processing device 20 can also acquire map data of the imaging target range from a geographic information system (not shown) via the network 22. The map data may be acquired in advance before imaging, or may be acquired after imaging.

[0054] The terminal device 24 may be a mobile information terminal such as a smartphone or a tablet terminal. The terminal device 24 includes a display 24A. The terminal device 24 may have the functionality of the remote controller 16. The terminal device 24 may also have the processing functionality of the image processing device 20.

[0055] [Configuration Example of Drone with Camera] Figure 2 is a block diagram that schematically illustrates an example of the electrical configuration of a drone 12 equipped with a camera 14. The drone 12 includes a Global Navigation Satellite System (GNSS) receiver 30, a barometric pressure sensor 32, a direction sensor 34, an inertial measurement unit (IMU) 36, and a motor 38. The IMU 36 includes a gyro sensor 362, an acceleration sensor 364, and a temperature sensor 366, and detects translational and rotational motion in three orthogonal axes. The motor 38 is a power source that rotates a rotor (not shown), and the drone 12 includes multiple motors 38 that drive multiple rotors.

[0056] The GNSS receiver 30 acquires position information including the latitude and longitude of the drone 12. The barometric pressure sensor 32 detects the barometric pressure in the drone 12. The drone 12 can acquire its altitude based on the barometric pressure detected using the barometric pressure sensor 32. Note that the term "acquire" includes the concept of generating information by data processing such as calculation. The latitude, longitude, and altitude of the drone 12 constitute position information of the drone 12 and the camera 14.

[0057] The orientation sensor 34 may be, for example, a geomagnetic sensor, and may detect the azimuth angle at which the lens of the camera 14 is facing.

[0058] The gyro sensor 362 detects a roll angle representing the angle of rotation about the roll axis, a pitch angle representing the angle of rotation about the pitch axis, and a yaw angle representing the angle of rotation about the yaw axis. The drone 12 acquires attitude information of the drone 12 based on the rotation angles acquired using the gyro sensor 362. Note that some or all of the sensors, such as the GNSS receiver 30, the barometric pressure sensor 32, the orientation sensor 34, and the IMU 36, may be located on the camera 14 side.

[0059] The drone 12 includes a processor 40, a storage device 42, and a communication interface 44. The storage device 42 may be a memory, an internal storage device, an external storage device, or a combination thereof. The processor 40 serves as a flight controller and performs various calculations necessary for flight control of the drone 12 based on sensor data obtained from various sensors.

[0060] The communication interface 44 is a communication unit that performs wireless communication with the remote controller 16, etc. The communication interface 44 may also include a communication terminal that supports wired communication. Furthermore, the drone 12 includes a battery and a battery charging terminal (not shown).

[0061] 3 is a block diagram showing an example of the hardware configuration of the image processing device 20. The image processing device 20 includes a processor 202, a computer-readable medium 204 which is a non-transitory tangible entity, a communication interface 206, and an input / output interface 208.

[0062] The processor 202 includes a central processing unit (CPU) and may include a graphics processing unit (GPU). The processor 202 is connected to a computer-readable medium 204, a communication interface 206, and an input / output interface 208 via a bus 210.

[0063] The computer-readable medium 204 includes a memory 212 serving as a primary storage device and a storage 214 serving as a secondary storage device. The computer-readable medium 204 may be, for example, a semiconductor memory, a hard disk drive (HDD), a solid state drive (SSD), or a combination of these. The computer-readable medium 204 stores various programs, data, and the like, including an image processing program and a display control program.

[0064] The image processing device 20 may include an input device 222 and a display device 224. The input device 222 and the display device 224 are connected to the bus 210 via the input / output interface 208. The input device 222 is configured by, for example, a keyboard, a mouse, a multi-touch panel, or other pointing device, or a voice input device, or an appropriate combination of these.

[0065] The display device 224 is configured by, for example, a liquid crystal display, an organic electro-luminescence (OEL) display, a projector, or an appropriate combination of these.

[0066] [Outline of Processing Functions of Image Processing Device 20] An outline of the processing executed by the processor 202 of the image processing device 20 is as follows. The processor 202 acquires an image captured by the camera 14 and information relating to the approximate position (approximate capturing position) and approximate attitude of the camera 14 when the image was captured. The information relating to the approximate position and approximate attitude of the camera 14 may be position information and attitude information indicated by sensor data obtained by sensors such as the GNSS receiver 30, barometric pressure sensor 32, orientation sensor 34, and IMU 36 mounted on the drone 12.

[0067] The position information and attitude information indicated by the sensor data of the drone 12 includes GNSS positioning errors, errors caused by the sensor, and the like, and indicates the approximate camera position (shooting position) and camera attitude (shooting attitude). The position information and attitude information at the time of shooting indicated by the sensor data of the drone 12 are referred to as camera position information and camera attitude information at the time of shooting. The camera information including the camera position information and camera attitude information at the time of shooting may be recorded as metadata such as tag information attached to the image file of the captured image, or may be recorded as a separate file linked to the image file.

[0068] The processor 202 detects an object from the acquired image. The object is, for example, a house. The object may be a building other than a house, or may be a road intersection, a manhole, or the like. The object may be a feature whose geographical location is included in map information of a geographic information system. It is also possible to handle multiple types of objects as objects, and the processor 202 may detect multiple different types of objects, such as houses and roads, from the image.

[0069] In the map information referenced by the image processing device 20, for example, each house is assigned a house ID (Identification) as an identification code for identifying the house, and location data indicating the respective positions of multiple specific points constituting the perimeter of the house is recorded, linked to the house ID. The location data for each specific point may be three-dimensional data of latitude, longitude, and altitude. In the case of Japan, map data including such geographic coordinate data can be obtained, for example, from basic map information provided by the Geospatial Information Authority of Japan. Alternatively, such map data can be obtained from the OpenStreetMap database.

[0070] The processor 202 uses map information about the geographic space including the area captured by the camera 14, and the camera position information and camera attitude information at the time of capture, to associate objects in the image with objects in the map information. The object in the map information refers to an object on the map (on the map information) shown in the map information. This association process converts the position of the object on the map shown in the map information (geospace coordinates of the object) into a position on the image (image coordinates) by perspective projection transformation using parameters identified from the camera position information and camera attitude information at the time of capture, and projects the converted position onto the image. Based on the positional relationship between the position of the object on the map projected onto the image and the object on the image, objects whose positions are close to each other are associated.

[0071] If the correspondence is based on transformation using camera information, there will be a large positional discrepancy between the position of the object in the map information projected onto the image and the position of the object in the corresponding image, so processor 202 estimates transformation parameters to convert geospatial coordinates into coordinates within the image so as to reduce the positional error between the two.

[0072] Furthermore, once the transformation parameters for converting geospatial coordinates into coordinates within the image are determined, it is possible to convert the coordinates within the image into geospatial coordinates by the inverse transformation. Therefore, estimating the transformation parameters for converting geospatial coordinates into coordinates within the image is understood to include the concept of estimating the transformation parameters for converting coordinates within the image into geospatial coordinates.

[0073] [Example of Image Processing Method] A more specific example of an image processing method according to this embodiment will now be described.

[0074] 4 is a flowchart showing an example of an image processing method executed by the image processing device 20. In step S11, the processor 202 acquires an image captured by the camera 14.

[0075] In step S12, the processor 202 acquires information regarding the approximate position and approximate attitude of the camera 14 when the image was captured. That is, the processor 202 acquires camera position information and camera attitude information at the time of capture that are linked to the captured image. The captured image is linked to the camera position information and camera attitude information at the time of capture. The camera position information may be position information obtained from the GNSS receiver 30 of the drone 12 and includes latitude, longitude, and altitude data. The altitude data in the camera position information 312 may be calculated based on data obtained from the barometric pressure sensor 32. The camera attitude information 314 includes azimuth angle, tilt angle, and roll angle data obtained from the orientation sensor 34 and the IMU 36. The tilt angle is the camera angle toward the ground and is synonymous with the "depression angle." The camera position information is approximate position information that includes measurement errors of GNSS positioning. The camera attitude information is approximate attitude information that includes errors in sensor data.

[0076] In step S13, the processor 202 detects an object in the acquired image. The process of detecting the object in the image may be, for example, a process using image recognition AI (artificial intelligence) that applies a trained model that has been trained to recognize the area of ​​the object in the image using an object detection algorithm that uses machine learning. Here, as an example, the processor 202 detects the object in the image using the image recognition AI and obtains the outline of the object in the image (see FIG. 5).

[0077] In step S14, the processor 202 performs a first association process to associate an object on the image detected from the image with an object on the map information. For example, the processor 202 performs perspective projection transformation on the position (geographical coordinates) of a representative point of the object on the map identified from the map information based on the camera position information and camera attitude information at the time of capture, and projects the position (geographical coordinates) of the representative point of the object on the map that is paired with the object on the image (see FIG. 5 ). The representative point of the object may be, for example, the center of gravity of the object. The center of gravity is an example of a representative point.

[0078] In step S15, the processor 202 calculates the center of gravity of the object on the image. For example, the processor 202 calculates the center of gravity of the object from the outer frame of the object on the image detected in step S13 (see FIG. 6).

[0079] In step S16, the processor 202 performs a second association process to associate the center of gravity of the object on the image with the center of gravity of the object on the map. The processor 202 calculates the center of gravity of the object on the map from the map information for the object on the map that has been associated with the object on the image by the first association process in step S14. The center of gravity of a house on the map can be calculated, for example, from the geospatial coordinates of multiple points that form the perimeter of the house included in the map information.

[0080] The coordinate system of the GIS data may be a geographic coordinate system expressed by latitude and longitude, or may be a coordinate system to which a map projection method such as the Universal Transverse Mercator (UTM) coordinate system is applied. When the map data is expressed as latitude and longitude data, the processor 202 preferably converts the map data including the latitude and longitude into Cartesian coordinate data. The processor 202 can obtain three-dimensional geospatial coordinates (X, Y, Z) from the map information.

[0081] The centroids of houses on the map calculated from map information are represented by geospatial coordinates, while the centroids of objects on the image are represented by image coordinates. The processor 202 estimates transformation parameters for a projection transformation that links three-dimensional geospatial coordinates with two-dimensional image coordinates from the relationship between the image coordinates and geospatial coordinates of the centroids of the corresponding objects.

[0082] When estimating the position and orientation of an object in three-dimensional space or within an image based on the correspondence between the position of the object in three-dimensional space and the position of the object in an image coordinate system, the position and orientation of the target object can generally be estimated by solving the PnP problem (Perspective n-Point problem).

[0083] The transformation parameters estimated in the image processing device 20 are as follows:

[0084]

[0085] In the equation, (X, Y, Z) represent three-dimensional coordinates in the world coordinate system. Here, they are geospatial coordinates. (u, v) represent the coordinates of a point projected onto the image plane, i.e., the coordinates within the image. (t1, t2, t3) and (r11, r12, r13, r21, r22, r23, r31, r32, r33) are the extrinsic parameters of the camera. (t1, t2, t3) represent translation parameters, and (r11, r12, r13, r21, r22, r23, r31, r32, r33) represent rotation parameters. (cx, cy) and (fx, fy) are the intrinsic parameters of the camera. (cx, cy) is the principal point, which may be, for example, the image center. (fx, fy) is the focal length expressed in pixels. The transformation matrix applied to the coordinate transformation between image coordinates and geospatial coordinates is called the camera matrix, and is expressed as the product of the intrinsic parameter matrix and the extrinsic parameter matrix. Once the intrinsic parameter matrix is ​​set, it can be used repeatedly as long as the focal length of the photographing optical system is not changed.

[0086] In this embodiment, the transformation parameters that are of interest are the elements of the extrinsic parameter matrix. As described above, when the extrinsic parameter matrix set from the camera position information and camera attitude information at the time of shooting is applied, the position error of the projective transformation is large, so the processor 202 estimates the elements of the extrinsic parameter matrix by the processes of steps S15 and S16 so as to minimize the position error of the projective transformation of the center of gravity of the corresponding object.

[0087] That is, the processor 202 performs a second matching process to correct the relationship between the position in the image coordinate system and the position in the geographical coordinate system so as to reduce the positional error between the center of gravity of the object in the image matched by the first matching process and the point obtained by projecting the center of gravity of the object on the map onto the image using perspective projection transformation (see Figure 7).

[0088] This second association process generates transformation parameters that realize coordinate transformation between geographical space coordinates and image coordinates with minimal positional error (step S17).

[0089] In step S18, the processor 202 stores the generated transformation parameters in the computer-readable medium 204. After step S18, the processor 202 ends the flowchart of FIG.

[0090] [Example of Object Detection Processing in Step S13] Figure 5 is an explanatory diagram showing an example of the processing in step S13. Figure 5 is an example image showing the detection result of detecting a house HS, which is an object, from a captured image IM1. Figure 5 shows an example of an image in which each of multiple houses included in the captured image IM1 is detected on a house-by-house basis, and a rectangular frame (bounding box) surrounding the area of ​​each house is superimposed on the captured image IM1. Here, the frame surrounding the object is referred to as an outer frame FR. Multiple houses HS are captured in the captured image IM1, and the processor 202 can acquire information about the outer frame FR surrounding each of the multiple houses HS detected from this image and display the outer frame FR indicating the area of ​​each house HS superimposed on the captured image IM1.

[0091] The outer frame FR of the object may be a rectangular bounding box output by an image recognition AI that performs object detection processing. Note that, although an unrotated rectangle is illustrated as the outer frame FR in Fig. 5, it may also be a rectangle that is rotated within the image plane to match the orientation of the house HS.

[0092] Furthermore, the image recognition AI that detects an object in an image is not limited to one configured to output a bounding box that surrounds the area of ​​the object in the image, but may also be one configured to perform segmentation processing that detects the area of ​​the object in the image on a pixel-by-pixel basis. In this case, the processor 202 acquires segmentation information that identifies the object on a pixel-by-pixel basis as information indicating the area of ​​the object in the image.

[0093] [Example of First Correspondence Processing in Step S14] Figure 6 is an explanatory diagram showing an example of the processing in step S14. Figure 6 schematically shows an example of the positional relationship between the outer frame FR of an object detected from a captured image and a projection point PT obtained by projecting the position of a representative point of the object on a map onto the captured image based on camera position information and camera attitude information at the time of capture. The representative point of the object on the map is, for example, the center of gravity of the object on the map calculated from map information. Figure 6 shows the positional relationship between the outer frame FR of each object and the projection point PT obtained by projecting the position of the representative point of the object on the corresponding map onto the captured image for seven objects in the image.

[0094] As shown in Fig. 6, when coordinate transformation is performed based on camera position information and camera attitude information at the time of shooting, there is a positional deviation (positional error) between the object on the image and the object on the map. In the first correspondence process of step S14, the processor 202 takes into account the positional error due to this projection transformation and corresponds the object on the map at the projection point PT projected at a position overlapping the area indicated by the outer frame FR of the object on the image with the object on the image surrounded by that outer frame FR. In the example shown in Fig. 6, seven pairs are obtained by corresponding the outer frames FR of each of the seven objects on the image with the projection points PT of the objects on the map.

[0095] In Figure 6, the outer frame FR of the object and the projection point PT projected at a position overlapping the area surrounded by the outer frame FR are paired, but it is sufficient to pair the outer frame FR of the object and the projection point PT if they are close within an allowable range in terms of their positional relationship.

[0096] It is also possible that the amount of positional deviation between the outer frame FR of an object on the image, the projection point PT of the object on the map, and the object on the image may differ between the central and peripheral parts of the image. Therefore, the allowable range of the positional relationship that serves as the criterion for determining a pair may be different between the central and peripheral parts of the image. For example, if the positional deviation in the peripheral part of the image is greater than the positional deviation in the central part, the allowable range of the positional relationship between the object on the peripheral part of the image and the corresponding projection point PT may be set larger than the allowable range in the central part. The central part of the image refers to the area near the center of the image. The peripheral part of the image refers to the peripheral area outside the central part of the image.

[0097] Furthermore, the result of the automatic association performed by the processor 202 may be displayed on the display device 224, and an instruction to manually correct the association may be received from the user. The user can input an instruction to change or add a corresponding pair from the input device 222, and the processor 202 associates the objects on the image with the objects on the map in accordance with the received instruction. The first association process may include processing in response to an instruction from the user.

[0098] [Example of Processing for Calculating Center of Gravity of Object in Step S15] Fig. 7 is an explanatory diagram showing an example of the processing in step S15. Fig. 7 schematically shows how the center of gravity G of an object is calculated from the outer frame FR of the object on the image. The processor 202 calculates the center of gravity G for each object detected from the image. The center of gravity G of the object on the image may be calculated as the center of gravity of the outer frame FR based on the outer frame FR of the object on the image. The processor 202 may calculate the center of gravity of the outer frame FR from the positions of the four corners of the outer frame FR of the object detected from the image, and acquire this as the center of gravity G of the object.

[0099] [Example of Second Correspondence Processing in Step S16] FIG. 8 is an explanatory diagram showing an example of the processing in step S16. FIG. 8 shows a conceptual diagram of the second correspondence processing. Based on the positional relationship between the center of gravity G of an object on an image and the projection point PT of the center of gravity of the object on the corresponding map, the processor 202 corrects the positional relationship between the two so as to reduce the positional deviation (positional error) between them. That is, based on the correspondence relationship between the center of gravity G of the object on the image expressed in intra-image coordinates (u, v) and the center of gravity of the object on the map expressed in three-dimensional geographic coordinates (X, Y, Z), the processor 202 estimates transformation parameters for performing perspective projection transformation so as to minimize the positional error (positional deviation) between the center of gravity G of the object on the image and the projection point PT obtained by projecting the center of gravity of the object on the corresponding map onto the image.

[0100] The processor 202 estimates optimal transformation parameters so that the sum of the distances (sum of errors) between the centers of gravity Gi of each of the multiple objects detected in the image and the projection points PTi of the corresponding object points on the map approaches a minimum. The subscript i is an index number for distinguishing between the multiple objects. The evaluation value used to evaluate the positional error between the multiple objects in the image and the corresponding objects on the map projected onto the image may be, for example, a simple sum or a weighted sum of the positional errors for each object.

[0101] 9 is a functional block diagram showing the functional configuration of the image processing device 20. The image processing device 20 includes an image acquisition unit 230, a camera information acquisition unit 232, a map information acquisition unit 236, an object detection unit 240, a center of gravity calculation unit 242, a first association unit 250, a second association unit 252, and a transformation parameter storage unit 256.

[0102] The image acquisition unit 230 acquires a captured image 300. The captured image 300 is an aerial image captured from the air by the camera 14. The camera information acquisition unit 232 acquires camera information including camera position information 312 and camera attitude information 314 at the time of capturing the captured image 300. The camera information acquisition unit 232 includes a camera position information acquisition unit 233 that acquires the camera position information 312, and a camera attitude information acquisition unit 234 that acquires the camera attitude information 314. The camera position information acquisition unit 233 acquires the camera position information 312 at the time of capturing the captured image 300.

[0103] The map information acquisition unit 236 acquires necessary map information from the map database 260. The map database 260 may include, for example, some or all of the GIS data among basic map information, location reference information, national land numerical information, map data attached to land registry documents, and city, ward, town, and village boundary data. The term "map data" includes the concept of GIS data.

[0104] The image processing device 20 may include a map database 260. In this case, the image processing device 20 includes a map data storage unit that stores the map database 260. The map data storage unit may be a storage area of ​​the computer-readable medium 204 in the image processing device 20, or may be a storage area of ​​an external storage device separate from the image processing device 20.

[0105] The object detection unit 240 detects an object from the captured image 300. The object detection unit 240 may be configured to include, for example, an image recognition AI module to which an object detection algorithm based on machine learning is applied. The object detection unit 240 detects an object from the captured image 300 and generates an outer frame 320 of the object.

[0106] The first correspondence unit 250 corresponds an outer frame 320 of an object on the image detected from the captured image 300 with the object on the map. The first correspondence unit 250 projects a point 326 of the object on the map into the image coordinate system of the captured image 300 by applying first transformation parameters 324 corresponding to a perspective projection transformation matrix determined from camera position information 312 and camera attitude information 314 at the time of capture acquired via the camera information acquisition unit 232, and corresponds the object based on the positional relationship between the projected point and the outer frame 320 of the object on the image. The point 326 of the object on the map may be, for example, the center of gravity of the object on the map obtained from map information. If the object is a house, the three-dimensional geographic coordinates of the center of gravity of the house can be calculated from the three-dimensional geographic coordinates of each of the multiple points constituting the perimeter of the house.

[0107] The first association unit 250 associates the object on the map projected at a position overlapping the area surrounded by the outer frame 320 of the object as the object on the map corresponding to the object on the image surrounded by the outer frame 320. In this way, a pair of the object on the image and the corresponding object on the map is estimated.

[0108] The center of gravity calculation unit 242 calculates the center of gravity G of the object on the image detected by the object detection unit 240. The center of gravity calculation unit 242 may calculate the center of gravity of the outer frame 320 of the object on the image detected by the object detection unit 240 as the center of gravity G of the object on the image.

[0109] The second correspondence unit 252 includes a parameter estimation unit 254 that corrects the positional relationship between the object on the image associated by the first correspondence unit 250 and the corresponding object on the map projected onto the image, and estimates transformation parameters to minimize the positional error between them. The parameter estimation unit 254 estimates second transformation parameters 328 that perform coordinate transformation based on the correspondence between the center of gravity G of the object on the image represented by coordinates within the image and a point 326 of the object on the map (here, the center of gravity G of the object on the map) represented by three-dimensional geographic coordinates, so as to minimize the positional error (positional deviation) between the center of gravity G of the object on the image and a projection point PT obtained by projecting the center of gravity of the corresponding object on the map onto the image. The second transformation parameters 328 are transformation parameters that achieve more accurate coordinate transformation than the first transformation parameters 324.

[0110] The second correspondence unit 252 may, for example, assign a parameter value based on the value of the first transformation parameter 324, convert the center of gravity of the object on the map into coordinates in the image using the transformation parameter (camera matrix) of the changed parameter value, evaluate the error (deviation) between the transformation result and the center of gravity of the object on the captured image, and search for the parameter value that minimizes the positional error, thereby estimating the second transformation parameter 328.

[0111] The second transformation parameters 328 generated by the second association unit 252 are stored in the transformation parameter storage unit 256. The transformation parameter storage unit 256 may be a storage area of ​​the computer-readable medium 204.

[0112] [Other Functions of Image Processing Device 20] In addition to the above-described processes, the image processing device 20 may also perform the following processes.

[0113] [1] Weighting Function in the First Correspondence Process In the first correspondence process, the outer frame FR of the object on the image may be associated with the projection point PT of the center of gravity of the object on the map so that the weight of the object in the center of the image is increased. Such weighting increases the detection accuracy of the object in the center of the image, and increases the alignment accuracy of the center of the image.

[0114] Alternatively, in the first correspondence process, the outer frame FR of the object on the image may be associated with the projection point PT of the center of gravity of the object on the map so as to weight the object in the peripheral part of the image more highly. Such weighting results in higher detection accuracy for the object in the peripheral part of the image, and higher alignment accuracy for the peripheral part of the image. Alignment accuracy refers to the accuracy of the correspondence between a position on the image (image coordinates) and a position in geographic space (geospatial coordinates).

[0115] [2] Weighting function in the second matching process When multiple objects are captured in the central and peripheral parts of an image, in the second matching process, the processor 202 may assign different weights to objects located in the central part of the image and objects located in the peripheral parts of the image to evaluate the overall position error in the image and estimate the transformation parameters so as to minimize the overall position error.

[0116] For example, the processor 202 may place emphasis on alignment in the center of the image, weight the evaluation of the position error for objects located in the center of the image and the position error for objects located in the periphery of the image, and determine an overall evaluation value that emphasizes the position error in the center.

[0117] In the second matching process, the position of the object on the image is matched with the position of the object on the map so as to give a higher weight to the object in the center of the image, thereby increasing the accuracy of alignment in the center of the image. Alternatively, in the second matching process, the position of the object on the image is matched with the position of the object on the map so as to give a higher weight to the object in the peripheral part of the image, thereby increasing the accuracy of alignment in the peripheral part of the image.

[0118] [3] Function of Using Objects in Each of the Corner and Center Regions of the Image In the second matching process, the processor 202 may estimate transformation parameters using objects in each of the corner and center regions of the image so as to minimize the overall position error within the image plane. The corner regions of the image refer to the regions near the four corners of a rectangular image region. If no objects exist in some of the corner and center regions, the transformation parameters may be estimated using existing objects so as to minimize the overall position error within the image plane.

[0119] [4] Function of Removing Outlier Points In the second matching process, the processor 202 may be configured to estimate transformation parameters while removing outlier points so as to minimize the positional error between each of multiple objects in the image and the corresponding object on the map. Outlier points refer to pairs of object points that have significantly larger positional deviations than the other points. By removing the outlier points and minimizing the positional error for the remaining objects, transformation parameters with high registration accuracy can be obtained.

[0120] [5] Function of displaying objects on a map on a captured image Preferably, the processor 202 is configured to convert the geospatial coordinates of an object on a map into coordinates within the image (coordinates within the camera's field of view) and display the object on the map superimposed on the captured image. More preferably, the processor 202 is configured to execute both a process of superimposing an object on a map transformed by applying the first transformation parameter 324 on the captured image and a process of superimposing an object on a map transformed by applying the second transformation parameter on the captured image.

[0121] By displaying the object on the map that has been transformed using the first transformation parameters 324 superimposed on the captured image, it is easy for the user to manually correct the association in the first association process.

[0122] By superimposing and displaying the object on the map that has been transformed using the second transformation parameter 328 on the captured image, the user can confirm whether the alignment using the estimated transformation parameter (second transformation parameter 328) is correct.

[0123] [6] Manual alignment adjustment function It is preferable that the processor 202 is configured to accept an operation to manually match an object on the image with a corresponding object on the map based on the result of overlaying the object on the map on the captured image, and to execute a process to match the object on the image with the object on the map in accordance with the user's operation instructions.

[0124] For example, the processor 202 may automatically match map data containing information about the location of houses to the captured image, and then accept operations to move polygons indicating the location of individual houses on the image, and may be equipped with a manual position adjustment function that fine-tunes the position to an even more optimal position in accordance with the user's operations.

[0125] [7] Cooperation with house cutout processing When highly accurate alignment between the captured image and the geospatial coordinates of map information is achieved, the area (partial image) of each house shown in the captured image can be cut out by matching it with the map information. The area of ​​each house may be cut out, for example, by a rectangular frame that encompasses the area of ​​the house. When cutting out a house, it is desirable to use the height data of the house in addition to the position data of points that make up the exterior perimeter of the house to determine the coordinates in the image for the shape of the roof, and to determine the entire area of ​​the house including the roof. The cut-out image of the house is saved, linked to a house ID that identifies the house.

[0126] [8] Collaboration with automatic house damage assessment function: Images of houses extracted from captured images are input to a processing unit of an automatic house damage assessment AI (Artificial Intelligence) that automatically determines the degree of damage to affected houses, and the degree of damage is assessed. This makes it possible to streamline damage survey work.

[0127] [Example of a damage assessment survey support system using image processing device 20] The image processing device 20 can be applied to systems for various purposes that link captured images with geospatial information. Below, as a specific application example of the image processing device 20, a damage assessment support system that serves as an information processing system useful for understanding the damage status of houses due to a disaster or the like will be described.

[0128] 10 is a flowchart showing an example of processing in a damage assessment investigation support system to which the image processing device according to the embodiment is applied. Note that, although the explanation here is given as an example of the operation by the processor 202 of the image processing device 20, some or all of the steps shown in FIG. 10 may be executed by a processor of an information processing device other than the image processing device 20.

[0129] In step S21, an image of the area to be surveyed is taken, and the processor 202 acquires the image taken by the camera 14 and camera information at the time of taking the image. Step S21 may be a step that includes the same processes as steps S11 and S12 in FIG.

[0130] In step S22, the processor 202 detects houses from the image and performs a process of cutting out the houses from the image. The "cutout process" may be understood as a process of extracting house regions. Step S22 may include a process similar to step S13 of FIG. 4. The processor 202 may cut out regions surrounded by the outlines of the houses detected from the image for each house. The processor 202 may also accept a designation of a house region to be cut out from the captured image. The user can specify the house to be cut out using a UI (User Interface) such as the input device 222. An operation to individually designate a target house may be accepted, or an area including multiple houses may be designated, and each of the multiple houses included in the designated area may be designated as a target house for the cutout process. In addition to an operation to designate an individual house or an operation to comprehensively designate multiple houses within the designated area, an operation menu such as "select all houses at once" may be provided to designate all houses in the captured image. The processor 202 executes the house segmentation process in accordance with the received instruction.

[0131] In order to grasp the geospatial information of each house in the image extracted in step S22, in step S23, the processor 202 associates the image with geographic coordinates. Note that the processing of step S23 may be executed before the processing of step S22 or may be executed in parallel with the processing of step S22.

[0132] Step S23 may include the same processes as steps S14 to S17 in Fig. 4. That is, in the process of step S23, correspondence is performed using the transformation parameters generated according to the flowchart in Fig. 4. This allows accurate correspondence of geographic coordinates to each house in the image, thereby achieving highly accurate georeferencing.

[0133] In step S24, the processor 202 determines the degree of damage to the house cut out from the image. This determination process may be performed by AI. The image of the cut-out house is input, for example, to an image recognition device (not shown), and the damage status of the house is automatically determined through image recognition. The image recognition device may be configured to use a trained model trained by machine learning. The trained model may be a deep learning model such as a neural network. The processing function of the image recognition device may be incorporated into the image processing device 20, or may be implemented in an image processing server or cloud server (not shown) connected via the network 22.

[0134] Steps S23 and S24 provide information on the degree of damage for each house on the image that is accurately linked to the geographic information.

[0135] By using the geographic information and damage level information for each house in the image thus obtained, the processor 202 can support the formulation of an efficient damage assessment survey plan in step S26. The formulation of the damage assessment survey plan includes the creation of a survey route and / or schedule for dispatching surveyors to the site to survey and confirm the damage situation there. After step S26, the processor 202 ends the flowchart of FIG. 10.

[0136] Fig. 11 is a functional block diagram of the image processing device 20 that executes the flowchart of Fig. 10. In Fig. 11, elements common to the configuration shown in Fig. 9 are assigned the same reference numerals, and duplicated explanations will be omitted. In addition to the configuration described in Fig. 4, the image processing device 20 includes an object extraction unit 244, a damage degree determination unit 246, a georeferencing unit 270, a geocoding unit 272, a damage certification survey plan formulation unit 280, a display image generation unit 282, and a display control unit 284.

[0137] The object cutout unit 244 performs a process of cutting out an image area of ​​an object in the image detected by the object detection unit 240 from the captured image 300. The object cutout unit 244 executes step S21 in Fig. 10 and performs a process of cutting out an area of ​​a house from the captured image 300.

[0138] The damage level determination unit 246 determines the damage level of a house from an image of the house cut out from the captured image 300 by the object cutout unit 244. The damage level determination unit 246 includes a damage level determination AI 247 that applies a learned model trained by machine learning. The damage level determination AI 247 is an AI module that receives an input of an image of the house, estimates the damage level of the house, and outputs the estimated damage level. The damage level determination AI 247 may be a classification model that classifies the input image into labels according to the damage level, or may be a regression model that infers a numerical value of the damage level.

[0139] The georeferencing unit 270 is a processing unit including the centroid calculation unit 242, the first association unit 250, and the second association unit 252 in Fig. 4, and executes step S23 in Fig. 10. The georeferencing unit 270 also applies the transformation parameters stored in the transformation parameter storage unit 256 to execute a coordinate transformation process between the image coordinate system and the geographic space coordinate system.

[0140] The geocoding unit 272 performs a process of associating geospatial coordinates with information such as addresses. The geocoding unit 272 performs a process of converting geospatial coordinates acquired via the georeferencing unit 270 or the map information acquisition unit 236 into geographic information such as addresses. The geocoding unit 272 can also convert geographic information such as addresses into geospatial coordinates.

[0141] The damage assessment survey plan formulation unit 280 performs processing to support the formulation of a damage assessment survey plan for a house based on the information on the damage level of the house obtained from the damage level determination unit 246 and the geospatial information linked to the house. The damage assessment survey plan formulation unit 280 supports the creation of a survey route and a survey schedule in accordance with instructions from the input device 222.

[0142] The display image generation unit 282 performs processing to generate an image to be displayed on the display device 224. The display image generation unit 282 applies the transformation parameters stored in the transformation parameter storage unit 256 to convert three-dimensional geographic coordinate data obtained from the map information into coordinates within the image, and from the transformation result, can generate a composite image for display in which map information aligned with the photographed image is superimposed on the photographed image. Furthermore, the display image generation unit 282 can generate various images, such as a display image for displaying the determination result of the damage degree determination unit 246 for each house, a display image for displaying the investigation route of the damage assessment investigation, and a display image for displaying the investigation schedule.

[0143] The display control unit 284 generates data for display on the display device 224. The display image, such as the composite image generated by the display image generation unit 282, is displayed on the display device 224 via the display control unit 284.

[0144] [Regarding the hardware configuration of each processing unit] The hardware structure of the processing units in the image processing device 20 that perform various processes, such as the image acquisition unit 230, camera information acquisition unit 232, map information acquisition unit 236, object detection unit 240, center of gravity calculation unit 242, object cut-out unit 244, damage level determination unit 246, first matching unit 250, second matching unit 252, georeferencing unit 270, geocoding unit 272, damage assessment survey plan formulation unit 280, display image generation unit 282 and display control unit 284, is, for example, various processors as shown below.

[0145] Various types of processors include CPUs, which are general-purpose processors that execute programs and function as various processing units, GPUs, which are processors specialized for image processing, programmable logic devices (PLDs), such as FPGAs (Field Programmable Gate Arrays), which are processors whose circuit configuration can be changed after manufacture, and dedicated electrical circuits, such as ASICs (Application Specific Integrated Circuits), which are processors with a circuit configuration designed specifically for executing specific processes.

[0146] A single processing unit may be configured with one of these various processors, or may be configured with two or more processors of the same or different types. For example, a single processing unit may be configured with multiple FPGAs, a combination of a CPU and an FPGA, or a combination of a CPU and a GPU. Multiple processing units may also be configured with a single processor. Examples of multiple processing units configured with a single processor include, first, a configuration in which a single processor is configured with a combination of one or more CPUs and software, as typified by computers such as client and server computers, and this processor functions as multiple processing units. Second, a configuration in which a processor is used to realize the functions of an entire system including multiple processing units on a single IC (Integrated Circuit) chip, as typified by a system-on-chip (SoC). In this way, the various processing units are configured with one or more of the above-mentioned various processors as a hardware structure.

[0147] Furthermore, the hardware structure of these various processors is, more specifically, an electric circuit made up of a combination of circuit elements such as semiconductor elements.

[0148] [Regarding the program that operates the computer] A program that causes a computer to realize some or all of the processing functions of the image processing device 20 can be recorded on a computer-readable medium, such as an optical disk, a magnetic disk, a semiconductor memory, or other tangible, non-transitory information storage medium, and the program can be provided through this information storage medium.

[0149] In addition, instead of providing the program by storing it on such a tangible, non-transitory computer-readable medium, it is also possible to provide the program signal as a download service using a telecommunications line such as the Internet.

[0150] Furthermore, some or all of the processing functions of the image processing device 20 may be realized by cloud computing, and may also be provided as a SaaS (Software as a Service) service.

[0151] Advantages of this embodiment The image processing device 20 according to this embodiment has the following advantages.

[0152] [1] It is possible to accurately link captured images with geospatial coordinates without the need to install ground control points (GCPs) in advance.

[0153] [2] The image processing device 20 can accurately link a house in an image with its geospatial information, enabling a quick understanding of the situation on the ground. This can also contribute to the quick issuance of disaster damage certificates.

[0154] [Variation 1] In the second matching process, the calculation method is not limited to minimizing the sum of the distances between the center of gravity G of the object on the image and the projection point PT obtained by projecting the center of gravity of the corresponding object on the map onto the image. For example, a calculation method may be applied in which the overlapping area between the object regions is maximized. In this case, the processor 202 may calculate, as the evaluation value, the sum of the areas of the regions enclosed by the object's outer frame FR detected from the image and the region of the projected figure obtained by projecting the outer periphery of the object on the map onto the image. The greater the degree of overlap (degree of overlap) between the object regions, the smaller the sum of the areas of the regions enclosed by the object's outer frame FR and the region of the projected figure obtained by projecting the outer periphery of the object on the map onto the image. Alternatively, the processor 202 may calculate, as the evaluation value, the sum of the areas of the overlapping areas between the corresponding object regions and correct the positional relationship between the object on the image and the object on the map so as to maximize this evaluation value.

[0155] The first association process is also not limited to the mode of calculating the representative points of the object, and association may be performed based on the shape and / or area of ​​the region of the object.

[0156] [Variation 2] The processing functions of the image processing device 20 may be realized by a plurality of computers or by cloud computing. The processing functions of the image processing device 20 may be implemented in the remote controller 16 and / or the terminal device 24.

[0157] [Variation 3] In the above embodiment, an example of processing a still image as a captured image has been described, but the camera 14 may also capture a video, and the image processing device 20 may extract some frames from the captured video and perform similar processing.

[0158] [Variation 4] In the above embodiment, an example of a process of projecting an object in map information onto an image and superimposing it on a captured image has been described. However, a process of projecting an object on an image onto a map and superimposing it on the map image may also be performed.

[0159] [Other Application Examples] In the above embodiment, an example is given of processing images captured by the camera 14 mounted on the drone 12, but the scope of application of the present disclosure is not limited to this example. For example, images captured using a camera installed at a high location overlooking the ground, such as on the roof of a building or on a steel tower, may be processed.

[0160] [Others] The present disclosure is not limited to the above-described embodiments, and various modifications are possible within the scope of the gist of the technical idea of ​​the present disclosure.

[0161] REFERENCE SIGNS LIST 10 Photographing image processing system 12 Drone 13 Gimbal head 14 Camera 16 Remote controller 16A Display 20 Image processing device 22 Network 24 Terminal device 24A Display 30 GNSS receiver 32 Barometric pressure sensor 34 Orientation sensor 36 IMU 38 Motor 40 Processor 42 Storage device 44 Communication interface 202 Processor 204 Computer readable medium 206 Communication interface 208 Input / output interface 210 Bus 212 Memory 214 Storage 222 Input device 224 Display device 230 Image acquisition unit 232 Camera information acquisition unit 233 Camera position information acquisition unit 234 Camera attitude information acquisition unit 236 Map information acquisition unit 240 Object detection unit 242 Center of gravity calculation unit 244 Object extraction unit 246 Damage degree determination unit 250 First correspondence unit 252 Second correspondence unit 254 Parameter estimation unit 256 Conversion parameter storage unit 260 Map database 270 Georeferencing unit 272 Geocoding unit 280 Damage certification survey plan formulation unit 282 Display image generation unit 284 Display control unit 300 Photographed image 312 Camera position information 314 Camera attitude information 320 Outer frame 324 First conversion parameter 326 Point of object on map 328 Second conversion parameter 362 Gyro sensor 364 Acceleration sensor 366 Temperature sensor 247 Damage degree determination AI IM Image IM1 Photographed image FR Outer frame HS House G Center of gravity PT Projection point S11-S11 Steps of image processing method S21-S26 Processing steps in the damage assessment survey support system

Claims

1. An image processing device having one or more processors, wherein the one or more processors perform the following processes: acquiring an image captured by a camera; acquiring camera information including information on the position and attitude of the camera when the image was captured; detecting an object from the image; using the camera information to associate the object in the image with an object on a map included in map information; and estimating transformation parameters for the coordinate transformation from the positional relationship between the associated object in the image and the object on the map so as to reduce positional errors of the object due to coordinate transformation between geospatial coordinates and coordinates in the image.

2. The image processing device of claim 1, wherein the process of detecting the object from the image includes a process of acquiring information indicating the area of ​​the object on the image, and the matching process includes a first matching process that matches the area of ​​the object on the image with the object on the map by conversion using the camera information.

3. The image processing device according to claim 2, wherein the one or more processors obtain a representative point of the object on the map from the map information, and the first matching process includes matching an area of ​​the object on the image with the representative point of the object on the map.

4. The image processing device of claim 3, wherein the first matching process includes matching the area of ​​the object on the image with the representative point of the object on the map based on the positional relationship between the point where the representative point of the object on the map is projected onto the image by transformation using the camera information and the area of ​​the object on the image.

5. The image processing device according to claim 3, wherein the representative point is a center of gravity.

6. The image processing device according to claim 2, wherein the information indicating the area of ​​the object on the image includes information on an outer frame surrounding the area of ​​the object.

7. The image processing device according to claim 2, wherein the information indicating the region of the object on the image includes segmentation information in which the region of the object is identified on a pixel-by-pixel basis.

8. The image processing device according to claim 2, wherein the first association process includes associating each of the plurality of objects shown in the image with the object on the map included in the map information.

9. The image processing device according to claim 2, wherein the first matching process includes matching the object in a central portion of the image so as to give a higher weight to the object.

10. The image processing device according to claim 2, wherein the first matching process includes matching the objects so as to give a higher weight to the objects in the peripheral portion of the image.

11. The image processing device according to claim 2, wherein the process of estimating the transformation parameters includes a second matching process that corrects the positional relationship between the object on the image that has been matched by the first matching process and the object on the map.

12. The image processing device described in claim 11, wherein the one or more processors further execute a process of calculating a representative point of the object on the image detected from the image, and the second matching process includes matching the coordinates in the image with the geospatial coordinates by correcting the positional relationship between the representative point of the object on the image matched by the first matching process and a point obtained by projecting the corresponding representative point of the object on the map onto the image.

13. The image processing device according to claim 11, wherein the second matching process includes estimating the transformation parameters using the object in each of the four corners and the central area of ​​the image so as to minimize the position error.

14. The image processing device according to claim 11, wherein the second matching process includes estimating the transformation parameters by using a plurality of the objects on the image and removing outlying points so as to minimize the position error.

15. The image processing device according to claim 1, wherein the one or more processors further execute a process of converting the object on the map into coordinates within the image of the image, and displaying the result of the conversion on the image in an overlaid manner.

16. The image processing device according to claim 15, further comprising a display device that displays the image and the conversion result in an overlapping manner.

17. The image processing device described in claim 15, wherein the one or more processors further accept an operation to manually match the object on the image with a corresponding object on the map based on the transformation result overlaid on the image, and execute a process to match the object on the image with the object on the map in accordance with the instructions of the accepted operation.

18. The image processing device according to claim 17, further comprising an input device for manually inputting the instruction.

19. The image processing device described in claim 1, wherein the image is an aerial image taken by the camera mounted on an aircraft, and the camera information is obtained from sensor data obtained by a sensor placed on at least one of the camera and the aircraft.

20. The image processing device according to claim 1, wherein the object includes at least one of a house, a manhole, and a road intersection.

21. The image processing device according to claim 1, wherein the object includes a house, and the one or more processors further execute a process of extracting the object, which is a house, from the image and determining the degree of damage to the extracted house.

22. An image processing method comprising the steps of: acquiring an image taken by a camera; acquiring camera information including information about the position and attitude of the camera when the image was taken; detecting an object from the image; using the camera information to associate the object on the image with an object on a map included in map information; and estimating transformation parameters for the coordinate transformation from the positional relationship between the associated object on the image and the object on the map so as to reduce positional errors of the object due to coordinate transformation between geospatial coordinates and coordinates in the image.

23. A program that enables a computer to implement the following functions: acquire an image taken by a camera; acquire camera information including information about the position and attitude of the camera when the image was taken; detect an object from the image; use the camera information to associate the object in the image with an object on a map included in map information; and estimate transformation parameters for the coordinate transformation from the positional relationship between the associated object on the image and the object on the map so as to reduce positional errors of the object due to coordinate transformation between geospatial coordinates and coordinates in the image.

24. A non-transitory computer-readable recording medium on which the program according to claim 23 is recorded.

Citation Information

Patent Citations

  • Map information collation and update system

    JP1993181411A

  • Method and device for updating map information

    JP1999328378A

  • Map information updating method and map updating device

    JP2000310940A

  • Image processing device, image processing method, and program

    WO2023047799A1

  • House state provision device and method

    WO2023167017A1