Learning device, inference device, image generation system, change detection system, image generation method, and image generation program
The technology generates simulated remote sensing images using a trained model based on land use data to address the challenges of capturing images during normal times, ensuring accurate change point extraction by accounting for land use and time variations.
Patent Information
- Application Number
- PCT/JP2024/039842
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-07-23
- Filing Date
- 2024-11-08
- Publication Date
- 2026-01-29
AI Technical Summary
Existing methods for calculating disaster locations from remote sensing images face challenges due to the impracticality of comprehensively capturing images during normal times and the generation of unsuitable images due to differences in image features, such as green spaces and farmland, which can lead to poor analysis results.
A technology that generates simulated remote sensing images using a trained model based on land use data, allowing for the inference of images without the need for comprehensive time and area capture, and accounting for differences in image features.
Enables the generation of suitable simulated images for analysis, facilitating accurate extraction of change points without the need for extensive image capture, and accounting for land use and time variations.
Smart Images

Figure JP2024039842_29012026_PF_FP_ABST
Abstract
Description
Learning device, inference device, image generation system, change detection system, image generation method, and image generation program
[0001] The present disclosure relates to a learning device, an inference device, an image generation system, a change detection system, an image generation method, and an image generation program.
[0002] In technology for calculating disaster locations (such as flood damage due to flooding or landslides) from remote sensing images, a commonly used method is to extract change points by calculating the difference between a disaster image and a normal image that includes the same points as the disaster locations. A disaster image is a remote sensing image taken after the disaster occurs (after the disaster occurs) and that includes the same points as the disaster locations. A normal image is a remote sensing image taken before the disaster occurs. Patent Document 1 discloses a technology for calculating the difference between two remote sensing images.
[0003] JP 2019-12400 A
[0004] Disaster images can be obtained by capturing new images after a disaster occurs. Meanwhile, normal-time images containing the same location as the disaster site have traditionally been searched for from past images (archived images). In this case, the disaster images and normal-time images must be captured using the same type of sensor (e.g., optical images, SAR (Synthetic Aperture Radar) images, or infrared images), and it is desirable that the corresponding features (resolution) be the same or similar. However, comprehensively capturing images of both domestic and international locations during normal times and continuously building an archive is not practical due to limited image capture opportunities. Therefore, there has been a problem in that it is sometimes impossible to obtain remote sensing images equivalent to past images including the disaster area, and it is therefore impossible to prepare normal-time images corresponding to the disaster images. Furthermore, even if normal-time images corresponding to the disaster images exist, only a combination of two images that is not suitable for analysis can be used, resulting in poor results due to significant differences in image features, such as green spaces and farmland, between the normal-time and disaster-time images. This problem can occur for reasons such as the date on which the normal-time image was taken being far removed from the time of the disaster, or the season on which the normal-time image was taken being different from the season on which the disaster occurred.
[0005] The present disclosure aims to provide a technology for extracting change points based on the difference between two remote sensing images that include the same location, which enables the simulated generation of remote sensing images so that it is not necessary to take remote sensing images comprehensively in terms of time and area during normal times, and so that images that are unsuitable for analysis are not generated due to differences in image features depending on the combination of land use status and time period.
[0006] The learning device according to the present disclosure includes a model generation unit that generates a trained model that uses reference land use data indicating the land use status in an area of interest as input and infers a simulated image, which is a simulated remote sensing image showing the area of interest, by learning the correspondence between each piece of land use data indicating the land use status in at least a portion of an area to be learned and each reference image in a reference image group consisting of one or more reference images, each of which is a remote sensing image showing an area corresponding to each piece of land use data.
[0007] According to the present disclosure, a trained model is generated that infers a simulated remote sensing image based on land use data indicating land use conditions in at least a portion of a training area and reference images indicating areas corresponding to the land use data. Here, by generating the trained model using reference images captured at desired times, the trained model can infer a simulated image corresponding to a desired combination of land use conditions and time. Therefore, by utilizing the present disclosure, in a technology for extracting change points based on the difference between two remote sensing images containing the same location, it is possible to generate simulated remote sensing images without the need to comprehensively capture remote sensing images in terms of time and area on a regular basis, and without generating images unsuitable for analysis due to differences in image features depending on the combination of land use conditions and time.
[0008] FIG. 1 is a diagram showing an example of the configuration of an image generation system 90 according to the first embodiment. FIG. 2 is a diagram explaining an overview of the image generation system 90 according to the first embodiment. FIG. 3 is a diagram showing an example of the configuration of a learning device 20 according to the first embodiment. FIG. 4 is a diagram showing an example of the configuration of an inference device 30 according to the first embodiment. (a) is a diagram explaining an overview of a shadow, and (b) is a diagram explaining problems with a shadow. FIG. 5 is a diagram showing an example of the hardware configuration of each device included in the image generation system 90 according to the first embodiment. A flowchart showing the operation of the learning device 20 according to the first embodiment. A diagram explaining the learning phase according to the first embodiment. A flowchart showing the operation of the inference device 30 according to the first embodiment. A diagram explaining the utilization phase according to the first embodiment. A diagram showing an example of the hardware configuration of each device included in the image generation system 90 according to a modification of the first embodiment. A diagram showing an example of the configuration of a change detection system 91 according to a modification of the first embodiment. A diagram showing an example of the configuration of a change detection device 40 according to a modification of the first embodiment.
[0009] In the description of the embodiments and drawings, the same elements and corresponding elements are given the same symbols. The description of elements given the same symbols is omitted or simplified as appropriate. Arrows in the drawings mainly indicate the flow of data or the flow of processing. Furthermore, "unit" may be interpreted as "circuit," "device," "equipment," "process," "step," "procedure," "processing," or "circuitry" as appropriate. The functions of each unit provided in each device may be realized by firmware, software, hardware, or a combination of these.
[0010] First Embodiment Hereinafter, the present embodiment will be described in detail with reference to the drawings.
[0011] ***Description of Configuration*** FIG. 1 shows an example configuration of an image generation system 90 according to this embodiment. As shown in FIG. 1, the image generation system 90 includes a learning device 20, an inference device 30, and a trained model storage unit 203. The learning device 20 and the inference device 30 may be configured integrally. The image generation system 90 is also called an observation image simulation generation system. Specific examples of remote sensing images include optical images, SAR (Synthetic Aperture Radar) images, and infrared images. The remote sensing images may be images captured from an aircraft or from an artificial satellite.
[0012] The learning device 20 is a device that executes the processing of the learning phase and generates a trained model by learning the correspondence between land use status and remote sensing images, or by retraining a trained model stored in the trained model storage unit 203 to generate a trained model. The trained model is a model that infers simulated remote sensing images of a specified area based on data indicating the land use status of the specified area. The trained model may infer remote sensing images based on other data related to the characteristics of the area, the characteristics of the time period, or the imaging conditions. The inferred remote sensing images may be used to extract change points (disaster locations) by calculating differences from disaster images that include points included in the remote sensing images. In this specification, remote sensing images may also be simply referred to as "images." Disaster images are remote sensing images captured after a disaster occurs (after the disaster occurs) and that include the same points as the disaster locations. Each period is one of multiple periods divided over a long period of time. Specific examples of each period include one of a number of periods in a year, one of the first, middle, and last parts of a month, or a week. The boundary between two consecutive periods may be determined based on the time at which an event that may cause an image difference is likely to occur in each region. The length of each period does not have to be constant. Each period may be information indicating a combination of whether or not each event that may cause an image difference has occurred in each region, rather than a period of time. Each period may be a period determined by time or time zone.
[0013] The inference device 30 is a device that executes processing in the inference phase, and infers remote sensing images using the trained model stored in the trained model storage unit 203.
[0014] FIG. 2 is a diagram illustrating an overview of the image generation system 90. The image generation system 90 may generate simulated remote sensing images of normal times (pre-disaster) at the same time and location as images captured at the time of a disaster from a land use map obtained from the analysis of remote sensing images. The image generation system 90 generates simulated remote sensing images according to specified conditions, without particular time or area limitations, based on a trained model generated by learning the correspondence between multiple remote sensing images limited in time and area and land use maps. The generated remote sensing images may be used for comparison with remote sensing images captured after a disaster. As shown in FIG. 2, when the image generation system 90 is used in Japan, the image generation system 90 samples the earth's surface throughout the country in normal times (pre-disaster) using a SAR satellite and generates a trained model based on the correspondence between the sampled SAR satellite images and data (land use maps) showing the land cover of the locations included in the sampled SAR satellite images. The image generation system 90 then inputs data showing land cover across the country into the trained model to generate a simulated SAR image showing the land surface across the entire country before the disaster.
[0015] 3 shows an example of the configuration of the learning device 20. The learning device 20 includes a data acquisition unit 201 and a model generation unit 202.
[0016] The data acquisition unit 201 acquires various data for training the trained model and passes the acquired various data to the model generation unit 202. The various data may include, as a specific example, a reference image DIN1 and land use data DIN2. The various data may include at least one of reference imaging conditions DIN3, time data DIN10, agricultural history data DIN11, and GIS (Geographic Information System) data DIN12. The reference image DIN1 is an actually captured remote sensing image. The land use data DIN2 indicates the land use status of at least a portion of the training target area. The training target area is an area in which actually captured remote sensing images that can be used as sampling data exist, and may be any defined area. As a specific example, the land use data DIN2 is data indicating a land use map. The land use map may be a commonly available one or may be generated from the reference image DIN1. The reference imaging conditions DIN3 may be used when generating the land use map. The land use data DIN2 may include data indicating detailed land use conditions, such as at least one of data indicating the type of crops in agricultural land, data indicating the type of plants in green spaces, and data indicating the height and shape of each building. The reference imaging conditions DIN3 consist of the conditions when the reference image DIN1 was captured. The reference imaging conditions DIN3 may include imaging parameters as each condition. The reference imaging conditions DIN3 may also be included as metadata in the reference image DIN1. The time data DIN10 is data indicating the time when the reference image DIN1 was captured. The agricultural history data DIN11 is data indicating agricultural history. The GIS data DIN12 may be any GIS data. Specific examples of the GIS data DIN12 include data indicating a DEM (Digital Elevation Model) or a DSM (Digital Surface Model). The elevation information indicated by the GIS data DIN12 may be used to analyze areas prone to flooding and areas unlikely to flooding.The GIS data DIN12 may include data indicating latitude and longitude, or data indicating the region or prefecture where the land is located. Because GIS data refers to geographic information, the land use data DIN2 is also included in the concept of GIS data in a broad sense. However, in this specification, land use data is described separately from GIS data. The reference image DIN1, land use data DIN2, and GIS data DIN12 input to the data acquisition unit 201 contain geospatial location information. The data acquisition unit 201 matches the location information to obtain coordinate correspondence between each piece of information, and then passes each piece of information to the model generation unit 202. In this case, the data acquisition unit 201 may convert the input data into data with a three-dimensional structure based on the location information, and then pass this data to the model generation unit 202 as a single piece of data.
[0017] The model generation unit 202 generates a trained model by learning the correspondence between each piece of land use data DIN2 and each reference image DIN1 in the reference image group. The reference image group consists of one or more reference images DIN1, each of which is a remote sensing image showing an area corresponding to each piece of land use data DIN2. Specifically, the area corresponding to the land use data DIN2 is the area targeted by the land use data DIN2, i.e., the area shown in the land use data DIN2. In this case, the trained model is a model that infers the simulated image DOUT using the reference land use data DIN4 as input. Specifically, the model generation unit 202 trains or retrains the trained model based on various data acquired by the data acquisition unit 201, and stores the trained model generated by training or retraining in the trained model storage unit 203. Specifically, the model generation unit 202 generates the trained model using machine learning or AI (artificial intelligence) technology. The model generation unit 202 trains a trained model that infers the reference image DIN1 from the land use data DIN2 based on the reference image DIN1 and the land use data DIN2 corresponding to the reference image DIN1. The model generation unit 202 may further learn the correspondence between the time data DIN10 indicating the time when each reference image DIN1 in the reference image group was captured and each reference image DIN1 in the reference image group, thereby generating, as a trained model, a model that infers a simulated remote sensing image DOUT corresponding to the time indicated by the input time data, using input time data indicating the time as an input. The input time data corresponds to the time data DIN10. The model generation unit 202 may further learn the correspondence between the agricultural history data DIN11 indicating the agricultural history in the area indicated by each reference image DIN1 of the reference image group and each reference image DIN1 of the reference image group, and thereby generate, as a trained model, a model that further inputs input agricultural history data indicating the agricultural history in the area of interest and infers a simulated remote sensing image corresponding to the input agricultural history data as a simulated image DOUT. The input agricultural history data corresponds to the agricultural history data DIN11. The area of interest does not need to be included in the area to be trained.The model generation unit 202 may further learn the correspondence between the GIS data DIN12 in the region indicated by each reference image DIN1 of the reference image group and each reference image DIN1 of the reference image group, thereby further inputting input GIS data that is the GIS data DIN12 in the region of interest, and generating a trained model that infers a simulated remote sensing image corresponding to the input GIS data as a simulated image DOUT.The model generation unit 202 may further learn the correspondence between the reference imaging conditions DIN3 consisting of the imaging conditions when each reference image DIN1 of the reference image group was captured, and each reference image DIN1 of the reference image group, thereby further inputting simulated imaging conditions DIN5 consisting of the imaging conditions of the simulated image DOUT, and generating a trained model that infers a simulated remote sensing image corresponding to the simulated imaging conditions as a simulated image DOUT.
[0018] The trained model storage unit 203 stores trained models.
[0019] 4 shows an example of the configuration of the inference device 30. The inference device 30 includes a data acquisition unit 301 and an inference unit 302.
[0020] The data acquisition unit 301 acquires various data for inferring remote sensing images and passes the acquired various data to the inference unit 302. Specific examples of the various data include reference land use data DIN4 and simulated imaging conditions DIN5. The various data may include at least one of time data DIN10, agricultural history data DIN11, and GIS data DIN12. The reference land use data DIN4 is similar to the land use data DIN2 and is data indicating the land use status in the area of interest. The area of interest is the area that is the target of inference for the remote sensing image. The simulated imaging conditions DIN5 are similar to the reference imaging conditions DIN3 and indicate various conditions related to the simulated image DOUT.
[0021] The inference unit 302 infers the simulated image DOUT using the various data acquired by the data acquisition unit 301 and the trained model. That is, the inference unit 302 infers the simulated image DOUT using the reference land use data DIN4 as input to the trained model. In this case, the trained model is a model generated by learning the correspondence between each piece of land use data DIN2 and each reference image DIN1 in the reference image group. The inference unit 302 may further input input time data indicating a time to the trained model and infer, as the simulated image DOUT, a simulated remote sensing image corresponding to the time indicated by the input time data. In this case, the trained model is a model generated by further learning the correspondence between each reference image DIN1 in the reference image group and time data DIN10 indicating the time when each reference image DIN1 in the reference image group was captured. The inference unit 302 may further use input agricultural history data indicating the agricultural history in the region of interest as input to the trained model, and infer a simulated remote sensing image corresponding to the input agricultural history data as simulated image DOUT. In this case, the trained model is a model generated by further learning the correspondence between agricultural history data DIN11 indicating the agricultural history in the region indicated by each reference image DIN1 of the reference image group and each reference image DIN1 of the reference image group. The inference unit 302 may further use input GIS data, which is GIS data DIN12 in the region of interest, as input to the trained model, and infer a simulated remote sensing image corresponding to the input GIS data as simulated image DOUT. In this case, the trained model is a model generated by further learning the correspondence between GIS data DIN12 in the region indicated by each reference image DIN1 of the reference image group and each reference image DIN1 of the reference image group. The inference unit 302 may further use simulated imaging conditions DIN5 consisting of each imaging condition of the simulated image DOUT as input to the trained model and infer a simulated remote sensing image corresponding to the simulated imaging conditions DIN5 as the simulated image DOUT. In this case, the trained model is a model generated by further learning the correspondence between the reference imaging conditions DIN3 corresponding to each reference image DIN1 of the reference image group and each reference image DIN1 of the reference image group.The reference imaging conditions DIN3 are made up of imaging conditions under which each reference image DIN1 in the reference image group was captured.
[0022] The simulated image DOUT is a simulated remote sensing image showing the region of interest, which corresponds to a normal image. A normal image is a remote sensing image captured before a disaster occurs or a remote sensing image that would be captured if a disaster did not occur.
[0023] Here, the appearance of a remote sensing image may vary depending on the imaging parameters used during capture. Therefore, it is considered important to use imaging parameters as inputs during the learning phase. The imaging parameters are numerical values that can be changed or set by the user capturing the remote sensing image. The reference imaging condition DIN3 may be general data that can affect the image and can be acquired or recognized by the user, regardless of whether the numerical value can be set. In this specification, concepts similar to imaging conditions may also be referred to as imaging parameters. When the reference image is an optical image, specific examples of imaging parameters include the pointing angle (also known as the off-nadir angle) and the imaging time. When TDI (Time Delay Integration) imaging is used, the imaging time is the total integration time (the product of the TDI stage number and the imaging period), which corresponds to the exposure time of a typical camera, and the TDI stage number is set as the imaging parameter. When the reference image is a SAR image, specific examples of imaging parameters include the angle of incidence, the imaging azimuth angle, and the radio wave irradiation time. Even for satellites with the same specifications, the resolution can vary depending on the pointing angle, angle of incidence, and other factors. Here, the name of the satellite capturing the image and the imaging mode are both imaging parameters. This is because the specifications are determined according to the satellite's name. As a specific example, a satellite named "ALOS-2" uses L-band radio waves, and the approximate range of resolution is determined according to the imaging mode used, with the detailed resolution being further determined by information such as the angle of incidence.
[0024] If the trained model is trained using the reference image DIN1 and land use data DIN2 without considering the reference imaging conditions DIN3, the accuracy of the simulated image DOUT generated by the inference device 30 may be low. As a specific example, when multiple satellite images captured at the same location corresponding to different modes are used as training data, if the training data includes imaging parameters, the model generation unit 202 can train the trained model along with the mode indicated by the imaging parameters. On the other hand, if the training data does not contain imaging parameters (e.g., when image data with a converted image format is used, or when only image data without metadata is used as training data), each image cannot be associated with the imaging parameters in the trained model. In such cases, it becomes necessary for the user to carefully examine the training data in advance, such as by generating training data by selecting only images with similar corresponding imaging parameters. Furthermore, the simulated image DOUT that can be created during inference is limited to images corresponding to limited imaging parameters.
[0025] Furthermore, if the reference image DIN1 has a wide area that corresponds to a (radar) shadow, the shadow will also be included in the image after map projection. When learning the correspondence between the reference image DIN1 containing many shadows and the land use data DIN2, the accuracy of the trained model may be reduced. A shadow is an area where radar waves do not hit due to the relationship between the unevenness of the ground surface, such as a tall building or mountain, and the radar's angle of incidence or azimuth angle, i.e., an area that is in shadow when captured. A shadow is generally displayed as black on the image, i.e., without reflection, with a low D / N value. Figure 5 is a diagram explaining a shadow. In Figure 5(a), a shadow is generated due to the positional relationship between the unevenness of the ground surface and the SAR satellite.
[0026] Figure 5(b) illustrates the problem of shadows. In Figure 5(b), parts of forests and lakes are covered by shadows, resulting in low D / N values. When a SAR image contains a shadow, if the shadow is used for training, the land use status corresponding to the shadow will be associated with an image that has different characteristics from the original characteristics of the image corresponding to the land use status. Therefore, to prevent such inaccurate associations, it is desirable not to use the shadow for training. Possible methods for preventing the association of shadows with land use map classifications when training by associating images with land use maps include automatically determining the shadow and not including it in the training data, or by assigning a low weight to the shadow for learning parameters commonly used in deep learning. Here, when the inference device 30 is able to use both GIS data including height data of buildings or land, etc., and imaging parameters for the simulated image DOUT, a configuration may be adopted in which the user can select whether or not to include shadows in the simulated image DOUT. However, from the perspective of comparing images before and after a disaster, it is preferable to create the simulated image DOUT with shadows by matching the simulated imaging conditions DIN5 as closely as possible to the imaging parameters of the disaster image. This is because actual SAR images contain shadows, and it is desirable that the shadow-related conditions be consistent when subtracting the simulated image DOUT from the disaster image. Furthermore, optical images may contain clouds (including thin clouds that obscure the ground surface but are transparent). Unlike shadows, clouds do not always reside in the same location. However, when the reference image DIN1 contains clouds, as with shadows in SAR images, there is a possibility that the land use situation corresponding to the cloudy area may be associated and learned with an image that has characteristics different from the inherent characteristics of the image corresponding to that land use situation. Therefore, a method similar to that for dealing with shadows may be used. Common cloud detection techniques may be used to extract cloud areas from within the image.
[0027] 6 shows an example of the hardware configuration of each device included in the image generation system 90 according to this embodiment. Each device is made up of a computer. Each device may also be made up of multiple computers.
[0028] As shown in the figure, each device is a computer equipped with hardware such as a processor 11, a memory 12, an auxiliary storage device 13, an input / output IF (Interface) 14, and a communication device 15. These pieces of hardware are connected as appropriate via signal lines 19.
[0029] The processor 11 is an integrated circuit (IC) that performs arithmetic processing and controls the hardware of a computer. Specific examples of the processor 11 include a central processing unit (CPU), a digital signal processor (DSP), or a graphics processing unit (GPU). Each device may include multiple processors that serve as the processor 11. The multiple processors share the role of the processor 11.
[0030] The memory 12 is typically a volatile storage device, specifically a random access memory (RAM). The memory 12 is also called a primary storage device or a main memory. Data stored in the memory 12 is saved in the secondary storage device 13 as needed.
[0031] The auxiliary storage device 13 is typically a non-volatile storage device, and specific examples thereof include a ROM (Read Only Memory), an HDD (Hard Disk Drive), or a flash memory. Data stored in the auxiliary storage device 13 is loaded into the memory 12 as needed. The memory 12 and the auxiliary storage device 13 may be configured integrally.
[0032] The input / output IF 14 is a port to which an input device and an output device are connected. Specific examples of the input / output IF 14 include a USB (Universal Serial Bus) terminal. Specific examples of the input device include a keyboard and a mouse. Specific examples of the output device include a display.
[0033] The communication device 15 is a receiver and a transmitter, and is specifically a communication chip or a NIC (Network Interface Card).
[0034] Each unit of each device may use the input / output IF 14 and the communication device 15 as appropriate when communicating with other devices.
[0035] The auxiliary storage device 13 stores an image generation program. The image generation program is a program that causes a computer to realize the functions of each unit of each device. The image generation program is loaded into the memory 12 and executed by the processor 11.
[0036] Data used when executing the image generation program and data obtained by executing the image generation program are stored in a storage device as appropriate. Each part of each device uses a storage device as appropriate. As a specific example, the storage device consists of at least one of the memory 12, the auxiliary storage device 13, a register in the processor 11, and a cache memory in the processor 11. Note that the terms "data" and "information" may have the same meaning. The storage device may be independent of the computer. The functions of the memory 12 and the auxiliary storage device 13 may be realized by other storage devices.
[0037] The image generation program may be recorded on a computer-readable non-volatile recording medium. Specific examples of the non-volatile recording medium include an optical disk and a flash memory. The image generation program may be provided as a program product.
[0038] ***Explanation of Operation*** The operation procedure of each device included in the image generation system 90 corresponds to an image generation method. Also, the program that realizes the operation of each device included in the image generation system 90 corresponds to an image generation program.
[0039] Fig. 7 is a flowchart showing an example of the operation of the learning device 20. Fig. 8 is a diagram illustrating the learning phase. The operation will be described with reference to Figs. 7 and 8.
[0040] (Step S101) The data acquisition unit 201 acquires various data for generating a trained model. As shown in Fig. 8 , in pre-processing, a land use map indicating land cover labels may be generated from the optical image.
[0041] (Step S102) The model generation unit 202 generates a trained model using various data and stores the generated trained model in the trained model storage unit 203. As a specific example, as shown in FIG. 8 , the model generation unit 202 learns the correspondence between the land cover labels indicated by the land use data DIN2 and the actual SAR images captured at each point in time. The SAR images contain information indicating the date and time of capture. Note that even for the same area, different land cover labels may be used at each point in time depending on changes in land use conditions.
[0042] Fig. 9 is a flowchart showing an example of the operation of the inference device 30. Fig. 10 is a diagram illustrating the utilization phase. The operation will be explained using Figs. 9 and 10.
[0043] (Step S111) The data acquisition unit 301 acquires various data for inferring the simulated image DOUT. As shown in Fig. 10, in pre-processing, a land use map showing land cover labels may be generated from the optical image.
[0044] (Step S112) The inference unit 302 infers a simulated image DOUT using various data and the trained model stored in the trained model storage unit 203. As shown in FIG. 10 , if an image captured close to the time of the disaster and at the same time as the disaster exists in the archive, the image in the archive can be used. Otherwise, the inference unit 302 uses the trained model to generate a simulated image DOUT captured close to the time of the disaster and at the same time as the disaster. The inference unit 302 may generate, as the simulated image DOUT, a SAR image of the region of interest that is an SAR image captured at the time of the disaster image, and that is an SAR image of the region of interest that would be expected if the disaster had not occurred. The generated SAR image may be used for comparison with an actual SAR image of the region of interest. The analysis pair is a pair of a remote sensing image captured after the disaster and a remote sensing image or a simulated image DOUT captured before the disaster.
[0045] ***Explanation of the Effects of Embodiment 1*** As described above, according to this embodiment, the simulated image DOUT is generated using a trained model, eliminating the need to comprehensively capture remote sensing images in terms of time and area during normal times. Furthermore, according to this embodiment, when the simulated image DOUT is inferred using the reference land use data DIN4, the time data DIN10, the agricultural history data DIN11, and the GIS data DIN12 as inputs, a remote sensing image can be generated that takes into account image features corresponding to the combination of land use status and time. Therefore, according to this embodiment, by using the simulated image DOUT to compensate for the lack of archived images, change points can be extracted based on two remote sensing images. Furthermore, this embodiment generates a simulated image DOUT for the same period as the disaster occurrence, i.e., it simulates normal-time images by taking into account differences in image features depending on the season. Therefore, by appropriately utilizing the simulated image DOUT according to this embodiment, it is possible to prevent areas other than the disaster area from being extracted as change points.
[0046] Here, when a land-use map shows classifications such as rice paddies or farmland, it is important to consider agricultural history. For example, in winter, rice paddies have characteristics similar to those of bare soil because the harvest is over. On the other hand, during rice planting season, the characteristics of rice paddies change as they are flooded. Furthermore, as the rice plants grow after planting, rice paddies have characteristics similar to those of grassland. Therefore, when a flooded area includes a rice-growing region, differences in the timing of image capture can affect the accuracy of detecting change points. For example, if a normal image is taken in winter and a disaster image is taken in summer, calculating the difference between these images can lead to the misidentification of flood damage in paddy fields, even when flood damage does not actually occur. Therefore, when estimating the extent of flood damage in land use areas where water availability varies depending on the season, it is important to generate images that simulate the appearance during the season at the time of the disaster. Furthermore, because the timing of watering rice paddies varies depending on the region, it is also important to consider the area where the area is located.
[0047] As another specific example, the appearance of a forest changes in both SAR images and optical images depending on whether deciduous trees have leaves (e.g., in the case of defoliation in winter). Therefore, it is desirable to reflect differences in the appearance of deciduous trees in the simulated image DOUT. As a specific example, the simulated image DOUT may reflect differences in the appearance of specific vegetation in areas where optical images exhibit different characteristics during a specific short period, such as during flowering or autumn (yellow) foliage. According to this embodiment, by inputting the agricultural history data DIN11 and the time data DIN10 in the utilization phase, the simulated image DOUT generated from the land use map can be made closer to an image under normal circumstances at the time of the disaster. Furthermore, according to this embodiment, by further associating and learning the imaging time and agricultural history in the learning phase, the simulated image DOUT can be generated in the utilization phase taking the disaster time and agricultural history into account.
[0048] In this embodiment, the change points between the simulated image DOUT generated from the land use map and the remote sensing image are assumed to be due to a disaster, but changes in land use can also be caused by human activities in addition to natural factors such as disasters. As a specific example, the simulated image DOUT may be generated for the purpose of detecting large-scale deforestation, the filling of a water body, or the construction or demolition of a building. According to this embodiment, it is possible to accurately detect large-scale change points between the simulated image DOUT and the remote sensing image. Therefore, this embodiment can also be applied to extracting changes other than disasters.
[0049] ***Other Configurations*** <Modification 1> Fig. 11 shows an example of the hardware configuration of each device included in an image generation system 90 according to this modification. Each device includes a processing circuit 18 instead of the processor 11, the processor 11 and memory 12, the processor 11 and auxiliary storage device 13, or the processor 11, memory 12, and auxiliary storage device 13. The processing circuit 18 is hardware that realizes at least a portion of the components included in each device. The processing circuit 18 may be dedicated hardware, or may be a processor that executes a program stored in the memory 12.
[0050] When processing circuitry 18 is dedicated hardware, processing circuitry 18 may be, for example, a single circuit, multiple circuits, a programmed processor, a parallel programmed processor, an ASIC (Application Specific Integrated Circuit), an FPGA (Field Programmable Gate Array), or a combination thereof. Each device may include multiple processing circuits that replace processing circuitry 18. The multiple processing circuits share the role of processing circuitry 18.
[0051] In each device, some functions may be realized by dedicated hardware, and the remaining functions may be realized by software or firmware.
[0052] The processing circuitry 18 is realized by, for example, hardware, software, firmware, or a combination of these. The processor 11, memory 12, auxiliary storage device 13, and processing circuitry 18 are collectively referred to as "processing circuitry." In other words, the functions of each functional component of each device are realized by the processing circuitry.
[0053] <Modification 2> Fig. 12 shows an example configuration of a change detection system 91 according to this modification. As shown in Fig. 12, the change detection system 91 includes a learning device 20, an inference device 30, a trained model storage unit 203, and a change detection device 40. The learning device 20 and the inference device 30 may be configured integrally. The functions of the learning device 20, the inference device 30, and the trained model storage unit 203 are the same as those of the devices in embodiment 1, and detailed description thereof will be omitted.
[0054] 13 shows an example of the configuration of the change detection device 40. The change detection device 40 includes an image adjustment unit 401 and a change detection unit 402.
[0055] The image adjustment unit 401 performs image processing on each of two or more images to facilitate detection of changes between the images. Specifically, the image adjustment unit 401 acquires a previous image DIN21 and a later image DIN22, applies image processing to the acquired images, and then passes the processed images to the change detection unit 402. The previous image DIN21 is a remote sensing image captured before the disaster or a simulated image DOUT generated by the inference device 30. The later image DIN22 is a remote sensing image captured after the disaster. The image adjustment unit 401 performs image processing as preprocessing for change detection, which facilitates the determination of changes. Specific examples of the image processing include alignment processing, image quality correction processing, noise removal processing, artifact removal processing, cloud detection processing, or a combination thereof. In alignment processing, when pixel-by-pixel comparison is performed using the position information held by the images, differences due to misalignment affect the changed portions. Therefore, detailed alignment may be achieved by matching the images. In the image quality correction process, since optical images may have different brightness, contrast, or color due to external factors such as weather, even when captured under the same conditions, image correction may be performed to adjust these. In the image quality correction process, noise removal may be performed to remove noise from the image using a smoothing filter or the like. False images (also called ambiguity) are a phenomenon unique to SAR images, in which an image appears to contain an object at a point where there is no actual object in the image. False images can also be suppressed by false image removal processing. Cloud detection processing detects cloud regions visible in optical images. The image adjustment unit 401 may use common cloud detection technology to determine cloud regions and exclude the determined cloud regions from areas for which changes are to be determined.
[0056] The change detection unit 402 receives two or more images, including the simulated image DOUT generated by the inference device 30, as input and detects changes between the images. Each image input to the change detection unit 402 may be an image that has undergone image processing. As a specific example, the change detection unit 402 compares the previous image DIN21 with the next image DIN22 to determine the difference, and creates a region as a change point based on the determined difference. The created change point is a region consisting of specific pixels in the image, or a polygon or other region surrounded by a polygon. The change detection unit 402 determines the change by setting a specific or variable threshold for the difference obtained by the comparison.
[0057] ***Other Embodiments*** Although the first embodiment has been described, it is possible to combine multiple parts of this embodiment and implement it. Alternatively, it is possible to implement this embodiment in part. In addition, various modifications may be made to this embodiment as necessary, and it is possible to implement it in any combination, either as a whole or in part. Note that the above-described embodiments are essentially preferred examples and are not intended to limit the scope of the present disclosure, its applications, and uses. The procedures described using flowcharts, etc. may be modified as appropriate.
[0058] Various aspects of the present disclosure are summarized below as appendices.
[0059] (Supplementary Note 1) A learning device including a model generation unit that generates a trained model that uses reference land use data indicating the land use status in an area of interest as input and infers a simulated image, which is a simulated remote sensing image showing the area of interest, by learning the correspondence between each piece of land use data indicating the land use status in at least a part of an area to be learned and each reference image in a reference image group consisting of one or more reference images, each of which is a remote sensing image showing an area corresponding to each piece of land use data.
[0060] (Supplementary Note 2) The learning device according to Supplementary Note 1, wherein the model generation unit further learns the correspondence between time data indicating the time when each reference image in the reference image group was captured and each reference image in the reference image group, and thereby generates, as the trained model, a model that uses input time data indicating the time as an input and infers, as the simulated image, a simulated remote sensing image corresponding to the time indicated by the input time data.
[0061] (Supplementary Note 3) The learning device according to Supplementary Note 1 or 2, wherein the model generation unit further learns the correspondence between agricultural history data indicating the agricultural history in the area indicated by each reference image in the reference image group and each reference image in the reference image group, and thereby generates, as the trained model, a model that further inputs input agricultural history data indicating the agricultural history in the area of interest and infers, as the simulated image, a simulated remote sensing image corresponding to the input agricultural history data.
[0062] (Supplementary Note 4) The learning device according to any one of Supplementary Notes 1 to 3, wherein the model generation unit further learns a correspondence between GIS (Geographic Information System) data in an area indicated by each reference image in the reference image group and each reference image in the reference image group, and thereby generates, as the trained model, a model that further inputs input GIS data that is GIS data in the area of interest and infers, as the simulated image, a simulated remote sensing image corresponding to the input GIS data.
[0063] (Supplementary Note 5) The learning device according to any one of Supplementary Notes 1 to 4, wherein the model generation unit further learns a correspondence between reference imaging conditions consisting of each imaging condition when each reference image in the reference image group was captured and each reference image in the reference image group, and generates, as the trained model, a model that infers, as the simulated image, a simulated remote sensing image corresponding to the simulated imaging conditions, using simulated imaging conditions consisting of each imaging condition of the simulated image as an input.
[0064] (Supplementary Note 6) An inference device comprising an inference unit that uses reference land use data indicating the land use status in an area of interest as input to a trained model and infers a simulated image that is a simulated remote sensing image that indicates the area of interest, wherein the trained model is a model generated by learning the correspondence between each piece of land use data indicating the land use status in at least a portion of the training area and each reference image of a reference image group consisting of one or more reference images that are each a remote sensing image that indicates an area corresponding to each piece of land use data.
[0065] (Supplementary Note 7) The trained model is a model generated by further learning the correspondence between time data indicating the time when each reference image in the reference image group was captured and each reference image in the reference image group, and the inference unit further uses input time data indicating the time as input to the trained model, and infers as the simulated image a simulated remote sensing image corresponding to the time indicated by the input time data.
[0066] (Supplementary Note 8) The trained model is a model generated by further learning the correspondence between agricultural history data indicating the agricultural history in the area indicated by each reference image of the reference image group and each reference image of the reference image group, and the inference unit further uses input agricultural history data indicating the agricultural history in the area of interest as input to the trained model, and infers a simulated remote sensing image corresponding to the input agricultural history data as the simulated image.
[0067] (Supplementary Note 9) The inference device according to any one of Supplementary Notes 6 to 8, wherein the trained model is a model generated by further learning a correspondence between GIS (Geographic Information System) data in an area indicated by each reference image of the reference image group and each reference image of the reference image group, and the inference unit further uses input GIS data, which is GIS data in the area of interest, as input to the trained model and infers a simulated remote sensing image corresponding to the input GIS data as the simulated image.
[0068] (Supplementary Note 10) The trained model is a model generated by further learning the correspondence between reference imaging conditions consisting of each imaging condition when each reference image of the reference image group was captured and each reference image of the reference image group, and the inference unit further uses simulated imaging conditions consisting of each imaging condition of the simulated image as input to the trained model, and infers a simulated remote sensing image corresponding to the simulated imaging conditions as the simulated image.
[0069] (Supplementary Note 11) An image generation system comprising: a learning device according to any one of Supplementary Notes 1 to 5; and an inference device according to any one of Supplementary Notes 6 to 10.
[0070] (Appendix 12) A change detection system comprising: a learning device according to any one of Appendices 1 to 5; an inference device according to any one of Appendices 6 to 10; and a change detection device, wherein the change detection device comprises a change detection unit that receives two or more images including the simulated image generated by the inference device as input and detects changes between the images.
[0071] (Appendix 13) The change detection device further includes an image adjustment unit that performs image processing on each of the two or more images so as to make it easier to detect changes between the images, and each image input to the change detection unit is an image that has been subjected to image processing.
[0072] 11 Processor, 12 Memory, 13 Auxiliary storage device, 14 Input / output IF, 15 Communication device, 18 Processing circuit, 19 Signal line, 90 Image generation system, 91 Change detection system, 20 Learning device, 201 Data acquisition unit, 202 Model generation unit, 203 Learned model storage unit, 30 Inference device, 301 Data acquisition unit, 302 Inference unit, 40 Change detection device, 401 Image adjustment unit, 402 Change detection unit, DIN1 Reference image, DIN2 Land use data, DIN3 Reference imaging conditions, DIN4 Reference land use data, DIN5 Simulated imaging conditions, DIN10 Time data, DIN11 Agricultural history data, DIN12 GIS data, DIN21 Previous time image, DIN22 Later time image, DOUT Simulated image.
Claims
1. A learning device having a model generation unit that generates a trained model that uses reference land use data indicating the land use status in an area of interest as input and infers a simulated image, which is a simulated remote sensing image showing the area of interest, by learning the correspondence between each piece of land use data indicating the land use status in at least a portion of the area to be learned and each reference image in a reference image group consisting of one or more reference images, each of which is a remote sensing image showing the area corresponding to each piece of land use data.
2. The learning device described in claim 1, wherein the model generation unit further learns the correspondence between time data indicating the time when each reference image in the reference image group was captured and each reference image in the reference image group, and generates as the trained model a model that uses input time data indicating the time as input and infers as the simulated image a simulated remote sensing image corresponding to the time indicated by the input time data.
3. The learning device described in claim 1 or 2, wherein the model generation unit further learns the correspondence between agricultural history data indicating the agricultural history in the area indicated by each reference image in the reference image group and each reference image in the reference image group, and thereby generates as the trained model a model that uses input agricultural history data indicating the agricultural history in the area of interest as input and infers a simulated remote sensing image corresponding to the input agricultural history data as the simulated image.
4. A learning device according to any one of claims 1 to 3, wherein the model generation unit further learns the correspondence between GIS (Geographic Information System) data in the area indicated by each reference image in the reference image group and each reference image in the reference image group, thereby generating, as the trained model, a model that uses input GIS data, which is GIS data in the area of interest, as input and infers, as the simulated image, a simulated remote sensing image corresponding to the input GIS data.
5. A learning device as claimed in any one of claims 1 to 4, wherein the model generation unit further learns the correspondence between reference imaging conditions consisting of the imaging conditions when each reference image in the reference image group was captured and each reference image in the reference image group, and generates as the trained model a model that uses simulated imaging conditions consisting of the imaging conditions of each simulated image as an input and infers a simulated remote sensing image corresponding to the simulated imaging conditions as the simulated image.
6. An inference device comprising an inference unit that uses reference land use data indicating the land use status in an area of interest as input to a trained model and infers a simulated image which is a simulated remote sensing image showing the area of interest, wherein the trained model is a model generated by learning the correspondence between each piece of land use data indicating the land use status in at least a portion of the training area and each reference image of a reference image group consisting of one or more reference images which are each a remote sensing image showing an area corresponding to each piece of land use data.
7. The inference device described in claim 6, wherein the trained model is a model generated by further learning the correspondence between time data indicating the time when each reference image in the reference image group was captured and each reference image in the reference image group, and the inference unit further uses input time data indicating the time as input to the trained model and infers as the simulated image a simulated remote sensing image corresponding to the time indicated by the input time data.
8. The inference device described in claim 6 or 7, wherein the trained model is a model generated by further learning the correspondence between agricultural history data indicating the agricultural history in the area indicated by each reference image in the reference image group and each reference image in the reference image group, and the inference unit further uses input agricultural history data indicating the agricultural history in the area of interest as input to the trained model and infers a simulated remote sensing image corresponding to the input agricultural history data as the simulated image.
9. An inference device according to any one of claims 6 to 8, wherein the trained model is a model generated by further learning the correspondence between GIS (Geographic Information System) data in the area indicated by each reference image in the reference image group and each reference image in the reference image group, and the inference unit further uses input GIS data, which is GIS data in the area of interest, as input to the trained model and infers a simulated remote sensing image corresponding to the input GIS data as the simulated image.
10. An inference device as described in any one of claims 6 to 9, wherein the trained model is a model generated by further learning the correspondence between reference imaging conditions consisting of each imaging condition when each reference image of the reference image group was captured and each reference image of the reference image group, and the inference unit further uses simulated imaging conditions consisting of each imaging condition of the simulated image as input to the trained model and infers a simulated remote sensing image corresponding to the simulated imaging conditions as the simulated image.
11. An image generation system comprising: a learning device according to any one of claims 1 to 5; and an inference device according to any one of claims 6 to 10.
12. A change detection system comprising: a learning device according to any one of claims 1 to 5; an inference device according to any one of claims 6 to 10; and a change detection device, wherein the change detection device comprises a change detection unit that receives as input two or more images including the simulated image generated by the inference device and detects changes between the images.
13. The change detection system according to claim 12, further comprising an image adjustment unit that performs image processing on each of the two or more images so as to make it easier to detect changes between the images, and each image input to the change detection unit is an image that has undergone image processing.
14. An image generation method in which a computer learns the correspondence between each piece of land use data indicating the land use status in at least a portion of a learning area and each reference image in a reference image group consisting of one or more reference images, each of which is a remote sensing image showing an area corresponding to each land use data, thereby generating a trained model that uses reference land use data indicating the land use status in an area of interest as input and infers a simulated image, which is a simulated remote sensing image showing the area of interest.
15. An image generation method in which a computer uses reference land use data indicating the land use status in an area of interest as input to a trained model to infer a simulated image which is a simulated remote sensing image showing the area of interest, wherein the trained model is a model generated by learning the correspondence between each piece of land use data indicating the land use status in at least a portion of the training area and each reference image in a reference image group consisting of one or more reference images which are each a remote sensing image showing an area corresponding to each piece of land use data.
16. An image generation program that causes a learning device, which is a computer, to execute a model generation process that generates a trained model that uses reference land use data that indicates the land use status in an area of interest as input and infers a simulated image that is a simulated remote sensing image that shows the area of interest, by learning the correspondence between each land use data that indicates the land use status in at least a portion of the learning area and each reference image in a reference image group consisting of one or more reference images that are each remote sensing images that show the area corresponding to each land use data.
17. An image generation program that causes an inference device, which is a computer, to execute an inference process to infer a simulated image, which is a simulated remote sensing image, showing an area of interest, using reference land use data showing the land use status in the area of interest as input to a trained model, wherein the trained model is a model generated by learning the correspondence between each piece of land use data showing the land use status in at least a portion of the training area and each reference image in a reference image group consisting of one or more reference images, each of which is a remote sensing image showing an area corresponding to each piece of land use data.
Citation Information
Patent Citations
Remote sensing image matching method based on deep learning and multi-subgraph matching
CN112861714A
Change detection system, change detection method, and change detection program
JP2024004521A
Information processing device, information processing method, and program
WO2020174649A1