Change detecting apparatus, change detecting method, and non-transitory computer-readable medium
The change detecting apparatus addresses the issue of angle discrepancies in image comparison by transforming and simulating images to a consistent view, enhancing the accuracy of change detection.
Patent Information
- Application Number
- PCT/JP2024/010805
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-03-19
- Publication Date
- 2025-09-25
AI Technical Summary
Existing change detection techniques fail to accurately account for differences in images captured from different angles, leading to inaccuracies in detecting changes in a target area over time.
A change detecting apparatus that encodes and transforms image features based on angle information to simulate images from a consistent viewing angle, allowing for accurate comparison and detection of changes by generating a simulated image from the second angle at the first time.
Enables precise change detection by aligning image features across different capture angles, improving accuracy in identifying changes in a target area.
Smart Images

Figure JP2024010805_25092025_PF_FP_ABST
Abstract
Description
CHANGE DETECTING APPARATUS, CHANGE DETECTING METHOD, AND NON-TRANSITORY COMPUTER-READABLE MEDIUM
[0001] The present disclosure generally relates to change detecting apparatus, change detecting method, and non-transitory computer-readable medium.
[0002] There are techniques to detect change in a particular place by comparing two images that capture the particular place at different times. NPL1 discloses a technique to compare two images X and Y respectively generated by different sensors Sx and Sy at different times Tx and Ty to detect change in the place captured on those two images. Specifically, the system of NPL1 applies a regression function to the image X to predict an image Y^ which would have been obtained if the sensor Sy captures that place at time Tx. Then, the images Y and Y^ are compared to detect change in the place from the time Tx to Ty.
[0003] NPL1: Luigi T. Luppino, Filippo M. Bianchi, Gabriele Moser, and Stian N. Anfinsen, "Remote sensing image regression for heterogeneous change detection", [online], July 31, 2018, [retrieved on 2024-2-21], retrieved from <arXiv, https: / / arxiv.org / pdf / 1807.11766.pdf>
[0004] NPL1 does not consider the difference of angles between sensors that capture the images to be compared. An objective of the present disclosure is to provide a novel technique to compare images to detect change in a place captured on the images.
[0005] The present disclosure provides a change detecting apparatus that comprises at least one memory that is configured to store instructions and at least one processor. The at least one processor is configured to execute the instructions to: acquire a first captured image, a second captured image, and angle information, the first captured image being generated by capturing a target place from a first angle at a first time, the second captured image being generated by capturing the target place from a second angle at a second time, the angle information indicating the first angle and the second angle; encode the first captured image into a first structural feature value, which represents features of a three-dimensional structure of the target place seen from the first angle at the first time; perform feature transformation on the first structural feature value based on the first angle and the second angle to compute a second structural feature value, which represents features of a three-dimensional structure of the target place seen from the second angle at the first time; generate, based on the second structural feature value, a simulated image of the target place seen from the second angle at the first time; and compare the simulated image and the second captured image to detect change in the target place from the first time to the second time.
[0006] The present disclosure further provides a change detecting method that is performed by a computer, comprises: acquiring a first captured image, a second captured image, and angle information, the first captured image being generated by capturing a target place from a first angle at a first time, the second captured image being generated by capturing the target place from a second angle at a second time, the angle information indicating the first angle and the second angle; encoding the first captured image into a first structural feature value, which represents features of a three-dimensional structure of the target place seen from the first angle at the first time; performing feature transformation on the first structural feature value based on the first angle and the second angle to compute a second structural feature value, which represents features of a three-dimensional structure of the target place seen from the second angle at the first time; generating, based on the second structural feature value, a simulated image of the target place seen from the second angle at the first time; and comparing the simulated image and the second captured image to detect change in the target place from the first time to the second time.
[0007] The present disclosure further provides a non-transitory computer readable storage medium storing a program. The program that causes a computer to execute: acquiring a first captured image, a second captured image, and angle information, the first captured image being generated by capturing a target place from a first angle at a first time, the second captured image being generated by capturing the target place from a second angle at a second time, the angle information indicating the first angle and the second angle; encoding the first captured image into a first structural feature value, which represents features of a three-dimensional structure of the target place seen from the first angle at the first time; performing feature transformation on the first structural feature value based on the first angle and the second angle to compute a second structural feature value, which represents features of a three-dimensional structure of the target place seen from the second angle at the first time; generating, based on the second structural feature value, a simulated image of the target place seen from the second angle at the first time; and comparing the simulated image and the second captured image to detect change in the target place from the first time to the second time.
[0008] According to the present disclosure, a novel technique to compare images to detect change in a place captured on the images.
[0009] Fig. 1 illustrates an overview of a change detecting apparatus.Fig. 2 is a block diagram showing an example of the functional configuration of the change detecting apparatus.Fig. 3 is a block diagram illustrating an example of the hardware configuration of a computer realizing the change detecting apparatus.Fig. 4 shows a flowchart illustrating an example flow of processes performed by the change detecting apparatus.Fig. 5 illustrates the angle information.Fig. 6 illustrates the first structural feature value as a feature cuboid.Fig. 7 illustrates an example of the feature transformation achieved using the transforming model.Fig. 8 illustrates an example of the training of the models in the change detecting apparatus.Fig. 9 illustrates another example of the training of the models in the change detecting apparatus.Fig. 10 illustrates an example of ways to extract features from a set of the first angle information and the second angle information.Fig. 11 is a block diagram showing another example of the functional configuration of the change detecting apparatus.Fig. 12 shows a flowchart illustrating another example flow of processes performed by the change detecting apparatus.Fig. 13 illustrates another example of the training of the models in the change detecting apparatus.
[0010] Example embodiments according to the present disclosure will be described hereinafter with reference to the drawings. The same numeral signs are assigned to the same elements throughout the drawings, and redundant explanations are omitted as necessary. In addition, predetermined information (e.g., a predetermined value or a predetermined threshold) is stored in advance in a storage unit to which a computer using that information has access unless otherwise described. In the present disclosure, a storage unit may be implemented with one or more storage devices, such as hard disks, solid-state drives (SSDs), or random-access memories (RAMs).
[0011] FIRST EXAMPLE EMBODIMENT <Overview> Fig. 1 illustrates an overview of a change detecting apparatus 2000. It is noted that Fig. 1 does not limit operations of the change detecting apparatus 2000, but merely show an example of possible operations of the change detecting apparatus 2000.
[0012] The change detecting apparatus 2000 is an apparatus that is configured to acquire a first captured image 10 and a second captured image 20 to be compared with each other. The first captured image 10 is generated by capturing a target place 30 at a first time. The second captured image 20 is generated by capturing the target place 30 at a second time. The second time is a point in time later than the first time. The change detecting apparatus 2000 detects change in the target place 30 from the first time to the second time based on the first captured image 10 and the second captured image 20.
[0013] The target place 30 is captured from different angles between the first captured image 10 and the second captured image 20. Specifically, the first captured image 10 is generated by a sensor 50 that captures the target place 30 from the first angle at the first time. On the other hand, the second captured image 20 is generated by a sensor 55 that captures the target place 30 from the second angle at the second time. It is noted that the sensor 55 may be different sensor from the sensor 50, or may be the same sensor as the sensor 50. In the latter case, the sensor 55 is the sensor 50 whose capturing angle (or antenna) is adjusted after the first time.
[0014] Due to the difference of the angle of the sensor 50 and the angle of the sensor 55 mentioned above, the first captured image 10 and the second captured image 20 can have difference in their images of the target place 30 even when there is no change in the target place 30 from the first time to the second time. Thus, it may be difficult to accurately detect change in the target place 30 from the first time to the second time by simply comparing the first captured image 10 and the second captured image 20.
[0015] To tackle this issue, the change detecting apparatus 2000 converts the first captured image 10 into a simulated image 80, which shows the target place 30 seen from the second angle at the first time, and compares the simulated image 80 and the second captured image 20 to detect change in the target place 30 from the first time to the second time. More specific operations of the change detecting apparatus 2000 may be as follows.
[0016] The change detecting apparatus 2000 acquires the first captured image 10, the second captured image 20, and angle information 40. The angle information 40 indicates the first angle and the second angle. The change detecting apparatus 2000 encodes the first captured image 10 into a first structural feature value 60. It is noted that the encoding of an image into a feature value is a process of extracting features from the image, such as the encoding performed by an autoencoder. The first structural feature value 60 is a value, such as a tensor, that represents features of the three-dimensional (3D) structure of the target place 30 captured on first captured image 10. Hereinafter, a feature value that represents features of the 3D structure of a place is generally called "structural feature value".
[0017] The change detecting apparatus 2000 performs feature transformation on the first structural feature value 60 to transform the first structural feature value 60 into a second structural feature value 70, based on the first angle and the second angle. The second structural feature value 70 is synthetically generated so as to represent features of the 3D structure of the target place 30 seen from the second angle at the first time. Detailed explanation of the feature transformation will be described later.
[0018] The change detecting apparatus 2000 decodes the second structural feature value 70 into the simulated image 80. It is noted that the decoding of a feature value into an image is a process of generating an image from the features, such as the decoding performed by an autoencoder. As mentioned above, the second structural feature value 70 is synthetically generated to represent the features of the 3D structure of the target place 30 seen from the second angle at the first time. Thus, the simulated image 80 is generated such that it includes an image of the target place 30 captured from the second angle at the first time.
[0019] The change detecting apparatus 2000 compares the simulated image 80 and the second captured image 20 to detect change in the target place 30 from the first time to the second time. Unlike the first captured image 10, the simulated image 80 includes the target place 30 seen from the angle same as the angle from which the target place 30 in the second captured image 20 is captured. Thus, the comparison between the simulated image 80 and the second captured image 20 can detect change in the target place 30 from the first time to the second time more accurately than the comparison between the first captured image 10 and the second captured image 20 does.
[0020] It is noted that the first captured image 10 may be an optical image or a radar image. When the first captured image 10 is an optical image, the sensor 50 is an optical camera that is configured to receive light to generate an optical image based on the received light. When the first captured image 10 is a radar image, the sensor 50 is a radar that is configured to transmit radio waves, receive reflection of the radio waves, and generate a radar image based on the received reflection of the radio waves. The sensor 50 may be installed on an artificial satellite to capture objects on the Earth, other planets, satellites, etc. An example of the radar used as the sensor 50 is a synthetic-aperture radar.
[0021] Similarly, the second captured image 20 may be an optical image or a radar image. When the second captured image 20 is an optical image, the sensor 55 is an optical camera that is configured to receive light to generate an optical image based on the received light. When the second captured image 20 is a radar image, the sensor 55 is a radar that is configured to transmit radio waves, receive reflection of the radio waves, and generate a radar image based on the received reflection of the radio waves. The sensor 55 may be installed on an artificial satellite to capture objects on the Earth, other planets, satellites, etc. An example of the radar used as the sensor 55 is a synthetic-aperture radar.
[0022] The change detecting apparatus 2000 is useful in and applicable to various situations. One of the situations in which the change detecting apparatus 2000 is useful is the situation where a disaster, such as an earthquake, has occurred. In this situation, it is preferable to ascertain the damage caused by the disaster as quickly as possible to handle the problems caused by the disaster (e.g., providing appropriate supports to the victims).
[0023] To ascertain the damage caused by the disaster, it is effective to detect change in the disaster-stricken area by comparing the images obtained from the satellites that capture that area, such as SAR (synthetic-aperture radar) images. Specifically, for each of the sections in the disaster-stricken area, the change in the section can be detected by comparing the image generated by capturing that section before the disaster and the image generated by capturing that section after the disaster.
[0024] Suppose that the images obtained from the same angle by the same satellite are used. In this case, it takes time (e.g., ten days) to obtain the images to be compared for the change detection since the satellite keeps moving. Thus, in order to quickly ascertain the damage in the disaster-stricken area, it is preferable to use the images obtained from different angles by adjusting the antennas of the same satellite or from different satellites that happen to fly close to the area of interest.
[0025] As mentioned above, the images obtained from different angles may have difference in the same place as each other even when there is no change in the place due to the change of viewing direction. Thus, it is effective to apply the change detecting apparatus 2000 to this situation since the change detecting apparatus 2000 can accurately detect change in the disaster-stricken area by comparing the images obtained from different angles with the feature transformation.
[0026] Hereinafter, more detailed explanation of the change detecting apparatus 2000 will be described.
[0027] <Example of Functional Configuration> Fig. 2 is a block diagram showing an example of the functional configuration of the change detecting apparatus 2000. The change detecting apparatus 2000 includes an acquiring unit 2020, an encoding unit 2040, a feature transforming unit 2060, a decoding unit 2080, and a detecting unit 2100. The acquiring unit 2020 acquires the first captured image 10, the second captured image 20, and the angle information 40. The encoding unit 2040 encodes the first captured image 10 into the first structural feature value 60. The feature transforming unit 2060 performs the feature transformation on the first structural feature value 60 by referencing to the angle information 40 to compute the second structural feature value 70. The decoding unit 2080 decodes the second structural feature value 70 into the simulated image 80. The detecting unit 2100 compares the simulated image 80 and the second captured image 20 to detect change in the target place 30 from the first time to the second time.
[0028] <Example of Hardware Configuration> The change detecting apparatus 2000 may be realized by one or more computers. Each of the one or more computers may be a special-purpose computer manufactured for implementing the change detecting apparatus 2000, or may be a general-purpose computer like a personal computer (PC), a server machine, or a mobile device.
[0029] The change detecting apparatus 2000 may be realized by installing an application in the computer. The application is implemented with a program that causes the computer to function as the change detecting apparatus 2000. In other words, the program is an implementation of the functional units of the change detecting apparatus 2000. There are various ways to acquire the program. For example, the program can be acquired from a storage medium (such as a DVD (digital versatile disc) or a USB (Universal Serial Bus) memory) in which the program is stored in advance. In another example, the program can be acquired by downloading it from a server machine that manages a storage medium in which the program is stored in advance.
[0030] Fig. 3 is a block diagram illustrating an example of the hardware configuration of a computer 1000 realizing the change detecting apparatus 2000. In Fig. 3, the computer 1000 includes a bus 1020, a processor 1040, a memory 1060, a storage device 1080, an input / output (I / O) interface 1100, and a network interface 1120.
[0031] The bus 1020 is a data transmission channel in order for the processor 1040, the memory 1060, the storage device 1080, and the I / O interface 1100, and the network interface 1120 to mutually transmit and receive data. The processor 1040 is a processer, such as a CPU (Central Processing Unit), GPU (Graphics Processing Unit), FPGA (Field-Programmable Gate Array), or a DSP (Digital Signal Processor). The memory 1060 is a primary memory component, such as a RAM or a ROM (Read Only Memory). The storage device 1080 is a secondary memory component, such as a hard disk, an SSD, or a memory card. The I / O interface 1100 is an interface between the computer 1000 and peripheral devices, such as a keyboard, mouse, or display device. The network interface 1120 is an interface between the computer 1000 and a network. The network may be a LAN (Local Area Network) or a WAN (Wide Area Network). The storage device 1080 may store the program mentioned above. The processor 1040 executes the program to realize each functional unit of the change detecting apparatus 2000.
[0032] The hardware configuration of the computer 1000 is not restricted to that shown in Fig. 3. For example, as mentioned-above, the change detecting apparatus 2000 may be realized by plural computers. In this case, those computers may be connected with each other through the network.
[0033] <Flow of Process> Fig. 4 shows a flowchart illustrating an example flow of processes performed by the change detecting apparatus 2000. The acquiring unit 2020 acquires the first captured image 10, the second captured image 20, and the angle information 40 (S102). The encoding unit 2040 encodes the first captured image 10 into the first structural feature value 60 (S104). The feature transforming unit 2060 performs the feature transformation on the first structural feature value 60 by referencing to the angle information 40 to compute the second structural feature value 70 (S106). The decoding unit 2080 decodes the second structural feature value 70 into the simulated image 80 (S108). The detecting unit 2100 compares the simulated image 80 and the second captured image 20 to detect change in the target place 30 from the first time to the second time (S110).
[0034] <Acquisition of Images and Angle Information: S102> The acquiring unit 2020 acquires the first captured image 10, the second captured image 20, and the angle information 40 (S102). Hereinafter, a set of the first captured image 10, the second captured image 20, and the angle information 40 is also called "the set of input data".
[0035] There are various ways to acquire the set of input data. For example, the set of input data is stored in advance in a storage unit in a manner that the change detecting apparatus 2000 can acquire the set of input data. In this case, the acquiring unit 2020 acquires the set of input data from this storage unit.
[0036] In another example, the set of input data is input by a user of the change detecting apparatus 2000. In this case, the change detecting apparatus 2000 may prompt the user to input the first captured image 10, the second captured image 20 and the angle information 40.
[0037] In some embodiments, each captured image is stored in a storage unit in association with the angle of the sensor at the time when the sensor generated the captured image. This means that the first captured image 10 and the second captured image are stored in the storage unit in association with the first angle and the second angle, respectively. In this case, the acquiring unit 2020 may acquire the first captured image 10 with the first angle and the second captured image 20 with the second angle from the storage unit, thereby acquiring the set of input data.
[0038] In some embodiments, there are two or more satellites each of which has a sensor to generate the captured images. In this case, each captured image can be stored in a storage unit in association with the identifier of the satellite on which the sensor having generated the captured image is installed. Furthermore, the history of locations (i.e., the orbit) of each satellite is recorded. Hereinafter, the history of locations of the satellite is called "satellite information". The satellite information may include the location of the satellite in association with time. In this case, the angle of the sensor at a particular time can be calculated based on the positional relationship between the captured place and the satellite corresponding to the sensor at that particular time.
[0039] When the satellite information can be used, the acquiring unit 2020 may acquire the first captured image 10 and the identifier of the corresponding satellite (i.e., the satellite on which the sensor 50 is installed). The acquiring unit 2020 determines the location of that satellite at the first time by looking up the location associated with the first time in the satellite information of that satellite. The acquiring unit 2020 computes the first angle based on the positional relationship between the determined location and the target place 30.
[0040] Similarly, the acquiring unit 2020 acquires the second captured image 20 and the identifier of the corresponding satellite (i.e., the satellite on which the sensor 55 is installed). The acquiring unit 2020 determines the location of that satellite at the second time by looking up the location associated with the second time in the satellite information of that satellite. The acquiring unit 2020 computes the second angle based on the positional relationship between the determined location and the target place 30.
[0041] In some embodiments, the satellite information may also include the angle of the satellite (i.e., the angle of the sensor installed on that satellite). In this case, the acquiring unit 2020 can acquire the angle of the sensor 50 at the first time by looking up the angle of the satellite associated with the first time in the satellite information of the satellite corresponding to the sensor 50. Similarly, the acquiring unit 2020 can acquire the angle of the sensor 55 at the second time by looking up the angle of the satellite associated with the second time in the satellite information of the satellite corresponding to the sensor 55.
[0042] The captured images to be acquired as the first captured image 10 and the second captured image 20 may be specified by the user or automatically determined by the change detecting apparatus 2000. In the latter case, for example, the change detecting apparatus 2000 prompts the user to enter a place of interest and time of interest. In this case, the change detecting apparatus 2000 determines whether the place of interest has substantial change before and after the time of interest (in other words, detects change in the place of interest before and after the time of interest). This means that the place of interest is handled as the target place 30. When it is required to ascertain the damage in the disaster-stricken area as mentioned above, a section of the disaster-stricken area and the time at which the disaster occurred are handles as the place of interest and the time of interest, respectively.
[0043] Suppose that each captured image is stored in a storage unit in association with location information (e.g., global positioning system (GPS) coordinates) that indicates the location of the place captured on the captured image and in association with the identifier of the corresponding satellite. In this case, as the first captured image 10, the acquiring unit 2020 acquires the captured image that meets conditions of: (1) including the place of interest; (2) being generated before the time of interest; and (3) having the latest generation time among the captured images meeting the conditions (1) and (2). In addition, as the second captured image 20, the acquiring unit 2020 acquires the captured image that meets conditions of: (1) including the place of interest; (2) being generated after the time of interest; and (3) having the earliest generation time among the captured images meeting the conditions (1) and (2). The first angle may be determined by, for example, looking up in the satellite information of the satellite corresponding to the sensor 50. Similarly, the second angle may be determined by, for example, looking up in the satellite information of the satellite corresponding to the sensor 55.
[0044] <As to Angle information 40> There are various ways to represent an angle of a sensor. In some embodiments, the angle of a sensor at a particular time may be represented by a pair of the incident angle of the sensor and the azimuth angle of the sensor at the particular time. In this case, the first angle may be represented by a pair of the incident angle of the sensor 50 and the azimuth angle of the sensor 50 at the first time. Similarly, the second angle may be represented by a pair of the incident angle of the sensor 55 and the azimuth angle of the sensor 55 at the second time.
[0045] Fig. 5 illustrates the angle information 40. In the example of Fig. 5, the angle information 40 includes a pair of the first angle 41 and the second angle 44. The first angle 41 includes a pair of the first incident angle 42 and the first azimuth angle 43. The second angle 44 includes a pair of the second incident angle 45 and the second azimuth angle 46. For the purpose of the clarity of the drawing, the sensor 55 is not depicted in Fig. 5.
[0046] The representation of the angle of the sensor is not limited to the pair of the incident angle and the azimuth angle of the sensor. For example, an zenith angle of the sensor can be used instead of the incident angle of the sensor.
[0047] <Encode of First Captured Image: S104> The encoding unit 2040 encodes the first captured image 10 into the first structural feature value 60 (S104). In some embodiments, the encoding unit 2040 includes an encoder, which is a machine learning-based model (e.g., a neural network), to perform the encoding on the first captured image 10. The encoder is configured to take an image as input and extracts features of the 3D structure of the place captured on the input image from the input image, thereby encoding the input image into a structural feature value of the input image.
[0048] The encoder may be trained in conjunction with one or more other models included in the change detecting apparatus 2000, such as a model performing the feature transformation or a model performing the decoding of the second structural feature value 70. The training of the encoder will be explained later in detail.
[0049] The structural feature value may be a set of cells that has a feature vector (i.e., value of features) and coordinates in a specific coordinate system. It is noted that, in this disclosure, a term "cell" is used to describe a pair of a value and coordinates. The feature vector at particular coordinates represents features of the 3D structure of a sub-region of the captured place corresponding to those coordinates. The set of cells may be represented by a cuboid of cells each of which indicates the feature vector at the corresponding coordinates of the cell. Hereinafter, this cuboid of cells is called "feature cuboid".
[0050] Hereinafter, unless otherwise stated, the structural feature value is described as being a feature cuboid. However, the techniques described in this disclosure can also be applied to cases where the structural feature value is represented by a form other than cuboid (e.g., a list of cells).
[0051] Fig. 6 illustrates the first structural feature value 60 as a feature cuboid. In Fig. 6, the first structural feature value 60 is represented as a feature cuboid in a first coordinate system 90, which is defined by a first azimuth-axis 92, a first range-axis 94, and a first incident-axis 96. The first range-axis 94 is an axis that is the direction perpendicular to flying direction of the sensor 50 (or the satellite).The first azimuth-axis 92 is an axis that is parallel to the flying direction of the sensor 50 (or the satellite) and is perpendicular to the first range-axis 94. The first range-axis 94 is an axis that is on the ground plane and perpendicular to the first azimuth-axis 92. The first incident-axis 96 is an axis that forms the first incident angle 42 from a direction opposite to the gravity direction, called "elevation direction".
[0052] The encoder 100 is the above mentioned encoder included in the encoding unit 2040. As illustrated by Fig. 6, the features of the 3D structure of the sub-region 12 of the first captured image 10 may be extracted as a sequence 62 of cells of the first structural feature value 60 along the first incident-axis 96.
[0053] <Feature Transformation: S106> The feature transforming unit 2060 performs the feature transformation on the first structural feature value 60 to compute the second structural feature value 70 (S106). In some embodiments, the feature transforming unit 2060 includes a machine learning-based model (e.g., a neural network) that performs the feature transformation on the first structural feature value 60. Hereinafter, this model is called "transforming model".
[0054] The transforming model is configured to take a first input angle, a second input angle, and an input structural feature value representing features of the 3D structure of a place seen from the first input angle. Conceptually, based on the first input angle and the second input angle, the transforming model converts an input structural feature value into an output structural feature value, which represents features of the 3D structure of the place seen from the second input angle.
[0055] Fig. 7 illustrates an example of the feature transformation achieved using the transforming model. In this example, the transforming model 120 includes a first model 121 and a second model 122. The first model 121 is configured to take a pair of an incident angle and an azimuth angle as input, and convert the input pair into a matrix. The second model 122 is configured to take a feature cuboid that is a concatenation of the input structural feature value, a matrix representing the first input angle, and a matrix representing the second input angle. The second model 122 is also configured to convert the input feature cuboid into the output structural feature value whose size is the same as the input structural feature value.
[0056] The feature transforming unit 2060 takes the first structural feature value 60, the first incident angle 42, the first azimuth angle 43, the second incident angle 45, and the second azimuth angle 46. The first incident angle 42 and the first azimuth angle 43 are fed into the first model 121, thereby being converted into a matrix 130. The second incident angle 45 and the second azimuth angle 46 are fed into the first model 121, thereby being converted into a matrix 140. The first structural feature value 60, the matrix 130, and the matrix 140 are concatenated into a feature cuboid 150. The feature cuboid 150 is fed into the second model 122, thereby being converted into the second structural feature value 70.
[0057] In some embodiments, the feature transformation may be achieved without using the machine learning-based model. For example, the feature transformation may be achieved using a matrix, called resampling matrix. The feature transformation using the resampling matrix R can be expressed as follows. Equation 1
[0058] In the equation (1), φ1 represents the first incident angle 42. φ2 represents the second incident angle 45. Δθ represents the second azimuth angle 46 minus the first azimuth angle 43. k represents the ratio of elevation-axis unit length to range-axis unit length. The matrix is applied to point vectors in the order (z,y,x) where z represents a vale of elevation-axis, y represents a value on range-axis and x represents a value on azimuth-axis.
[0059] In some embodiments, the feature transformation may be achieved using both the resampling matrix and a machine learning-based model. For example, the feature transforming unit 2060 applies the resampling matrix to the first structural feature value 60, and then input the resulting value into a machine learning-based model (e.g., a neural network), thereby obtaining the second structural feature value 70.
[0060] <Decode of Second Structural Feature Value: S108> The decoding unit 2080 decodes the second structural feature value 70 into the simulated image 80 (S108). In some embodiments, the decoding unit 2080 includes a decoder, which is a machine learning-based model (e.g., a neural network) to perform the decoding on the second structural feature value 70. The decoder is configured to take a structural feature value as input and convert the input structural feature value into an output image that includes a place whose features of the 3D structure are represented by the input structural feature value.
[0061] The decoder may be trained in conjunction with one or more other models included in the change detecting apparatus 2000, such as the encoder, the transforming model 120, or both. The training of the decoder will be explained later in detail.
[0062] <Change Detection: S110> The detecting unit 2100 compares the simulated image 80 and the second captured image 20 to detect change in the target place 30 from the first time to the second time (S110). For example, the detecting unit 2100 computes the pixel-wise difference (e.g., distance) between the simulated image 80 and the second captured image 20 to generate a dissimilarity map. The dissimilarity map may be a gray scale image in which each pixel represents a degree of difference between the corresponding pixel in the simulated image 80 and the corresponding pixel in the second captured image 20.
[0063] The detecting unit 2100 may perform a further process on the dissimilarity map. For example, the detecting unit 2100 computes a sum of the pixel values of the dissimilarity map, and determine whether there is a significant change in the target place 30 based on the sum of the pixel values. Specifically, the detecting unit 2100 determines that there is the significant change in the target place 30 when the sum of the pixel values of the dissimilarity map is larger than a predefined threshold Th1. On the other hand, the detecting unit 2100 determines that there is not the significant change in the target place 30 when the sum of the pixel values of the dissimilarity map is not larger than the predefined threshold Th1.
[0064] In another example, the detecting unit 2100 may counts the number of pixels in the dissimilarity map whose value is larger than a predefined threshold Th2, and determine whether there is a significant change in the target place 30 based on the counted number. This counted number may represent the size of the area that has a significant change. Specifically, the detecting unit 2100 determines that there is the significant change in the target place 30 when the counted number is larger than a predefined threshold Th3. On the other hand, the detecting unit 2100 determines that there is not the significant change in the target place 30 when the counted number is not larger than the predefined threshold Th3.
[0065] To eliminate the subtle differences that do not represent substantial changes, the detecting unit 2100 may apply thresholding to the dissimilarity map, thereby removing differences below a certain threshold. Hereinafter, the dissimilarity map after the thresholding is called "change map". For example, the change map is a binary image in which a white pixel represents that the corresponding part of the target place 30 has a substantial change while a black pixel represents that the corresponding part of the target place 30 does not have a substantial change.
[0066] The further processes performed on the dissimilarity map mentioned above can also be applied to the change map. It is noted that, regarding the count of the pixels, the detecting unit 2100 counts the number of white pixels of the change map to determine the size of the area having significant changes.
[0067] <Output of Result> The change detecting apparatus 2000 may output information (hereinafter, called "output information") that includes one or more pieces of data representing the result of the processing performed by the change detecting apparatus 2000. For example, the output information may include the dissimilarity map, the change map, or both. Additionally or alternatively, the output information may include the result of the determination of whether there is a substantial change in the target place 30 from the first time to the second time.
[0068] It is noted that the change detecting apparatus 2000 may take two or more pieces of the input dataset (i.e., two or more sets of the first captured image 10, the second captured image 20, and the angle information 40). In this case, the change detecting apparatus 2000 processes each input dataset and generates the output information for each input dataset.
[0069] <Training of Models> Hereinafter, the training of models used in the change detecting apparatus 2000 will be explained in detail. Fig. 8 illustrates an example of the training of the models in the change detecting apparatus 2000. In Fig. 8, the change detecting apparatus 2000 includes three models: the encoder 100, the transforming model 120, and the decoder 160. The decoder 160 is the decoder included in the decoding model 2080 mentioned above. Hereinafter, an apparatus that performs the training of the models to be used in the change detecting apparatus 2000 is called "training apparatus".
[0070] The training apparatus acquires a training data 200, which includes a first training image 210, a second training image 220, a first training angle 230, and a second training angle 240. The first training image 210 includes a specific place that is captured from the first training angle 230. The second training image 220 includes the specific place that is captured from the second training angle 240. The first training image 210 and the second training image 220 are prepared such that the captured place has no substantial change between them.
[0071] The training apparatus inputs the first training image 210 into the encoder 100. The encoder 100 encodes the first training image 210 into the first structural feature value 60. The first structural feature value 60 is fed into the transforming model 120. The transforming model 120 takes the first structural feature value 60, the first training angle 230, and the second training angle 240 as input, and converts the first structural feature value 60 into the second structural feature value 70. The second structural feature value 70 is fed into the decoder 160. The decoder 160 decodes the second structural feature value 70 into the simulated image 80.
[0072] The training apparatus uses the simulated image 80 and the second training image 220 to compute a dissimilarity score, which is a score represents a degree of change in the captured place from the time at which the first training image 210 is generated to the time at which the second training image 220 is generated. For example, the training apparatus computes the pixel-wise difference between the simulated image 80 and the second training image 220 to generate the dissimilarity map. Then, the training apparatus computes, as the dissimilarity score, a statistical value of the pixel values of the dissimilarity map (e.g., an average of the pixel values of the dissimilarity map).
[0073] As mentioned above, the first training image 210 and the second training image 220 are prepared such that the captured place has no substantial change between them. Thus, the more accurately the models (i.e., the encoder 100, the transforming model 120, and the decoder 160) work, the closer the dissimilarity score is to zero. Therefore, the training apparatus updates trainable parameters in the encoder 100, the transforming model 120, and the decoder 160 based on the dissimilarity score such that the dissimilarity score becomes closer to zero.
[0074] In some embodiments, as mentioned above, the feature transformation can be achieved with the resampling matrix without the transforming model 120. Fig. 9 illustrates another example of the training of the models in the change detecting apparatus 2000. Unlike the example shown by Fig. 8, the feature transformation is performed using the resampling matrix. In this case, the training apparatus updates the encoder 100 and the decoder 160 based on the dissimilarity score.
[0075] The training apparatus uses multiple pieces of the training data 200 to repeatedly update the models. Multiple pieces of the training data 200 may be generated using database of the captured images, called "image database".
[0076] For example, the change detecting apparatus 2000 retrieve a plurality of captured images from the image database. Then, the change detecting apparatus 2000 divides the captured images into groups based on the captured location. For each group, the change detecting apparatus 2000 makes the list of the captured images by sorting them in the order of the generation time.
[0077] The two captured images adjacent to each other in the list are highly likely to have no substantial change in the captured place since their generation times are close to each other. Thus, for each of the list, the training apparatus makes pairs of the two captured images adjacent to each other in the list in order to use them as pairs of the first training image 210 and the second training image 220. For each pair of the first training image 210 and the second training image 220, the training apparatus determines the first training angle 230 and the second training angle 240 by using, for example, the satellite information mentioned above.
[0078] SECOND EXAMPLE EMBODIMENT Fig. 10 illustrates another overview of a change detecting apparatus 2000. It is noted that Fig. 10 does not limit operations of the change detecting apparatus 2000, but merely show an example of possible operations of the change detecting apparatus 2000.
[0079] The change detecting apparatus 2000 of the second example embodiment is different from the change detecting apparatus 2000 of the first example embodiment in that not only the first captured image 10 but also the second captured image 20 are encoded. Specifically, the change detecting apparatus 2000 of the second example embodiment extracts features of the 3D structure of the target place 30 captured on the second captured image 20, which is the target place 30 seen from the second angle at the second time, to compute a third structural feature value 170.
[0080] After the second structural feature value 70 is computed, the change detecting apparatus 2000 performs feature fusion on the second structural feature value 70 and the third structural feature value 170, thereby generating a fused feature value 180. The change detecting apparatus 2000 decodes the fused feature value 180 into the simulated image 80.
[0081] Examples of the advantageous points of the second example embodiment are the following. By incorporating the second captured image 20 in the change detecting apparatus 2000 (especially the encoder 100, the feature fusion, and the decoder 160), the model can learn to reconstruct the simulated image 80 that is capable of compensating for variations in lighting, shadow, and other environmental factors that could either conceal or falsely represent changes between the input images. Consequently, in addition to alleviate the structural changes induced by the different sensor capturing angles, the change detecting apparatus of the second example embodiment further mitigates the occurrence of false alarms or noise in the dissimilarity map that caused by the discrepancies in color or illumination aspects.
[0082] Hereinafter, more detailed explanation of the change detecting apparatus 2000 of the second example embodiment will be described.
[0083] <Example of Functional Configuration> Fig. 11 is a block diagram showing another example of the functional configuration of the change detecting apparatus 2000. In addition to the functional units shown by Fig. 3, the change detecting apparatus 2000 shown by Fig. 11 further includes a feature fusing unit 2120. The feature fusing unit 2120 performs the feature fusion on the second structural feature value 70 and the third structural feature value 170 to compute the fused feature value 180.
[0084] In the change detecting apparatus 2000 of the second example embodiment, the encoding unit 2040 further encodes the second captured image 20 into the third structural feature value 170. When the encoding unit 2040 is implemented with the encoder 100, the encoder 100 is also used to encode the second captured image 20 into the third structural feature value 170.
[0085] <Example of Hardware Configuration> The change detecting apparatus 2000 of the second example embodiment may have the same hardware configuration as the change detecting apparatus 2000 of the first example embodiment. The storage unit 1080 of the second example embodiment stores the program implementing the functional units of the second example embodiment.
[0086] <Flow of Process> Fig. 12 shows a flowchart illustrating another example flow of processes performed by the change detecting apparatus 2000. The acquiring unit 2020 acquires the first captured image 10, the second captured image 20, and the angle information 40 (S202). The encoding unit 2040 encodes the first captured image 10 into the first structural feature value 60 (S204). The feature transforming unit 2060 performs the feature transformation on the first structural feature value 60 by referencing to the angle information 40 to compute the second structural feature value 70 (S206).
[0087] The encoding unit 2040 encodes the second captured image 20 into the third structural feature value 70 (S208). The feature fusing unit 2120 performs the feature fusion on the second structural feature value 70 and the third structural feature value 170 to compute the fused feature value 180 (S210).
[0088] The decoding unit 2080 decodes the fused feature value 180 into the simulated image 80 (S212). The detecting unit 2100 compares the simulated image 80 and the second captured image 20 to detect change in the target place 30 from the first time to the second time (S214).
[0089] It is noted that the flow of processes performed by the change detecting apparatus 2000 of the second example embodiment is not limited to that shown by Fig. 12. For example, the encoding of the second captured image 20 (S208) may be performed an arbitrary timing between S202 and S210: e.g., performed in parallel with the encoding of the first captured image 10 (S204) or the feature transformation (S206).
[0090] <Feature Fusion: S210> The feature fusing unit 2120 performs the feature fusion on the second structural feature value 70 and the third structural feature value 170 to compute the fused feature value 180 (S210). There are various ways to perform the feature fusion. For example, the feature fusion is achieved using a machine learning-based model (e.g., a neural network), called "fusion model". The feature fusion can alternatively be implemented as arithmetic operations for the amalgamation of the input feature values. These operations may include, but are not limited to, summation, subtraction, or a weighted addition mechanism.
[0091] In some implementation, the fusion model is implemented as an attention network. In this case, for example, the fusion model uses the third structural feature value 170 to compute attention weights and applies the attention weights to the second structural feature value 70, thereby outputting the fused feature value 180.
[0092] <Training of Models> Fig. 13 illustrates another example of the training of the models in the change detecting apparatus 2000. In Fig. 13, the change detecting apparatus 2000 includes four models: the encoder 100, the transforming model 120, the decoder 160, and the fusion model 190. The fusion model 190 is the fusion model (e.g., an attention network) included in the feature fusing unit 2120 mentioned above.
[0093] In the training shown by Fig. 13, the second training image 220 is fed into the encoder 100, thereby being encoded into the third structural feature value 170. The fusion model 190 takes the second structural feature value 70 and the third structural feature value 170 as input, and outputs the fused feature value 180. The fused feature value 180 is fed into the decoder 160, thereby being decoded into the simulated image 80. The training apparatus uses this simulated image 80 and the second training image 220 to compute the dissimilarity score. The training apparatus updates trainable parameters of the encoder 100, the transforming model 120, the decoder 160, and the fusion model 190 based on the dissimilarity score.
[0094] It is noted that when the feature transformation is achieved using the resampling matrix, the training apparatus updates trainable parameters of the encoder 100, the decoder 160, and the fusion model 190 based on the dissimilarity score.
[0095] The program can be stored and provided to a computer using any type of non-transitory computer readable media. Non-transitory computer readable media include any type of tangible storage media. Examples of non-transitory computer readable media include magnetic storage media (such as floppy disks, magnetic tapes, hard disk drives, etc.), optical magnetic storage media (e.g., magneto-optical disks), CD-ROM (compact disc read only memory), CD-R (compact disc recordable), CD-R / W (compact disc rewritable), and semiconductor memories (such as mask ROM, PROM (programmable ROM), EPROM (erasable PROM), flash ROM, RAM (random access memory), etc.). The program may be provided to a computer using any type of transitory computer readable media. Examples of transitory computer readable media include electric signals, optical signals, and electromagnetic waves. Transitory computer readable media can provide the program to a computer via a wired communication line (e.g., electric wires, and optical fibers) or a wireless communication line.
[0096] Although the present disclosure is explained above with reference to example embodiments, the present disclosure is not limited to the above-described example embodiments. Various modifications that can be understood by those skilled in the art can be made to the configuration and details of the present disclosure within the scope of the invention.
[0097] The whole or part of the example embodiments disclosed above can be described as, but not limited to, the following supplementary notes. <Supplementary notes> (Supplementary Note 1) A change detecting apparatus comprising: at least one memory that is configured to store instructions; and at least one processor that is configured to execute the instructions to: acquire a first captured image, a second captured image, and angle information, the first captured image being generated by capturing a target place from a first angle at a first time, the second captured image being generated by capturing the target place from a second angle at a second time, the angle information indicating the first angle and the second angle; encode the first captured image into a first structural feature value, which represents features of a three-dimensional structure of the target place seen from the first angle at the first time; perform feature transformation on the first structural feature value based on the first angle and the second angle to compute a second structural feature value, which represents features of a three-dimensional structure of the target place seen from the second angle at the first time; generate, based on the second structural feature value, a simulated image of the target place seen from the second angle at the first time; and compare the simulated image and the second captured image to detect change in the target place from the first time to the second time. (Supplementary Note 2) The change detecting apparatus according to supplementary note 1, wherein the generation of the simulated image includes decoding the second structural feature value into the simulated image. (Supplementary Note 3) The change detecting apparatus according to supplementary note 1, wherein the at least one processor is configured to further execute: encoding the second captured image into a third structural feature value, which represents features of a three-dimensional structure of the target place seen from the second angle at the second time; and fusing the second structural feature value and the third structural feature value into a fused feature value, and wherein the generation of the simulated image includes decoding the fused feature value into the simulated image. (Supplementary Note 4) The change detecting apparatus according to supplementary note 3, wherein the fusion of the second structural feature value and the third structural feature value includes applying attention based on the third structural feature value to the second structural feature value to generate the fused feature value. (Supplementary Note 5) The change detecting apparatus according to supplementary note 1, wherein the at least one memory is configured to further store a transforming model, which is a model configured to: take an input structural feature value representing features of a three-dimensional structure of a particular place seen from a first input angle, a first input angle, and a second input angle; and output an output structural feature value representing features of a three-dimensional structure of the particular place seen from the second input angle, the performing of the feature transformation includes: inputting the first structural feature value, the first angle, and the second angle to the transforming model as the input structural feature value, the first input angle, and the second input angle, respectively; and acquiring the output structural feature value from the transforming model as the second structural feature value. (Supplementary Note 6) The change detecting apparatus according to supplementary note 5, wherein training data used to train the transforming model is generated by performing: sorting the plurality of images that capture a same place in an order of a generation time of the image; selecting two of the sorted images that are adjacent to each other; and generating the training data including the selected images as images equivalent to the first captured image and the second captured image. (Supplementary Note 7) The change detecting apparatus according to supplementary note 1, wherein the detection of change in the target place includes: generating a dissimilarity map, which is an image each of whose pixel indicates a difference between a corresponding pixel of the simulated image and a corresponding pixel of the second captured image; and determining whether there is change in the target place from the first time to the second time based on pixel values of the dissimilarity map. (Supplementary Note 8) A change detecting method performed by a computer, comprising: acquiring a first captured image, a second captured image, and angle information, the first captured image being generated by capturing a target place from a first angle at a first time, the second captured image being generated by capturing the target place from a second angle at a second time, the angle information indicating the first angle and the second angle; encoding the first captured image into a first structural feature value, which represents features of a three-dimensional structure of the target place seen from the first angle at the first time; performing feature transformation on the first structural feature value based on the first angle and the second angle to compute a second structural feature value, which represents features of a three-dimensional structure of the target place seen from the second angle at the first time; generating, based on the second structural feature value, a simulated image of the target place seen from the second angle at the first time; and comparing the simulated image and the second captured image to detect change in the target place from the first time to the second time. (Supplementary Note 9) The change detecting method according to supplementary note 8, wherein the generation of the simulated image includes decoding the second structural feature value into the simulated image. (Supplementary Note 10) The change detecting method according to supplementary note 8, further comprising: encoding the second captured image into a third structural feature value, which represents features of a three-dimensional structure of the target place seen from the second angle at the second time; and fusing the second structural feature value and the third structural feature value into a fused feature value, and wherein the generation of the simulated image includes decoding the fused feature value into the simulated image. (Supplementary Note 11) The change detecting method according to supplementary note 10, wherein the fusion of the second structural feature value and the third structural feature value includes applying attention based on the third structural feature value to the second structural feature value to generate the fused feature value. (Supplementary Note 12) The change detecting method according to supplementary note 8, wherein the computer includes a transforming model, which is a model configured to: take an input structural feature value representing features of a three-dimensional structure of a particular place seen from a first input angle, a first input angle, and a second input angle; and output an output structural feature value representing features of a three-dimensional structure of the particular place seen from the second input angle, the performing of the feature transformation includes: inputting the first structural feature value, the first angle, and the second angle to the transforming model as the input structural feature value, the first input angle, and the second input angle, respectively; and acquiring the output structural feature value from the transforming model as the second structural feature value. (Supplementary Note 13) The change detecting method according to supplementary note 12, wherein training data used to train the transforming model is generated by performing: sorting the plurality of images that capture a same place in an order of a generation time of the image; selecting two of the sorted images that are adjacent to each other; and generating the training data including the selected images as images equivalent to the first captured image and the second captured image. (Supplementary Note 14) The change detecting method according to supplementary note 8, wherein the detection of change in the target place includes: generating a dissimilarity map, which is an image each of whose pixel indicates a difference between a corresponding pixel of the simulated image and a corresponding pixel of the second captured image; and determining whether there is change in the target place from the first time to the second time based on pixel values of the dissimilarity map. (Supplementary Note 15) A non-transitory computer-readable medium storing a program that causes a computer to execute: acquiring a first captured image, a second captured image, and angle information, the first captured image being generated by capturing a target place from a first angle at a first time, the second captured image being generated by capturing the target place from a second angle at a second time, the angle information indicating the first angle and the second angle; encoding the first captured image into a first structural feature value, which represents features of a three-dimensional structure of the target place seen from the first angle at the first time; performing feature transformation on the first structural feature value based on the first angle and the second angle to compute a second structural feature value, which represents features of a three-dimensional structure of the target place seen from the second angle at the first time; generating, based on the second structural feature value, a simulated image of the target place seen from the second angle at the first time; and comparing the simulated image and the second captured image to detect change in the target place from the first time to the second time. (Supplementary Note 16) The medium according to supplementary note 15, wherein the generation of the simulated image includes decoding the second structural feature value into the simulated image. (Supplementary Note 17) The medium according to supplementary note 15, wherein the program causes the computer to further execute: encoding the second captured image into a third structural feature value, which represents features of a three-dimensional structure of the target place seen from the second angle at the second time; and fusing the second structural feature value and the third structural feature value into a fused feature value, and wherein the generation of the simulated image includes decoding the fused feature value into the simulated image. (Supplementary Note 18) The medium according to supplementary note 17, wherein the fusion of the second structural feature value and the third structural feature value includes applying attention based on the third structural feature value to the second structural feature value to generate the fused feature value. (Supplementary Note 19) The medium according to supplementary note 15, wherein the program includes a transforming model, which is a model configured to: take an input structural feature value representing features of a three-dimensional structure of a particular place seen from a first input angle, a first input angle, and a second input angle; and output an output structural feature value representing features of a three-dimensional structure of the particular place seen from the second input angle, the performing of the feature transformation includes: inputting the first structural feature value, the first angle, and the second angle to the transforming model as the input structural feature value, the first input angle, and the second input angle, respectively; and acquiring the output structural feature value from the transforming model as the second structural feature value. (Supplementary Note 20) The medium according to supplementary note 19, wherein training data used to train the transforming model is generated by performing: sorting the plurality of images that capture a same place in an order of a generation time of the image; selecting two of the sorted images that are adjacent to each other; and generating the training data including the selected images as images equivalent to the first captured image and the second captured image. (Supplementary Note 21) The medium according to supplementary note 15, wherein the detection of change in the target place includes: generating a dissimilarity map, which is an image each of whose pixel indicates a difference between a corresponding pixel of the simulated image and a corresponding pixel of the second captured image; and determining whether there is change in the target place from the first time to the second time based on pixel values of the dissimilarity map.
[0098] 10 first captured image 12 sub-region 20 second captured image 30 target place 40 angle information 41 first angle 42 first incident angle 43 first azimuth angle 44 second angle 45 second incident angle 46 second azimuth angle 50 sensor 55 sensor 60 first structural feature value 62 sequence 70 second structural feature value 80 simulated image 90 first coordinate system 92 first azimuth-axis 94 first range-axis 96 first incident-axis 100 encoder 120 transforming model 121 first model 122 second model 130 matrix 140 matrix 150 feature cuboid 160 decoder 170 third structural feature value 180 fused feature value 190 fusion model 200 training data 210 first training image 220 second training image 230 first training angle 240 second training angle 1000 computer 1020 bus 1040 processor 1060 memory 1080 storage device 1100 input / output interface 1120 network interface 2000 change detecting apparatus 2020 acquiring unit 2040 encoding unit 2060 feature transforming unit 2080 decoding unit 2100 detecting unit 2120 feature fusing unit
Claims
1. A change detecting apparatus comprising: at least one memory that is configured to store instructions; and at least one processor that is configured to execute the instructions to: acquire a first captured image, a second captured image, and angle information, the first captured image being generated by capturing a target place from a first angle at a first time, the second captured image being generated by capturing the target place from a second angle at a second time, the angle information indicating the first angle and the second angle; encode the first captured image into a first structural feature value, which represents features of a three-dimensional structure of the target place seen from the first angle at the first time; perform feature transformation on the first structural feature value based on the first angle and the second angle to compute a second structural feature value, which represents features of a three-dimensional structure of the target place seen from the second angle at the first time; generate, based on the second structural feature value, a simulated image of the target place seen from the second angle at the first time; and compare the simulated image and the second captured image to detect change in the target place from the first time to the second time.
2. The change detecting apparatus according to claim 1, wherein the generation of the simulated image includes decoding the second structural feature value into the simulated image.
3. The change detecting apparatus according to claim 1, wherein the at least one processor is configured to further execute: encoding the second captured image into a third structural feature value, which represents features of a three-dimensional structure of the target place seen from the second angle at the second time; and fusing the second structural feature value and the third structural feature value into a fused feature value, and wherein the generation of the simulated image includes decoding the fused feature value into the simulated image.
4. The change detecting apparatus according to claim 3, wherein the fusion of the second structural feature value and the third structural feature value includes applying attention based on the third structural feature value to the second structural feature value to generate the fused feature value.
5. The change detecting apparatus according to claim 1, wherein the at least one memory is configured to further store a transforming model, which is a model configured to: take an input structural feature value representing features of a three-dimensional structure of a particular place seen from a first input angle, a first input angle, and a second input angle; and output an output structural feature value representing features of a three-dimensional structure of the particular place seen from the second input angle, the performing of the feature transformation includes: inputting the first structural feature value, the first angle, and the second angle to the transforming model as the input structural feature value, the first input angle, and the second input angle, respectively; and acquiring the output structural feature value from the transforming model as the second structural feature value.
6. The change detecting apparatus according to claim 5, wherein training data used to train the transforming model is generated by performing: sorting the plurality of images that capture a same place in an order of a generation time of the image; selecting two of the sorted images that are adjacent to each other; and generating the training data including the selected images as images equivalent to the first captured image and the second captured image.
7. The change detecting apparatus according to claim 1, wherein the detection of change in the target place includes: generating a dissimilarity map, which is an image each of whose pixel indicates a difference between a corresponding pixel of the simulated image and a corresponding pixel of the second captured image; and determining whether there is change in the target place from the first time to the second time based on pixel values of the dissimilarity map.
8. A change detecting method performed by a computer, comprising: acquiring a first captured image, a second captured image, and angle information, the first captured image being generated by capturing a target place from a first angle at a first time, the second captured image being generated by capturing the target place from a second angle at a second time, the angle information indicating the first angle and the second angle; encoding the first captured image into a first structural feature value, which represents features of a three-dimensional structure of the target place seen from the first angle at the first time; performing feature transformation on the first structural feature value based on the first angle and the second angle to compute a second structural feature value, which represents features of a three-dimensional structure of the target place seen from the second angle at the first time; generating, based on the second structural feature value, a simulated image of the target place seen from the second angle at the first time; and comparing the simulated image and the second captured image to detect change in the target place from the first time to the second time.
9. The change detecting method according to claim 8, wherein the generation of the simulated image includes decoding the second structural feature value into the simulated image.
10. The change detecting method according to claim 8, further comprising: encoding the second captured image into a third structural feature value, which represents features of a three-dimensional structure of the target place seen from the second angle at the second time; and fusing the second structural feature value and the third structural feature value into a fused feature value, and wherein the generation of the simulated image includes decoding the fused feature value into the simulated image.
11. The change detecting method according to claim 10, wherein the fusion of the second structural feature value and the third structural feature value includes applying attention based on the third structural feature value to the second structural feature value to generate the fused feature value.
12. The change detecting method according to claim 8, wherein the computer includes a transforming model, which is a model configured to: take an input structural feature value representing features of a three-dimensional structure of a particular place seen from a first input angle, a first input angle, and a second input angle; and output an output structural feature value representing features of a three-dimensional structure of the particular place seen from the second input angle, the performing of the feature transformation includes: inputting the first structural feature value, the first angle, and the second angle to the transforming model as the input structural feature value, the first input angle, and the second input angle, respectively; and acquiring the output structural feature value from the transforming model as the second structural feature value.
13. The change detecting method according to claim 12, wherein training data used to train the transforming model is generated by performing: sorting the plurality of images that capture a same place in an order of a generation time of the image; selecting two of the sorted images that are adjacent to each other; and generating the training data including the selected images as images equivalent to the first captured image and the second captured image.
14. The change detecting method according to claim 8, wherein the detection of change in the target place includes: generating a dissimilarity map, which is an image each of whose pixel indicates a difference between a corresponding pixel of the simulated image and a corresponding pixel of the second captured image; and determining whether there is change in the target place from the first time to the second time based on pixel values of the dissimilarity map.
15. A non-transitory computer-readable medium storing a program that causes a computer to execute: acquiring a first captured image, a second captured image, and angle information, the first captured image being generated by capturing a target place from a first angle at a first time, the second captured image being generated by capturing the target place from a second angle at a second time, the angle information indicating the first angle and the second angle; encoding the first captured image into a first structural feature value, which represents features of a three-dimensional structure of the target place seen from the first angle at the first time; performing feature transformation on the first structural feature value based on the first angle and the second angle to compute a second structural feature value, which represents features of a three-dimensional structure of the target place seen from the second angle at the first time; generating, based on the second structural feature value, a simulated image of the target place seen from the second angle at the first time; and comparing the simulated image and the second captured image to detect change in the target place from the first time to the second time.
16. The medium according to claim 15, wherein the generation of the simulated image includes decoding the second structural feature value into the simulated image.
17. The medium according to claim 15, wherein the program causes the computer to further execute: encoding the second captured image into a third structural feature value, which represents features of a three-dimensional structure of the target place seen from the second angle at the second time; and fusing the second structural feature value and the third structural feature value into a fused feature value, and wherein the generation of the simulated image includes decoding the fused feature value into the simulated image.
18. The medium according to claim 17, wherein the fusion of the second structural feature value and the third structural feature value includes applying attention based on the third structural feature value to the second structural feature value to generate the fused feature value.
19. The medium according to claim 15, wherein the program includes a transforming model, which is a model configured to: take an input structural feature value representing features of a three-dimensional structure of a particular place seen from a first input angle, a first input angle, and a second input angle; and output an output structural feature value representing features of a three-dimensional structure of the particular place seen from the second input angle, the performing of the feature transformation includes: inputting the first structural feature value, the first angle, and the second angle to the transforming model as the input structural feature value, the first input angle, and the second input angle, respectively; and acquiring the output structural feature value from the transforming model as the second structural feature value.
20. The medium according to claim 19, wherein training data used to train the transforming model is generated by performing: sorting the plurality of images that capture a same place in an order of a generation time of the image; selecting two of the sorted images that are adjacent to each other; and generating the training data including the selected images as images equivalent to the first captured image and the second captured image.
Citation Information
Patent Citations
Image capturing method and image capturing apparatus
JP2021096805A
Object amount calculation device and object amount calculation method
WO2020179438A1