Classifying objects with respect to their angles

The method addresses the challenge of varying image data by using a machine learning model to classify objects based on angle of incidence and image parameters, improving classification accuracy in imaging systems.

JP7785180B2Active Publication Date: 2025-12-12アイサイ オサケユキチュア
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
JP2024536031
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2021-12-17
Filing Date
2022-11-30
Publication Date
2025-12-12
Estimated Expiration
2042-11-30

Smart Images

  • Figure 0007785180000005
    Figure 0007785180000005
  • Figure 0007785180000006
    Figure 0007785180000006
  • Figure 0007785180000007
    Figure 0007785180000007
Patent Text Reader

Abstract

1. A computer-implemented method for classifying objects in an image, the method comprising: receiving image data associated with the image; receiving angle of incidence data, the angle of incidence data indicating an angle of incidence at which the image data is collected by a detector; and classifying one or more objects in the image as belonging to one of one or more categories using a machine learning model, the classification of the one or more objects in the image being based on the angle of incidence data and respective values ​​of one or more parameters of the image data.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] FIELD OF THE INVENTION This application relates to a method for classifying objects in an image, and in particular to a method for classifying objects in an image based at least in part on the angle of incidence at which associated image data is collected. [Background technology]

[0002] In many imaging systems, the qualitative and quantitative properties of the collected imaging data can vary significantly based on the angle of incidence at which the image data is collected. For example, in an active radar system configured for remote sensing, the collected image data includes radar backscatter signals received at a detector. The intensity, phase, and other characteristics of the radar backscatter signals are highly dependent on both the angle of incidence at which the image data is collected and the optical properties of the target imaged by the radar signal. Such optical properties may include reflectivity, transmittance, and other properties that are themselves strongly dependent on the angle of incidence.

[0003] In satellite-based radar imaging competitions, image data may be collected from a range of incidence angles. In such systems, two images of the same imaging object may differ significantly in both quantity and quality due to the differences in the incidence angles at which the respective image data associated with each image is collected. This variability in image data can present significant challenges when attempting to accurately analyze and / or compare images, such as when attempting to classify objects or regions within the images.

[0004] Based on the above considerations, the inventors have devised the claimed invention.

[0005] The embodiments described below are not limited to implementations that address any or all of the drawbacks of known methods discussed above. Summary of the Invention

[0006] This Summary is provided to introduce a selection of concepts in a simplified form. These concepts are further described below in the Detailed Description. This Summary is not intended to identify key features or essential features of the claimed subject matter, nor is it intended for use in determining the scope of the claimed subject matter. Modifications and alternative features used to facilitate the practice of the invention and / or to achieve a substantially similar technical effect are deemed to be within the scope of the invention(s) disclosed herein.

[0007] Generally, the present disclosure provides a method for classifying objects in images that takes into account each angle of incidence at which image data associated with each image is collected.

[0008] The invention is defined as set forth in the following claims.

[0009] In a first aspect of the present invention, there is provided a computer-implemented method for classifying objects in an image, the method comprising: receiving image data associated with the image; receiving angle of incidence data, the angle of incidence data indicating an angle of incidence at which image data was collected by a detector; and using a machine learning model to classify one or more objects in the image as belonging to one of one or more categories, wherein the classification of the one or more objects in the image is based on the angle of incidence data and values ​​of one or more parameters of the image data.

[0010] Thus, classification of objects by this method may be sensitive to differences in image data caused by changes in the angle of incidence at which the image data is collected.

[0011] In another aspect of the invention, there is provided a computing device configured to perform any of the methods disclosed herein.

[0012] In another aspect of the present invention, a computer-readable storage medium is provided that includes instructions that, when executed by a computer, cause the computer to perform any of the methods disclosed herein.

[0013] In another aspect of the present invention, there is provided a computer program comprising instructions which, when executed by a computer, cause the computer to carry out any of the methods disclosed herein.

[0014] The methods described herein can be implemented by software in machine-readable form on a tangible storage medium, e.g., in the form of a computer program including computer program code means adapted to perform all the steps of any of the methods described herein when the program is run on a computer, and the computer program can be embodied on a computer-readable medium. Examples of tangible (or non-transitory) storage media include magnetic disks, thumb drives, memory cards, etc., but do not include propagating signals. The software may be adapted to run on a parallel or serial processor such that the method steps can be performed in any suitable order or simultaneously.

[0015] This application recognizes that firmware and software are individually tradable commodities with value. It is designed to include software that operates or controls on "dumb" or standard hardware to perform a desired function. It is also intended to include software that "describes" or defines hardware configurations, such as HDL (Hardware Description Language) software that designs silicon chips or configures general-purpose programmable chips to perform desired functions.

[0016] The features and embodiments described herein may be combined as appropriate as would be apparent to one skilled in the art, and may be combined with any aspect of the invention unless expressly specified that such a combination is not possible, or unless one skilled in the art would understand that such a combination is not possible. Hereinafter, an embodiment of the present invention will be described by way of example with reference to the drawings. [Brief explanation of the drawings]

[0017] [Figure 1] 1 is a flowchart illustrating a method for classifying objects in an image. [Figure 2] 1 illustrates an example of a satellite capturing images of an imaging object from different angles of incidence. [Figure 3a] 10a-10d show example scenarios of imaging signals reflecting from the surface of an imaging object. [Figure 3b] 10a-10d show example scenarios of imaging signals reflecting from the surface of an imaging object. [Figure 3c] 10a-10d show example scenarios of imaging signals reflecting from the surface of an imaging object. [Figure 3d] 10a-10d show example scenarios of imaging signals reflecting from the surface of an imaging object. [Figure 4] 1 is a flowchart illustrating a method for training a CNN to classify one or more objects in an image based on incidence angle data and one or more parameters of the image data. [Figure 5] 4 shows a schematic of how the incident angle data is augmented to generate a training dataset for training the CNN in FIG. [Figure 6] 1 illustrates a method for classifying objects in an image according to the claimed invention. [Figure 7]We show the performance of the CNN in identifying water with and without incident angle information. [Figure 8] An example of a SAR image containing both water and land areas is shown below. [Figure 9] The results of processing SAR images without incident angle information are shown. [Figure 10] 1 shows the results of processing a SAR image containing angle of incidence information. Common reference numbers are used in the figures to denote the same or similar features. DETAILED DESCRIPTION OF THE INVENTION

[0018] Embodiments of the present invention are described below by way of example only. These examples represent the best mode of practicing the invention currently known to the applicant, and are not the only way to accomplish this. The description sets forth the functions of the examples and the sequence of steps for constructing and operating the examples. However, the same or equivalent functions and sequences may be implemented by different examples.

[0019] FIG. 1 is a flow chart illustrating a method for classifying objects in an image.

[0020] In act S100, image data associated with an image is received.

[0021] The angle of incidence data indicates the angle of incidence at which image data is collected by the detector.

[0022] In some embodiments, the incident angle data may be received simultaneously with the image data, in other words, operations S100 and S102 may be performed simultaneously.

[0023] In this way, the accuracy of the incident angle data may be increased so that the incident angle data truly represents the incident angle when the detector collects the image data.

[0024] Additionally or alternatively, the image data may include metadata, such as information about the time and / or date the image data was collected, e.g., a timestamp associated with the image data. In such an example, receiving the angle of incidence data may include receiving the angle of incidence data based on the timestamp and other metadata related to the time and / or date the image data was collected. In this manner, analysis of the image may be delayed until a time convenient for a user of the methods disclosed herein. For example, a user may select an image to process from a stack of images collected by a detector. The user may receive the angle of incidence data based on metadata associated with the image data, e.g., a timestamp, as opposed to the angle of incidence data associated with the time the pre-stored image data was received. In this case, operations S100 and S102 may be performed at different times.

[0025] In examples involving satellite imaging, the metadata may include information such as a satellite identifier number, i.e., an identifier for which of multiple satellites the detector that collected the image data is on board. This information may also indicate the hardware on board the satellite. For example, identifying the satellite may provide information about the characteristics of the detector and / or the satellite itself. In some examples, this information may be included directly in the metadata.

[0026] In some embodiments, the incidence angle data may be derived from a local incidence angle model that is based on a geometric model of the Earth and one or more state vectors that indicate the position of the detector relative to the geometric model of the Earth.

[0027] The angle of incidence data may be a sufficiently accurate approximation without the need to obtain precise angle of incidence measurements, which are difficult and expensive.

[0028] In some embodiments, the geometric model of the Earth may be an ellipsoidal model.

[0029] Alternatively, the geometric model of the Earth may be a spherical model, a "real Earth" model, or based on a digital elevation model (DEM). Each model provides a different approach to determining the elevation of objects on the Earth's surface, based on a different model of the Earth's shape. In practice, each geometric model strikes a balance between the computational cost of determining shape and elevation and the accuracy of the model. In other words, the more complex the model (e.g., a "real Earth" model), the more accurately the shape of the Earth's surface can be determined, but the computational cost of modeling the elevation and precise coordinates of objects increases. In contrast, less complex models (e.g., spherical models) are significantly faster and cheaper to process, but offer lower accuracy due to their simplified model of the Earth's shape. In practice, we have found that ellipsoidal models offer a good balance by reducing the computational cost associated with processing based on geometric models without sacrificing processing accuracy beyond an acceptable level.

[0030] In another operation S104, a machine learning model is used to classify one or more objects in the image into one of one or more categories based on the angle of incidence data and the respective values ​​of the one or more parameters of the image data.

[0031] In some embodiments, classifying one or more objects in an image may include classifying each of a plurality of pixels of the image as belonging to one of one or more categories.

[0032] In this way, each pixel may be classified according to one or more categories, allowing for fine detail in the classification, and allowing for classification at a resolution that matches the resolution of the images collected by the detector.

[0033] In some embodiments, classifying one or more objects in an image may include classifying one or more objects in a portion of the image.

[0034] In this way, computing and processing resources may be optimized by restricting object classification to be performed only within a target region within the image. This section or region of interest of the image may be determined automatically by a machine learning model based on one or more parameters used to train the model, or based on one or more parameters selected by a user of the method and input into the machine learning model. Additionally or alternatively, a user of the method may manually select the section or region of interest within which the classification should be performed.

[0035] In some embodiments, the one or more parameters of the image data may include one or more of an intensity value for each of the plurality of pixels, a color channel value for each of the plurality of pixels, and / or phase information for each of the plurality of pixels.

[0036] For example, the intensity value of each of the plurality of pixels may indicate the intensity of the imaging signal received at each pixel of the detector. The intensity values ​​may correspond to grayscale values ​​indicating signal strength, for example, in an 8-bit format allowing for grayscale values ​​ranging from 0 to 255, i.e., 256 grayscales. Alternatively, the grayscale values ​​may be encoded in a smaller format, for example, a 4-bit format, allowing for grayscale values ​​ranging from 0 to 15, or in a larger format, for example, a 16-bit format, allowing for grayscale values ​​ranging from 0 to 65535. Generally, when grayscale values ​​are encoded in an n-bit format, the grayscale values ​​may range from 0 to 255. n-1 The range may be:

[0037] Additionally or alternatively, the one or more parameters of the image data may include respective channel values ​​for a plurality of pixels. In some examples, such as in radar imaging, the channel values ​​may be polarization channels, with each channel configured to be sensitive to a different radar signal depending on its polarization. For example, in many radar imaging systems, including satellite-based radar imaging systems, imaging signals collected by a detector may be backscattered from one or more imaging objects. The strength of the backscattered signal may depend on many factors, such as the angle of incidence and the polarization of the imaging signal. In addition, the backscattered imaging signal may include components of one or more polarizations different from the polarization of the incident imaging signal.

[0038] In such an example, the polarization channels may include HH channels, i.e., channels that are particularly sensitive to radar signals transmitted and received with "horizontal" polarization; VV channels, i.e., channels that are particularly sensitive to radar signals transmitted and received with "vertical" polarization; HV channels, i.e., channels that are particularly sensitive to radar signals transmitted with "horizontal" polarization and received with "vertical" polarization; and VH channels, i.e., channels that are particularly sensitive to radar signals transmitted with "vertical" polarization and received with "horizontal" polarization. Additionally or alternatively, the polarization channels may include channels that are particularly sensitive to radar signals transmitted and / or received with circular or elliptically polarized light, or any combination of linear, circular, and elliptically polarized light.

[0039] Additionally, the strength of the backscattered signal often depends on the wavelength of the imaging signal. In all practical applications, the imaging signal has a bandwidth. In other words, the image signal is not completely monochromatic, but rather contains signals with a range of wavelengths. In such cases, there may be variations in the received signal based on the wavelengths of the various components of the signal. Thus, in such instances, the channel values ​​may encode information based on the variation in backscattered intensity with respect to the signal wavelength.

[0040] In other examples, such as in an optical system, the wavelength of the imaging signal may correspond to the color of the signal, where a range of wavelengths corresponds to a range or mixture of colors, and in such cases the channel values ​​may be color channel values ​​for multiple pixels that encode information based on the variation of backscattering intensity with respect to signal wavelength.

[0041] The color channels may be classified using any suitable channel scheme. For example, the color channels may be based on an additive structure, such as RGB channels, HSL channels, or HSV channels (also known as HSB channels). Additionally or alternatively, the color channels may be based on a subtractive structure, such as CMYK channels. Additionally or alternatively, the color channels may be based on luma and chromaticity components, such as YUV channels.

[0042] In some examples, the one or more parameters may include phase information for each of a plurality of pixels. The relative phase information between pixels may provide information indicative of the relative height or orientation of an object imaged at different pixels based on the phase difference between the pixels. The relative phase information may be particularly relevant to processing images imaged by fast-moving detectors, such as detectors on satellites, because the orientation of an object can significantly affect the strength of a backscattered signal due to the orientation of reflective surfaces.

[0043] The phase values ​​may be encoded separately from the intensity values ​​described above. In another example, the phase and intensity values ​​may be encoded as complex data values, i.e., data encoded as complex numbers having real and imaginary parts. The magnitude, or modulus, of each complex data value may correspond to a respective intensity value, and the argument of each complex data value may correspond to a respective phase value.

[0044] In some embodiments, the detector may be mounted on a satellite in Earth orbit.

[0045] FIG. 2 shows an example of a satellite 210 capturing images of an imaging object 222 from different angles of incidence.

[0046] As the satellite 210 orbits the Earth 220, the detector may be oriented to capture several images of the same imaging object 222. In some examples, the entire satellite 210 may rotate to change the orientation of the detector. Additionally or alternatively, the detector may rotate relative to the platform on which it is carried, in conjunction with the satellite or aircraft, to adjust the orientation of the detector. However, motion of the satellite 210 may change the angle of incidence θ at which the detector collects image data. In some examples, the angle of incidence may vary significantly between successively imaged images.

[0047] In the example shown in FIG. 2 , the angles of incidence vary as satellite 210 collects image data from three exemplary points in orbit around Earth 220. For example, when the satellite is at a first location along its orbit, the angle of incidence of the imaging signal reflected from imaging object 222 is a first exemplary incidence. In the example shown in FIG. 2 , the first angle of incidence is approximately −30°. When the satellite is at a second location along its orbit, the angle of incidence of the imaging signal reflected from imaging object 222 is a second exemplary incidence. In the example shown in FIG. 2 , the second angle of incidence is approximately 10°. When the satellite is at a third location along its orbit, the angle of incidence of the imaging signal reflected from imaging object 222 is a third exemplary incidence. In the example shown in FIG. 2 , the third angle of incidence is approximately +30°.

[0048] 2 shows only one imaging object 222 for illustrative purposes, each image collected by the detectors onboard the satellite 210 may collect image data associated with multiple imaging objects 222. There may be complete, partial, or no overlap between the imaging objects 222 imaged in each of the different images. In other words, each image collected by the detectors onboard the satellite 210 may include image data associated with all of the same imaging objects 222, each image collected by the detectors onboard the satellite 210 may include image data associated with portions of the same imaging object 222 and / or different imaging objects 222, or each image collected by the detectors onboard the satellite 210 may include image data associated with different imaging objects 222.

[0049] In some examples of satellite-based imaging, the satellite 210 does not pass directly overhead the imaging object 222. In other words, the detector may capture images offset from the satellite's orbital path, i.e., the detector may not capture images of the area directly below (nadir) the satellite 210. Additionally, as seen in FIG. 2 , the angle of incidence may be a combination of the cross-track angle (perpendicular to the satellite's direction of travel) between the satellite 210 and the imaging object 222 and the along-track angle between the satellite 210 and the imaging object 222.

[0050] The orbital period of the satellite 210 may be 6 hours or less, 12 hours or less, 18 hours or less, 24 hours or less, 36 hours or less, 48 ​​hours or less, or 72 hours or less.

[0051] In some embodiments, one or more state vectors indicative of detector positions (based at least in part on the incidence angle data) may be based on a model of the orbital path of satellite 210 .

[0052] It may be possible to determine the incidence angle data without requiring precise GPS measurements of the satellite's 210 position at the time the image data was collected by the detector. Instead, given sufficiently accurate initialization coordinates, the satellite's future orbit may be modeled over the satellite's orbital lifetime. For example, if the semi-major axis radius of the satellite's 210 orbit around Earth 220 is provided along with coordinates of the satellite's 210 position and velocity at a given time, it may be possible to determine a model of the satellite's 210 orbit. This model may be based, for example, on a Keplerian orbit model or similar. Additionally, the model of the satellite's orbit may be updated (periodically and / or as needed), for example, if the satellite 210 undergoes maneuvers that change its orbital path.

[0053] In some embodiments, the image may be a synthetic aperture radar image.

[0054] Synthetic aperture radar (SAR) imagery is typically based on transmitting a radar imaging signal and then detecting and recording the backscattered signal received by a detector that is reflected and / or scattered from one or more imaging objects 222. The strength of the backscattered signal depends on the angle of incidence θ at which the image data is collected. The backscattered strength may also depend on other factors, such as the wavelength of the imaging signal, the reflectivity, transmittance, and orientation of each imaging object 222. Considering each of these factors is important for accurately classifying imaging objects according to one or more categories. As one of the primary factors affecting the strength of the backscattered signal, it is crucial to consider the angle of incidence to improve classification accuracy.

[0055] In some embodiments, one or more categories may include one or more geographic features.

[0056] In some embodiments, the one or more geographic features may include one or more of water, ice, cultivated land, forested areas, and / or man-made structures.

[0057] In satellite-based or aerial imaging applications, detectors may capture images of a geographic environment. These images may be imaged for environmental, academic, military, or other purposes. When these images are based on SAR imaging technology, the collected image data may be based on backscattered signals reflected from various geographic features within the geographic environment. The strength of the backscattered signals may be strongly dependent on the nature of the geographic feature from which the signal is reflected.

[0058] Figure 3a shows a scenario where an imaging signal reflects off the surface of a smooth imaging object.

[0059] When a radar signal reflects off a relatively smooth surface, the surface acts like a mirror and the signal may be specularly reflected from the surface, as seen in Figure 3a. If the specular reflection is strong, backscattering will be minimal, and therefore the strength of the backscattered signal collected by the imaging detector will be very low.

[0060] Figure 3b shows the scenario of an imaging signal reflected from the surface of a rough imaging object.

[0061] In contrast to specular reflection from a smooth surface, when a radar signal reflects from a relatively rough surface, the surface causes the signal to be diffusely reflected from the surface. When a signal is diffusely reflected, it is reflected in all directions, as seen in Figure 3b. This reflection can be nearly isotropic or relatively anisotropic. For example, a surface that is intermediate between "rough" and "smooth" may reflect a radar signal in a generally specular manner, but with a diffuse component. Thus, the angular spectrum of the reflected signal may resemble a cone whose central axis is along the line of specular reflection. The half-angle of this cone increases with the roughness of the reflecting surface.

[0062] FIG. 3c illustrates the scenario of an imaging signal reflecting from the surface of an imaging object whose surface roughness is intermediate between "rough" and "smooth," as discussed above.

[0063] In radar reflection analysis, the terms "smooth" and "rough" refer to surface smoothness and surface roughness. Surface smoothness or roughness may be quantified, for example, in terms of the Rayleigh criterion. For example, the surface roughness, d, of a surface may be determined as the root mean square roughness height from a reference plane, i.e., the root mean square deviation of the object surface from a plane defining the average position of the surface. Relatively rough surfaces will have higher surface roughness values ​​than relatively smooth surfaces. When an image signal of wavelength λ is incident on a surface at an angle of incidence θ, the surface may be considered smooth if:

[0064]

number

[0065] Alternatively, a surface may be considered rough if:

[0066]

number

[0067] Alternatively, a surface may be considered intermediate between a rough and a smooth surface if:

[0068]

number

[0069] It is therefore clear that the relative smoothness or roughness of a surface depends not only on the surface properties but also on the wavelength and angle of incidence of the imaging (radar) signal.

[0070] For example, if the terrain is water, the degree of surface roughness depends on how relatively still the water surface is. For example, if a detector captures an image of relatively calm water, the effective surface will be relatively smooth, resulting in specular reflection of the radar signal from the water surface. Thus, the backscatter signal strength from calm water will be stronger at steeper angles of incidence, i.e., when the detector is closer to directly overhead than at shallower angles of incidence, i.e., when the detector is closer to the horizon relative to the imaging plane. In contrast, if a detector captures an image of rough water, the effective surface will be rougher, thereby increasing the range of angles of incidence over which a strong backscatter signal will be captured.

[0071] Similarly, if the terrain is ice, the ice surface for radar imaging will be smooth, resulting in specular reflection of the radar.

[0072] In contrast, if the geographic feature is forested, the surface roughness (caused by treetops) results in a large degree of diffuse reflection, making the reflected radar signal nearly isotropic. Therefore, the strength of the backscattered signal from a forested area depends only slightly on the angle of incidence.

[0073] If the geographic feature is arable land, the surface roughness caused by the presence of crops on flat land may result in a reflected radar signal that is primarily dominated by specular reflection but also includes a diffuse reflection component that broadens the angular spectrum of the reflected signal into a cone. In other words, if the imaging signal is a radar signal, arable land may represent a surface that is intermediate between a rough and a smooth surface, as discussed above.

[0074] FIG. 3d illustrates a scenario in which an imaging signal experiences corner reflections from man-made structures, such as buildings.

[0075] A special case of specular reflection may occur if the geographic feature is an artificial structure. Artificial structures are typically constructed with walls (or other vertical elements) perpendicular to their surface. This means that the incident radar signal may experience so-called corner reflections: when a radar signal reflects off two smooth surfaces that are perpendicular to each other, the signal is reflected twice (once from each surface) and reflected back towards the detector. In such cases, the strength of the reflected radar signal collected at the detector will be very high, higher than the diffuse reflection signal collected from the backscattered signal from the rough surfaces.

[0076] The machine learning model may be trained to understand the relationship between object features (such as surface roughness) and the angle of incidence, and how that relationship affects the signal received by the detector, thereby training the machine learning model to more accurately classify the imaging object 222 based on the angle of incidence data indicative of the angle of incidence and the value of one or more parameters of the image data, such as the signal intensity collected for each pixel of the image, as described above.

[0077] In some examples, it may be necessary to capture images of the imaging object 222 from multiple different angles of incidence to enable more accurate classification of the imaging object 222. For example, more accurate classification may be achievable if images of the imaging object 222 are available from a relatively shallow angle, a relatively steep angle, and an intermediate angle between the shallow and steep angles. The strength of the signal strength across the range of angles of incidence may enable more accurate classification.

[0078] For purely illustrative purposes, examples of lookup tables for classifying "rough" water, "calm" water, ice, cultivated land, forest land, and man-made structures are included below for "shallow" incidence angles, "steep" incidence angles, and "intermediate" incidence angles, which are angles between shallow and steep. It should be noted that the tables below are qualitative and are included to illustrate some of the logic that machine learning models learn during training.

[0079] [Table 1] Table 1: Logic lookup table showing the relationship between signal strength and angle of incidence for various geographic features.

[0080] As can be seen from Table 1 above, some graphical features may exhibit a very similar relationship between the signal strength collected by the image detector and the angle of incidence. In such cases, the machine learning model may be further trained based on metadata, such as the GPS coordinates of the imaging object 222, that provides contextual information to the machine learning model, such as to distinguish between rough water and forested areas or calm water and ice.

[0081] FIG. 4 illustrates a method for training a CNN to classify one or more objects in an image based on incident angle data and one or more parameters of the image data.

[0082] In act S400, the machine learning model receives image data associated with a plurality of training images.

[0083] The incident angle data may be similar to the incident angle model described above for incident angle data received in the manner shown in FIG.

[0084] Further, in operations S404-S412, a training data patch is generated for each received training image. As shown in operation S406, generating a training patch for a given training image may include concatenating incident angle data associated with the given training image and image data associated with the given training image.

[0085] In other words, the machine learning model is trained to classify images by receiving image data associated with a plurality of training images; for each training image, receiving angle of incidence data indicating the angle of incidence at which image data associated with each training image was collected by a detector; and for each training image, generating a training data patch, wherein generating the training data patch includes concatenating the associated angle of incidence data to the associated image data.

[0086] In this way, the machine learning model can learn how the quantitative and / or qualitative properties of the image data are linked and associated with the angle of incidence data. In other words, rather than the machine learning model learning how to process each of these data sets individually, the model is constrained to learn how the two data sets are linked together and alter each other to produce the final overall image.

[0087] In another operation S408, the received angle of incidence data is expanded for each training image to generate angle range data for each training image.

[0088] In another operation S412, for each training image, the associated angle range data is concatenated with the associated image data as part of generating the training data patch.

[0089] In other words, in some embodiments, generating a training data patch further includes expanding the associated incidence angle data to generate angle range data, the angle range data indicating a range of angles that the machine learning model is trained to recognize as similar to the incidence angle indicated by the incidence angle data, and concatenating the angle range data to the associated image data.

[0090] Increasing the incidence angle data to angle range data can make the training of machine learning models more robust. In particular, this expansion can prevent machine learning models from learning false positive conclusions. For example, if the training data identifies images collected at an incidence angle of 21.79° as images of water, future images captured at 21.79° may be mistakenly identified as images of water. Increasing the incidence angle data to generate angle range data significantly reduces the risk of machine learning models learning such false positive conclusions, making the trained machine learning model more robust and reliable.

[0091] For each incident angle value of the incident angle data, increasing the incident angle data may include generating a respective incident angle range distributed about the incident angle value. The incident angle range may be uniformly distributed on either side of the incident angle value. For example, the incident angle range may be 0.05 degrees or more on either side of the incident angle value, 0.1 degrees or more on either side of the incident angle value, 0.25 degrees or more on either side of the incident angle value, 0.5 degrees or more on either side of the incident angle value, or 1 degree or more on either side of the incident angle value.

[0092] The range of incidence angles may not be uniformly distributed for each incidence angle value. In other words, the distribution defining the angles within the range may be non-uniform. For example, the distribution may be symmetrical with respect to the incidence angle value, or may be a normal distribution. In another example, the distribution may be asymmetrical with respect to the incidence angle value. For example, the distribution may be biased toward angles greater than the incidence angle value or angles smaller than the incidence angle value.

[0093] In one example, the range of incidence angles may be distributed according to a normal distribution with a standard deviation of 0.25 degrees or less. In such an example, 95% of the angle values ​​within the range of incidence angles are within 0.5 degrees on either side of the incidence angle value.

[0094] In another operation S404, for each training image, the associated incidence angle data is projected onto a two-dimensional patch of the same size as the associated image data before being concatenated with the associated image data.

[0095] Similarly, another operation S410 is to project, for each training image, the associated angular range data onto a two-dimensional patch of the same size as the associated image data before concatenating with the associated image data.

[0096] In other words, in some embodiments, generating the training data patch may further include projecting the associated incident angle data and / or angle range data onto a two-dimensional patch of the same size as the associated image data before concatenation.

[0097] In this way, the incidence angle data, angle range data, and image data may all be combined into a single training patch that can be passed to, for example, a convolutional neural network, where the convolution operations performed by a CNN can train a machine learning model to recognize relationships and interlinks between the three data sets.

[0098] FIG. 5 is a schematic diagram showing how the incidence angle data is augmented to generate a training data set for training a machine learning model.

[0099] In some embodiments, the local incidence angle model may comprise a function 502 that defines the relationship between the incidence angle and the distance in a first direction within the image.

[0100] Each image may be based on a coordinate system having a primary direction and two mutually orthogonal directions, such as in the case where the detector is mounted on a satellite 210. For example, if the detector is mounted on a satellite 210, a first direction along the image may correspond to the satellite's range direction. In satellite-based imaging, the range direction corresponds to a direction perpendicular to the satellite's motion, i.e., the range direction is "cross-track" across the satellite's orbit. A second direction is orthogonal to the first direction and corresponds to the satellite's azimuth direction. In satellite-based imaging, the azimuth direction is "along-track" parallel to the satellite's orbit.

[0101] By determining the angle of incidence for each pixel of an image based on distance in only one direction within the image, the computational cost of determining the angle of incidence may be significantly reduced.

[0102] In some embodiments, function 502 may be a polynomial determined based on a geometric model of the Earth and a model of the detector trajectory.

[0103] In some embodiments, the polynomial may be a cubic polynomial.

[0104] In other examples, the function may be a linear function, or a quadratic, quartic, quintic, or higher degree polynomial. Making the function a polynomial may reduce the computational cost of determining the polynomial, for example, relative to a trigonometric function, because polynomial functions generally require significantly fewer terms to calculate than trigonometric, hyperbolic, exponential, or other functions that require quasi-infinite Taylor or Maclaurin expansions to determine (e.g., accurately calculating a trigonometric, hyperbolic, or exponential function may require calculating hundreds or thousands of terms).

[0105] In practice, cubic functions may prove to be particularly suitable, as they allow the definition of relatively smooth functions with less "kinks" than would be present in higher-order polynomials and with more control over slope variation than quadratic or linear functions.

[0106] The range of angles may be defined by a distribution function such that the "original" angle of incidence is the most likely angle, and the angles at the ends of the range are the most likely angles. The distribution function may be symmetric, e.g., the distribution may be approximately normal. Alternatively, the distribution function may be asymmetric with respect to the "original" angle of incidence, e.g., the distribution may be approximately beta.

[0107] The expansion function 504 may be projected onto a two-dimensional patch 506 the same size as the image data associated with each training image, as described above in connection with Figure 4. This two-dimensional patch 506 may be concatenated with the associated image data to generate the training data patch.

[0108] A number of training data patches may then be compiled to form a training dataset 508 for training a machine learning model.

[0109] In some embodiments, the machine learning model may be trained based at least in part on an elevation model of the terrain imaged by the detector.

[0110] In this manner, a machine learning model may be trained based on training data that includes information from the elevation model, such that both the machine learning model and the user are able to recognize the overall scene geometry of the terrain imaged by the detector, which may be further based on the classified objects, the incidence angle data, and, in some examples, a determination of the overall orientation of one or more classified objects relative to the Earth and / or the detector surface.

[0111] In some embodiments, the machine learning model may include an artificial neural network.

[0112] In some embodiments, the artificial neural network may be a convolutional neural network.

[0113] FIG. 6 shows a schematic diagram of a convolutional neural network (CNN) 600 that may be used to classify one or more objects in an image as belonging to one of one or more categories based on incident angle data and one or more parameters of the image data.

[0114] Convolutional neural networks (CNNs) may be particularly useful in the context of the methods described herein due to their ability to analyze images by convolving neighboring pixels and removing artifacts that can cause the CNN to mislearn. In particular, CNNs are particularly well-suited for systems where the data to be analyzed, or the data on which the network is trained, includes data patches that include concatenations of various data types due to their ability to convolve multiple data sets into a single factor (or a reduced number of factors) for analysis and / or processing, as in training the machine learning models described herein (as described below).

[0115] In the illustrative example shown in FIG. 6, CNN 600 may receive multiple input data sets 610 in this example. These input data sets are input to a set of nodes 620 according to the data links (represented by straight lines) shown in FIG. 6. The nodes 620 make up a first layer of CNN 600. In some examples, there may be as many nodes 620 as there are inputs 610. The first layer of nodes 620 convolves the received inputs 610 into a smaller number of nodes 630 that form a second layer of CNN 600. In the example shown in FIG. 6, there are six nodes 620 in the first layer of CNN 600 and five nodes 630 in the second layer of CNN 600.

[0116] In some examples, there may be more convolutional layers of CNN 600. For example, in the example shown in Figure 6, the (five) nodes 630 in the second layer of CNN 600 convolve data with a smaller number (four) of nodes 640 that form the third layer of CNN 600. These (four) nodes 640 in the third layer of CNN 600 convolve data with an even smaller number (three) of nodes 650 that form the fourth layer of CNN 600.

[0117] In some examples, after the data has been convolved into the desired number of nodes (in this case, three), the convolved data is fed to further layers of the neural network, which may operate similarly to a feedforward network, for example. In the example shown in FIG. 6, the convolved data is passed from (three) nodes 650 in the fourth layer of the CNN to a set of (four) feedforward nodes 660 that define the fifth layer of the CNN 600. In some examples, the (four) feedforward nodes 660 in the fifth layer of the CNN 600 forward propagate the data to at least one further set of (four) feedforward nodes 670 in the sixth or higher layer of the CNN 600. Finally, the data may be output from the final (sixth) layer of the (four) feedforward nodes 670 to multiple output points 680 (three in this example).

[0118] In one example, a CNN was developed to detect water in SAR images. The example CNN has a 2D convolutional input layer with two input channels fed through 64 filters, resulting in 64 output channels with a kernel size of 3x3. The CNN also has an activation layer that uses a Swish function to enable the network to learn to recognize water in complex images. Other functions, such as the rectified linear unit (ReLU) function, could also be used. In the example, the normalization layer utilizes group normalization using four groups. Other normalization methods, such as batch normalization, layer normalization, and instance normalization, could also be used.

[0119] Using the above CNN, we conducted an experiment to demonstrate how the exact same model, trained on the same images and incidence angles, can better detect features such as water in SAR images when incidence angle data is also provided. The test was performed on a semantic segmentation CNN autoencoder trained to segment regions containing water. The model had approximately 40 million learnable parameters and was trained using focal loss, a form of weighted cross-entropy. The training dataset consisted of 225,000 512x512 pixel patches taken from 543 SAR images acquired by an X-band SAR satellite operated by ICEYE Oy in Espoo, Finland. The validation set consisted of 25,000 similar patches. The validation set was never seen by the model during training.

[0120] Two experiments were performed to see how well the model could mimic the validation set by accurately detecting water in the right locations. In this case, water was chosen as a feature because SAR images can have strong nonlinearities when imaged from low incidence angles. In both experiments, the model was trained to segment water using the same training set and the same random seed. In one experiment, the model was given a validation set of images but no incidence angle information. In the second experiment, the model was provided with both the validation image data and the associated incidence angle information.

[0121] Figure 7 shows the performance of both experiments by plotting validation loss against time. Validation loss is a measure of how well the model can mimic the validation set. Trace 701 shows the validation loss for Experiment 1, where validation images are provided to the model but the incident angle data associated with each image is not. Trace 702 shows the validation loss for Experiment 2, where the model is provided with both the incident angle information associated with the validation images. Experiment 2 showed that providing incident angle information improved the performance of the CNN. Trace 702 shows an improvement of approximately 10–20% at every time step, terminating with a 25% advantage (0.028 vs. 0.035) after training. When Experiment 1 reached the early stopping criterion, training was terminated due to a plateau in training loss.

[0122] Figure 8 shows a SAR image 800 imaged from the west coast of Florida around Baywood Lake. This image contains numerous water features. The area indicated by 801 is ocean near the coastline. In addition to the ocean, the image also contains numerous small lakes, including Baywood Lake 802, Hunters Lake 803, and a series of small lakes 804, shown in the top center of the image. Area 805 represents water further offshore. This image was imaged at a relatively low central incidence angle of approximately 16 degrees. In some instances, such as SAR images obtained by the ICEYE satellite, there can be a strong contrast inversion near this angle. When imaged from an incidence angle of 20 degrees or greater, water typically appears dark unless the water surface is quite rough (such as when disturbed by wind or waves). However, in images imaged from an incidence angle less than 20 degrees, such as image 800 in Figure 8, the water can appear relatively bright, as indicated by the bright areas. This particular image was chosen to determine whether a model could interpret this contrast variation to accurately identify areas covered by water despite their bright colors.

[0123] Figure 9 shows image 900 after processing by a CNN presented with image 800 from Figure 8 without incident angle data and asked to determine which areas are covered by water. Image 900 represents the probability determined by the CNN that a particular pixel is water, with bright white indicating a high probability of classifying the pixel as water and dark areas indicating a low probability that the area is water. In image 900, the CNN does a fairly good job of identifying near-shore area 801 as water. However, Baywood Lake 802, Hunters Lake 803, and Small Lake 804 all appear very dark, meaning the likelihood that these areas are water as assigned by the CNN is still very low.

[0124] Figure 10 shows image 1000, generated by the same CNN, with the probability that a given pixel and region is water, but this time with the benefit of knowing the angle of incidence (~16°) at which image 800 was imaged. Comparing Figures 9 and 10, both models successfully identify ocean 801 near the coast as water, with image 1000 likely showing a slightly higher probability of those regions being water (brighter white) than image 900. Both images 900 and 1000 appear to exhibit similar issues with identifying open ocean areas further offshore (e.g., area 805). This may be due to wind or wave turbulence in water bodies further from the coast. However, image 1000 shows that the CNN, with the benefit of knowing the angle of incidence, is much better at identifying small lakes located inland from the coast. Baywood Lake 802, Hunters Lake 803, and Small Lake 804 all appear very dark in image 900 but very bright in image 1000. This indicates that the CNN assigned a much higher probability that these lakes were water (which, of course, they were). Thus, as an example, we can show that a CNN trained on incidence angle data, and provided with incidence angle data for a particular image, while still not perfect, can do a significantly better job of segmenting regions with specific features than a CNN that does not provide the same incidence angle information.

[0125] For clarity, the above description has described embodiments of the present invention with reference to a single user, although it should be understood that in practice the system may be shared by multiple users, and may even be shared by many users simultaneously.

[0126] The above embodiments may be fully automatic, and in some instances, a user or operator of the system may manually instruct some steps of the method to be performed.

[0127] In embodiments described herein, the system may be implemented as any form of computing and / or electronic device. Such a device may include one or more processors, which may be a microprocessor, a controller, or any other suitable type of processor that processes computer-implementable instructions that control the operation of the device to collect and record routing information. In some examples, for example, when using a system-on-chip architecture, the processor may include one or more fixed function blocks (also called accelerators) that implement portions of the methodology in hardware (rather than software or firmware). Platform software, including an operating system or any other suitable platform software, may be provided in the computing-based device to enable application software to run on the device.

[0128] The various functions described herein may be implemented in hardware, software, or any combination thereof. If implemented in software, the functions may be stored on or transmitted over a computer-readable medium as one or more instructions or code. Computer-readable media may include, for example, computer-readable storage media. Computer-readable storage media may include volatile or nonvolatile, removable or non-removable media implemented in any method or technology for storing information such as computer-readable instructions, data structures, program modules, or other data. Computer-readable storage media may be any available storage medium accessible by a computer. By way of example, and not limitation, such computer-readable storage media may include RAM, ROM, EEPROM, flash memory or other storage devices, CD-ROM or other optical storage disks, magnetic disk storage devices or other magnetic storage devices, or any other medium accessible by a computer that can be used to carry or store desired program code in the form of instructions or data structures. As used herein, optical disk and diskette include optical disks (CDs), laser disks, optical discs, digital versatile disks (DVDs), floppy disks, and Blu-ray disks (BDs). Also, propagated signals are not included within the scope of computer-readable storage media. Computer-readable media also includes communication media, which includes any medium that facilitates transmission of a computer program from one place to another. A connection may be, for example, a communications medium. For example, if software is transmitted from a website, server, or other remote source using coaxial cable, fiber optic cable, twisted pair, DSL, or wireless technologies such as infrared, radio, and microwave, it is included within the definition of communications media. Combinations of the above should also be included within the scope of computer-readable media.

[0129] Alternatively or additionally, the functions described herein may be performed, at least in part, by one or more hardware logic components, such as, but not limited to, field programmable gate arrays (FPGAs), programmable application specific integrated circuits (ASICs), programmable application specific standard products (ASSPs), systems on chips (SOCs), complex programmable logic devices (CPLDs), etc.

[0130] Although illustrated as a single system, it should be understood that the computing device may be a distributed system, whereby, for example, several devices may communicate over network connections and may jointly perform tasks described as being performed by the computing devices.

[0131] Although illustrated as a local device, it should be appreciated that the computing device may be located remotely and accessed via a network or other communications link (e.g., using a communications interface).

[0132] As used herein, the term "computer" refers to any device having processing capability that enables the execution of instructions. Those skilled in the art will recognize that such processing capability may be incorporated into many different devices, and thus the term "computer" includes PCs, servers, mobile phones, personal digital assistants, and many other devices.

[0133] Those skilled in the art will recognize that storage devices storing program instructions may be distributed across a network. For example, a remote computer may store an example of a process written as software. A local or terminal computer may access the remote computer and download some or all of the software to execute the program. Alternatively, a local computer may download pieces of software as needed, or may execute some software instructions at a local terminal and some at a remote computer (or computer network). Those skilled in the art will also recognize that, by utilizing conventional techniques known to those skilled in the art, all or some of the software instructions may be executed by dedicated circuitry, such as a DSP, a programmable logic array, or the like.

[0134] It should be understood that the benefits and advantages described above may relate to one embodiment or to several embodiments. The embodiments are not limited to those that solve any or all of the problems noted or that have the benefits and advantages noted. Variations are to be considered within the scope of the present invention.

[0135] A reference to "an" or "an" item means one or more of those items. The term "comprising" is used herein to mean including identified method steps or elements, but these steps or elements do not comprise an exclusive list and a method or device may include additional steps or elements.

[0136] As used herein, the terms "component" and "system" are intended to include a computer-readable data store comprised of computer-implementable instructions that, when executed by a processor, cause the computer to perform a particular function. The computer-executable instructions may include routines, functions, etc. It should also be understood that a component or system may be located on a single device or distributed across multiple devices.

[0137] Moreover, as used herein, the word "exemplary" is intended to mean "serving as an example or instance of something."

[0138] Furthermore, to the extent the term "comprising" is used in the detailed description or claims, it is intended that the term have the same inclusiveness as the term "comprising," as the term "comprising" is interpreted as a transitional term within the claims.

[0139] Additionally, the operations described herein may include computer-implementable instructions, which may be implemented by one or more processors and / or stored on one or more computer-readable media. Computer-executable instructions may include routines, subroutines, programs, threads of execution, etc. Additionally, the results of the operations of these methods may be stored on a computer-readable medium, displayed on a display device, and / or the like.

[0140] Although the ordering of steps in the methods described herein is exemplary, these steps may be performed in any suitable order, or simultaneously where appropriate. Furthermore, steps may be added to or substituted into any method, or single steps may be deleted from any method, without departing from the scope of the subject matter described herein. Aspects of any of the above implementations may be combined with aspects of any other implementations described to form further implementations without losing the desired effect.

[0141] What has been described above includes examples of one or more embodiments. Of course, for purposes of describing the above aspects, it is not possible to describe every possible modification and variation of the above-described devices or methods, but those skilled in the art will recognize that many further modifications and arrangements of the various aspects are possible. Accordingly, the described aspects are intended to include all such modifications, variations, and variations that fall within the scope of the appended claims.

Claims

1. 1. A computer-implemented method for classifying objects in an image, comprising: receiving image data associated with the image; receiving angle of incidence data, the angle of incidence data indicating an angle of incidence at which image data is collected by a detector; using a machine learning model to classify one or more objects in the image as belonging to one of one or more categories, wherein the classification of the one or more objects in the image comprises: the incident angle data; and the step is performed based on values ​​of one or more parameters of the image data, The machine learning model is receiving image data associated with a plurality of training images; receiving, for each training image, angle of incidence data indicative of the angle of incidence at which image data associated with each said training image was collected by a detector; generating training data patches for each training image, said generating training data patches comprising: augmenting the associated incidence angle data to generate angle range data, the angle range data indicating a range of angles that the machine learning model is trained to recognize as similar to the incidence angle indicated by the incidence angle data; and and concatenating the angular range data with the associated image data. It is trained to classify images, augmenting the associated incidence angle data includes, for each incidence angle value in the incidence angle data, generating a respective range of incidence angles distributed about the incidence angle value by using a function that defines the respective range of angles for each incidence angle.

2. The step of generating training data patches comprises: The computer-implemented method of claim 1 , further comprising projecting the associated incidence angle data and / or the angle range data into a two-dimensional patch having the same size as the associated image data before concatenating.

3. The computer-implemented method of claim 1 , wherein the machine learning model is trained based at least in part on an elevation model of terrain imaged by a detector.

4. 2. The computer-implemented method of claim 1, wherein the incidence angle data is derived from a local incidence angle model that is based on a geometric model of the Earth and one or more state vectors that indicate a detector's position relative to the geometric model of the Earth.

5. The computer-implemented method of claim 4 , wherein the geometric model of the Earth is an ellipsoidal model.

6. The computer-implemented method of claim 4 or 5, wherein the local incidence angle model includes a function that defines a relationship between the incidence angle and a distance in a first direction within the image.

7. The computer-implemented method of claim 6 , wherein the function is a polynomial determined based on a geometric model of the Earth and a model of the detector trajectory.

8. 2. The computer-implemented method of claim 1, wherein classifying the one or more objects in the image comprises classifying each of a plurality of pixels of the image as belonging to one of one or more categories.

9. The computer-implemented method of claim 1 , wherein classifying one or more objects in the image comprises classifying one or more objects in a section of the image.

10. The computer-implemented method of claim 1 , wherein the angle of incidence data is received simultaneously with the image data.

11. the detector is mounted on a satellite in Earth orbit; The computer-implemented method of any one of claims 4 to 5, wherein the one or more state vectors indicative of the detector positions are based on a model of the satellite's orbital path.

12. The one or more parameters of the image data are: an intensity value for each of the plurality of pixels; the color channel values ​​of each of the plurality of pixels, and / or The computer-implemented method of claim 1 , including one or more of the phase information for each of the plurality of pixels.

13. The computer-implemented method of claim 1 , wherein the image is a synthetic aperture radar image.

14. The computer-implemented method of claim 1 , wherein the one or more categories include one or more geographic features.

15. The one or more geographic features include: water, ice, cultivated land, forested areas, and / or The computer-implemented method of claim 14 , including one or more of the following:

16. The computer-implemented method of claim 1 , wherein the machine learning model comprises an artificial neural network.

17. 17. The computer-implemented method of claim 16, wherein the artificial neural network is a convolutional neural network.

18. A computer apparatus comprising a processor for executing the method of claim 1.

19. A computer readable medium comprising logic that, when executed by a computer, causes the computer to perform the method of claim 1.

20. A computer program comprising instructions which, when executed by a computer, cause the computer to carry out the method of claim 1.

Citation Information

Patent Citations

  • Arctic sea ice classification method based on middle-method ocean satellite scatterometer

    CN113095375A