Neural network-based analysis of images captured under different visibility conditions
By adaptively selecting low-resolution ANN architectures for edge devices based on visibility conditions, the computational demands of ANN-based image analysis are reduced, ensuring efficient performance in foggy or smoggy environments.
Patent Information
- Application Number
- JP2025101891
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-06-27
- Filing Date
- 2025-06-18
- Publication Date
- 2026-02-03
AI Technical Summary
ANN-based image analysis solutions require significant computational resources, making them unsuitable for edge devices with limited resources, especially when multiple tasks need to be performed under varying visibility conditions like fog and smog.
Adaptive selection of ANN architectures based on visibility conditions, using low-resolution architectures for reduced visibility scenarios to conserve computational resources, and optionally adjusting image resolution to match the selected architecture.
Reduces computational complexity and resource usage while maintaining accurate image analysis results, particularly in low-visibility conditions, by selecting appropriate ANN architectures and adjusting image resolution.
Smart Images

Figure 2026016306000001_ABST
Abstract
Description
[Technical Field]
[0001] The present disclosure relates generally to the field of image analysis using artificial neural networks (ANNs). In particular, the present disclosure relates to ANN-based analysis of images of scenes captured under different visibility conditions, such as those caused by different levels of fog and / or smog. [Background technology]
[0002] Artificial neural networks (ANNs) have proven useful for a variety of machine learning tasks related to the analysis / processing of images, such as both still and video image content. Examples of tasks suitable for ANN-based analysis include, for example, classification, segmentation, and / or detection of objects in images, depth analysis of images, etc.
[0003] However, ANN-based machine learning solutions are often demanding in terms of the computational power required to implement them, both for training and actual inference. This can prevent such solutions from being implemented, for example, on edge devices, which often have more limited computational resources, especially when multiple solutions for performing different specific tasks need to be implemented on the same device. U.S. Patent No. 11,447,151 B2 discloses a system for detecting objects in a scene under rainy weather conditions. U.S. Patent No. 10,586,132 B2 discloses a system for autonomous vehicle operation that detects and classifies pedestrians, traffic signs, and other vehicles based on environmental conditions around the vehicle.
[0004] The present disclosure seeks to further develop such ANN-based machine learning solutions for image analysis and mitigate their above-mentioned drawbacks. Summary of the Invention
[0005] To the above-mentioned ends, the present disclosure proposes improved methods, devices, computer programs and computer program products for analyzing images of a scene, taking into account that visibility conditions in the scene may change over time, as defined by the attached independent claims. Various embodiments are defined by the attached dependent claims.
[0006] According to a first aspect of the present disclosure, there is provided a (computer-implemented) method for analyzing images of a scene captured under different visibility conditions. The method includes obtaining images of the scene captured by one or more cameras. For each of the images, the method further includes: i) obtaining an indication of the actual or assumed visibility in the scene at the time the image was captured; ii) selecting an ANN architecture from a plurality of ANN architectures based on the indicated visibility, the ANN architectures each being trained for image analysis / processing but configured to accommodate different input image resolutions; and iii) analyzing the image using the selected ANN architecture.
[0007] The proposed solution improves upon current technology in that it allows for the determination of which particular ANN architecture to use based on the visibility conditions of the scene. As a result, in scenes where the visibility conditions are such that greater complexity is not required and / or useful, the use of a more complex high-resolution ANN architecture may be avoided, and computational resources may be freed up and made available for the performance of other tasks by instead choosing to use a less complex low-resolution ANN architecture. In other words, as will be described in more detail later in this specification, the proposed solution takes advantage of the fact that there may be visibility conditions where the use of a high-resolution ANN architecture does not provide any substantial benefit, but instead allows for the use of a low-resolution ANN architecture for the same purpose of reducing computational complexity.
[0008] Optionally, the method includes, prior to operation iii), changing the resolution of the image to match the resolution of the selected ANN architecture. For example, visibility conditions of the scene may result in selecting an ANN architecture that is trained to operate on images having a lower resolution than the resolution of the images from the camera, and therefore the images from the camera may, for example, be downsampled to match such lower resolution.
[0009] Selecting which ANN architecture to use includes selecting a low-resolution ANN architecture for lower visibility and a high-resolution ANN architecture for higher visibility. The low-resolution ANN architecture is configured so that its operation consumes fewer computer processing resources and / or time than the high-resolution ANN architecture. For example, a "low-resolution" ANN architecture as used herein may have fewer input neurons than a "high-resolution" ANN architecture but may still be trained to perform the same type of task. For example, if the task is object detection, reducing the resolution due to low visibility can still produce accurate results, as image resolution is often more important for detecting smaller objects, e.g., objects farther away from the camera. In low or lower visibility conditions, such objects are likely to be obscured anyway, e.g., by fog and / or smog, and therefore would not be detected even using a higher-resolution ANN architecture. Therefore, the proposed solution allows for the use of a low-resolution ANN architecture to produce approximately the same results, but with less computational resource consumption (and / or in a shorter amount of time).
[0010] In some embodiments, analyzing the image may include performing at least one of object detection, object classification, object segmentation, depth analysis, and keypoint detection within the image, all of which operations are expected to be such that during low visibility conditions, the use of a high-resolution ANN architecture provides little or no benefit over the use of a lower-resolution architecture.
[0011] In some embodiments, obtaining an indication of visibility may include detecting a level of fog and / or smog in the scene. Fog and / or smog are often a cause of reduced visibility, particularly in larger cities and other environments where, for example, surveillance cameras are frequently found and used, where smog is already or becoming a problem. Fog and / or smog can reduce the range of the camera, as more distant objects are only partially or completely obscured by the fog and / or smog. Thus, the level of fog and / or smog in the scene correlates well with the range of the camera.
[0012] In some embodiments, detecting the level of fog and / or smog in a scene may include assessing contrast and / or edges in the image. Additional particles in the air due to fog and / or smog may cause additional scattering of light, resulting in reduced contrast and reduced high frequency content in the image, which may be suitably assessed using, for example, edge detection methods.
[0013] In some embodiments, detecting the level of fog and / or smog in a scene may involve using an ANN architecture trained for this purpose.
[0014] In some embodiments, obtaining an indication of visibility may include mapping the detected level of fog and / or smog to object detection performance. This may be performed by approximating how far various levels of fog and / or smog can realistically be seen, as evaluated, for example, in a laboratory environment, and / or by assessing at what level of fog and / or smog object detection begins to detect, for example, a person or other object of interest. Such experiments may be performed over time for different (naturally or artificially) occurring levels of fog and / or smog, and the results may be stored and analyzed to derive the approximate range of the camera as a function of the level of fog and / or smog.
[0015] In some embodiments, obtaining an indication of visibility may include making a prediction based on previously determined fog and / or smog level patterns. For example, historically recorded and / or estimated fog and / or smog levels may be stored and analyzed to make a prediction regarding future fog and / or smog levels, such as based on time series analysis. Other examples may include, for example, detecting that fog and / or smog is more likely to be present during certain times of day, such as during the morning, during the afternoon, etc., and making assumptions regarding current fog and / or smog levels based on the current time of day, etc.
[0016] In some embodiments, obtaining an indication of visibility may include using current and / or historical weather data for the scene. For example, there may be weather forecasts that include estimates of fog and / or smog levels, or even estimates regarding visibility distances, often provided, for example, by METAR data used, for example, by aircraft pilots. Other examples may include mapping how other weather parameters, such as humidity, temperature, wind speed, etc., correlate with fog and / or smog, and using such correlations to predict fog and / or smog levels based on such other parameters.
[0017] In some embodiments, obtaining an indication of visibility may include receiving data from one or more sensors configured for this purpose, such as a fog detector, a smog detector, a visibility detector, etc. Such detectors may, for example, be provided as part of a camera and / or located at one or more other positions and / or at one or more other locations within the scene.
[0018] According to a second aspect of the present disclosure, there is provided a camera, such as a surveillance camera. The camera includes at least one image sensor for capturing an image of a scene (a "scene" may be defined, for example, as everything within the camera's field of view FOV at the time the image is captured, or, for example, as a geographic location and / or area). The camera further includes processing circuitry configured to perform the method of the first aspect.
[0019] In some examples, the processing circuitry may be further configured to perform any embodiment of the method of the first aspect.
[0020] According to a third aspect of the present disclosure, there is provided a computer program comprising computer code which, when executed on processing circuitry of a device such as a camera, causes the device to perform the method of the first aspect.
[0021] In some examples, the computer code may further be such as to cause a device to perform any embodiment of the method of the first aspect.
[0022] According to a fourth aspect of the present disclosure, there is provided a computer program product including a computer-readable storage medium having stored thereon the computer program of the third aspect (or any embodiment thereof). As used herein, a computer-readable storage medium may, for example, be non-transitory and provided as, for example, a hard disk drive (HDD), a solid-state drive (SSD), a USB flash drive, an SD card, a CD / DVD, and / or any other storage medium capable of non-transitory storage of data. In other embodiments, the computer-readable storage medium may be transitory and correspond, for example, to signals (electrical, optical, mechanical, etc.) present on a communications link, wire, or similar means of signal transfer, in which case the computer-readable storage medium is naturally a data carrier rather than a data storage entity.
[0023] Other objects and advantages of the present disclosure will be apparent from the following detailed description, drawings, and claims. Within the scope of the present disclosure, it is contemplated that, for example, all features and advantages described with reference to the method of the first aspect also relate to, apply to, and may be used in combination with the camera of the second aspect, the computer program of the third aspect, and the computer program product of the fourth aspect, and vice versa.
[0024] Exemplary embodiments are described below with reference to the accompanying drawings. [Brief explanation of the drawings]
[0025] [Figure 1] 1A-1D show schematic images of an exemplary scene captured under different visibility conditions. [Figure 2] 1A-1D are functional block diagrams illustrating various examples of devices according to the present disclosure. [Figure 3]1 is a flowchart illustrating various examples of methods for image analysis according to the present disclosure. [Figure 4A] 1A-1C are diagrams illustrating schematic components and functional modules of various example devices according to the present disclosure. [Figure 4B] 1A-1C are diagrams illustrating schematic components and functional modules of various example devices according to the present disclosure. [Figure 5] 1A-1C are diagrams that schematically illustrate examples of computer programs, computer program products, and computer-readable storage media according to the present disclosure. DETAILED DESCRIPTION OF THE INVENTION
[0026] In the drawings, like reference numerals are used for like elements unless otherwise specified. Unless explicitly stated to the contrary, the drawings show only the elements necessary to illustrate the exemplary embodiments, while other elements may be omitted or only suggested for clarity. As shown in the figures, the sizes (absolute or relative) of elements and regions may be exaggerated or understated compared to their true values for illustrative purposes, and are thus provided to illustrate the general structure of the embodiments.
[0027] FIG. 1 shows a collection 100 of images 101, 102, and 103 of the same scene captured under different visibility conditions. In image 101, visibility is good, with no visible fog or smog present in the scene. The scene includes three exemplary objects: a first person 110, a second person 112, and a vehicle 114. Of the three objects, the first person 110 is located closest to the camera, followed by the second person 112 and then the vehicle 114. In the scene captured by first image 101, visibility distance d1 indicates the range of the camera under the current visibility conditions, here extending from the camera toward infinity. As indicated by the dashed portion of the arrow for d1, the range of the camera (due to clear visibility conditions) does not have a clearly defined far end and may be considered sufficient to capture the entire depth of the scene.
[0028] As envisioned herein, image 101 may be analyzed / processed, for example, to perform object detection and provide respective bounding boxes 120, 122, and 124 around each of objects 110, 112, and 114, for example, as shown in the figure.
[0029] The detection range of an object detection algorithm may depend on the resolution of the image, increasing with increasing image resolution and decreasing with decreasing image resolution. As used herein, the "detection range" of an object detection algorithm may be defined as the distance into a scene at which the algorithm successfully detects objects it was trained on. As a result, detecting objects further away from the camera may require a higher resolution input image, while detecting objects closer to the camera may require a lower resolution input image. In image 101, detecting more distant objects, such as a vehicle 114, will likely require the use of an ANN architecture trained to operate on high-resolution images.
[0030] In Image 102, there is moderate fog and / or smog 130 within the scene, which reduces the visibility distance to distance d2 < d1. As a result, only the first person 110 and the second person 112 are clearly visible here, and the object detection algorithm may struggle or fail to detect objects that are farther from the camera than d2, such as vehicle 114.
[0031] In Image 103, there is more severe fog and / or smog 132 within the scene, which reduces the visibility distance to distance d3 < d2. In this example, only the first person 110 is still visible, and the object detection algorithm may struggle or fail to detect vehicle 114 and the second person 112, both of which are (at least partially) hidden within the fog and / or smog 132.
[0032] As will be described in more detail with reference to FIGS. 2 and 3 here, the present disclosure contemplates improving the efficiency of image analysis by taking into account that while some visibility conditions can enable reducing the computational complexity of image analysis, the same or similar results can still be obtained without such reduction. In particular, this may involve using a computationally inexpensive / low computational load ANN architecture that is trained to operate on images with lower resolution under visibility conditions such as those of Images 102 and 103 where the visibility distance d i is smaller than that of Image 101.
[0033] 2 shows a functional block diagram of an example device 200 contemplated herein, and FIG. 3 shows a flowchart of an example method 300 performed by such a device. Device 200 acquires multiple images 212 of a scene captured by one or more cameras 210 (e.g., as part of operation S310 of method 300). The images 212 may be received, for example, directly from the cameras or from some other entity that possesses such images. Device 200 may also form part of one of the cameras 210, in which case the images 212 are received internally, for example, from one or more image sensors of a camera configured to capture such images of the scene. Of course, operation S310 may involve receiving only a single image 212, or at least fewer than all images 212, at a time, and repeating subsequent operations of the method for each such image.
[0034] For each of the images 212, the device obtains (e.g., as part of operation S320 of method 300) an indication of the actual or assumed visibility in the scene at the time the image was captured. The indication may be generated, for example, by the visibility estimation block (or module) 220, which may base such determination on data found in the image 212 itself and / or from input from one or more other entities 226, for example.
[0035] Based on the indicated visibility of the scene when each image was captured, i.e., based on the expected visibility distance d per image, the device selects (e.g., as part of operation S330 of method 300) from among a plurality 240 of ANN architectures 240-1, 240-2, ..., 240-N (N is an integer indicating the total number of such ANN architectures) that are trained for the same image analysis task but that are trained (and configured) to operate on input images of different resolutions. Such selection may be performed, for example, by implementing a demultiplexing function (denoted by block 230) that is controlled based on the output from visibility estimation block 220. For example, the indicated / estimated visibility distance d may be sorted into one of a plurality of visibility categories, and each such category may be associated with one of the plurality of ANN architectures 240. For example, there may be one category for "good visibility," one for "medium visibility," one for "poor visibility," or categories with more or less granularity, as long as there are at least two categories of different visibility (or visibility distances).
[0036] Device 200 provides functionality for changing the resolution of the image (e.g., as part of operation S325 of method 300) as needed to match the resolution at which the selected ANN architecture has been trained to perform the image analysis. As shown in FIG. 2, this may be implemented, for example, as one or more downsampling blocks 250-1, 250-2, . . . , 250-N, although not all such blocks shown may be included.
[0037] As envisioned herein, in one example, ANN architecture 240-1 may be trained / configured to operate at the highest input image resolution, ANN architecture 240-N may be trained / configured to operate at the lowest input image resolution, while any remaining ANN architectures may be trained / configured to operate at one or more intermediate input image resolutions between the highest and lowest input image resolutions. As used herein, "highest" and "lowest" should not be understood at an absolute level, but only as an indication that the resolution is, for example, the highest or lowest of multiple different resolutions at which multiple ANN architectures 240 are trained / configured. For example, if the input resolution of ANN architecture 240-1 matches the resolution of image 212, downsampling block 250-1 may not be necessary.
[0038] As an illustrative example, the resolution of the input image (when expressed as the number of pixels used to capture the scene) may be X pixels wide and Y pixels high, i.e., X×Y, and the ith ANN architecture of ANN architectures 240 has a resolution of α i Width in X pixels and β i Y pixel height, i.e. α i X×β i It may be trained / configured to work for an input resolution of Y, where α i and β i are scaling factors, and the i-th downsampling block 250-i may therefore be configured to provide a corresponding downsampling of the image. Preferably, these factors are such that the resulting resolution in each dimension is an integer. In other examples, if the factors are not, an integer may be obtained, for example, by rounding a non-integer resolution to, for example, the nearest integer value. In some examples, it may be that α1 > α2 > . . . > α N and β1>β2>···>β N For example, for each i, α i =β iOf course, other examples are possible, for example, at least α1X × β1Y > α2X × β2Y > > α N X×β N Y, etc.
[0039] As generally used herein, the "resolution" of an image refers to the number of pixels present in the image, with a higher resolution image using more pixels to represent an object with a higher level of detail, and a lower resolution image using fewer pixels to represent the same object but with a lower level of detail.
[0040] As generally used herein, a "high-resolution ANN architecture" may, for example, include a larger number of input neurons or be configured in some way that is more difficult to implement (computationally). Similarly, a "low-resolution ANN architecture" may include a smaller number of input neurons or be configured in some way that is easier to implement (computationally). For example, an ANN architecture configured for object detection may have a stack of convolutional layers that take in an image of a particular resolution as input and then successively reduce the resolution to gather more and more semantic context. A high-resolution ANN architecture may, for example, use a larger first convolutional layer in such a stack and / or include a larger number of layers in the stack. Similarly, a low-resolution ANN architecture may, for example, use a smaller first convolutional layer in the stack and / or include a smaller number of layers in the stack. As a result, because there are fewer neurons, fewer connections, and therefore fewer weights to evaluate when implementing a low-resolution ANN architecture, a low-resolution ANN architecture is likely to be easier to implement and use in terms of computational power.
[0041] For example, in the case of a convolutional neural network (CNN) used for object detection, the computational complexity of the network depends on the number of filters, the filter dimensions, and the input dimensions. A convolution operation may have a complexity of O(XYPQRS), where X and Y are the input dimensions mentioned above, P and Q are the filter dimensions, and R and S are the filter strides. Therefore, it can be seen that reducing the input dimensions also reduces the computational complexity of the network, thus freeing up computational resources that can be used for other tasks.
[0042] Device 200 is then further configured to analyze the images using the selected ANN architecture (e.g., as part of operation S340 of method 300), such that different ANN architectures are used for images capturing the scene under different visibility conditions, or at least while the visibility conditions are classified into different visibility categories.
[0043] If the image analysis is or includes, for example, object detection, device 200 may be configured so that the highest-resolution ANN architecture 240-1 is selected for high-visibility conditions, such as those classified as a “good visibility” category, and the lowest-resolution ANN architecture 240-N is selected for low-visibility conditions, such as those classified as a “poor visibility” category. As previously described herein, this may have the advantage that the higher resolution often required to detect more distant objects in a scene is no longer needed because such objects are likely obscured due to poor visibility anyway, and ANN architecture 240-N may be implemented more cheaply in terms of computational resources while still being able to provide the same or similar results. The computational resources thus freed may instead be used for other tasks, thus improving the overall efficiency of the device. Other exemplary image analysis tasks where increased levels of smog and / or fog help to obscure objects that require high-resolution images to detect and therefore may benefit from the proposed solution include, for example, object classification, object segmentation, depth analysis, keypoint detection, and the like.
[0044] Obtaining an indication of the visibility of the scene at the time each image was captured may be performed in many ways. For example, information may be provided (e.g., to visibility estimation block 220) from one or more sensors (indicated by dashed block 226) configured to detect, for example, smog, fog, overcast, rain, snow, blizzard, mist, smoke, air pollution, and / or additional particles of typical size and / or concentration in the air. Such sensors may be provided in the scene as part of a camera, as part of device 200, etc. Based on input from such sensors 226, block 220 can draw a conclusion regarding the actual visibility in the scene and control block 230 so that the correct ANN architecture is used.
[0045] In other examples, one or more ANN architectures specifically trained to identify, for example, fog and / or smog in images and / or to determine the level of fog and / or smog may be used to provide an indication of scene visibility. To reduce the amount of computation required to run such an architecture, it may be sufficient to estimate visibility only occasionally, rather than for each new image, such that, for example, previously indicated visibility is assumed to be valid for one or more subsequently captured images. For example, device 200 may be configured to run such an architecture, for example, every minute, every hour, every day, or more frequently, or less frequently, etc.
[0046] In some examples, device 200 may be configured to obtain an indication of visibility by analyzing data contained in one or more images 212 themselves, such as by looking at contrast and / or edges. For example, when visibility is low due to the increased presence of additional particles in the air, such particles may cause additional scattering of light, resulting in an image appearing blurrier than during clear visibility conditions. This may cause a decrease in image contrast and a reduction in high frequency content in the image, which may be appropriately assessed using, for example, edge detection methods.
[0047] In some examples, to determine how visibility (distance) of a scene depends on, for example, the level of fog and / or smog, device 200 may include a mapping between fog and / or smog levels and visibility distance. Such a mapping may be obtained, for example, by performing controlled experiments in a laboratory environment, where, for example, the range of a camera or human eye may be studied for different levels of fog and / or smog, and conclusions drawn regarding the relationship between visibility distance and fog and / or smog levels.
[0048] In other examples, the results of object detection may be analyzed to ascertain the distance at which the algorithm used begins to detect the object and then correlate this distance with the estimated level of fog and / or smog in the scene. For example, for an object in the scene that is known to be detectable and remain stationary in the image, it may be determined when the object detection algorithm begins (or stops) being able to detect the object, and the prevailing level of fog and / or smog may be noted and associated with that object and, for example, the known distance between the object and the camera. If there are multiple such objects at different distances from the camera, the process may be repeated to generate a mapping.
[0049] The mapping itself may involve the use of more or less sophisticated processes, such as the use of any suitable model for estimating how one or more dependent variables depend on one or more predictor / independent variables, such as, for example, a linear regression model, a Gaussian process regression based on a covariance kernel, etc., or the training of a neural network or other machine learning model for such purpose. For example, in the case of a model that is also able to output one or more confidence intervals, device 200 may be configured to select a high-resolution ANN architecture even for indicated lower visibility distances when the uncertainty of such an indication is indicated as high (e.g., above a threshold), e.g., in order to avoid missing detecting more distant objects that are not obscured by fog and / or smog, or at least not with sufficient certainty.
[0050] In some examples, the indication of visibility may be obtained by making a prediction based on previously determined patterns of fog and / or smog levels. For example, if it is established that fog and / or smog is likely to be present during a particular time of day, a particular day of the week, a particular week of the year, etc., then assumptions can be made regarding visibility in a scene simply based on the particular time each image was captured, etc.
[0051] Thus, as an example, device 200 may be configured such that, for image 101, ANN architecture 240-1 is used to analyze the image to also detect the more distant vehicle 114. For image 102, higher resolution is not needed because the vehicle 114 is likely obscured by fog and / or smog 130 anyway, so device 200 may instead select lower resolution ANN architecture 240-2. For image 103, because only the first person 110 is visible and, for example, the second person 112 and vehicle 114 are likely obscured by more severe fog and / or smog 132, device 200 may instead select an even lower resolution ANN architecture, which may be more efficient to implement and still provide the same amount of results.
[0052] In other examples, the visibility of the scene at the time each image was captured may instead or additionally be derived from meteorological data, such as data indicating one or more of temperature, humidity, rainfall, and snowfall, e.g., as part of one or more weather forecasts associated with the area in which the scene occurs. For example, if such meteorological data indicates that it was raining or snowing heavily and / or that there was fog and / or smog when the image was captured, visibility may be indicated as low, e.g., with a corresponding low visibility distance, and device 200 may, for example, select a low-resolution ANN architecture for analyzing the image. Similarly, if the data instead indicates that visibility is good, device 200 may instead select a high-resolution ANN architecture to detect objects that are further away from the camera and are less likely to be obscured, e.g., by fog and / or smog. Meteorological data may also, or instead, include predictive data for one or more future times beyond the capture time of the image, and such data may be used to estimate what visibility will be like even when the image was captured. The weather data may in some instances include, for example, METAR data (or the like) used by aircraft pilots, which may include, for example, indicated visibility distances or at least indicated visibility categories from which the visibility distance of a scene may be derived.
[0053] As envisaged herein, low visibility may also be caused by factors other than fog and / or smog, such as rain, snow, blizzards, dust storms, hailstorms, tornadoes, flying debris, and / or by a lack of sunlight or other light, such as, for example, during evening or night hours, or any other conditions that make it more difficult to detect objects at greater distances, and therefore it may be appropriate to reduce the computational complexity of the ANN architecture used so as not to waste computational effort where there is little or no benefit due to low visibility conditions.
[0054] As generally envisioned herein, a mapping may be created, for example, between an estimated level of fog and / or smog, e.g., level L, and a visibility distance (e.g., camera range) D, e.g., a mapping f:L→D.
[0055] In other examples, device 200 may be configured to analyze only a portion of the image, e.g., a portion of the image that is likely to contain closer objects, e.g., by discarding portions of the image that are likely to contain more distant objects, and then provide only the remaining portion of the image as input to the selected ANN architecture.
[0056] 4A schematically illustrates a further example of a device 400 for performing the methods contemplated herein, i.e., a device (such as a camera) configured to perform the method 300 described with reference to FIG. 3. The device 400 includes at least a processor (or “processing circuitry”) 410 and, optionally, a memory 412. A “processor” or “processing circuitry” as used herein may be, for example, any combination of one or more suitable central processing units (CPUs), multiprocessors, microcontrollers (μCs), digital signal processors (DSPs), application specific integrated circuits (ASICs), field programmable gate arrays (FPGAs), graphics processing units (GPUs), etc., capable of executing software instructions stored in the memory 412. The memory 412 may be external to the processor 410 or internal to the processor 410. A “memory” as used herein may be any combination of random access memory (RAM) and read-only memory (ROM), or any other type of memory capable of storing instructions. Memory 412 includes (i.e., stores) instructions that, when executed by processor 410, cause device 400 to perform a method described herein (i.e., method 300 or any embodiment thereof). Device 400 may further include one or more additional items 414 that may, in some circumstances, be useful in performing a method. In some exemplary embodiments, device 400 may be, for example, a (video) camera, such as a (video) surveillance camera, and additional items 414 may then include, for example, an image sensor and one or more lenses for focusing light from a scene onto the image sensor, for example, so that the surveillance camera can capture images of the scene as part of performing a contemplated method. Additional items 414 may also include various other electronic components necessary, for example, to properly operate the image sensor and / or lenses as desired, for example, to capture a scene.Performing the method within the surveillance camera can be useful in that the processing moves to the "edge," i.e., closer to where the actual scene was captured, compared to performing image analysis elsewhere (such as on a more centralized processing server). Further items 414 may also include one or more sensors for detecting / estimating visibility, e.g., fog and / or smog levels, e.g., any of sensors 226 shown in FIG. 2.
[0057] The device 400 may be connected to a network, for example, so that results from executing methods can be transmitted to a user. To this end, the device 400 may include a network interface 416, which may be, for example, a wireless network interface (e.g., supporting Wi-Fi, e.g., as defined by IEEE 802.11 or any successor standard) or a wired network interface (e.g., supporting Ethernet, e.g., as defined by IEEE 802.3 or any successor standard). The network interface 416 may also support any other wireless standard capable of transferring encoded video, such as, for example, Bluetooth. The various components 410, 412, 414, and 416 (if present) may be connected via one or more communication buses 420 so that these components can communicate with each other and exchange data as necessary.
[0058] Device 400 may be a surveillance camera mounted or mountable on a building, for example, in the form of a PTZ camera, or a fisheye camera capable of providing a wider perspective of a scene, or any other type of surveillance / reconnaissance camera. Device 400 may also be, for example, a body camera, action camera, dash cam, etc. suitable for mounting on people, animals, and / or various vehicles, etc. Device 400 may also be, for example, a smartphone or tablet that a user can carry and that can capture a scene. In any such example of device 400, it is contemplated that device 400 may include all necessary components (if any) other than those already described herein, so long as device 400 is still capable of performing method 300 or any embodiment thereof contemplated herein. Various components of device 400 may, in some examples, implement various ANN architectures / entities described herein, such as a plurality of 240, and may be further configured to implement various functional blocks (e.g., 220, 230) to select which ANN architecture to use for processing an image based on an estimated visibility distance in the image and process the image using the selected ANN architecture.
[0059] 4B illustrates one or more embodiments of device 400 with respect to several functional / computing blocks 410a-410d. Each such block 410a-410d is responsible for performing a function according to a particular operation of method 300, as illustrated in the flowchart of FIG. 3. For example, one such functional block 410a may be configured to acquire input images from at least one camera (operation S310), another block 410b may be configured to acquire an indication of actual or assumed visibility in the scene (operation S320), another block 410c may be configured to select which one of a plurality of ANN architectures to use based on the visibility (distance) of the scene (operation S330), and another block 410d may be configured to analyze / process each image using the ANN architecture selected for that image (operation S340). The device 400 may optionally include one or more further functional blocks 410e, such as, for example, a block for performing downscaling of the image to match the resolution / dimensions of the input of the selected ANN architecture (operation S325).
[0060] Generally speaking, each of the functional modules 410a-e may be implemented in hardware or software. Preferably, one or more or all of the functional modules 410a-e may be implemented by the processing circuitry 410, possibly in cooperation with a storage medium / memory 412 and / or a communication interface 416. Accordingly, the processing circuitry 410 may be configured to fetch instructions provided by the functional modules 410a-e from the memory 412 and execute these instructions, thereby performing any operations of the method 300 performed by / in the device 400 disclosed herein.
[0061] 5 schematically illustrates a computer program product 510 including a computer-readable means / storage medium 530. The computer storage medium 530 may store a computer program 520 that may cause the processor 410 of the device 400 and entities and devices operatively coupled thereto, such as the communication interface 416 and the memory 412, to perform the method 300 according to the embodiments described herein, for example, with reference to FIGS. 1, 2, and 3. The computer program 520 and / or the computer program product 510 may thus provide means for performing any of the operations of the method 300 performed by the device 400 disclosed herein.
[0062] 5, computer program product 510 and computer-readable storage medium 530 are shown as optical discs, such as CDs (compact discs) or DVDs (digital versatile discs) or Blu-Ray discs. Computer program product 510 and computer-readable storage medium 530 may be embodied as memory, such as random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM), or electrically erasable programmable read-only memory (EEPROM), and more particularly as non-volatile storage media in devices, such as external memory, such as USB (universal serial bus) memory, or flash memory, such as compact flash memory. Thus, while computer program 520 is shown here schematically as tracks on the depicted optical disc, computer program 520 may be stored in any manner suitable for computer program product 510 and computer-readable storage medium 530.
[0063] Summarizing all of the above, the present disclosure improves upon current technology by providing a solution that takes into account visibility distance (i.e., camera range) in a scene and adaptively selects a low-resolution ANN architecture for image analysis / processing when visibility is reduced, e.g., due to fog and / or smog, thereby freeing up computational resources for other tasks. This may be particularly useful, for example, in edge devices where processing resources are more limited than, e.g., servers. Additional advantages of the proposed solution include the fact that high-resolution ANN architectures may be more prone to generating false positives (e.g., as part of object detection) in, e.g., foggy and / or smoggy images compared to low-resolution ANN architectures. This may be particularly true, for example, for high-resolution ANN architectures that are not specifically trained to detect objects in foggy and / or smoggy conditions; therefore, additional image details / data provided in high-resolution images of a scene may confuse such networks. Thus, using a low-resolution ANN architecture to operate on a low-resolution version of an image, in addition to being more computationally efficient, may also, for example, reduce the number of such false positives.
[0064] Although the features and elements may be described above in particular combinations, each feature or element may be used alone without the other features and elements, or in various combinations with or without the other features and elements. Additionally, variations to the disclosed embodiments may be understood and effected by those skilled in the art in practicing the claimed invention, from a study of the drawings, the disclosure, and the appended claims.
[0065] In the claims, the words "comprise" and "include" do not exclude other elements, and the indefinite articles "a" or "an" do not exclude a plurality. The mere fact that certain features are recited in mutually different dependent claims does not indicate that a combination of these features cannot be used to advantage. [Explanation of symbols]
[0066] A collection of 100 images captured under different visibility conditions 101 Images captured with good / high visibility 102 Images captured in medium visibility 103 Images captured with poor / low visibility 110, 112, 114 object 120, 122, 124 Object detection / bounding boxes 130, 132 Light fog and / or smog and heavy fog and / or smog 200 devices 210 Camera 212 images from the camera 220 Visibility Estimation Block 226 One or more sensors / further items 230 Demultiplexer / ANN Architecture Selector 240 Multiple Different ANN Architectures 250 Image Downscaling Block 300 ways S310-S340 Method Operation 400 devices 410 Processing Circuit 410a-e Function Blocks / Modules 412 memory 414 More Items 416 Communication Interface 420 communication bus 510 Computer Program Products 520 Computer Programs 530 Computer-readable storage medium d Visibility distance / camera range
Claims
1. 1. A computer-implemented method (300) for artificial neural network (ANN)-based analysis of images of a scene captured under different visibility conditions, comprising: Obtaining (S310) images (212) of a scene captured by one or more cameras (210); For each of said images, i) obtaining an indication of the actual or assumed visibility in the scene at the time the image was captured; ii) selecting an ANN architecture from a plurality of ANN architectures based on the indicated visibility, the ANN architectures each being trained for image analysis but configured to accommodate different input image resolutions; iii) analyzing the image using the selected ANN architecture, including, if necessary, changing the resolution of the image to match the resolution of the selected ANN architecture; Including, The method (300), wherein selecting the ANN architecture includes selecting a low-resolution ANN architecture for lower visibility and selecting a high-resolution ANN architecture for higher visibility, the low-resolution ANN architecture being configured such that its operation consumes less computer processing resources and / or time than the high-resolution ANN architecture.
2. The method of claim 1 , wherein the analyzing comprises performing at least one of object detection, object classification, object segmentation, depth analysis, and keypoint detection within the image.
3. The method of claim 1 or 2, wherein obtaining the indication of the visibility comprises detecting a level of fog and / or smog within the scene.
4. The method of claim 3 , wherein detecting a level of fog and / or smog in the scene comprises assessing contrast and / or edges in the image.
5. The method of claim 3 or 4, wherein obtaining the indication of the visibility comprises using a mapping of detected fog and / or smog levels to visibility distances.
6. The method of any one of claims 3 to 5, wherein obtaining the indication of the visibility comprises making a prediction based on previously determined fog and / or smog level patterns.
7. The method of any preceding claim, wherein obtaining the indication of the visibility comprises use of current and / or historical weather data associated with the scene.
8. at least one image sensor for capturing images of the scene; For each of a plurality of images captured by the at least one image sensor, i) obtaining an indication of the actual or assumed visibility in the scene at the time the image was captured; ii) selecting an artificial neural network (ANN) architecture from a plurality of ANN architectures (240) based on the indicated visibility, the ANN architectures each being trained for image analysis but configured to accommodate different input image resolutions; iii) analyzing the image using the selected ANN architecture, including, if necessary, changing the resolution of the image to match the resolution of the selected ANN architecture; a processing circuit (410) configured to perform A camera (210, 400) comprising: The camera (210, 400), wherein selecting the ANN architecture includes selecting a low-resolution ANN architecture for lower visibility and selecting a high-resolution ANN architecture for higher visibility, the low-resolution ANN architecture being configured such that its operation consumes less computer processing resources and / or time than the high-resolution ANN architecture.
9. The camera of claim 8, wherein the processing circuitry is further configured to perform the method of any one of claims 2 to 7.
10. When executed on a processing circuit (410) of a device (200, 400), such as a camera, the device acquiring an image (212) of a scene; For each of said images, i) obtaining an indication of the actual or assumed visibility in the scene at the time the image was captured; ii) selecting an artificial neural network (ANN) architecture from a plurality of ANN architectures (240) based on the indicated visibility, the ANN architectures each being trained for image analysis but configured to accommodate different input image resolutions; iii) analyzing the image using the selected ANN architecture, including, if necessary, changing the resolution of the image to match the resolution of the selected ANN architecture; A computer program (520) comprising computer code for causing A computer program (520) in which selecting the ANN architecture includes selecting a low-resolution ANN architecture for lower visibility and selecting a high-resolution ANN architecture for higher visibility, the low-resolution ANN architecture being configured such that its operation consumes less computer processing resources and / or time than the high-resolution ANN architecture.
11. The computer program of claim 10, wherein the computer code is further such as to cause the device to perform the method (300) of any one of claims 2 to 9.
12. A computer program product (510) comprising a non-transitory computer-readable storage medium (530) having stored thereon a computer program (520) according to claim 10 or 11.