System and method for determining image capture degradation of a camera sensor
Through high-frequency multi-scale fusion transformation technology, the problem of image degradation is solved, and image quality improvement and user experience improvement is achieved.
Patent Information
- Application Number
- CN202211362585.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2022-04-25
- Filing Date
- 2022-11-02
- Publication Date
- 2025-05-23
- Estimated Expiration
- 2042-11-02
AI Technical Summary
The prior art is difficult to effectively identify and cope with degradation in camera image capture, resulting in a decline in image quality and affecting subsequent algorithms and vehicle occupants' use.
By using high-frequency multi-scale fusion transformation technology, a frequency layer is generated, image capture degradation is identified, and occlusion is removed by cleaning the system or degradation areas are ignored, generating a notification to prompt the user.
Effectively identify and respond to camera image degradation, improve image quality, ensure images are available for vehicle systems, and provide cleaning and notifications to improve user experience.
Smart Images

Figure CN116962662B_ABST
Abstract
Description
[0001] introduction
[0002] The present disclosure relates to systems and methods for determining image capture degradation of a camera, and more particularly to systems and methods for determining image capture degradation of a camera using a high frequency multi-scale fusion transform. Summary of the invention
[0003] In some embodiments, the present disclosure relates to a method for determining image capture degradation of a camera sensor. The method includes: capturing a series of image frames over time by a camera of a vehicle via one or more sensors. The method includes: generating a latent image from a series of image frames captured over time by a camera of a vehicle using a processing circuit. The latent image represents the temporal and / or spatial differences over time among the series of image frames. In one embodiment, the latent image is generated by determining the pixel dynamic range of the series of images. In another embodiment, the latent image is generated by determining the gradient dynamic range of the series of images. In another embodiment, the latent image is generated by determining the temporal difference of each pixel of the series of images. In another embodiment, the latent image is generated by determining the average gradient of the series of images. In some embodiments, the image gradient is determined by applying a Sobel filter or a bilateral filter. The method includes: using a processing circuit and generating multiple frequency layers based on the latent image. Each frequency layer in the frequency layer corresponds to a frequency-based decomposition of the latent image at a corresponding scale and frequency. In some embodiments, the method uses a high frequency fusion transform to generate the frequency layer. In some embodiments, the method performs a high frequency fusion transform at a single scale. In other embodiments, the method performs a high frequency fusion transform at multiple scales.The method includes: using a processing circuit, identifying image capture degradation based on multiple frequency layers.
[0004] In some embodiments, generating a plurality of frequency layers comprises determining a plurality of frequency-based coefficients for each of a plurality of scales centered at each of a plurality of locations in a later image. In some embodiments, the frequency-based coefficients correspond to a plurality of spectral sub-band frequencies. In some embodiments, each frequency layer in the plurality of frequency layers contains a frequency-based coefficient for a corresponding one of the spectral sub-band frequencies. In some embodiments, identifying image capture degradation comprises selecting a subset of the coefficients based on a frequency magnitude threshold.
[0005] In some embodiments, the frequency layers are determined by performing a high frequency multi-scale fusion transform on the latent image.
[0006] In some embodiments, generating the plurality of frequency layers further comprises: selecting a subset of coefficients based on their frequency. The method comprises: sorting the subset of frequency-based coefficients with respect to magnitude. The method comprises: normalizing the sorted subset of frequency-based coefficients to generate the plurality of layers.
[0007] In some embodiments, the camera captures the series of image frames at a sample frequency and the sample frequency is determined based on the vehicle speed. In some embodiments, when the vehicle speed is below a predetermined threshold, the image frame is excluded from the series of image frames.
[0008] In some embodiments, the method includes adjusting a frequency magnitude threshold.
[0009] In some embodiments, the method includes determining whether an occlusion exists based on the identified image capture degradation. The method includes applying a fluid to a face of the camera using a vehicle washing system in response to determining that an occlusion exists.
[0010] In some embodiments, the method includes generating a notification on the display device indicating image capture degradation.
[0011] In some embodiments, the method includes: ignoring one or more regions of one or more image frames based on image degradation.
[0012] In some embodiments, the present disclosure relates to a system for determining image capture degradation. The system includes a camera system and a control circuit. The camera is configured to capture an image sequence. The control circuit is coupled to the camera and is configured to:
[0013] A latent image is generated from a series of image frames captured by a camera. The latent image represents the temporal and / or spatial differences among the series of image frames over time. In one embodiment, the latent image is generated by determining the pixel dynamic range of the series of images. In another embodiment, the latent image is generated by determining the gradient dynamic range of the series of images. In another embodiment, the latent image is generated by determining the time difference of each pixel of the series of images. In another embodiment, the latent image is generated by determining the average gradient of the series of images. In some embodiments, the image gradient is determined by applying a Sobel filter or a bilateral filter. The control circuit generates multiple frequency layers based on the latent image. Each frequency layer in the frequency layer corresponds to a frequency-based decomposition of the latent image at a corresponding scale and frequency. In some embodiments, the control circuit uses a high-frequency fusion transform to generate the frequency layer. In some embodiments, the control circuit performs a high-frequency fusion transform at a single scale. In other embodiments, the control circuit performs a high-frequency fusion transform at multiple scales. The control circuit uses a processing circuit to identify image capture degradation based on multiple frequency layers.
[0014] In some embodiments, the camera is integrated into the vehicle and the camera captures the series of image frames at a sample frequency based on the speed of the vehicle.
[0015] In some embodiments, image frames are excluded from the latent image when the capture is performed when the speed of the vehicle is below a predetermined threshold.
[0016] In some embodiments, the control circuit ignores the camera output.
[0017] In some embodiments, the system includes a cleaning system that applies a fluid to the face of the camera.
[0018] In some embodiments, the system includes a display device configured to display a notification indicating an occlusion event.
[0019] In some embodiments, the present disclosure relates to a non-transitory computer readable medium. The non-transitory computer readable medium includes program instructions for image capture degradation. In some embodiments, the program instructions cause a computer processing system to perform steps including the following:
[0020] A series of image frames are captured by a camera. The step also includes: using a processing circuit to generate a latent image from a series of image frames captured by a camera of the vehicle over time. The latent image represents the temporal and / or spatial differences among the series of image frames over time. In one embodiment, the latent image is generated by determining the pixel dynamic range of the series of images. In another embodiment, the latent image is generated by determining the gradient dynamic range of the series of images. In another embodiment, the latent image is generated by determining the temporal difference of each pixel of the series of images. In another embodiment, the latent image is generated by determining the average gradient of the series of images. In some embodiments, the image gradient is determined by applying a Sobel filter or a bilateral filter. The step further includes: using a processing circuit and generating multiple frequency layers based on the latent image. Each frequency layer in the frequency layer corresponds to a frequency-based decomposition of the latent image at a corresponding scale and frequency. In some embodiments, the step further includes: generating the frequency layer using a high frequency fusion transform. In some embodiments, the step includes: performing a high frequency fusion transform at a single scale. In other embodiments, the step includes performing a high frequency fusion transform at multiple scales. The step includes: using a processing circuit to identify image capture degradation based on multiple frequency layers. BRIEF DESCRIPTION OF THE DRAWINGS
[0021] The present disclosure according to one or more various embodiments is described in detail with reference to the following drawings. The drawings are provided for illustrative purposes only and only show typical or exemplary embodiments. These drawings are provided to facilitate understanding of the concepts disclosed herein, and these drawings should not be considered as limitations on the breadth, scope or applicability of these concepts. It should be noted that for clarity and ease of description, these drawings are not necessarily drawn to scale.
[0022] Figure 1 depicts a top view of an illustrative vehicle having several cameras according to some embodiments of the present disclosure;
[0023] Figure 2 depicts a diagram of exemplary output from a camera according to some embodiments of the present disclosure;
[0024] Figure 3 depicts a system diagram of an illustrative system for determining image capture degradation of a camera sensor according to some embodiments of the present disclosure;
[0025] Figure 4 depicts a flow chart of an illustrative process for generating a latent image according to some embodiments of the present disclosure;
[0026] FIG. 5A to FIG. 5B depicts an illustrative process for applying a high frequency multi-scale transform according to some embodiments of the present disclosure;
[0027] FIG. 6A to FIG. 6B depicts exemplary regions according to some embodiments of the present disclosure;
[0028] Figure 7 depicts a flow chart of an illustrative process for determining image capture degradation of a camera sensor according to some embodiments of the present disclosure; and
[0029] Figure 8 Depicted is a flow chart of an illustrative process for managing image capture degradation and response of a camera sensor, according to some embodiments of the present disclosure. DETAILED DESCRIPTION
[0030] Image degradation may occur for a variety of reasons, such as, for example, dust accumulation above the camera lens, bird droppings, objects placed on or near the camera, and environmental factors, such as the general direction of a camera pointing toward a strong light source. Additionally, image degradation may be caused by camera blur, fogging, or other obstructions that may cause degradation of the images captured by the camera. Such image degradation reduces the quality of the images and may render them unusable for use by other algorithms or by vehicle occupants. The systems and methods of the present disclosure are directed to determining which portions of an image frame are degraded and responding to image degradation.
[0031] Figure 1 A top view of an exemplary vehicle 100 with several cameras according to some embodiments of the present disclosure is shown. As shown, vehicle 100 includes cameras 101, 102, 103, and 104, but it should be understood that according to the present disclosure, the vehicle may include any suitable number of cameras (e.g., one camera, more than one camera). In addition, although the present disclosure may show, discuss, or describe cameras, any image capture device may be implemented without departing from the contemplated embodiments. For example, any device that generates an actinic, digital, or analog representation of the environment (including those captured by a camera, video camera, infrared camera, radar device, or lidar device) may be used, and may be implemented according to the techniques described herein without departing from the contemplated embodiments.
[0032] Panel 150 shows a cross-sectional view of a camera exhibiting an occluder. In the depicted exemplary embodiment, the occluder covers portion 152 of the camera, while portion 151 is not covered (e.g., but portion 151 may be affected by, for example, the occluder). The occluder may completely cover portion 152, and may effectively cover at least some portion of portion 151 (e.g., due to uneven distribution of reflected light from the occluder). The occluder may accumulate on the camera and may continue to exist for some time (e.g., fall off, dissipate, or remain for an extended period of time). In some embodiments, the systems and methods of the present disclosure involve determining which portions of an image are degraded (e.g., caused by an occluder), and responding to the degradation by clearing the occluder, ignoring the image exhibiting degradation, modifying image processing for output from the camera, generating a notification of degradation and / or occluder, any other suitable function, or any combination thereof. Although the present disclosure discusses embodiments in which an occluder obscures a portion of a camera and may therefore cause image degradation, contemplated embodiments include those in which the entire camera's field of view is obscured by the occluder or the image is completely corrupted.
[0033] Figure 2 A diagram depicts an exemplary output 200 from a camera according to some embodiments of the present disclosure. As shown, the output 200 includes a plurality of captured images 201-205 indexed by time (e.g., the images are sequential). Although images are shown and described, any photochemical, digital, or analog representation may be used, including those captured from a video camera, a video camera, an infrared camera, a radar device, or a lidar device, and may be implemented without departing from the contemplated embodiments.
[0034] A partition grid is applied to images 201 to 205 to define regions, and a point 210 in the partition grid is shown. In some embodiments, point 210 corresponds to a single pixel of image 201. Region 211 corresponds to a location of the partition grid. The partition grid includes N×M points, and region 211 may correspond to a specific number of pixels corresponding to each point (e.g., 7×7 pixels, 9×9 pixels, or any other A×B pixel set). For example, images 201 to 205 may each include (N*A)×(M*B) pixels, which are grouped into N×M regions, each region including A×B pixels. In some embodiments, the regions do not overlap. For example, each pixel may be associated with a single region (e.g., together with other pixels). In other embodiments, the regions may at least partially overlap. For example, at least some pixels may be associated with more than one region (e.g., adjacently indexed regions). In other embodiments, the regions do not overlap and are spaced apart. For example, at least some pixels do not need to be associated with any region (e.g., adjacently indexed regions). According to the present disclosure, any suitable regions may be used, overlapping or non-overlapping, or spaced or non-spaced, or combinations thereof. In addition, regions of different sizes (eg, different dimensions) may be implemented without departing from the contemplated embodiments.
[0035] In some embodiments, the output of one camera or more than one camera may be analyzed to determine whether any particular image or region of an image is degraded.The partition grid need not be rectangular and may include gaps, blanks, irregularly arranged dots, arrays, or combinations thereof.
[0036] Figure 3 A system diagram of an exemplary system 300 for determining image capture degradation of a camera sensor according to some embodiments of the present disclosure is depicted. As shown, the system 300 includes a transformation engine 310, a degradation map engine 320, a smoothing engine 330, a response engine 340, reference information 350, preference information 360, and a memory storage device 370. Those skilled in the art will readily appreciate that the illustrated arrangement of the system 300 may be modified in accordance with the present disclosure. For example, components may be functionally combined, separated, added, reduced, modified in functionality, omitted, or otherwise modified in accordance with the present disclosure. The system 300 may be implemented as a combination of hardware and software, and may include, for example, control circuitry (e.g., for executing computer-readable instructions), memory, a communication interface, a sensor interface, an input interface, a power supply (e.g., a power management system), any other suitable components, or any combination thereof. To illustrate, the system 300 is configured to generate a latent image, perform a frequency-based transform on the generated latent image, create an activation map based on the transform, process the activation map, generate a degradation map, and generate or produce a suitable response based on the degradation map, or any other process therein.
[0037] The transform engine 310 is configured to create a latent image from a series of images, pre-process the latent image, create multiple image layers by performing a frequency-based transform on the latent image, create an activation map based on the multiple image layers, and perform further processing (e.g., post-processing) on the activation map.
[0038] The transform engine 310 may utilize any frequency-based transform to create multiple image layers. For example, the transform engine 310 may utilize a discrete cosine transform (DCT) to express a finite sequence of data points (e.g., image information) according to the sum of cosine functions oscillating at different frequencies. Although the present disclosure discusses the use of a discrete cosine transform, any type of transform may be implemented without departing from the contemplated embodiments. For example, the following transforms may be implemented without departing from the contemplated embodiments: a binomial transform, a discrete Fourier transform, a fast Fourier transform, a discrete Hartley transform, a discrete sine transform, a discrete wavelet transform, a Hadamard transform (or a Walsh-Hadamard transform), a fast wavelet transform, a Hankel transform, a discrete Chebyshev transform, a finite Legendre transform, a spherical harmonic transform, an irrational basis discrete weighted transform, a number theoretic transform, and a Stirling transform, or any combination thereof. In addition, different types of discrete cosine transforms may be implemented without departing from the contemplated embodiments, including type-I DCT, type-II DCT, type-III DCT, type-IV DCT, type-V DCT, type-VI DCT, type-VII DCT, type-VIII DCT, multi-dimensional type-II DCT (MD DCT-II) and multi-dimensional type-IV DCT (MD-DCT-IV), or any combination thereof.
[0039] The transformation engine 310 may consider a single image (e.g., a group of one image), multiple images, reference information, or a combination thereof. For example, the images may be captured at 5 frames / second to 10 frames / second, or any other suitable frame rate. In another example, a group of images may include ten images, less than ten images, or more than ten images for analysis by the transformation engine 310. In some embodiments, the transformation engine 310 applies pre-processing to each image in the group of images to prepare the image for processing. For example, the transformation engine 310 may brighten one or more images or portions thereof in the captured images, darken one or more images or portions thereof in the captured images, color shift one or more images in the captured images (e.g., in color schemes, from color to grayscale, or other mapping, etc.), crop the image, resize the image, adjust the aspect ratio of the image, adjust the contrast of the image, adjust the contrast of the image, perform any other suitable processing to prepare the image, or any combination thereof. Additionally, transformation engine 310 may change processing techniques based on the output of transformation engine 310, degradation graph engine 320, smoothing engine 330, response engine 340, output 390, reference information 350, preference information 360, or any combination thereof.
[0040] In some embodiments, the transform engine 310 subsamples each image by dividing the image into regions according to a grid (e.g., forming an array of regions that constitute the image as a whole). For illustration, with reference to the subsampled grid, the transform engine 310 selects a small neighborhood for each center pixel (e.g., A×B pixels), resulting in N×M regions. For example, for illustration, N and M can be positive integers that can be equal to each other but do not have to be equal to each other (e.g., a region can be a square of 7×7 pixels or 8×8 pixels; or alternatively, 10×6 pixels).
[0041] In some embodiments, the transformation engine 310 generates a latent image by receiving a plurality of images from a camera or alternatively images stored in a storage device (e.g., memory storage device 370). The plurality of images includes a series of images captured by, for example, a camera (e.g., camera 102) attached to the vehicle. In such examples, the series of images contains visual information related to the vehicle's surroundings, such as roads, road conditions, signs, other vehicles, etc. According to the techniques and embodiments shown and described in the present disclosure, the latent image contains information related to the temporal and / or spatial differences among the series of images from which the latent image is generated.
[0042] The smoothing engine 330 is configured to smooth the output of the degradation map engine 320. In some embodiments, the smoothing engine 330 takes the degradation map from the degradation map engine 320 as input and determines a smoothed degradation map, which can be the same as the output of the degradation map engine 320, but need not be the same. To illustrate, the degradation map engine 320 can identify image degradation (e.g., caused by an occluder) or the removal of an occluder relatively quickly (e.g., from frame to frame or during a few frames). The smoothing engine 330 smoothes the transition to ensure some confidence in the state change (e.g., from degraded to non-degraded and / or from occluded to non-occluded, and vice versa). For example, the smoothing engine 330 can increase the delay in the state change (e.g., occluded-non-occluded or degraded-non-degraded), reduce the frequency state changes (e.g., prevent short time scale fluctuations in the state), increase the confidence in the transition, or a combination thereof. In some embodiments, the smoothing engine 330 applies the same smoothing to each transition direction. For example, the smoothing engine 330 may implement the same algorithm and its same parameters regardless of the direction of the state change (e.g., occluded to unoccluded, or unoccluded to occluded). In some embodiments, the smoothing engine 330 applies different smoothing to each transition direction. For example, the smoothing engine 330 may determine the smoothing technique or its parameters based on the current state (e.g., the current state may be "degraded", "occluded", or "unoccluded"). The smoothing engine 330 may apply statistical techniques, filters (e.g., moving averages or other discrete filters), any other suitable techniques for smoothing the output of the degradation map engine 320, or any combination thereof. To illustrate, in some embodiments, the smoothing engine 330 applies Bayesian smoothing to the output of the degradation map 320. In some embodiments, more smoothing is applied to the transition from occluded to unoccluded than to the transition from unoccluded to occluded. As shown, the smoothing engine 330 may output a degradation map 335 that corresponds to the smoothed degradation map values for each region. As shown, for example, black in the degradation mask 335 corresponds to degraded areas, and white in the degradation mask 335 corresponds to non-degraded or non-occluded areas. For example, as depicted, the bottom of the camera is exhibiting image degradation that may be caused by an occluder.
[0043] The response engine 340 is configured to generate an output signal based on the state determined by the degradation graph engine 320 and / or the smoothing engine 330. The response engine 340 may provide the output signal to an auxiliary system, an external system, a vehicle system, any other suitable system, their communication interfaces, or any combination thereof. In some embodiments, the response engine 340 provides the output signal to a cleaning system (e.g., a washing system) to spray water or other liquids on the camera face (e.g., or enable mechanical cleaning such as a wiper) to remove obstructions that cause degradation. In some embodiments, the response engine 340 provides the output signal to a notification system, or otherwise includes a notification system to generate a notification. For example, the notification may be displayed on a display screen, such as a touch screen of a smartphone, a screen of a vehicle console, any other suitable screen, or any combination thereof. In another example, the notification may be provided as an LED light, a console icon, or other suitable visual indication. In further examples, a screen configured to provide a video feed from the classified camera feed may provide a visual indication, such as a warning message, a highlighted area of the video feed corresponding to image degradation or camera occlusion, any other suitable indication superimposed on the video or otherwise presented on the screen, or any combination thereof. In some embodiments, the response engine 340 provides an output signal to an imaging system of the vehicle. For example, a vehicle may receive images from multiple cameras to determine environmental information (e.g., road information, pedestrian information, traffic information, location information, path information, proximity information), and may therefore alter how the image is processed in response to image degradation.
[0044] In some embodiments, as shown, the response engine 340 includes one or more settings 341, which may include, for example, notification settings, degradation thresholds, predetermined responses (e.g., the type of output signal generated in response to the degradation mask 335), any other suitable settings for affecting any other suitable process, or any combination thereof.
[0045] In an illustrative example, the system 300 (e.g., its transform engine 310) may receive a set of images from a camera output (e.g., repeatedly at a predetermined rate). The transform engine 310 generates a latent image from the set of images. The transform engine 310 may perform one or more pre-processing techniques on the latent image. The transform engine 310 performs a high frequency multi-scale fusion transform on the latent image, thereby generating a plurality of frequency layers, each frequency layer corresponding to a frequency-based decomposition of the latent image. The transform engine 310 processes the plurality of frequency layers to generate an activation map corresponding to the frequency with the largest coefficient among the plurality of frequency layers. The transform engine 310 may apply a post-processing technique to the activation map. The activation map is output to the degradation map engine 320. The smoothing engine 330 receives the degradation map from the degradation map engine 320 to generate a smoothed degradation map. As more images are processed over time (e.g., by the transform engine 310 and the degradation map engine 320), the smoothing engine 330 manages a changing degradation mask 335 (e.g., based on the smoothed degradation map). Thus, the output of smoothing engine 330 is used by response engine 340 to determine a response to a determination that an image captured from a camera is degraded due to, for example, the camera being at least partially obscured or unobstructed. Response engine 340 determines an appropriate response based on settings 341 by generating output signals to one or more auxiliary systems (e.g., a cleaning system, an imaging system, a notification system).
[0046] Figure 4 A diagram of an exemplary process for generating a latent image 430 according to some embodiments of the present disclosure is shown. Process 400 may be performed by one or more processes or techniques described herein, such as transformation engine 310. Latent image generator 410 receives a plurality of images, for example, from camera 102. Although with respect to Figure 4 Only a single camera 102 is described, but any number of cameras may be used without departing from the contemplated embodiments. In addition, the latent image generator 410 may receive input images from a memory storage device, such as the memory storage device 370. As shown, the latent image generator 410 receives images 402A, 402B, and 402C from the camera 102. Although only three images (402A to 402C) may be shown and described, any number of images may be used, up to and including image 402N. Images 402A to 402N are a series of images captured over a period of time and may be received from the camera 102. For example, a camera 102 mounted to a moving vehicle 100 and oriented in the direction of travel (e.g., facing forward) generates a series of images 404A to 404C. The exemplary images 404A to 404C depict the scenery around the vehicle 100 as it crosses a road. In addition, the system may utilize vehicle speed information 442 obtained from, for example, a vehicle speed sensor 424. Latent image generator 410 may use various techniques to generate latent image 430, including but not limited to pixel dynamic range, gradient dynamic range, and pixel absolute difference.
[0047] The pixel dynamic range (or "PDR") utilizes the total variation of pixels within one time frame over a series of images and may be represented by way of example as follows:
[0048]
[0049] Wherein, "k" is an image index having a value from 1 to the number of images in the image sequence (e.g., 1 to N). The dynamic range feature captures activities occurring at locations among the images 404A to 404C with respect to time. In some embodiments, the activity is captured by determining the minimum and maximum values at each location {i, j} among the group of images 402A to 402N. For illustration, for each group of images (e.g., the group of images 402A to 402N), a single maximum and a single minimum are determined for each location {i, j} (e.g., at each pixel). In some embodiments, the dynamic range is determined as the difference between the maximum and minimum values, and indicates the amount of change that occurs for the region (e.g., corresponding to the group of images 402A to 402N) within a time interval. The system can utilize vehicle speed information 422 generated from, for example, a vehicle speed sensor 424 to determine whether the vehicle moves as the input image is captured. For illustration, if the region is degraded (e.g., due to the camera being partially blocked), the difference between the maximum and minimum values will be relatively small or even zero (i.e., not relatively large). That is, the area of the latent image that may be degraded will have little change over time. To further illustrate, the dynamic range feature can also help identify whether the area is degraded, especially in low light conditions (e.g., at night) when most of the image content is black. In some embodiments, the system may select all pixels in the area or may subsample the pixels in the area. For example, in some cases, selecting fewer pixels allows sufficient performance to be retained while minimizing the computational load. In an illustrative example, the system may determine an average value for each area of each image in an image sequence (e.g., images 404A to 404C) to generate a sequence of average values for each area of a partitioned grid. The system determines the difference between the maximum and minimum values of the sequence of average values for each position or area of the partitioned grid. Using pixel dynamic range technology, the latent image generator 410 can output a pixel dynamic range map 444 that can be used as a latent image 430.
[0050] In another exemplary embodiment, process 400 may determine one or more gradient values to be used as a latent image, also referred to as a gradient dynamic range ("GDR"). GDR represents the dynamic range of the input image gradients (e.g., images 402A to 402N) over a period of time. In contrast to the PDR metric that captures temporal variation, GDR allows for some spatial information to be considered. To capture spatial variation, the system determines the image gradients (e.g., or other suitable difference operators) using any suitable technique, such as, for example, a Sobel operator (e.g., a 3×3 matrix operator), a Prewitt operator (e.g., a 3×3 matrix operator), a Laplacian operator (e.g., gradient divergence), a Gaussian gradient technique, any other suitable technique, or any combination thereof. To illustrate, the system determines the range of gradient values at each region (e.g., at any pixel location or pixel group) over time (e.g., for a set of images) to determine the variation of the gradient metric. Thus, the gradient dynamic range captures spatiotemporal information. In such embodiments, the gradient or spatial difference determination captures spatial variation, while the dynamic range component captures temporal variation. In the illustrative example, the system can determine the gradient difference by determining the gradient value of each region of each image in the series of images to generate a gradient value sequence for each region, and determining, for each corresponding gradient value sequence, the difference among the gradient values in the corresponding gradient value sequence. In this way, the system determines the gradient difference (e.g., gradient dynamic range) across the series of images, and can output, for example, the gradient dynamic range averaged over a period of time. In the illustrative example, the process 400 can treat the images 404A to 404C as input images and output a gradient average map 446 that can be used as the latent image 430.
[0051] In addition to implementing PDR and GDR techniques, process 400 may also apply pixel absolute difference (or "PAD") techniques. In such examples, process 400 may determine the difference by capturing frame-to-frame changes (e.g., the inverse of the frame rate) that occur in a scene within a very short time interval as a temporal feature. For example, when considering two consecutive image frames, the absolute difference between the two frames (e.g., the difference in averages) may capture the change. In an illustrative example, the system may determine the difference by determining the average of each region of a first image to generate a first set of averages, determining the average of each region of a second image to generate a second set of averages (e.g., the second image is temporally adjacent to the first image), and determining the difference between each average in the first set of averages and the corresponding average in the second set of averages (e.g., to generate an array of difference values). In an illustrative example utilizing the PAD technique, process 400 may treat images 404A to 404C as input images and output a temporal difference map 442 that may be used as a latent image 430.
[0052] In some embodiments, process 400 may combine one or more of the aforementioned techniques to generate latent image 430. For example, process 400 may utilize images 404A-404C to output temporal difference map 442, dynamic range map 444, and gradient mean map 446. Additionally, the system may perform one or more processes to combine some or all of output maps 440 to generate latent image 430.
[0053] Figure 5A A flow chart of an exemplary process 500 for determining image capture degradation of a camera sensor using a high frequency multi-scale fusion transform (HiFT) according to some embodiments of the present disclosure is depicted. HiFT is used to perform frequency domain analysis to find regions of latent images having high frequency content and low frequency content. For illustration, a transform (e.g., DCT) is applied to express a spatial domain image (e.g., input image 504) as a linear combination of cosine functions of different frequencies. In this way, regions of latent images containing high frequency content are identified, indicating that those regions may not experience image degradation, and conversely, regions of latent images containing low frequency content indicate that those regions may be experiencing image degradation. Image degradation may be caused by, for example, a camera being partially occluded. At step 502, a latent image (e.g., latent image 430) generated using one or more of the techniques described herein may be applied as input image 504. For example, input image 504 may be embodied by a latent image generated from a series of images captured by camera 102 by applying, for example, PDR, GDR, or PAD techniques. Although input image 504 may be shown and described as a latent image (eg, latent image 430 ), the input image may be any image without departing from contemplated embodiments.
[0054] Applications such as Figure 5A and Figure 5B, the latent image 430 is divided into an area including A×B blocks. In some embodiments, each block contains a single pixel. To illustrate such embodiments, a 7×7 area contains 7×7 pixels (i.e., forty-nine pixels). In other embodiments, each block contains multiple pixels. To illustrate such embodiments, each block may contain, for example, four pixels (e.g., 2×2 pixels), and the corresponding 7×7 area contains 196 pixels (forty-nine blocks, each containing four pixels). In addition, although each area can be shown and described as square (i.e., A=B), A and B can be any integer without departing from the envisioned embodiment. In addition, the latent image 520 can be divided into areas of different sizes. In such embodiments, three areas of different sizes can be applied, each area reflecting a scale (or resolution). For example, area 522 includes 5×5 pixels centered at pixel {i, j}, area 524 includes 7×7 pixels centered at pixel {i, j}, and area 526 includes 9×9 pixels centered at pixel {i, j}. Although three scales having resolutions of 5x5, 7x7, and 9x9, respectively, are shown and described, any number of scales having any resolutions may be implemented without departing from the contemplated embodiments.
[0055] At steps 506A to 506C, transforms are applied to each region at scale 1, scale 2, and scale 3, respectively, to express those spatial domain signals as linear combinations of cosine functions of different frequencies. For example and as shown at step 506B, region 524 includes 7×7 blocks, each corresponding to one pixel of latent image 520. Thus, region 524 contains 7×7 pixels centered at pixel {i, j}. The 7×7 region defines scale 1. The value of each pixel is related to a visual parameter such as brightness. In such embodiments, for example, a pixel value of 0 corresponds to a black pixel, and a pixel value of 255 corresponds to a white pixel, and all values in between correspond to varying shades of gray. At steps 506A and 506C, transforms are similarly applied to region 524 (at scale 1) and region 526 (at scale 3), respectively. In this way, process 500 provides a multi-scale (i.e., at scales 1 to 3) approach to determining camera occlusion.
[0056] Applying a transform (e.g., a DCT transform) to each A×B region approximates each of those regions by A×B cosine functions, each having a coefficient (or magnitude) corresponding to the contribution of that particular function to the region as a whole. As shown by the frequency matrix visualization 532, the approximated cosine waves increase in frequency from left to right (i.e., in the x-direction) and from top to bottom (i.e., in the y-direction). The resulting frequency matrix contains A×B spectral subbands, each of which includes transform coefficients related to how much its corresponding cosine frequency contributes to the region. As shown, the highest frequency spectral subband is located in the lower right corner of the decomposition 530, and conversely, the lowest frequency spectral subband is located in the upper left corner.
[0057] At steps 508A to 508C, all frequencies except high frequency coefficients are filtered. The presence of high frequency content in a region indicates that the region may not experience image degradation. Therefore, by filtering the low frequency content and the medium frequency content, the regions containing high frequency content are isolated, thereby indicating which regions are experiencing image degradation and which regions are original. Although 28 spectral subbands are shown as constituting high frequency content, any number of spectral subbands may be considered as high frequency content without departing from the contemplated embodiment. In addition, the number of spectral subbands identified as high frequency can be changed by, for example, the input or output of the transform engine 310, the degradation map 320, the smoothing engine 330, the response engine 340, the output 390, or a combination thereof.
[0058] At step 512, the spectral subbands are classified according to their corresponding frequencies. A plurality of output frequency layers are generated, each frequency layer including all magnitudes of a particular spectral subband. Thus, each frequency layer represents an activation map for a particular frequency. Figure 5A In the illustrated exemplary embodiment, 117 output layers 510 are generated as a result of applying HiFT to the input image 504 at three different scales. Each output layer 510 represents a specific frequency, and the intensity of the brightness depicted in each layer represents the magnitude of the coefficient at each point (e.g., pixel) of the layer. For example, layer 1 represents the lowest frequency decomposition caused by applying DCT to the input image 504. As shown, the brighter areas of layer 1 represent locations with larger magnitudes of the lowest frequencies. In contrast, the darker areas of layer 1 represent locations with the lowest magnitudes of the lowest frequency cosine function. In such embodiments, the black portion of layer 1 (having the lowest magnitude) represents an area of the input image 502 that is not affected by the lowest frequency decomposition; on the other hand, the white (or brighter) portion of layer 1 represents an area of the input image 502 that is affected by the lowest frequency decomposition. In such examples, the brighter the area of layer 1, the greater the impact of the lowest frequency, and conversely, the darker the area, the smaller the impact of the lowest frequency contribution to the input image 504.
[0059] At step 514, the regions with maximum activations for each layer are selected and aggregated. In one embodiment, the output frequency layers 510 are compared, and the maximum coefficient value at each location is used to create an output layer 516. In such an embodiment, each location (e.g., each pixel) of each layer is compared to the corresponding location of all other layers. The frequency corresponding to the highest coefficient value is added to the output layer 516. In this way, the frequency corresponding to the highest coefficient value is selected and added to the output layer 516. The resulting output layer 516 includes a mixture of the highest activations of each layer at each frequency, and represents the highest frequency content at each location within the input image 504.
[0060] Fig. 6A and Figure 6B Various sizes and orientations of regions employed in exemplary HiFT according to some embodiments of the present disclosure are depicted. Figure 5B Depicted are region 522 including 5×5 blocks, region 524 including 7×7 blocks, and region 526 including 9×9 blocks. When applying the exemplary HiFT, for example at step 506A, the latent image 520 may be divided into a plurality of regions 522, each region 522 including 5×5 blocks, and each block containing one pixel. The entire latent image 520 is divided in this manner so that the entire latent image 520 is divided into regions.
[0061] like Figure 6B As depicted in FIG. 5 , the input image 504 is decomposed into a plurality of regions 524, each region comprising 7×7 blocks. Figure 6B Only six regions are depicted and described, but any number of regions may be implemented without departing from the envisioned embodiment. In one exemplary embodiment as shown in panel 602, the input image 504 may be evenly decomposed, where each block (or pixel) is contained within a single region. In another exemplary embodiment as shown in panel 604, the input image 504 may be decomposed into a plurality of overlapping regions 524. Although each region 524 is shown as being overlapped by two blocks (or pixels), any number of overlaps may be implemented without departing from the envisioned embodiment. In another exemplary embodiment as shown in panel 606, the input image 504 may be decomposed into a plurality of regions 524, such that each region 524 is separated by one or more blocks (or pixels). Although each region 524 is shown as being separated by two blocks (or pixels), the regions 524 may be separated by any number of blocks (or pixels) without departing from the envisioned embodiment.
[0062] Figure 7 A block diagram depicts an exemplary method 700 for determining image capture degradation using a high frequency multi-scale fusion transform (HiFT) according to some embodiments of the present disclosure. In some embodiments, the process 700 is performed by, for example, Figure 3 6. In some embodiments, process 700 is an application implemented on any suitable hardware and software that can be integrated into a vehicle, communicate with the vehicle's systems, include a mobile device (e.g., a smartphone application), or a combination thereof.
[0063] At step 702, the system generates a latent image. A series of images captured by, for example, camera 102 is processed to indicate temporal and / or spatial changes among the series of images. In one embodiment, the pixel dynamic range of the series of images is determined, thereby generating a latent image, which includes the total variation of each pixel within a certain time frame (e.g., a time frame corresponding to the duration of the capture of the series of images). In another embodiment, the gradient dynamic range of the series of images is determined, thereby generating a latent image, which includes the dynamic range of the image gradient of the series of images. In such embodiments, the image gradient can be the output of a Sobel filter on the series of images. In this way, the resulting latent image includes the spatiotemporal information of the series of images. In another embodiment, the latent image is generated by determining the temporal difference of corresponding pixels on the series of images. In such embodiments, the value of each pixel of the resulting latent image corresponds to the temporal variation experienced by the pixel on the series of images.
[0064] At step 704, the system divides the latent image into a plurality of regions. Each region contains A×B blocks, where A and B can be any integer greater than zero. In some embodiments, the regions are of the same size (i.e., the same resolution). In other embodiments, the system divides the latent image into regions of different sizes (i.e., different resolutions). To illustrate such embodiments, the system divides the latent image into regions of three different resolutions, such as 5×5 blocks, 7×7 blocks, and 9×9 blocks, each containing one pixel.
[0065] At step 706, the system determines the coefficients based on the frequency. The system performs a transform, e.g., a discrete cosine transform (DCT), on each region. The DCT decomposes each region into spectral subbands, each of which has a frequency and a coefficient. The coefficient (or magnitude) of each spectral subband indicates the effect of its corresponding frequency on the decomposed region. The system divides the spectral subbands of each region into high-band frequencies, mid-band frequencies, and low-band frequencies. The system filters the low-band frequencies and mid-band frequencies, thereby leaving only the high-band frequencies.
[0066] At step 708, the system then generates a plurality of frequency layers. Each frequency layer corresponds to a spectral subband frequency. In an illustrative example in which the system decomposes the latent image into a region comprising 7×7 blocks (or pixels), the decomposition produces a 7×7 matrix comprising 49 cosine functions (or spectral subbands), each of which has a frequency coefficient (or magnitude). After filtering the low-band frequencies and the mid-band frequencies, 28 high-band frequencies are retained. Then, the system then generates 28 frequency layers, each corresponding to a high-band frequency in the 28 remaining high-band frequencies and including the coefficient (magnitude) of the frequency.
[0067] At step 710, the frequency layers are aggregated into a single layer that includes the highest coefficients of the multiple layers. In one embodiment, the layers with the highest activation (i.e., highest coefficient) are aggregated using, for example, max pooling. In such embodiments, each coefficient in each layer is compared to the other coefficients at the corresponding position. In this way, the system identifies the frequency with the highest activation at each position (e.g., at each pixel) among the multiple layers. The resulting activation map contains the highest frequency with the highest coefficient.
[0068] At step 712, the activation map is filtered. In one embodiment, a local filter entropy is applied to the activation map. Entropy is a static measure of randomness and is applied as a local filter entropy to characterize the texture of the image (i.e., the density of high frequency content) by providing information about the local variability of the intensity values of pixels in the image. In the case of an image with dense texture (i.e., experiencing high frequency content), the result of the local filter entropy will be low. Conversely, in the case of an image experiencing sparse texture (i.e., experiencing low frequency content), the result of the local filter entropy will be high. To illustrate, when a local filter entropy is applied to an activation map, regions with little content will produce high entropy values and regions with more content will produce low entropy values. In this way, the system determines what regions of the activation map may be experiencing image degradation (by producing high values) and which regions may not experience image degradation (by producing low values). In some embodiments, edge-aware smoothing techniques (e.g., guided filters or domain transform edge-preserving recursive filters) may be used to filter the output of the local filter entropy.
[0069] Figure 8 A flow chart of an exemplary process 800 for determining image capture degradation according to some embodiments of the present disclosure is shown. In some embodiments, process 800 or aspects thereof may be combined with any of the exemplary steps of processes 300, 500, or 700.
[0070] At step 802, the system generates an output signal. For example, step 802 can be the same as step 514 of process 500 of FIG. 5. In another example, step 802 can be the same as step 514 of process 500 of FIG. Figure 7The system may generate an output signal and provide it to, for example, an auxiliary system, an external system, a vehicle system, a controller, any other suitable system, their communication interfaces, or any combination thereof.
[0071] At step 804, the system generates a notification. In some embodiments, the system provides an output signal to a display system to generate a notification. For example, the notification may be displayed on a display screen, such as a touch screen of a smartphone, a screen of a vehicle console, any other suitable screen, or any combination thereof. In another example, the notification may be provided as an LED light, a console icon, a visual indication such as a warning message, a highlighted area corresponding to a degraded video feed, a message (e.g., a text message, an email message, an on-screen message), any other suitable visual or sound indication, or any combination thereof. For illustration, panel 850 shows a message superimposed on a display of a touch screen (e.g., a smartphone or a vehicle console) indicating that the right rear (RR) camera (e.g., camera 104) is blocked by 50%. For further illustration, the notification may provide instructions to a user (e.g., a driver or vehicle occupant) to clean the camera, ignore images from the camera that have experienced degradation, or otherwise factor degradation into images from the camera.
[0072] At step 806, the system causes the camera to be cleaned. In some embodiments, the system provides an output signal to a cleaning system (e.g., a cleaning system) to spray water or other liquids on the camera face (e.g., or enable mechanical cleaning such as a wiper) to remove obstructions that contribute to image degradation. In some embodiments, the output signal causes the wiper motor to reciprocate the wiper across the camera lens. In some embodiments, the output signal causes the liquid pump to start and pump a cleaning fluid toward the lens (e.g., as a spray from a nozzle coupled to the pump through a pipe). In some embodiments, the output signal is received by a cleaning controller that controls the operation of a cleaning fluid pump, a wiper, or a combination thereof. For illustration, panel 860 shows a pump and a wiper configured to clean the camera lens. The pump sprays a cleaning fluid toward the lens to remove or otherwise dissolve / soften the obstruction, and the wiper rotates across the lens to mechanically remove the obstruction.
[0073] At step 808, the system modifies the image processing. In some embodiments, the system provides the output signal to the imaging system of the vehicle. For example, the vehicle may receive images from multiple cameras to determine environmental information (e.g., road information, pedestrian information, traffic information, location information, path information, proximity information), and thus may change how to process the image in response to image degradation. For illustration, panel 870 shows an image processing module that takes images from four cameras as input (e.g., but any suitable number of cameras may be implemented, including one camera, two cameras, or more than two cameras). As shown in panel 870, one of the four cameras experiences image degradation caused by an occlusion (e.g., indicated by an "×"), while the other three cameras do not experience degradation (e.g., indicated by a check mark). In some embodiments, the image processing module may ignore output from a camera that exhibits image degradation, ignore a portion of an image from a camera that exhibits occlusion, reduce the weight or importance associated with a camera that exhibits degradation, consider any other suitable modification of the overall output of a camera that exhibits degradation, or a combination thereof. The determination of whether to modify image processing may be based on the degree of degradation (e.g., the relative amount of occluded pixels to total pixels), the shape of the degradation (e.g., a sharply skewed aspect ratio such as a stripe occlusion may be less likely to trigger modification than a more square aspect ratio), which camera was identified as capturing the image exhibiting degradation, the time of day or night, user preferences (e.g., included as a threshold or other reference in the reference information), or a combination thereof.
[0074] In some embodiments, at step 808, the system ignores a portion of the camera's output. For example, the system may ignore or otherwise exclude during analysis a portion of the camera's output that corresponds to the degradation mask. In further examples, the system may ignore quadrants, halves, sectors, windows, any other suitable collection of pixels having a predetermined shape, or any combination thereof based on the degradation mask (e.g., the system may map the degradation mask to a predetermined shape and then size and arrange the shape accordingly to indicate the portion of the camera's output to ignore).
[0075] The foregoing merely illustrates the principles of the present disclosure, and various modifications may be made by those skilled in the art without departing from the scope of the present disclosure. The above embodiments are presented for the purpose of illustration and not limitation. The present disclosure may also take many forms other than those explicitly described herein. Therefore, it should be emphasized that the present disclosure is not limited to the methods, systems, and instruments explicitly disclosed, but is intended to include variations and modifications thereof, which are within the spirit of the following claims.
Claims
1. A method for determining image capture degradation, the method include: capturing a series of image frames over time by a camera of the vehicle via one or more sensors; generating, using processing circuitry, a latent image from the series of image frames captured by the camera, the latent image representing temporal or spatial differences among the series of image frames over time; Using processing circuitry, applying a high frequency multi-scale fusion transform and generating a plurality of frequency layers based on the latent image, wherein each frequency layer corresponds to a frequency-based decomposition of the latent image at a corresponding scale and frequency; and Using the processing circuit, identifying image capture degradation of the camera based on the plurality of frequency layers, including generating an activation map by aggregating the plurality of frequency layers into a single layer including highest coefficients of the plurality of frequency layers; applying a local filter entropy filter to the activation map to filter the activation map; identifying state transitions of the filtered activation map and smoothing the state transitions to ensure confidence in the state changes; outputting a degradation map corresponding to the smoothed degradation map values for each region; Image capture degradation of the camera is identified based on the degradation map.
2. The method according to claim 1, wherein the plurality of frequency layers are generated include: A plurality of frequency-based coefficients are determined for each of a plurality of scales centered at each of a plurality of locations in the latent image, wherein the plurality of frequency-based coefficients correspond to a plurality of spectral sub-band frequencies, and wherein each of the plurality of frequency layers includes frequency-based coefficients for a corresponding one of the spectral sub-band frequencies, and wherein identifying image capture degradation further comprises selecting a subset of the coefficients based on a frequency magnitude threshold prior to generating an activation map by aggregating the plurality of frequency layers into a single layer including highest coefficients of the plurality of frequency layers.
3. The method of claim 2, wherein generating a plurality of frequency layers further comprises: include: Selecting a subset of coefficients based on their frequencies; a subset of said frequency-based coefficients for categorizing said magnitude; as well as The classified subsets of frequency-based coefficients are normalized to generate multiple layers. 4 . The method of claim 1 , wherein the camera captures the series of image frames at a sampling frequency and wherein the sampling frequency is determined based on a vehicle speed.
5. The method of claim 1, wherein the latent image is generated based on a pixel dynamic range, wherein the pixel dynamic range is determined by a difference in the value of each pixel in the series of image frames.
6. The method of claim 1, wherein the latent image is generated based on a gradient dynamic range, wherein the gradient dynamic range is determined by differences in image gradients in the series of image frames. The method of claim 6 , wherein the image gradient is determined based on a Sobel filter. The method of claim 6 , wherein the image gradient is determined based on a bilateral filter.
9. The method according to claim 1, further comprising: include: identifying an image frame captured by the camera when the vehicle is traveling at a speed below a predetermined threshold; as well as excluding the identified image frame from the series of image frames; Wherein when the vehicle speed is lower than a predetermined threshold, the image frame is excluded from the series of image frames.
10. The method according to claim 2, further comprising: include: The frequency magnitude threshold is adjusted based on one or more frequency layers of the plurality of frequency layers.
11. The method according to claim 1, further comprising: include: determining whether an occlusion exists based on the identified image capture degradation; as well as In response to determining that an occlusion exists, a fluid is applied to the face of the camera.
12. The method according to claim 1, further comprising: include: Causes generation of a notification indicating image capture degradation.
13. The method according to claim 1, further comprising: include: One or more regions of one or more of the image frames are ignored based on the identified image capture degradation.
14. A system for determining image capture degradation: camera; a control circuit coupled to the camera and configured to: capturing a series of image frames over time by the camera via one or more sensors; generating a latent image from the series of image frames captured by the camera, the latent image representing temporal or spatial differences among the series of image frames over time; Applying a high frequency multi-scale fusion transform and generating a plurality of frequency layers based on the latent image, each frequency layer corresponding to a frequency-based decomposition of the latent image at a corresponding scale and frequency; as well as identifying image capture degradation of the camera based on the plurality of frequency layers, comprising generating an activation map by aggregating the plurality of frequency layers into a single layer including highest coefficients of the plurality of frequency layers; applying a local filter entropy filter to the activation map to filter the activation map; identifying state transitions of the filtered activation map and smoothing the state transitions to ensure confidence in the state changes; outputting a degradation map corresponding to the smoothed degradation map values for each region; Image capture degradation of the camera is identified based on the degradation map.
15. The system of claim 14, wherein the camera is integrated into a vehicle and wherein the camera captures the series of image frames at a sample frequency based on a speed of the vehicle.
16. The system of claim 15, wherein the control circuit is further configured to: identifying an image frame captured by the camera when the vehicle is traveling at a speed below a predetermined threshold; and The identified image frame is excluded from the image frames.
17. The system of claim 14, wherein the control circuit ignores camera output.
18. The system of claim 14, further comprising a cleaning system, wherein the cleaning system applies a fluid to a face of the camera.
19. The system of claim 14, further comprising a display device configured to display a notification indicating an occlusion event.
20. A non-transitory computer readable medium, the non-transitory computer readable medium include: Program instructions for determining image capture degradation, which when executed cause a computer processing system to perform steps including: capturing a series of image frames over time by a camera of the vehicle via one or more sensors; generating, using control circuitry, a latent image from the series of image frames captured by the camera, the latent image representing temporal or spatial differences among the series of image frames over time; Using a control circuit, applying a high frequency multi-scale fusion transform and generating a plurality of frequency layers based on the latent image, each frequency layer corresponding to a frequency-based decomposition of the latent image at a corresponding scale and frequency; as well as Using a control circuit, identifying image capture degradation of the camera based on the plurality of frequency layers, including generating an activation map by aggregating the plurality of frequency layers into a single layer including highest coefficients of the plurality of frequency layers; applying a local filter entropy filter to the activation map to filter the activation map; identifying state transitions of the filtered activation map and smoothing the state transitions to ensure confidence in the state changes; outputting a degradation map corresponding to the smoothed degradation map values for each region; Image capture degradation of the camera is identified based on the degradation map.
Citation Information
Patent Citations
Dirt detection device and method for vehicle-mounted camera
CN112261403A
Methods and apparatus relating to detection and / or indicating a dirty lens condition
US20160004144A1