Systems and methods for determining camera occlusion

By using partitioned grid and classifier technology, the occlusion area of ​​the camera is accurately identified, which solves the problem of image quality degradation caused by camera occlusion and achieves effective occlusion response and image processing.

CN115205522BActive Publication Date: 2025-10-31RUIWEIAN INTELLECTUAL PROPERTY HLDG CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202111631743.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2021-04-08
Filing Date
2021-12-29
Publication Date
2025-10-31
Estimated Expiration
2041-12-29

AI Technical Summary

Technical Problem

Existing technologies struggle to effectively determine camera obstruction, leading to decreased image quality and rendering the images unusable for vehicle occupants.

Method used

By applying a partitioned grid to spatially divide the image set, spatial, temporal, and combined features are extracted. A classifier is used to determine the occlusion region, and a smoothing technique is used to generate an output signal in response to the occlusion.

Benefits of technology

It accurately identifies areas obstructed by the camera, ensuring image quality, and can remove or ignore obstructed images, generate obstruction notifications, and improve the effectiveness of image processing.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115205522B_ABST
    Figure CN115205522B_ABST
Patent Text Reader

Abstract

A system for determining vehicle camera occlusion is configured to apply a partitioned grid with multiple locations to an image sequence to form multiple regions. The system determines at least one spatial feature corresponding to the partitioned grid and at least one temporal feature corresponding to the partitioned grid. The system generates a classification sequence for each of the multiple locations based on the at least one spatial feature, the at least one temporal feature, and reference information. The system applies a smoothing technique to determine a subset of occluded regions in the classification sequence and generates an output signal based on the subset of regions. The output can be provided to an output system to clean the camera lens, notify a user or vehicle of occlusion, or modify image processing.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates to systems and methods for determining camera occlusion, and more specifically, to systems and methods for determining camera occlusion using an occlusion classifier based on image features. Summary of the Invention

[0002] In some embodiments, this disclosure relates to a method for determining camera occlusion. The method includes, for example, spatially partitioning an image set by applying a partitioned grid with multiple locations to form multiple regions for each image. In some embodiments, each image in a time-indexed image sequence is partitioned. The method includes determining at least one spatial feature corresponding to the partitioned grid and at least one temporal feature corresponding to the partitioned grid. The method includes generating a classification sequence for each of the multiple locations based on the at least one spatial feature, the at least one temporal feature, and reference information. In some embodiments, the reference information includes thresholds, restrictions, instruction sets, or other reference information for classifying each region of the image. The method includes applying a smoothing technique to determine a subset of occluded regions in the classification sequence. For example, in some embodiments, the regions determined to be occluded are generally occlusion masks corresponding to effectively occluded regions of the camera. The method includes generating an output signal based on the subset of regions.

[0003] In some embodiments, the at least one spatial feature includes a scale feature. The method includes: determining a sequence of scale sizes at each location in each image; determining a range metric for each scale size of the scale size sequence at each location to generate a set of range metrics; and determining differences between the sets of range metrics. For example, in some embodiments, the method includes determining an array of scale features corresponding to a region array (e.g., specified by a partitioned grid).

[0004] In some embodiments, the at least one temporal feature includes an average feature. The method includes: for each image, determining a corresponding average metric at each location corresponding to more than one region to generate a sequence of average metrics; and determining the differences between the sequences of average metrics. For example, in some embodiments, the method includes determining an array of average features corresponding to an array of regions (e.g., specified by a partitioned grid).

[0005] In some embodiments, the at least one temporal feature includes a difference feature. The method includes: determining an average value for each region of a first image to generate a first set of average values; determining an average value for each region of a second image to generate a second set of average values; and determining the difference between each average value in the first set of average values ​​and a corresponding average value in the second set of average values. The second image is temporally adjacent to the first image. For example, in some embodiments, the method includes determining an array of difference features corresponding to an array of regions (e.g., specified by a partitioned grid).

[0006] In some embodiments, the at least one temporal feature includes a range feature. The method includes: determining an average value for each region of each image in the image sequence to generate a sequence of average values ​​for each location of the partitioned grid; and determining the difference between the maximum and minimum values ​​of the sequence of average values ​​for each location of the partitioned grid. For example, in some embodiments, the method includes determining a range feature array corresponding to a region array (e.g., specified by the partitioned grid).

[0007] In some embodiments, the at least one temporal feature includes gradient features. The method includes: determining gradient values ​​for each region of each image in the image sequence to generate a gradient value sequence for each location of the partitioned grid; and for each corresponding gradient value sequence, determining the difference between the gradient values ​​of the corresponding gradient value sequences. For example, in some embodiments, the method includes determining an array of gradient features corresponding to an array of regions (e.g., specified by a partitioned grid).

[0008] In some embodiments, determining the at least one spatial feature and the at least one temporal feature includes determining range features, gradient features, difference features, scale features, and average features.

[0009] In some embodiments, the reference information includes a reference value, and generating the output signal includes: determining the degree of occlusion, determining whether the degree of occlusion exceeds the reference value, and identifying a response if the occlusion exceeds the reference value, wherein the output signal indicates the response. For example, in some embodiments, the number or score of regions classified as occluded is compared with a threshold to determine the degree of occlusion. In another instance, in some embodiments, the number or score of regions classified as occluded is equivalent to the degree of occlusion.

[0010] In some embodiments, the output signal is configured to cause the image processing module to ignore the camera's output. In some embodiments, generating the output signal includes generating a notification on a display device indicating the degree of occlusion. In some embodiments, the output signal is configured to cause a washing system to apply liquid to the surface of the camera.

[0011] In some embodiments, applying smoothing techniques to determine the subset of regions includes: determining a smoothing metric based on the current classification of each of the plurality of locations; determining a smoothed classification value sequence based on the smoothing metric and the classification sequence; and determining a new classification based on the smoothed classification value sequence.

[0012] In some embodiments, this disclosure relates to a system for determining camera occlusion. The system includes a camera system, control circuitry, and an output interface. The camera system is configured to capture a sequence of images. The control circuitry is coupled to the camera system and configured to: apply a partitioned grid comprising multiple locations to each image of the image sequence to form multiple regions for each image, wherein each image of the image sequence is time-indexed. The control circuitry is further configured to: determine at least one spatial feature corresponding to the partitioned grid and at least one temporal feature corresponding to the partitioned grid; generate a classification sequence for each of the multiple locations based on the at least one spatial feature, the at least one temporal feature, and reference information; and apply a smoothing technique to determine a subset of occluded regions in the classification sequence. The output interface is configured to generate an output signal based on the subset of regions. Attached Figure Description

[0013] This disclosure is described in detail with reference to the following accompanying drawings, based on one or more embodiments. The drawings are provided for illustrative purposes only and depict only typical or exemplary embodiments. These drawings are provided to facilitate understanding of the concepts disclosed herein and should not be construed as limiting the breadth, scope, or applicability of these concepts. It should be noted that these drawings are not necessarily drawn to scale for clarity and ease of illustration.

[0014] Figure 1 A top view of an illustrative vehicle having several cameras according to some embodiments of the present disclosure is shown;

[0015] Figure 2 A diagram illustrating illustrative output from a camera according to some embodiments of the present disclosure is shown;

[0016] Figure 3 A system diagram of an illustrative system for managing camera occlusion and responses according to some embodiments of the present disclosure is shown;

[0017] Figure 4 A diagram illustrating a set of images for feature extraction according to some embodiments of the present disclosure is shown;

[0018] Figure 5 A flowchart illustrating an illustrative process for managing camera occlusion and responses according to some embodiments of the present disclosure is shown;

[0019] Figure 6 A flowchart illustrating an illustrative process for managing classifications according to some embodiments of this disclosure is shown;

[0020] Figure 7 A graph illustrating the illustrative response of a smooth classifier according to some embodiments of the present disclosure is shown;

[0021] Figure 8 A graph illustrating the relationship between time and smoothing measures according to some embodiments of the present disclosure is shown;

[0022] Figure 9 A block diagram illustrating an illustrative smoothing technique according to some embodiments of the present disclosure is shown;

[0023] Figure 10 A block diagram illustrating occlusion and extraction features according to some embodiments of the present disclosure is shown;

[0024] Figure 11 A block diagram illustrating an illustrative smoothing technique according to some embodiments of the present disclosure is shown;

[0025] Figure 12 A graph illustrating the illustrative response of a smooth classifier according to some embodiments of the present disclosure is shown;

[0026] Figure 13 A graph illustrating the illustrative response of a smooth classifier according to some embodiments of the present disclosure is shown;

[0027] Figure 14 A graph illustrating the illustrative relationship between time and smoothing metrics according to some embodiments of the present disclosure; and

[0028] Figure 15 A graph illustrating two illustrative confidence measures according to some embodiments of the present disclosure is shown. Detailed Implementation

[0029] Camera occlusion can occur due to various reasons, such as dust accumulation on the camera lens, bird droppings, or objects placed on the camera. Occlusion can degrade image quality, rendering it unusable for other algorithms or by vehicle occupants. The systems and methods disclosed herein relate to determining which portions of an image frame are occluded and responding to that occlusion.

[0030] Figure 1 A top view of an illustrative vehicle 100 having a plurality of cameras according to some embodiments of the present disclosure is shown. As shown, vehicle 100 includes cameras 101, 102, 103 and 104, but it should be understood that the vehicle may include any suitable number of cameras according to the present disclosure (e.g., one camera, more than one camera).

[0031] Panel 150 shows a cross-sectional view of a camera displaying occlusion. The occlusion covers a portion 152 of the camera, while a portion 151 remains uncovered (e.g., although portion 151 may be affected by the occlusion). The occlusion may completely cover portion 152 and may effectively cover at least some of portion 151 (e.g., uneven distribution of light reflected from the occlusion). The occlusion may be temporary on the camera and may persist for a period of time (e.g., fall off, dissipate, or remain). In some embodiments, the systems and methods of this disclosure involve determining which portions of the camera are occluded and responding to the occlusion by clearing the occlusion, ignoring an image showing the occlusion, modifying image processing for output from the camera, generating an occlusion notification, any other suitable function, or any combination thereof.

[0032] Figure 2A diagram illustrating illustrative output 200 from a camera according to some embodiments of the present disclosure is shown. As shown, output 200 includes a plurality of time-indexed captured images 201-205 (e.g., the images are sequential). A partition grid (point 210 of the partition grid is shown) is applied to images 201-205 to define regions. Region 211 corresponds to a location on the partition grid. The partition grid comprises N x M points, and region 211 may correspond to a specific number of pixels (e.g., 11 x 11 pixels, 10 x 10 pixels, or any other A x B pixel set), with pixels corresponding to each point. For example, images 201-205 may each comprise (N*A) x (M*B) pixels, which are divided into N x M regions, each region comprising A x B pixels. In some embodiments, these regions do not overlap. For example, each pixel may be associated with a single region (e.g., and other pixels). In some embodiments, these regions at least partially overlap. For example, at least some pixels may be associated with more than one region (e.g., adjacent indexed regions). In some embodiments, these regions do not overlap and are spaced apart. For example, at least some pixels do not need to be associated with any region (e.g., adjacent index regions). According to this disclosure, any suitable regions, whether overlapping or non-overlapping, spaced or not spaced, or a combination thereof, can be used. The systems and methods of this disclosure can be applied to determine the features of each region of each image, two adjacent images, a set of images, or a combination thereof, and therefore can include any suitable number of different feature types. For example, one or more features can be extracted for each region of any or every image (e.g., for any of the five possible values, any one of images 201-205). In another instance, one or more features can be extracted for each location of a partitioned grid by comparing two adjacent images (e.g., for four possible values ​​at each location, images 201 and 202, 202 and 203, 203 and 204, or 204 and 205). In another instance, one or more features can be extracted for each location of a partitioned grid by comparing image sets 201-205 (e.g., to generate a feature value for each location). In some embodiments, the locations of the partitioned grid are used to identify an occlusion mask corresponding to a set of locations corresponding to an occlusion state. In some embodiments, the output of one or more cameras can be analyzed to determine feature values. The partitioned grid need not be rectangular and may include gaps, spaces, irregularly arranged points, arrays, or combinations thereof. The partitioned grid corresponds to a set of indexes for the feature values, which may, but does not necessarily, correspond to an array. For example, in some embodiments, while the partitioned grid may be applied to a set of images, the feature values ​​obtained by indexing the points of the partitioned grid do not need to be directly applied to regions of a single image. For illustration, in some embodiments, the partitioned grid provides a structure for determining feature values ​​rather than for dividing a particular image.To further illustrate, the partitioned grid may include a data structure for storing feature values, which is indexed by the spatial locations of regions corresponding to the image set.

[0033] Figure 3 A system diagram of an illustrative system 300 for managing camera obstruction and response according to some embodiments of the present disclosure is shown. As shown, system 300 includes a feature extractor 310, a classifier 320, a smoothing engine 330, a response engine 340, reference information 350, preference information 360, and a memory storage device 370. It should be understood that the arrangement of the illustrated system 300 can be modified according to the present disclosure. For example, components can be combined, separated, added, removed, modified, omitted, or otherwise modified according to the present disclosure. System 300 can be implemented as a combination of hardware and software and can include, for example, control circuitry (e.g., for executing computer-readable instructions), memory, a communication interface, a sensor interface, an input interface, a power supply (e.g., a power management system), any other suitable components, or any combination thereof. For illustration, system 300 is configured to extract features from a set of images, classify images or regions thereof, evaluate the classification (e.g., smooth the classification), and generate or provoke an appropriate response to the classification or a change in classification.

[0034] Feature extractor 310 is configured to determine one or more features of an image set to determine spatial features, temporal features, spatial-temporal features, image information, any other suitable information, or any combination thereof. Feature extractor 310 may consider a single image (e.g., a set of images), multiple images, reference information, or combinations thereof to determine features. For example, images may be captured at a frame rate of 5-10 frames per second or any other suitable frame rate. In another example, a set of images may include ten images, fewer than ten images, or more images for analysis by feature extractor 310. In some embodiments, feature extractor 310 applies preprocessing to each image in the image set to prepare the image for segmentation and feature extraction. For example, feature extractor 310 may brighten an image or a portion thereof, darken an image or a portion thereof, shift the image color (e.g., between color schemes, from color to grayscale or other mappings), crop an image, scale an image, adjust the aspect ratio of an image, adjust the contrast of an image, perform any other suitable processing to prepare the image, or any combination thereof. In some embodiments, the feature extractor 310 subsamples each image by dividing the image into regions according to a grid (e.g., forming an array of regions that collectively constitute the image). For illustration, referring to the subsampling grid, the feature extractor 310 selects a small neighborhood for each center pixel (e.g., N x M pixels), thereby producing N*M regions (e.g., N*M values ​​for some features of each image). For example, for illustration, N and M can be positive integers that may, but do not necessarily, be equal to each other.

[0035] In some embodiments, the feature extractor 310 determines spatial features by considering a single image and determining a set of feature values ​​for each region of that image (e.g., N*M feature values ​​for each image). Therefore, the feature extractor 310 can store spatial feature values ​​of the image; compare feature values ​​of the image and a reference image or value, and store comparison metrics; scale, normalize, or otherwise modify the determined feature values; or combinations thereof. Spatial features include any suitable features determined based on regions of a single image, such as scaling features, gradient features, minimum / maximum values, average values, object recognition, any other suitable features indicating spatial variations of the image (e.g., regions thereof), or any combination thereof.

[0036] In some embodiments, the feature extractor 310 determines temporal features by considering multiple images (e.g., a consecutive set of images) and determining a set of feature values ​​(e.g., N*M feature values ​​per image) for the set of images or regions of the images. Therefore, the feature extractor 310 may: store the temporal feature values ​​of the set or subset of images; compare the feature values ​​of a first image and a second subsequent image (e.g., compare adjacent indexed images) and store the comparison metric; scale, normalize, or otherwise modify the determined feature values; or combinations thereof. Temporal features include any suitable features determined based on comparisons of multiple images, such as dynamic range, differences (e.g., changes in average, minimum, or maximum values), any other suitable features indicating temporal changes between images (e.g., regions of images), or any combination thereof.

[0037] In some embodiments, the feature extractor 310 determines spatial-temporal features (“combined features”) by considering multiple images (e.g., a consecutive set of images) and determining a set of feature values ​​(e.g., N*M feature values ​​for each image set in one or more image sets). Therefore, the feature extractor 310 may: store combined feature values ​​for a set or subset of images; compare feature values ​​of a first image and a second subsequent image (e.g., compare adjacent indexed images) to determine combined features and store the combined feature values; compare feature values ​​of more than two images (e.g., variations in some index intervals of the images); or combinations thereof. Combined features include any suitable features determined based on comparisons of regions of each of the multiple images, such as dynamic range, differences (e.g., variations in average, minimum, and maximum values), any other suitable features indicating spatial and temporal variations between regions of the images, or any combination thereof.

[0038] Classifier 320 is configured to determine a classification corresponding to a partitioned grid. For example, each location (e.g., each region) of the partitioned grid is classified into one of a series of states. In some embodiments, the system identifies both occluded and unoccluded states and determines whether each region corresponds to an occluded state or an unoccluded state. For illustration, the set of regions in an occluded state forms an occlusion mask and may correspond to, for example, the degree of physical occlusion by a camera. In some embodiments, the system identifies more than two states, such as unoccluded, occluded, and damaged states. For example, the system may determine an intermediate state between an occluded state and an unoccluded state (e.g., a damaged state). In some embodiments, classifier 320 retrieves or otherwise accesses reference information to determine the classification. For example, classifier 320 may retrieve thresholds, parameter values ​​(e.g., weights), algorithms (e.g., computer-implemented instructions), offset values, or combinations thereof from memory. In some embodiments, classifier 320 applies an algorithm to the output (e.g., feature values) of feature extractor 310 to determine the classification. For example, the classifier 320 can apply least squares determination, weighted least squares determination, support vector machine (SVM) determination, multilayer perceptron (MLP) determination, any other suitable classification technique, or any combination thereof.

[0039] In some embodiments, classifier 320 performs classification for each frame captured (e.g., each image), and thus updates the classification as each new image becomes available. In some embodiments, classifier 320 performs classification based on a set of images, and thus can determine the classification for each frame (e.g., at a frequency equal to the frame rate) or at a lower frequency (e.g., every ten frames or other suitable frequency). In some embodiments, classifier 320 performs classification at a downsampling frequency less than the frame rate, such as a predetermined frequency (e.g., by time or number of frames).

[0040] As shown, classifier 320 can retrieve or otherwise access settings 321, which may include, for example, classification settings, classification thresholds, predetermined classifications (e.g., two or more categories to which a region may belong), any other suitable settings for classifying regions of an image, or any combination thereof. Classifier 320 can apply one or more settings of 321 to classify regions of an image, locations corresponding to partitioned grids, or both, based on features extracted by feature extractor 310. As shown, classifier 320 includes a selector 322 configured to select from classifications, classification schemes, classification techniques, or combinations thereof. For example, in some embodiments, selector 322 is configured to select from a predetermined set of categories based on the output of feature extractor 310. In some embodiments, selector 322 is configured to select from a predetermined set of classification schemes such as occluded / unoccluded, occluded / partially occluded / unoccluded, any other suitable scheme, or any combination thereof. In some embodiments, selector 322 is configured to select from a predetermined set of classification techniques, such as least squares, weighted least squares, support vector machine (SVM), multilayer perceptron (MLP), any other suitable technique, or any combination thereof.

[0041] Smoothing engine 330 is configured to smooth the output of classifier 320. In some embodiments, smoothing engine 330 takes the classification from classifier 320 (e.g., for each region) as input and determines a smoothed classification that may, but does not have to, be the same as the output of classifier 320. For illustration, classifier 320 may identify occlusion or occlusion removal relatively quickly (e.g., frame-to-frame, or over several frames). Smoothing engine 330 smooths this transition to ensure some level of confidence in the state change. For example, smoothing engine 330 may increase the latency of state changes (e.g., occlusion-to-unoccluded), reduce the frequency of state changes (e.g., prevent short-timescale fluctuations in the state), increase the confidence of the transition, or a combination thereof. In some embodiments, smoothing engine 330 applies the same smoothing to each transition direction. For example, smoothing engine 330 may implement the same algorithm and its same parameters regardless of the direction of the state change (e.g., occlusion to unoccluded, or unoccluded to occlusion). In some embodiments, smoothing engine 330 applies different smoothing to each transition direction. For example, the smoothing engine 330 can determine the smoothing technique or its parameters based on the current state (e.g., the current state could be "occluded" or "unoccluded"). The smoothing engine 330 can apply statistical techniques, filters (e.g., moving averages or other discrete filters), any other suitable techniques for smoothing the output of the classifier 320, or any combination thereof. For illustration, in some embodiments, the smoothing engine 330 applies Bayesian smoothing to the classification of the classifier 320. In some embodiments, more smoothing is applied to the transition from occluded to unoccluded compared to the transition from unoccluded to occluded. As shown, the smoothing engine 330 can output an occlusion mask 335 corresponding to the smoothed classification value for each region. As shown, for example, black in the occlusion mask 335 corresponds to occlusion, while white in the occlusion mask 335 corresponds to unocclusion (e.g., the bottom of the camera shows occlusion).

[0042] The response engine 340 is configured to generate an output signal based on state transitions determined by the smoothing engine 330. The response engine 340 may provide the output signal to an auxiliary system, external system, vehicle system, any other suitable system, its communication interface, or any combination thereof. In some embodiments, the response engine 340 provides the output signal to a cleaning system (e.g., a washing system) to spray water or other liquids on the camera surface (e.g., or activate a mechanical cleaner such as a wiper) to remove obstructions. In some embodiments, the response engine 340 provides the output signal to a notification system or otherwise includes a notification system to generate a notification. For example, the notification may be displayed on a display screen, such as a smartphone touchscreen, a vehicle console screen, any other suitable screen, or any combination thereof. In another instance, the notification may be provided as an LED light, a console icon, or other suitable visual indicator. In yet another instance, a screen configured to provide video feeds from categorized camera feeds may provide visual indicators, such as warning messages, a highlighted area corresponding to an obstructed video feed, any other suitable indication overlaid on the video or otherwise presented on the screen, or any combination thereof. In some embodiments, the response engine 340 provides an output signal to the vehicle's imaging system. For example, the vehicle may receive images from multiple cameras to determine environmental information (e.g., road information, pedestrian information, traffic information, location information, path information, proximity information), and thus may change the way the images are processed in response to occlusion or lack of occlusion.

[0043] In some embodiments, as shown, the response engine 340 includes one or more settings 341, which may include, for example, notification settings, occlusion thresholds, predetermined responses (e.g., the type of output signal generated in response to occlusion mask 335), any other suitable settings for influencing any other suitable process, or any combination thereof.

[0044] In an illustrative example, system 300 (e.g., its feature extractor 310) can receive a set of images from camera output (e.g., repeated at a predetermined rate). Feature extractor 310 applies a partitioned grid and determines a set of feature values ​​corresponding to the grid, each feature corresponding to one or more images. Feature extractor 310 can determine one or more spatial features, one or more temporal features, one or more combined features (e.g., spatial-temporal features), any other suitable information, or any combination thereof. The extracted features are output to classifier 320, which classifies each location of the partitioned grid into one or more states (e.g., occluded or unoccluded). Smoothing engine 330 receives classification and historical classification information from classifier 320 to generate a smooth classification. As more images are processed over time (e.g., by feature extractor 310 and classifier 320), smoothing engine 330 (e.g., based on smooth classification) manages changing occlusion masks 335. Therefore, the output of smoothing engine 330 is used by response engine 340 to determine a response to the determination that the camera is at least partially occluded or unoccluded. The response engine 340 generates an output signal to one or more auxiliary systems (e.g., a washing system, an imaging system, a notification system) and determines an appropriate response based on settings 341.

[0045] Figure 4 A diagram illustrating an illustrative image set for feature extraction according to some embodiments of the present disclosure is shown. As shown, image set 400 includes five images denoted by {l, m, n, o, p}. An illustrative point of a partitioned grid is shown, denoted by {i, j}, where the partitioned grid includes N x M locations (e.g., where i is in a set of integers 1:N and j is in a set of integers 1:M). Image set 400 spans a certain time interval (e.g., equal to four times the frame rate multiplied by the time elapsed between image l and image p). In an illustrative example, Figure 4 The features disclosed in the context can be derived from Figure 3 The feature extractor 310 is determined.

[0046] Referring to Figure 400, the system can determine one or more average features (MFs) corresponding to a partitioned grid. For illustration, if the occluding material is clay covering only a portion of the image sensor, the unoccluded portion may experience fading depending on where the surrounding light originates (e.g., the direction of illumination). This could be because light reflects off the clay surface and brightens the unoccluded portion on the image sensor. In some cases, abrupt changes in illumination can cause a classifier (e.g., classifier 320) to give false positives. Therefore, the system can use small neighborhoods around each location of the partitioned grid and compute an average for each neighborhood. In some embodiments, the system can determine the average for all images in a set of images. For example, the average should be nearly identical for all these images because they were captured sequentially under assumed similar illumination. In an illustrative example, the system can determine the average feature (e.g., a temporal feature) by generating a sequence of average metrics for each image, corresponding to a specific average metric at each location corresponding to more than one region. The system then determines the differences between the sequences of average metrics to generate the average feature. In some embodiments, the differences include, for example, the difference between the maximum and minimum values ​​of the average metric for each location or a reduced subset of locations. In some embodiments, the difference includes, for example, the variance of each location or a reduced subset of locations, such as the standard deviation (e.g., the mean relative to the mean measure).

[0047] Referring to figure 420, the system can determine one or more difference features, such as pixel absolute difference (PAD). The system can determine the difference as a purely temporal feature by capturing frame-to-frame changes in a scene occurring within very short time intervals (e.g., the reciprocal of the frame rate). For example, when considering two consecutive image frames, the absolute difference between the two frames (e.g., the difference in averages) can capture this difference. In an illustrative example, the system can determine the difference features by determining the average value of each region of a first image to generate a first set of average values, determining the average value of each region of a second image to generate a second set of average values ​​(e.g., the second image is temporally adjacent to the first image), and determining the difference between each average value in the first set of average values ​​and the corresponding average value in the second set of average values ​​(e.g., to generate an array of difference feature values).

[0048] Referring to Figure 410, the system can determine one or more scale features (SFs) (e.g., as spatial features) corresponding to a partitioned grid. The system can determine scale features to capture spatial variations across various length scales (e.g., regions of various sizes, corresponding to various numbers of pixels). The system can determine scale features by identifying small neighborhoods (e.g., windows) of each region (e.g., any particular pixel or group of pixels) and determining the range within that neighborhood (e.g., maximum value minus minimum value, or any other suitable metric indicating the range) to capture differences in the scene at that scale. The system then changes the window size and repeats the determination of the ranges for various window sizes. In an illustrative example, the system can determine scale features by determining a sequence of scale sizes at each location in each image, determining a range metric for each scale size in the sequence of scale sizes at each location to generate a set of range metrics, and determining the differences between the sets of range metrics. For example, the system can determine the range of the set of range metrics, such as the difference between the maximum and minimum values ​​in the set of range metrics. In another example, the system can identify one or more scales (e.g., feature values ​​corresponding to scale sizes).

[0049] Figure 430 illustrates a technique for determining dynamic range features. For example, dynamic range features may include pixel dynamic range (PDR), which is a temporal feature. This dynamic range feature captures activity occurring at locations with respect to time. In some embodiments, activity is captured by determining the minimum and maximum values ​​in the image set 400 at each location {i,j}. To illustrate, for each image set (e.g., image set 400), a single maximum and a single minimum value are determined for each location. In some embodiments, the dynamic range is determined as the difference between the maximum and minimum values ​​and indicates the amount of change that region undergoes within a time interval (e.g., corresponding to image set 400). To illustrate, if a region is occluded, then the difference between the maximum and minimum values ​​will not be relatively large. To further illustrate, dynamic range features can also help identify whether a region is occluded, particularly during nighttime when most image content is black. In some embodiments, the system may select all pixels in a region, or may subsample the pixels of a region. For example, in some cases, selecting fewer pixels still allows for sufficient performance. In an illustrative example, the system may determine the average value for each region of each image in an image sequence to generate a sequence of average values ​​for each location in a partitioned grid. The system then determines the difference between the maximum and minimum values ​​of the average sequence at each location of the partitioned grid (e.g., the partitioned grid may be subsampled to obtain a smaller or larger number of values ​​and resolution).

[0050] Referring to Figure 430, the system can determine one or more gradient features, also known as gradient dynamic range (GDR). While the system captures temporal variations when determining the gradient dynamic range (PDR) metric, GDR allows for the consideration of some spatial information. To capture spatial variations, the system uses any suitable technique to determine the image gradient (e.g., or other suitable difference operators), such as the Sobel operator (e.g., a 3x3 matrix operator), the Prewitt operator (e.g., a 3x3 matrix operator), the Laplacian operator (e.g., gradient divergence), the Gaussian gradient technique, any other suitable technique, or any combination thereof. The system determines the range of gradient values ​​at each region (e.g., any pixel location or group of pixels) over time (e.g., for an image set) to determine the variation in the gradient metric. Therefore, the gradient feature is a spatial-temporal feature. For illustration, gradient or spatial difference determines the capture of spatial variations, while the dynamic range component captures temporal variations. In an illustrative example, the system can determine gradient features by determining the gradient values ​​of each region of each image in an image sequence to generate a gradient value sequence for each location of a partitioned grid, and (for each corresponding gradient value sequence) determining the difference between the gradient values ​​of the corresponding gradient value sequences.

[0051] In some embodiments, calculating the maximum and minimum values ​​is computationally inexpensive and therefore can be used to determine (e.g., over time) ranges. Other values, such as the mean, standard deviation, variance, or others, could be used, but these might incur additional computational requirements compared to determining the minimum and maximum values ​​(e.g., their difference). For example, using the minimum and maximum values ​​can be very efficient (e.g., compared to other methods that require more computation).

[0052] Figure 5 A flowchart illustrating an illustrative process 500 for managing camera occlusion and response according to some embodiments of the present disclosure is shown. In some embodiments, process 500 is provided by, for example... Figure 3-4 The illustrative system and technology are system implementations of any of these. In some embodiments, process 500 is an application implemented on any suitable hardware and software, which may be integrated into the vehicle and communicate with the vehicle's systems, including mobile devices (e.g., smartphone applications), or combinations thereof.

[0053] At step 502, the system applies a partitioning grid to each image (e.g., of an image sequence). The partitioning grid includes a set of indexes that may correspond to multiple locations forming multiple regions in each image. The partitioning grid may include an array, a set of indexes, a mask, or a combination thereof, and may be a regular (e.g., rectangular array), an irregularly arranged set of indexes, an index set corresponding to irregularly sized regions (e.g., smaller regions at the periphery), overlapping regions, a set of indexes for spaced-out regions, or a combination thereof. For example, the partitioning grid defines the locations where each image is subdivided into regions, allowing each region to be analyzed. In another instance, the partitioning grid may include N x M locations, where N and M are positive integers greater than one, resulting in an array of regions. In another instance, in some embodiments, applying the partitioning grid includes retrieving or accessing pixel data for regions of one or more images in step 504. In some embodiments, each image in the image sequence is indexed temporally. For example, images may include video feed frames captured by a camera and stored as single images indexed in the order of capture (e.g., at a suitable frame rate). In an illustrative example, the system may apply a partitioned grid to each image to produce an array of regions for each image (e.g., an N*M product), and thus for K frames. In another illustrative example, for K frames and a specific feature (e.g., one or more features), the system may determine K*N*M values ​​corresponding to regions in each image (e.g., for each feature type, where more than one feature type may exist). In some embodiments, each region corresponds to a pixel, a set of pixels (e.g., AxB pixels in each region), or a combination thereof. In some embodiments, the partitioned grid includes an index array for indexing regions of an image or set of images.

[0054] At step 504, the system determines at least one spatial feature, at least one temporal feature, or a combination thereof corresponding to a partitioned grid. For illustration, step 504 can be... Figure 3 Feature extractor 310 is executed, and may include, for example Figure 4 Any descriptive feature described in the context. Features may include, for example, range features, gradient features, difference features, scale features, mean features, statistical features (e.g., mean, standard deviation, variance), any other suitable feature, or any combination thereof. For illustration, features may include one or more mean features (MF), one or more difference features, one or more scale features, one or more dynamic range features, one or more gradient features, maximum value, minimum value, any other suitable feature, or any combination thereof.

[0055] At step 506, the system generates a classification sequence for each of a plurality of locations based on the features from step 504. In some embodiments, for example, the classification sequence for each location is based on at least one spatial feature, at least one temporal feature, at least one combined feature, reference information, or a combination thereof. In some embodiments, the system trains a classifier to predict whether a given feature descriptor belongs to the “occluded” or “unoccluded” category. In some embodiments, the system implements the trained classifier to take feature values ​​as input and return a classification for each region (e.g., corresponding to a partitioned grid). For illustration, the output of the classifier may include an array corresponding to the partitioned grid, wherein each location in the array includes one of two values ​​(e.g., a binary system of occluded or unoccluded). In some embodiments, the system can generate a new set of classification values ​​corresponding to the partitioned grid at a rate the same as or slower than the frame rate (e.g., for every K images, such as ten images or any other suitable integer). In some embodiments, in response to an event (e.g., a trigger from an algorithm or controller) or a combination thereof, the system generates a new set of classification values ​​corresponding to the partitioned grid at a predetermined frequency.

[0056] At step 508, the system applies a smoothing technique to determine a subset of occluded regions in the classification sequence. In some embodiments, for example, the system implements Bayesian smoothing to smooth the classification. Because the classification in step 506 is discrete (e.g., outputting an occluded or unoccluded binary classifier for each region), the classification can exhibit fluctuations (e.g., when the classifier generates false positives). The system smooths the output of step 506 to address such fluctuations and establish confidence in any possible state changes. In some embodiments, the system monitors historical classifications (e.g., previous classifications for each region) and the current classification output of step 506. In some such embodiments, the system weights the historical and current values ​​to determine a smoothed output (e.g., which may be a state in the same state category as in step 506). For example, if the system determines a category in the occluded and unoccluded categories for each region at step 506, then at step 508, the system determines a smoothed classification value from the occluded or unoccluded categories. The system smooths the classifier output to establish confidence in the classification prediction before outputting a signal indicating a change in state.

[0057] At step 510, the system generates an output signal based on a subset of regions. For example, the system may determine the state change from unoccluded to occluded at step 508, and thus generate an output signal. In some embodiments, the system identifies all regions with occluded states, and collectively identifies this set of regions as an occlusion mask corresponding to an occlusion (e.g., the occluded area of ​​a camera). Figure 6The illustrative steps of process 600 provide an example of output based on smooth classification. The system may generate an output signal based on the state transition determined at step 508. The system may provide the output signal to, for example, an auxiliary system, an external system, a vehicle system, any other suitable system, its communication interface, or any combination thereof. In some embodiments, the system provides the output signal to a cleaning system (e.g., a washing system) to spray water or other liquids on the camera surface (e.g., or to activate a mechanical cleaner such as a wiper) to remove obstructions. In some embodiments, the system provides the output signal to a notification system or otherwise includes a notification system to generate a notification. For example, the notification may be displayed on a display screen, such as a smartphone touchscreen, a vehicle console screen, any other suitable screen, or any combination thereof. In another example, the notification may be provided as an LED light, a console icon, or other suitable visual indicator. In another example, a screen configured to provide video feeds from classified camera feeds may provide visual indicators, such as warning messages, a highlighted area corresponding to the obstructed video feed, any other suitable indication overlaid on the video or otherwise presented on the screen, or any combination thereof. In some embodiments, the system provides an output signal to the vehicle's imaging system. For example, the vehicle can receive images from multiple cameras to determine environmental information (e.g., road information, pedestrian information, traffic information, location information, path information, proximity information), and thus can change the way the images are processed in response to occlusion or lack of occlusion.

[0058] In an illustrative example, the system can implement process 500 as a feature extractor, a classifier, and a smoother. The feature extractor can determine various features (e.g., one or more features, such as five different feature descriptors extracted from a set of images). For illustration, the feature extractor creates feature descriptors for points corresponding to image regions. Feature descriptors (features) can include a set of values ​​(e.g., an array of digits) that “describe” the region undergoing classification. The classifier takes the extracted features as input and determines a label prediction for each region. The smoother (e.g., a Bayesian smoother) smooths the classifier prediction over time to mitigate false positives produced by the classifier, thereby increasing confidence in state changes (e.g., to improve detection performance).

[0059] In another illustrative example, process 500 may depend at least in part on the movement of the vehicle to predict whether pixels are occluded or not. In some embodiments, some temporal scene variations improve classification performance. For example, if the vehicle is stationary, feature descriptors may begin to produce fault values ​​or other values ​​that are difficult to characterize (e.g., dynamic values ​​will all be close to 0 due to the lack of change in the scene). In some embodiments, to address low scene variability, the system may integrate a vehicle state estimator to determine whether the vehicle is sufficiently moving.

[0060] Figure 6 A flowchart illustrating an illustrative process 600 for managing classifications according to some embodiments of the present disclosure is shown. In some embodiments, process 600 or aspects thereof may be combined with any illustrative steps of process 500.

[0061] At step 602, the system generates an output signal. For example, step 602 can be combined with... Figure 5 The process 500 is the same as step 510. The system can generate an output signal and provide the output signal to, for example, an auxiliary system, an external system, a vehicle system, a controller, any other suitable system, its communication interface, or any combination thereof.

[0062] At step 604, the system generates a notification. In some embodiments, the system provides an output signal to a display system to generate the notification. For example, the notification may be displayed on a display screen, such as a smartphone touchscreen, a vehicle console screen, any other suitable screen, or any combination thereof. In another instance, the notification may be provided as an LED light, a console icon, a visual indicator (e.g., a warning message), a highlighted area corresponding to the obstructed video feed, a message (e.g., a text message, an email message, an on-screen message), any other suitable visual or auditory indication, or any combination thereof. For illustration, diagram 650 shows a message overlaid on a display of a touchscreen (e.g., a smartphone or vehicle console) indicating that the right rear (RR) camera is 50% obstructed. To further illustrate, the notification may provide a user (e.g., a driver or vehicle occupant) with an indication to clear the camera, ignore the image from the camera, or otherwise account for the obstruction when considering the image from the camera.

[0063] At step 606, the system cleans the camera. In some embodiments, the system provides an output signal to a cleaning system (e.g., a washing system) to spray water or other liquid onto the camera surface (e.g., or activate a mechanical cleaner such as a wiper) to remove obstructions. In some embodiments, the output signal causes a wiper motor to reciprocate across the camera lens. In some embodiments, the output signal activates a liquid pump and pumps cleaning fluid (e.g., as a spray from a nozzle connected to the pump via a free pipe) toward the lens. In some embodiments, the output signal is received by a cleaning controller that controls the operation of the cleaning fluid pump, the wiper, or a combination thereof. For illustration, figure 660 shows a pump and a wiper configured to clean a camera lens. The pump sprays cleaning fluid toward the lens to remove or otherwise dissolve / soften obstructions, while the wiper rotates across the lens to mechanically remove the obstructions.

[0064] At step 608, the system modifies the image processing. In some embodiments, the system provides an output signal to the vehicle's imaging system. For example, the vehicle may receive images from multiple cameras to determine environmental information (e.g., road information, pedestrian information, traffic information, location information, path information, proximity information), and thus may change the way the images are processed in response to occlusion or lack of occlusion. For illustration, figure 670 shows an image processing module that takes images from four cameras as input (e.g., but any suitable number of cameras may include, for example, one, two, or more). As shown in figure 670, one of the four cameras is occluded (e.g., indicated by an "x"), while the other three cameras are not occluded (e.g., indicated by a checkmark). In some embodiments, the image processing module may ignore the output from the occluded camera, ignore a portion of the image from the occluded camera, reduce the weight or validity associated with the occluded camera, consider any other suitable modifications to the entire output of the occluded camera, or a combination thereof. Determining whether to modify image processing can be based on the degree of occlusion (e.g., the fraction of total pixels that are occluded), the shape of the occlusion (e.g., a large skewed aspect ratio, such as striped occlusion, is less likely to trigger modification than a more square aspect ratio), which camera is identified as displaying occlusion, daytime or nighttime, user preferences (e.g., included in reference information as a threshold or other reference), or a combination thereof.

[0065] In some embodiments, at step 608, the system ignores a portion of the camera output. For example, the system may ignore or otherwise exclude portions of the camera output corresponding to an occlusion mask during analysis. In another instance, the system may ignore a quarter, half, a sector, a window, any other suitable set of pixels with a predetermined shape, or any combination thereof, based on the occlusion mask (e.g., the system may map the occlusion mask to a predetermined shape, and then accordingly set and arrange the shape to indicate the portion of the camera output to be ignored).

[0066] Figure 7 A graph 700 illustrates the illustrative response of a smooth classifier according to some embodiments of the present disclosure. Figure 8 A graph 800 illustrates the illustrative relationship between time and smoothness metrics according to some embodiments of the present disclosure. Figure 9 A block diagram illustrating a smoothing technique according to some embodiments of the present disclosure is shown. In some embodiments, Figure 5 The process 500 may include Figure 8 Smoothing measures and Figure 9 The technique can be used to generate data for graph 700 (e.g., where the horizontal axis corresponds to time and the vertical axis corresponds to classification). Trace 701 represents the classifier output (e.g., from step 506 of process 500). Tracees 702 and 703 show the smoothed classification generated using process 500 (e.g., with different smoothing measures). When the classifier (e.g., Figure 3 When the classifier 320 changes its state from -1 (e.g., "occluded") to 1 (e.g., "unoccluded" or "normal") due to Bayesian smoothing, the smoothed classifier (e.g., Figure 3 The smoothing engine (330) does not immediately change the decision. The smoothing classifier delays state transitions for a period of time to build confidence in the true state of the region (e.g., corresponding to one or more pixels in that region). If the classifier repeatedly classifies a region as class "A," this slowly increases the confidence that the pixel's true label is class "A." Figure 7As shown, the smoothing classifier applies a non-uniform bias in the Bayesian smoothing technique. For example, to transition from state -1 to state 1, far fewer samples (e.g., frames or time) are required compared to transitioning from state 1 to state -1. This is because the classifier can be configured to consider the pixels most likely to be unoccluded (e.g., state 1 as shown). For example, if the classifier predicts an "unoccluded" label, it may be expected to quickly obtain confidence, while if a region is classified as occluded, it may be expected to wait longer before declaring that the region is occluded (e.g., transitioning to an "occluded" state). In some embodiments, each region may have associated parameter values ​​(e.g., α, β, or both), which may but do not have to be the same as other regions. For example, regions along edges, in the middle, at the top, at the bottom, or other regions corresponding to any other part of the camera may be associated with different parameter values.

[0067] like Figure 8-9 As shown, the parameter controlling the shape of the transition curve is "α" (also called "momentum"), which affects the time required for the state transition. The parameter α can be determined, for example, using... Figure 8 The formula shown in Figure 800 determines the value of α. The larger the value of α, the longer it takes to determine whether a complete change from the first state to the second state (e.g., occluded to unoccluded, or vice versa). In another instance, the parameter α can have different values ​​depending on the state at the time of classification. For illustration, if the current state is "unoccluded," the value of α can be greater than the value when the current state is "occluded," causing the response to take longer (e.g., more frames and more confidence) to transition to the "occluded" state. The trace shown in Figure 800 illustrates the parameter "β" over time. The parameter β can be determined based on empirical data, model fitting, functional relationships, or combinations thereof. The parameter α can be determined from β using, for example, the equation shown in Figure 800 or any other suitable relationship (e.g., another function, lookup table, or algorithm). The coefficients a, b, and c can be stored in reference information, can depend on the current state (e.g., occluded or unoccluded), and can include any suitable values. In the illustrative example, coefficients a, b, and c can have values ​​of -0.0028, 0.6962, and 1.9615, respectively. For example, see reference... Figure 8 In the graph 800, if a 60-second state transition elapsed time is required (e.g., before changing the predicted label), the smooth classifier can use a β value of 4.6.

[0068] like Figure 9As shown, parameter α is used to influence the "momentum" of the current state and thus the transition to another state. A sequence of block masks updated using parameter α at each time step t0, t1, t2, etc., is illustrated. Therefore, the classification value for any region of a particular mask can be a value corresponding to an intermediate value at a state or state transition. When the occlusion mask is updated, a smoothing classification approaches that classification and can eventually reach the classification value within the time interval (e.g., depending on the value of parameter β). Thus, using a smoothing classifier prevents false triggers or other fluctuations in classification from causing excessive state transitions. In some embodiments, parameter β is based on the time of day (e.g., a proxy for brightness). For example, a smoothing classifier may require less time to build confidence in the predicted label during the daytime compared to nighttime when the image may be less bright or display less contrast. The enhancement of the classifier output occurs over time in, for example, a non-linear manner to increase confidence and reduce false positives.

[0069] Figure 10 A block diagram illustrating illustrative images of occlusion and extraction features according to some embodiments of the present disclosure is shown. Image 1000 is subsequently captured from a camera displaying occlusion roughly corresponding to the left half. In some embodiments, the system may determine the extracted image generated by determining feature values ​​for each region or pixel and assigning visual indications to those values. For example, a maximum image 1010 and a minimum image 1011 may be determined by determining a minimum and a maximum value for each region or pixel, and then (i) reconstructing the maximum image 1010 by combining the maximum values ​​and (ii) reconstructing the minimum image 1011 by combining the minimum values. For illustration, the maximum image 1010 and the minimum image 1011 are not themselves corresponding images of image 1000, but rather constitute the pixels of image 1000 based on a minimum / maximum value classification for each region or pixel. The system may determine a difference image 1020 by subtracting the maximum image 1010 and the minimum image 1011 by region or by pixel. Similarly, difference image 1020 is not itself an image of image 1000, but is constructed by combining the differences between the maximum and minimum values ​​at each region or pixel (e.g., the maximum image 1010 and the minimum image 1011). For example, the left side of 1020 is shown as almost completely black, representing small differences in the corresponding region, while the right side of 1020 shows grayscale variations, representing larger differences and variations in the corresponding region. When occlusion is present, the expected differences between images are smaller in occluded regions because the variation in incident light on the camera is reduced, while when there is no occlusion, the expected image changes over time and therefore the expected differences are larger. Furthermore, the system may require some vehicle movement or changes in the captured images to distinguish between clockwise regions and static scenes (e.g., occlusion can be more easily identified if the vehicle and lighting are not static).

[0070] Figure 11 A block diagram illustrating an illustrative smoothing technique according to some embodiments of the present disclosure is shown. Figure 11 As shown, the parameter α is used to influence the "momentum" of the current state, and thus the transition to another state, such as... Figure 9 The same description is used in the context of [previous description]. A sequence of block masks updated using parameter α at each time step t0, t1, t2, etc., is shown. The top row corresponds to the classifier output (e.g., the instantaneous occlusion mask), and the bottom row corresponds to the smooth classifier output (e.g., the predicted occlusion mask). The occlusion mask is updated at each time step, where the smooth classification changes as the confidence in the state change increases. The enhancement of the classifier output occurs over time in, for example, a non-linear manner to increase confidence and reduce false positives. For example, as shown, regions that are classified the same as previously classified as later are more likely to change their classification as time progresses (e.g., from t0 to t1 to t2). Furthermore, as shown, regions near the boundary between occlusion and unocclusion (e.g., the boundary of the occlusion mask) exhibit relatively more value variation due to less consistent classifier values, temporal variation, or both.

[0071] Figure 12 A graphical representation of the illustrative response of a smoothing classifier according to some embodiments of the present disclosure is shown. Trace 1201 represents the classifier output (e.g., from step 506 of process 500). Trace 1202 illustrates the use of process 500 (e.g., using... Figure 13-14 The smooth classification is generated by the parameters described in the context. When the classifier (e.g., Figure 3 When the classifier 320 changes its state between -1 (e.g., "occluded") and 1 (e.g., "unoccluded" or "normal") due to Bayesian smoothing, the smoothed classifier (e.g., Figure 3 The smoothing engine (330) does not immediately change the decision. The smoothing classifier delays state transitions for a period of time to build confidence in the true state of the region (e.g., corresponding to one or more pixels in the region). Figure 12 As shown, the smoothing classifier applies a non-uniform bias in the Bayesian smoothing technique. For example, the number of samples (e.g., frames or time) required to transition from state -1 to state 1 is much smaller than the number of transitions from state 1 to state -1. Figure 12As illustrated in the example, the transition from -1 to 1 takes approximately 30 seconds, while the transition from 1 to -1 takes approximately 70 seconds. This is because the classifier can be configured to consider the pixels most likely to be unoccluded (e.g., state 1 as shown in the figure). For example, if the classifier predicts an "unoccluded" label, it might be expected to quickly obtain confidence, while if the region is classified as occluded, it might be expected to wait longer before declaring the region occluded (e.g., transitioning to an "occluded" state). In some embodiments, the difference in transition time can be achieved by using varying α and / or β values ​​(e.g., where the parameters are determined based on the current state).

[0072] Figure 13 A graph 1300 illustrates the illustrative response of a smooth classifier according to some embodiments of the present disclosure. Figure 14 A graph 1400 illustrates the illustrative relationship between time and smoothing metrics according to some embodiments of the present disclosure. In some embodiments, Figure 5 The process 500 may include Figure 14 The smoothing metric is used to generate data for graph 1300 (e.g., where the x-axis corresponds to time and the y-axis corresponds to classification). Traces 1301 and 1302 show smoothed classifications generated using process 500 (e.g., with different smoothing metrics). When the classifier (e.g., Figure 3 When the classifier 320 changes its state from -1 (e.g., "occluded") to 1 (e.g., "unoccluded" or "normal") due to Bayesian smoothing, the smoothed classifier (e.g., Figure 3 The smoothing engine 330 does not immediately change the decision. The smoothing classifier delays state transitions for a period of time to build confidence in the true state of the region (e.g., corresponding to one or more pixels in the region). As shown, trace 1301 exhibits a transition of approximately 25 seconds, while trace 1302 exhibits a transition of approximately 35 seconds, because each trace corresponds to... Figure 14 The different values ​​of β and α included. Figure 14 The expression for determining α and β is shown, where “previous prediction” corresponds to the previous classification value (e.g., depending on the current state).

[0073] Figure 15Graphical representations of two illustrative confidence measures according to some embodiments of this disclosure are shown. Graph 1500 shows a non-linear, entropy-based confidence measure, while graph 1550 shows a piecewise linear confidence measure. In some embodiments, the smoothed classifier value for each region can be used to determine the probability (e.g., the probability may be equal to the smoothed classifier value). In some embodiments, the probability value can be used to determine the confidence value. For example, the curves of graphs 1500 and 1550, any other suitable confidence measure relationship, or any combination thereof, can be used to determine the confidence value. For example, as shown, as the probability decreases below 0.5 or increases above 0.5, the confidence value increases (e.g., in opposite polarities). For example, higher probabilities above 0.50 tend to have higher confidence in one state, and lower probabilities below 0.50 tend to have higher confidence in another state.

[0074] The above description is merely an example of the principles of this disclosure, and various modifications can be made by those skilled in the art without departing from the scope of this disclosure. The above embodiments are presented for illustrative purposes and not for limitation. This disclosure can also take many forms besides those expressly described herein. Therefore, it is important to emphasize that this disclosure is not limited to the explicitly disclosed methods, systems, and devices, but is intended to include variations and modifications of the invention within the spirit of the following claims.

Claims

1. A method for determining camera occlusion, the method comprising: A partitioned grid, comprising multiple locations, is applied to each image in the image sequence to form multiple regions for each image, wherein each image in the image sequence is indexed in time; Determine at least one spatial feature corresponding to the partitioned grid and at least one temporal feature corresponding to the partitioned grid, wherein: One or more of at least one spatial feature or at least one temporal feature corresponds to a pixel feature, the feature value of which is compared with a dynamic value range, and The range of this dynamic value is modified based on the calculated changes in pixel values ​​in adjacent regions of the partitioned grid. Generate a binary classification sequence for each of the plurality of locations based on the at least one spatial feature and the at least one temporal feature; Applying smoothing techniques to the classification sequence to determine the occluded subset of regions; and The output signal is generated based on the subset of the region.

2. The method according to claim 1, wherein the at least one spatial feature includes a scale feature. And the determination of the scale features includes: At each location in each image, determine the sequence of scale sizes; At each location, determine a range metric for each scale size of the scale size sequence to generate a set of range metrics; and Determine the differences between the range metric sets.

3. The method of claim 1, wherein the at least one time feature comprises an average feature, and wherein determining the average feature comprises: For each image, the corresponding average metric at each location corresponding to more than one region is determined to generate an average metric sequence; as well as Determine the differences between the average metric sequences.

4. The method of claim 1, wherein the at least one time feature includes a difference feature, and wherein determining the difference feature comprises: Determine the average value of each region in the first image to generate a first set of average values; The average value of each region of the second image is determined to generate a second set of average values, wherein the second image is temporally adjacent to the first image; as well as Determine the difference between each average in the first set of averages and the corresponding average in the second set of averages.

5. The method of claim 1, wherein the at least one time feature includes a range feature, and wherein determining the range feature comprises: Determine the average value of each region in each image of the image sequence to generate an average value sequence for each location of the partitioned grid; as well as Determine the difference between the maximum and minimum values ​​of the average sequence at each location of the partitioned grid.

6. The method of claim 1, wherein the at least one time feature comprises a gradient feature, and wherein determining the gradient feature comprises: Determine the gradient value of each region in each image of the image sequence to generate a gradient value sequence at each location of the partitioned grid; as well as For each corresponding gradient value sequence, determine the difference between the gradient values ​​of the corresponding gradient value sequence.

7. The method of claim 1, wherein determining the at least one spatial feature and the at least one temporal feature comprises determining: Range characteristics; Gradient features; Difference characteristics; Scale characteristics; as well as Average characteristics.

8. The method of claim 1, wherein generating the binary classification sequence is further based on reference information including reference values, and wherein generating the output signal comprises: Determine the degree of occlusion; Determine whether the degree of occlusion exceeds the reference value; as well as If the occlusion exceeds the reference value, a response is identified, wherein the output signal indicates the response.

9. The method of claim 1, wherein the output signal is configured to cause the image processing module to ignore the output of the camera.

10. The method of claim 1, wherein generating the output signal includes generating a notification on the display device indicating the degree of occlusion.

11. The method of claim 1, wherein the output signal is used to modify image processing in response to the occluded subset of regions.

12. The method of claim 1, wherein applying the smoothing technique to the binary classification sequence to determine the subset of regions comprises: A smoothness metric is determined based on the current binary classification of each of the plurality of locations; A smoothed binary classification value sequence is determined based on the smoothness metric and the binary classification sequence; as well as A new binary classification is determined based on the smoothed binary classification value sequence.

13. A system for determining camera occlusion, the system comprising: A camera system for capturing image sequences; Control circuitry coupled to the camera system, wherein the control circuitry is configured to: A partitioned grid, comprising multiple locations, is applied to each image in the image sequence to form multiple regions for each image, wherein each image in the image sequence is indexed by time. Determine at least one spatial feature corresponding to the partitioned grid and at least one temporal feature corresponding to the partitioned grid, wherein: One or more of at least one spatial feature or at least one temporal feature corresponds to a pixel feature, the feature value of which is compared with a dynamic value range, and The range of this dynamic value is modified based on the calculated changes in pixel values ​​in adjacent regions of the partitioned grid. Based on the at least one spatial feature, the at least one temporal feature, and reference information, a binary classification sequence is generated for each of the plurality of locations, and Applying smoothing techniques to the classification sequence to determine the occluded subset of regions; and An output interface is configured to generate an output signal based on the subset of regions.

14. The system of claim 13, wherein the at least one spatial feature comprises a scale feature, and wherein, The control circuit is configured to determine the scale feature in the following manner: At each location in each image, determine the sequence of scale sizes; At each location, a range metric is determined for each scale size of the scale size sequence to generate a set of range metrics; as well as Determine the differences between the range metric sets.

15. The system of claim 13, wherein the at least one time feature comprises an average feature, and wherein, The control circuit is configured to determine the average characteristic in the following manner: For each image, the corresponding average metric at each location corresponding to more than one region is determined to generate an average metric sequence; as well as Determine the differences between the average metric sequences.

16. The system of claim 13, wherein the at least one time feature includes a difference feature, and wherein, The control circuit is configured to determine the difference characteristics in the following manner: Determine the average value of each region in the first image to generate a first set of average values; The average value of each region of the second image is determined to generate a second set of average values, wherein the second image is temporally adjacent to the first image; as well as Determine the difference between each average in the first set of averages and the corresponding average in the second set of averages.

17. The system of claim 13, wherein the at least one time feature includes a range feature, and wherein, The control circuit is configured to determine the range characteristics in the following manner: Determine the average value of each region in each image of the image sequence to generate an average value sequence for each location of the partitioned grid; as well as Determine the difference between the maximum and minimum values ​​of the average sequence at each location of the partitioned grid.

18. The system of claim 13, wherein the at least one temporal feature comprises a gradient feature, and wherein, The control circuit is configured to determine the gradient feature in the following manner: Determine the gradient value of each region in each image of the image sequence to generate a gradient value sequence at each location of the partitioned grid; as well as For each corresponding gradient value sequence, determine the difference between the gradient values ​​of the corresponding gradient value sequence.

19. The system of claim 13, wherein the control circuitry is configured to determine the at least one spatial feature and the at least one temporal feature by determining the following: Range characteristics; Gradient features; Difference characteristics; Scale characteristics; as well as Average characteristics.

20. The system of claim 13, wherein the control circuitry is further configured to generate the binary classification sequence based on reference information including reference values, and wherein the output interface is configured to generate the output signal in such a manner as follows: Determine the degree of occlusion; Determine whether the degree of occlusion exceeds the reference value; and If the occlusion exceeds the reference value, a response is identified, wherein the output signal indicates the response.

Citation Information

Patent Citations

  • Camera blockage detection for autonomous driving systems

    DE102018127738A1