Diverse redundancy for covering vehicle cameras
Patent Information
- Application Number
- EP2024714830
- Authority / Receiving Office
- EP · EP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-03-23
- Filing Date
- 2024-03-15
- Publication Date
- 2026-01-28
AI Technical Summary
Automated vehicle systems face inaccuracies and contamination issues due to pixel errors and local contamination in outward-facing vehicle cameras, which can lead to incorrect evaluation of sensor information and impact driver assistance and autonomous driving functions.
A redundant system combining frame-based and event-based cameras, where a computing unit generates and transforms camera images into a common format to detect deviations, utilizing machine learning methods and neural networks for error detection and adaptation, ensuring mutual monitoring and increased safety through diverse sensor properties.
The system effectively mitigates uncertainties and errors in visual environment detection, enhancing traffic safety by providing redundant protection and accurate data aggregation, especially in challenging lighting conditions, and enabling robust error detection and response.
Smart Images

Figure EP2024057067_26092024_PF_FP
Abstract
Description
[0001] Various redundancies to protect vehicle cameras
[0002] The invention relates to a system for a vehicle for visual environment detection, as well as a vehicle with such a system.
[0003] The following information is generally known prior art in the field of specialized cameras and therefore does not require citation of a specific prior art document: Frame-based cameras and event-based cameras are known. An event-based camera, also known as a neuromorphic camera, silicon retina, or dynamic vision sensor (DVS), comprises an image sensor that responds to local brightness changes. Event-based cameras do not capture images with a shutter like frame-based cameras. Instead, the pixels of an event-based camera are asynchronous and independent. They remain essentially silent but respond independently of each other to brightness changes in the environment of the event-based camera.For this purpose, each pixel stores a reference brightness and compares it with the new brightness level. If the brightness difference exceeds a threshold, this pixel sets its reference level to the new brightness value and generates a so-called "event." Event-based cameras therefore process local relative light intensity changes directly and send these information packets as events to the connected control devices. Here, these can be temporally and / or spatially aggregated into so-called event images. In contrast to frame-based cameras, event-based cameras typically only provide binary information about relative light intensity changes, optionally also providing information about the direction of change, i.e.An increase or decrease in relative light intensity changes, known as polarity, is present, but compared to frame-based cameras, they tend to have a higher temporal resolution and typically a higher dynamic range. This allows object movements in their surroundings to be tracked very precisely. WO 2022 / 258430 A1 relates in this context to a method for characterizing a dynamic image sensor (DVS), comprising: converting individual images of a scene generated by an image sensor into scene events based on conversion parameters that characterize the DVS.
[0004] Determining a difference between the converted events and events corresponding to the scene detected by the DVS; and adjusting the conversion parameters to reduce the difference between the converted and detected events.
[0005] The properties of event-based cameras mentioned above, particularly with regard to frame-based cameras, which are commonly used in the state of the art, can be used for automated driving applications, particularly for vehicles on public roads, to perform useful environmental detection. However, pixel errors and / or local contamination of outward-facing vehicle cameras can negatively impact vehicle functions, particularly driver assistance systems and (semi-)autonomous driving functions, as the sensor information from the vehicle cameras may then be incorrectly or at least incorrectly evaluated in the connected control units.
[0006] The object of the invention is to make visual environment detection of an automated vehicle more secure.
[0007] The invention is based on the features of the independent claims. Advantageous developments and refinements are the subject of the dependent claims.
[0008] A first aspect of the invention relates to a system for a vehicle for visual environmental detection, comprising a frame-based camera for generating a camera image and an event-based camera for generating event-based image information, and comprising a computing unit which is designed to generate an event image from image information of the event-based camera, and to transform at least one of the camera image and the event image for a common point in time of environmental detection by the frame-based camera and by the event-based camera such that the camera image and the event image associated with the point in time of environmental detection are present in a common format, and to compare the camera image and the event image in the common format for deviations from one another and to determine the presence of an error if a measure for the deviations is exceeded via a predetermined comparison condition.
[0009] Sensor elements of the event-based camera initially provide image information. This information is then converted into a respective event image. This conversion can take place within the event-based camera itself if it is appropriately equipped and has the necessary computing power—in this case, a computing module of the event-based camera is part of the computing unit. However, this conversion can also take place within the encapsulated computing unit of the event-based camera, functioning as a central control unit.
[0010] The computing unit is designed to generate an event image from image information from the event-based camera. In general, the respective "event images" are temporally aggregated event information (e.g., a histogram of all events or a temporal average of the events over a specific period of time for each pixel of the event-based camera, the polarity of the last event of each pixel, etc.). For this purpose, incoming image information with events is temporally and / or spatially aggregated into the respective event images in the computing unit. Local aggregation is an aggregation of events for individual pixel positions or regions around a pixel position, and temporal aggregation can weight and take into account the time difference between the event and the current time, e.g., with exponential temporal memory loss (e.g., so-called "eligibility traces"), etc.For example, with a temporal-spatial aggregation, a histogram can be created for each pixel of the event-based camera (e.g., 5 events at position xi, yi in the last 5 seconds), whereby the events of the event-based camera are summed over a fixed period of time dt (possibly time-weighted), thus generating an event image in the journals dt. For example, a temporal aggregation can be the number of all events of the event-based camera (e.g., 128 events in the last 5 seconds), whereby the events of the event-based camera are recorded over a fixed period of time dt (possibly time-weighted). For example, a spatial aggregation can be the number of all events of the event-based camera at a position of the sensor of the event-based camera or a region of the sensor at the current time t.Furthermore, in a temporal-spatial aggregation, a function / map can be trained and used with machine learning methods to aggregate a number of recent events and their times on each pixel.
[0011] In the context of this document, the term visual environment detection also includes visual interior detection of a vehicle interior.
[0012] According to the invention, two different systems are used for the visual detection of the surroundings of an automated vehicle, namely an event-based camera and a frame-based camera, which enable the same or at least similar functions for the automated vehicle (e.g., object detection, gesture recognition, etc.) and can thus be designed redundantly. The event-based camera typically processes local relative light intensity changes directly and asynchronously, which are aggregated to form the event images. The frame-based camera (e.g., CCD, APS), on the other hand, typically processes the information about the light intensity on the sensor on an image- or frame-based basis with a fixed or adaptive frame rate. Mutual function monitoring of the other camera system is possible, however, since the sensor information can be mutually translated or mapped, at least approximately.
[0013] An advantageous effect of the invention is that uncertainties such as pixel errors and / or local contamination, and thus limitations in the visual detection of a vehicle's surroundings, can be detected and mitigated early on through diverse redundancy. This is achieved by securing an associated surroundings detection system through the use of different (so-called "diverse") visual detection methods. Traffic safety is thus increased through mutual redundant protection of cameras with systemically different sensor properties. The particular advantage of mutual monitoring of a frame-based camera and an event-based camera lies in the presence of diverse redundancy, which, in contrast to identical redundancy, offers significantly increased safety.Identical redundancy would be present if, for example, two frame-based cameras or two event-based cameras were used and their results were compared. While this would allow for simpler implementation, as no transformation between the different sensor results would be necessary, it is precisely the diversity that is exploited in the diverse redundancy to visually capture the surroundings and / or the interior of the vehicle using systemically different approaches, and thus to better detect system-related errors in one of the sensor types. Especially with the combination of a frame-based camera and an event-based camera, not only is redundancy achieved through the use of different systems, there are also fundamentally different system properties with regard to the application on a vehicle, for example:The event-based camera has a higher temporal resolution and a lower latency and a higher bandwidth (the so-called "dynamic range"), i.e. in the present case of the application for environment recognition for a vehicle, situations with lighting conditions that are difficult to cover for frame-based cameras, such as tunnel driving, driving at night, driving in bright sunshine on an ice lake and a dazzling sun, can be easily captured by the event-based camera.
[0014] Possible uses for the visual information about the surroundings and / or the interior of the vehicle thus secured are:
[0015] - Object recognition and object detection;
[0016] - Semantic segmentation of the environment (pixel-accurate assignment of object identity, e.g. vehicle, road, tree, ... to a respective pixel)
[0017] - Depth determination (image-based estimation of the distance to objects in front of the vehicle)
[0018] - Driver monitoring (beat rate, pupils, direction of gaze, etc.);
[0019] It can also be planned that the frame-based camera is primarily used in lighting conditions that are good for it, and the event-based camera is used to ensure functionality in poor lighting conditions or when fast objects are moving (the relatively high temporal resolution of the event-based camera can be used here). In this scenario, both camera systems are largely redundant, but also complement each other at the edge of their capabilities.
[0020] According to an advantageous embodiment, the common format is the format of the camera image, and the computing unit is designed to transform the event image into the format of the camera image, or, the common format is the format of the event image, and the computing unit is designed to transform the camera image into the format of the event image.
[0021] Advantageously, an aggregated event image of the event-based camera is generated from a sequence of camera images of the frame-based camera by means of a corresponding transformation.
[0022] In particular, machine learning methods (e.g., using a convolutional neural network (CNN)) can be used to learn individual models that map the camera image or sequence of camera images from a frame-based camera installed in the vehicle into the aggregated event image of the event-based camera installed in the vehicle. For this purpose, a model is learned using training data from the camera images of a frame-based camera installed in the vehicle and the aggregated event image of the event-based camera installed in the vehicle using supervised learning. Alternatively or additionally, a mapping using a method known in the field for converting camera images using logarithmic scaling of the image intensity information, difference calculation, and threshold crossing can be used, for example, as a starting point for training the CNN.
[0023] In addition, generic models can be created during vehicle development, i.e., based on the types of frame-based / event-based cameras installed in the vehicle and their intended orientations within the vehicle. These generic models can be retrained in the user's vehicle, i.e., customized to account for or map variations in the cameras and their orientations, i.e., the individual cameras installed in the vehicle that may deviate slightly from the intended camera specifications and orientation.
[0024] It is also possible to generate one or more camera images (such as those from the frame-based camera) from aggregated event images of the event-based camera using a correspondingly different transformation. These reverse transformations for transforming an event image into a camera image of a frame-based camera are also known in the state of the art; see, for example, the publication: Rebecq et al. "High speed and high dynamic range video with an event camera" IEEE Transactions on Pattern Analysis and Machine Intelligence, vol. 43, no. 6, pp. 1964-1980, 2021.
[0025] The transformation of the sensor information from an event-based camera for comparability with the camera images from a frame-based camera is carried out as follows: As mentioned at the beginning, event-based cameras typically process local relative light intensity changes directly and forward these information packets to the corresponding modules of the processing unit. In the processing unit, incoming events from the event-based camera are, in particular, as explained above, temporally and / or spatially aggregated into event images. Frame-based cameras (e.g., of the "Charged Coupled Device" (CCD) or "Active Pixel Sensor" (APS) type) process the information about the light intensity on the sensor on an image- or frame-by-image basis with a fixed or, if necessary, adaptive frame rate. During the exposure of the incident light to this sensor, local light intensities are measured and integrated over this period.This enables mutual monitoring of the other camera's functions, as the sensor information can be mutually translated or mapped, at least approximately. Thus, both directions of mutual conversion are possible.
[0026] According to a further advantageous embodiment, the computing unit is designed to carry out two comparisons for deviations by means of transformation into two different common formats, wherein the first common format is the format of the camera image, and wherein the second common format is the format of the event image, wherein the computing unit is designed to determine the presence of an error case in the case of a) when at least one measure for the deviations of the comparisons is exceeded beyond the predetermined comparison condition, or alternatively b) when an average of the measures for the deviations, weighted across the comparisons, is exceeded beyond the predetermined comparison condition.
[0027] According to a further advantageous embodiment, the computing unit is designed to adapt the weightings of the weighted mean in the case of application of b) to the quality of the respective transformation used.
[0028] According to a further advantageous embodiment, the predefined comparison condition comprises a limit value. According to a further advantageous embodiment, the predefined comparison condition comprises a context-dependent limit value, wherein the computing unit is designed to determine the level of the limit value depending on the situation.
[0029] Preferably, the thresholds are context-dependent, for example to take into account the poorer dynamic range of the frame-based camera, e.g. the threshold is chosen higher in the dark than in daylight.
[0030] Furthermore, the comparison condition preferably comprises a respective sub-threshold value for each pixel individually, so that inaccuracies at the sensor edges or in areas with different detection areas FOV (abbreviation for "field of view") of the event-based camera and the frame-based camera can be taken into account, for example by setting the respective sub-threshold value of a pixel to infinity when the FOV does not overlap.
[0031] According to a further advantageous embodiment, the computing unit is designed to carry out a pairwise comparison of pixel information when comparing for deviations in order to be able to detect pixel errors and / or contamination of the frame-based camera and / or the event-based camera.
[0032] According to a further advantageous embodiment, the computing unit is designed to i) control an output unit of the vehicle to output a warning to the driver or user of the vehicle when the presence of an error is detected, and / or ii) control a vehicle subsystem using the respective camera image and / or the respective event image to switch off or switch to an emergency control program.
[0033] According to a further advantageous embodiment, the computing unit is designed to determine a mean square deviation of individual pixels and / or deviations in a limited local spread, e.g. weighted with a 2D Gaussian distribution centered around the individual pixels, for comparing the deviations.
[0034] According to a further advantageous embodiment, the computing unit is designed to apply a first artificial neural network for the transformation in the case of the transformation of the camera image of the frame-based camera into the format of the event image for the transformation, and to apply a second artificial neural network for the transformation in the case of the transformation of the event image into the format of the camera image.
[0035] According to a further advantageous embodiment, the computing unit is designed to aggregate detected deviations over a predetermined period of time. This ensures that spontaneous and gradual deterioration can be distinguished, allowing an error response to be adjusted accordingly, e.g., a lower threshold is used by the computing unit for spontaneous deterioration than for gradual deterioration.
[0036] Preferably, the processing unit does not react to the first threshold violation by detecting an error, but rather logs subsequent threshold violations and aggregates them (e.g., in a temporal histogram). The evaluation of the aggregated information can distinguish between gradual deterioration (e.g., pixel errors) and spontaneous deterioration (e.g., due to contamination), thus improving the robustness of error detection.
[0037] A further aspect of the invention relates to a vehicle with a system as described above and below.
[0038] Advantages and preferred developments of the proposed vehicle result from an analogous and analogous transfer of the statements made above in connection with the proposed system.
[0039] Further advantages, features, and details will become apparent from the following description, which – where appropriate with reference to the drawings – describes at least one embodiment in detail. Identical, similar, and / or functionally equivalent parts are provided with the same reference numerals.
[0040] It shows:
[0041] Fig. 1: A vehicle with a multi-redundant system for visual environment detection according to an embodiment of the invention. The illustrations in the figure are schematic and not to scale.
[0042] Fig. 1 shows a bird's-eye view of a vehicle with a system for visual environment detection, which has a frame-based camera 1 for generating a camera image and an event-based camera 3 for generating event-based image information. A computing unit 5 of the vehicle generates a respective event image from the image information of the event-based camera 3. This occurs at a predetermined frequency at repeated time intervals, so that a respective video is effectively generated by both the frame-based camera and the event-based camera. After the computing unit 5 has compiled the image information from the event-based camera into a respective event image, the computing unit 5 has access to an event image and a camera image, each relating to a respective one of many points in time during the ongoing detection of the vehicle's environment.In the next step, the processing unit 5 transforms the camera image and / or the event image so that both are in a comparable format. In the example of Fig. 1, this may occur twice and in both directions, so that the camera image is transformed once so that the camera image is in the format of the event image, and the event image is transformed once so that it is in the format of the camera image. This enables the processing unit 5 to compare the original event image with the camera image transformed into the format of the event image in a first comparison, and the original camera image with the event image transformed into the format of the camera image in a second comparison for deviations from each other.Thus, various sensor systems (frame-based cameras and event-based cameras) are used systematically to collect redundant environmental detection data, and these data are compared bidirectionally and pixel-by-pixel. If a certain pixel deviation exceeds a predefined comparison condition, the presence of an error is determined, and an appropriate response is initiated depending on which vehicle subsystem is using visual environmental detection. For non-safety-critical vehicle subsystems, a simple warning can be issued to the driver; for safety-critical systems, an emergency control mode of the vehicle subsystem in use can be activated in addition to the warning.Although the invention has been illustrated and explained in detail by preferred embodiments, the invention is not limited by the disclosed examples, and other variations may be derived therefrom by those skilled in the art without departing from the scope of the invention. It is therefore clear that a multitude of variations exist. It is also clear that exemplary embodiments are truly only examples and should not be construed as limiting the scope, possible applications, or configuration of the invention in any way.Rather, the preceding description and the description of the figures enable the person skilled in the art to implement the exemplary embodiments in concrete terms, whereby the person skilled in the art, with knowledge of the disclosed inventive concept, can make various changes, for example with regard to the function or the arrangement of individual elements mentioned in an exemplary embodiment, without departing from the scope of protection defined by the claims and their legal equivalents, such as further explanations in the description.
Claims
Patent claims 1. A system for a vehicle for visual environmental detection, comprising a frame-based camera (1) for generating a camera image and an event-based camera (3) for generating event-based image information, and comprising a computing unit (5) which is designed to generate an event image from image information of the event-based camera (3), and to transform at least one of the camera image and the event image for a common point in time of environmental detection by the frame-based camera (1) and the event-based camera (3) such that the camera image and the event image associated with the point in time of environmental detection are present in a common format, and to compare the camera image and the event image in the common format for deviations from one another and to determine the presence of an error if a measure for the deviations is exceeded above a predetermined comparison condition.
2. System according to claim 1, wherein the common format is the format of the camera image, and the computing unit (5) is designed to transform the event image into the format of the camera image, or, wherein the common format is the format of the event image, and the computing unit (5) is designed to transform the camera image into the format of the event image.
3. System according to claim 1, wherein the computing unit (5) is designed to carry out two comparisons for deviations by means of transformation into two different common formats, wherein the first common format is the format of the camera image, and wherein the second common format is the format of the event image, wherein the computing unit (5) is designed to determine the presence of an error case in the case of a) when at least one measure for the deviations of the comparisons exceeds the predetermined comparison condition, or alternatively tiv b) when a weighted average of the measures for the deviations over the specified comparison condition is exceeded.
4. System according to one of the preceding claims, wherein the predetermined comparison condition comprises a context-dependent limit value, wherein the computing unit (5) is designed to determine the level of the limit value depending on the situation.
5. System according to one of the preceding claims, wherein the computing unit (5) is designed to carry out a pairwise comparison of pixel information when comparing for deviations in order to be able to detect pixel errors and / or contamination of the frame-based camera (1) and / or the event-based camera (3).
6. System according to one of the preceding claims, wherein the computing unit (5) is designed, when the presence of an error is determined, to i) control an output unit of the vehicle to output a warning to the driver or user of the vehicle, and / or ii) control a vehicle subsystem using the respective camera image and / or the respective event image to switch off or switch to an emergency control program.
7. System according to one of the preceding claims, wherein the computing unit (5) is designed to determine a mean square deviation of individual pixels and / or deviations in a limited local spread for comparing the deviations.
8. System according to one of the preceding claims, wherein the computing unit (5) is designed to apply a first artificial neural network for the transformation in the case of the transformation of the camera image of the frame-based camera (1) into the format of the event image for the transformation, and to apply a second artificial neural network for the transformation in the case of the transformation of the event image into the format of the camera image.
9. System according to one of the preceding claims, wherein the computing unit (5) is designed to aggregate determined deviations over a predetermined period of time.
10. Vehicle with a system according to one of the preceding claims.