Visual environment detection system and vehicle
A dual-camera system with frame-based and event-based cameras and a computing unit for error detection improves safety in automated vehicles by ensuring redundancy and accurate sensor information evaluation.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-03-15
- Publication Date
- 2026-03-12
AI Technical Summary
Existing event-based cameras in automated vehicles are prone to pixel defects and local dirt, which can lead to inaccurate sensor information evaluation, affecting driver assistance and automated driving functions.
A vehicle visual environment detection system combining frame-based and event-based cameras, with a computing unit to generate and compare images in a common format, detecting deviations and errors to ensure redundancy and improve safety.
The system enhances safety by early recognition and mitigation of uncertainties through diversity redundancy, improving road safety by detecting errors in visual environment detection systems.
Smart Images

Figure 2026508782000001_ABST
Abstract
Description
[Technical Field]
[0001] The present invention relates to a vehicle system for visual environment detection, and to a vehicle equipped with such a system. [Background technology]
[0002] The following information is generally known prior art in the field of specialized cameras, and therefore does not require citation of specific prior art documents. Frame-based cameras and event-based cameras are known. Event-based cameras, also known as neuromorphic cameras, silicon retina cameras, or dynamic vision sensors (DVS), contain image sensors that respond to local brightness (luminance) changes. Unlike frame-based cameras, event-based cameras do not capture images using a shutter. Instead, the pixels of an event-based camera are asynchronous and independent. They are essentially silent (still), but react independently to changes in brightness around the event-based camera. To do this, each pixel stores a reference brightness and compares it with each new brightness level. If the brightness difference exceeds a threshold, the pixel sets its reference level to the new brightness value and generates a phenomenon, known as an "event." Thus, event-based cameras directly process relative changes in local light intensity and transmit these information packets as events to a connected control device. These can then be aggregated temporally and / or spatially into so-called event images. In contrast to frame-based cameras, event-based cameras typically provide only binary information about the relative change in light intensity, and optionally also provide information about the direction of the change, i.e., an increase or decrease in the relative change in light intensity, known as polarity. However, event-based cameras tend to have a higher temporal resolution and typically a wider dynamic range compared to frame-based cameras, which allows them to track the movement of objects in their environment very accurately.
[0003] In this context, WO 2022 / 258430 relates to a method for characterizing a dynamic image sensor (DVS), comprising transforming individual images of a scene produced by the image sensor into events of the scene based on transformation parameters that characterize the DVS, determining the difference between the transformed events and events corresponding to the scene recognized by the DVS, and adapting the transformation parameters to reduce the difference between the transformed events and the detected events.
[0004] The properties of the event-based cameras mentioned at the outset, taking into account the frame-based cameras commonly used in the prior art, can be used to perform purposeful environment detection, in particular for applications in automated driving of vehicles in public transport. However, pixel defects and / or local dirt on the outward-facing vehicle camera can have a negative effect on vehicle functions, in particular driver assistance systems and (semi-)automated driving functions, since the sensor information of the vehicle camera can possibly be incorrectly or at least inaccurately evaluated in the connected control unit. Summary of the Invention [Problem to be solved by the invention]
[0005] An object of the present invention is to make visual environment detection for automated vehicles safer. [Means for solving the problem]
[0006] The invention emerges from the features of the independent claims. Advantageous developments and embodiments are the subject of the dependent claims.
[0007] A first aspect of the present invention relates to a vehicle visual environment detection system comprising: a frame-based camera for generating camera images and an event-based camera for generating event-based image information; and a computing unit designed to generate an event image from the image information of the event-based camera, and to transform at least one of the camera image and the event image so that at a common point in time of environment detection by the frame-based camera and the event-based camera (the point in time when the surrounding environment is recognized simultaneously), the camera image and the event image related to the time of environment detection exist in a common format, and to compare the camera image and the event image in the common format for the presence or absence of deviations between them, and to determine the presence of an error if the degree of deviation exceeds a predetermined comparison condition.
[0008] The sensor elements of an event-based camera initially provide image information. This information is then converted into event images. This conversion can be performed by the event-based camera itself if it is adequately equipped and has sufficient computing power, in which case the event-based camera's computing module is part of the computing unit. However, this conversion can also be performed by a computing unit encapsulated and implemented within the event-based camera and functioning as a central control unit.
[0009] The computational unit is designed to generate event images from image information of an event-based camera. Generally, each "event image" is time-aggregated event information (e.g., a histogram of all events over a specific period for each pixel of the event-based camera, or the time-averaged value of events, the polarity of the last event for each pixel, etc.). To this end, the computational unit aggregates the input image information along with the events in time and / or space to generate each event image. Spatial aggregation is the aggregation of events for an individual pixel location or the region around a pixel location, while temporal aggregation can consider the time difference between the current time point t0 and the event, weighted by, for example, exponential temporal memory loss (e.g., so-called "eligibility trace"). For example, in temporal-spatial aggregation, a histogram of each pixel of the event-based camera can be created (e.g., events at location x1, y1 over the past 5 seconds), and the events of the event-based camera are summed over a fixed period dt (which may be time-weighted), thus generating an event image at time step dt. For example, temporal aggregation could be the total number of events in an event-based camera (e.g., 128 events in the last 5 seconds), where events are detected over a fixed period dt (which may be temporally weighted). Spatial aggregation could be the total number of events in an event-based camera at the sensor's location or within the sensor's region at the current time t0. Furthermore, for temporal and spatial aggregation, a function / image can be trained using machine learning techniques to aggregate the number of recent events on each pixel and their time points.
[0010] In this specification, the term visual environment detection also includes visual interior detection relating to the interior of a vehicle.
[0011] According to the present invention, two different types of systems for visual environment detection in automated vehicles, namely event-based cameras and frame-based cameras, are used, enabling the same or at least similar functions (e.g., object detection, gesture recognition, etc.) for automated vehicles, and thus can be designed redundantly. Event-based cameras typically process local relative changes in light intensity directly and asynchronously, which are aggregated into event images. In contrast, frame-based cameras (e.g., CCD, APS, etc.) typically process information about light intensity on the sensor in an image-based or frame-based manner using a fixed or adaptive frame rate. However, since the sensor information can be converted or mapped to each other, at least approximately, each camera system can monitor the interoperability of the other camera systems.
[0012] An advantageous effect of the present invention is that diversity redundancy allows for early recognition and mitigation of uncertainties, such as pixel defects and / or localized contamination, and thus limitations in the vehicle's visual environment detection, by protecting the associated environment detection systems through the use of different (so-called "diverse") visual environment detection techniques. This improves road safety through the systematic mutual redundancy of cameras with different sensor characteristics. A particular advantage of the mutual monitoring of frame-based and event-based cameras is the existence of diversity redundancy, which significantly improves safety. For example, identity redundancy occurs when two frame-based cameras or two event-based cameras are used and their results are compared with each other. This may enable simpler implementation because no conversion between the results of different sensors is required. It is precisely this heterogeneity that is exploited in diversity redundancy, allowing for systematically different approaches to visually detecting the vehicle's surroundings and / or interior, thereby enabling better recognition of errors caused by one of the sensor types. In particular, when combining frame-based and event-based cameras, not only is there redundancy due to the use of different systems, but in this case the system characteristics are fundamentally different for vehicular applications, so that, for example, event-based cameras have higher time resolution, lower latency, and wider bandwidth (so-called "dynamic range"), i.e., in this case of an application for vehicular environment recognition, the event-based camera can easily detect situations with exposure conditions that are difficult to meet with a frame-based camera, such as driving through a tunnel, driving at night, driving in bright sunlight on a frozen lake, and driving under glaring sunlight.
[0013] Possible uses of visual information about the environment and / or interior of a vehicle protected in this way include: -Object recognition and object detection, - Semantic segmentation of the surroundings (accurate pixel-by-pixel assignment of object IDs to each pixel, e.g. vehicles, roads, trees, etc.), -Depth detection (image-based estimation of the distance to objects in front of the vehicle), - Driver monitoring (blinking, pupils, gaze direction, etc.).
[0014] Furthermore, frame-based cameras can be primarily used in lighting conditions favorable to them, while event-based cameras can be intended to ensure functionality in lighting conditions unfavorable to frame-based cameras or when dealing with fast-moving objects (where the relatively higher temporal resolution of event-based cameras can be utilized). In this scenario, both camera systems are largely redundant but additionally complement each other at the limits of performance.
[0015] In a favorable embodiment, the common format is the camera image format, and the computing unit is designed to convert event images to the camera image format, or the common format is the event image format, and the computing unit is designed to convert camera images to the event image format.
[0016] Advantageously, the aggregated event image of an event-based camera is generated from a sequence of camera images of a frame-based camera through a corresponding transformation.
[0017] In particular, machine learning techniques (e.g., by applying a "Convolutional Neural Network," or CNN for short) can be used to learn an individual model that maps a camera image or a sequence of camera images from a frame-based camera mounted on a vehicle to an aggregated event image from an event-based camera mounted on a vehicle. To this end, the model is learned by supervised learning using training data from the camera images from the frame-based camera mounted on the vehicle and the aggregated event images from the event-based camera mounted on the vehicle. Alternatively or additionally, the images may be used for mapping using methods known in the art for transforming camera images by logarithmic scaling, differencing, and thresholding of image intensity information, for example, as a starting point for training a CNN.
[0018] Additionally, generic models can be created during vehicle development based on the type of frame-based / event-based cameras installed on the vehicle and their orientation on the vehicle. These generic models can be retrained or personalized on the user's vehicle to account for or map variations in cameras and camera orientations, i.e., individual cameras installed on the vehicle may deviate slightly from the intended camera specifications and orientations.
[0019] Furthermore, it is also possible to generate one or more camera images (from a frame-based camera) from the aggregated event images of the event-based camera using correspondingly different transformations. These inverse transformations for converting event images into camera images of a frame-based camera are also known in the prior art, see for example the publication Rebecq et al., "High speed and high dynamic range video with an event camera," IEEE Transactions on Pattern Analysis and Machine Intelligence, vol. 43, no. 6, pp. 1964-1980, 2021.
[0020] The sensor information of an event-based camera is converted, for example, so that it can be compared with the camera image of a frame-based camera as follows: As mentioned at the beginning, the event-based camera typically processes relative changes in local light intensity directly and transfers these information packets to a corresponding module of the computing unit. In the computing unit, the input events of the event-based camera are temporally and / or spatially aggregated into an event image, as described above. Frame-based cameras (e.g., charge-coupled devices (CCDs) or active pixel sensors (APSs)) process information about the light intensity on the sensor on an image or frame basis at a fixed or, in some cases, adaptive frame rate. During the exposure of the sensor to incident light, the local light intensity is measured and integrated over this period. This allows the sensor information to be converted or mapped, at least approximately, to one another, thereby enabling mutual functional monitoring of the respective cameras. Mutual conversion is therefore possible in both directions.
[0021] According to a further advantageous embodiment, the calculation unit is designed to carry out two comparisons for the presence or absence of deviations by conversion into two different common formats, the first common format being in the format of a camera image and the second common format being in the format of an event image, and the calculation unit is designed to determine the presence of an error occurrence if a) at least one degree of deviation of the comparison exceeds a predetermined comparison condition, or alternatively b) if the average value of the weighted deviation measures based on the comparison exceeds a predetermined comparison condition.
[0022] According to a further advantageous embodiment, the calculation unit is designed to adapt the weights of the weighted average when applying b) to the quality of the respective transformation used.
[0023] According to a further advantageous embodiment, the predetermined comparison condition comprises a limit value.
[0024] In a more advantageous embodiment, the predetermined comparison conditions include context-dependent limits, and the computational unit is designed to determine the level of the limits depending on the situation.
[0025] Preferably, the threshold is context-dependent to take into account, for example, that the dynamic range of a frame-based camera is not very good; for example, the threshold is selected to be higher in low light than in sunlight.
[0026] More preferably, the comparison conditions include respective sub-thresholds for each pixel individually, so that inaccuracies at the sensor edges or in areas with different detection regions FOV (abbreviation for "field of view") of the event-based camera and the frame-based camera can be taken into account, for example by setting the respective sub-thresholds for the pixels to infinity in non-overlapping FOVs.
[0027] In a more advantageous embodiment, the computing unit is designed to perform a pairwise comparison of pixel information in the case of a comparison for the presence or absence of deviations, so that it can detect pixel defects and / or stains in a frame-based camera and / or an event-based camera.
[0028] According to a further advantageous embodiment, the calculation unit is designed, if the presence of an error occurrence is determined, to i) control an output unit of the vehicle to issue a warning to a driver or user of the vehicle, and / or ii) control a vehicle subsystem that uses the respective camera image and / or the respective event image to shut down or switch over to an emergency control program.
[0029] In a more advantageous embodiment, the computing unit is designed to determine the mean squared deviation of individual pixels and / or the deviation in a limited local spread, weighted by, for example, a two-dimensional Gaussian distribution centered on individual pixels, in order to compare the deviations.
[0030] In a more advantageous embodiment, the computing unit is designed to use a first artificial neural network to convert camera images of a frame-based camera into event image format, and a second artificial neural network to convert event images into camera image format.
[0031] In a more advantageous embodiment, the computing unit is designed to aggregate the determined deviations over a predetermined period. This makes it possible to distinguish between spontaneous degradation and slow-progressing degradation, and to adapt the error response accordingly. For example, the computing unit uses a lower threshold for spontaneous degradation than for slow-progressing degradation.
[0032] Preferably, the computing unit does not react to the initial threshold exceedance by recognizing an error, but rather records and aggregates (e.g., into a time histogram) any further threshold exceedances. By evaluating the aggregated information, it is possible to distinguish between slowly progressing degradation (e.g., pixel defects) and spontaneously occurring degradation (e.g., due to dirt), thereby improving the robustness of error recognition.
[0033] Another aspect of the present invention relates to a vehicle equipped with the system described above and described below.
[0034] Advantageous and preferred developments of the proposed vehicle result from a similar and sensible application of the explanations given above in connection with the proposed system.
[0035] Other advantages, features and details will become apparent from the following description in which at least one embodiment is set forth in detail, with reference to the drawings where appropriate, and in which identical, similar and / or functionally identical parts are designated by the same reference numerals. [Brief explanation of the drawings]
[0036] [Figure 1] 1 is a diagram of a vehicle equipped with a diversely redundantly protected system for visual environment detection according to an exemplary embodiment of the present invention; DETAILED DESCRIPTION OF THE INVENTION
[0037] The illustrations in the figures are schematic and not to scale.
[0038] FIG. 1 shows a bird's-eye view of a vehicle equipped with a visual environment detection system having a frame-based camera 1 for generating camera images and an event-based camera 3 for generating event-based image information. A computation unit 5 on the vehicle generates respective event images from the image information of the event-based camera 3. This occurs at time intervals that are repeated at a predetermined frequency, so that respective images are effectively generated by both the frame-based camera and the event-based camera. After the image information of the event-based camera is organized into respective event images by the computation unit 5, the event image and camera image for each of many time points in the vehicle's continuous environment detection are each available to the computation unit 5. In a next step, the computation unit 5 converts the camera image and / or the event image so that both are in a comparable format. In the example of FIG. 1, this can be done twice in both directions: once to convert the camera image so that it is in the format of an event image, and again to convert the event image so that it is in the format of a camera image. The computing unit 5 can then compare the original event image with the converted camera image in a first comparison, and the converted event image in a second comparison, to see if there are any deviations between them. Thus, systematically different sensor systems (frame-based cameras and event-based cameras) are used for the redundant and diverse collection of environmental detection data, and these various collected data are checked against each other bidirectionally and pixel by pixel. If the degree of pixel deviation exceeds a predetermined comparison condition, an error is detected and a corresponding reaction is initiated depending on which vehicle subsystem uses visual environmental detection. In non-safety-critical vehicle subsystems, a simple warning can be issued to the vehicle driver, while in safety-critical systems, in addition to a warning, an emergency control mode of the vehicle subsystem used can be activated.
[0039] Although the present invention has been illustrated and described in detail with reference to preferred exemplary embodiments, the present invention is not limited to the disclosed examples, and those skilled in the art can derive other variations therefrom without departing from the scope of protection of the present invention. Therefore, it is clear that there are many possible variations. Likewise, it is clear that the exemplary embodiments are merely examples in nature and should not be understood in any way as limiting, for example, the scope of protection, applicability, or configuration of the present invention. Rather, the foregoing description and illustrations enable those skilled in the art to specifically implement the exemplary embodiments, and in so doing, those skilled in the art, having learned the disclosed inventive concept, can make various modifications, for example, with respect to the function or arrangement of individual elements listed in the exemplary embodiments, without departing from the scope of protection defined by the claims and their legal equivalents, e.g., the broad description in the specification. [Prior art documents] [Patent documents]
[0040] [Patent Document 1] International Publication No. 2022 / 258430
Claims
1. A vehicle visual environment detection system, A frame-based camera (1) for generating camera images, An event-based camera (3) for generating event-based image information, The system comprises a computing unit (5) configured to generate an event image from the image information of the event-based camera (3), The calculation unit (5) is designed to convert at least one of the camera image and the event image so that at a common point in time of environment detection by the frame-based camera (1) and the event-based camera (3), the camera image and the event image associated with the time of environment detection exist in a common format, and to compare the camera image and the event image in the common format for their deviation from each other, and to determine the presence of an error if the degree of the deviation exceeds a predetermined comparison condition.
2. The common format is the format of the camera image, and the computing unit (5) is designed to convert the event image to the format of the camera image, or the common format is the format of the event image, and the computing unit (5) is designed to convert the camera image to the format of the event image. The system of claim 1 .
3. The calculation unit (5) is designed to perform two comparisons of deviations by converting to two different common formats, the first common format being the format of the camera image and the second common format being the format of the event image, and the calculation unit (5) is designed to determine the presence of an error if a) at least one degree of the deviation in the comparison exceeds a predetermined comparison condition, or alternatively, b) the mean of a weighted deviation scale based on the comparison exceeds a predetermined comparison condition. The system of claim 1 .
4. The predetermined comparison conditions include context-dependent limits, and the calculation unit (5) is designed to determine the level of the limits depending on the situation. The system according to any one of claims 1 to 3.
5. 5. The system according to claim 1, wherein the calculation unit (5) is designed to detect pixel defects and / or contamination of the frame-based camera (1) and / or the event-based camera (3) by performing a pairwise comparison of pixel information in the case of the comparison for deviation.
6. the computing unit (5) is designed to, if it determines that an error has occurred, i) control an output unit of the vehicle to issue a warning to a driver or user of the vehicle, and / or ii) control a vehicle subsystem that uses the respective camera image and / or the respective event image to shut down or switch to an emergency control program, The system according to any one of claims 1 to 5.
7. the calculation unit (5) is designed to determine the mean square deviation of individual pixels and / or the deviation in a limited local area in order to compare said deviations, A system according to any one of claims 1 to 6.
8. the computing unit (5) is designed to apply a first artificial neural network when converting the camera image of the frame-based camera (1) into the format of the event image, and to apply a second artificial neural network when converting the event image into the format of the camera image; A system according to any one of claims 1 to 7.
9. said calculation unit (5) being designed to aggregate the determined deviations over a predetermined period of time; A system according to any one of claims 1 to 8.
10. A vehicle equipped with a system according to any one of claims 1 to 9.
Citation Information
Patent Citations
Method and apparatus for characterizing a dynamic vision sensor
WO2022258430A1