Detection with improved image quality for recognizing obstacles in the surroundings of a vehicle
The method of generating fusion images with short exposure times and neural network processing addresses image quality issues in ADAS systems, improving sensitivity and reducing data transmission, ensuring effective ADAS performance under poor visibility.
Patent Information
- Application Number
- PCT/EP2025/063891
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-05-29
- Filing Date
- 2025-05-20
- Publication Date
- 2025-12-04
AI Technical Summary
Existing camera-based advanced driver-assistance systems (ADAS) face challenges in maintaining image quality under poor visibility conditions, such as at night, due to motion blur and low signal-to-noise ratio, especially when using long exposure times, and high data transmission requirements for image fusion methods.
A method involving capturing a sequence of images with short exposure times, selecting a predefined number of consecutive images, and generating a fusion image with a virtual exposure time that combines the advantages of short and long exposure times, using a neural network for spatial alignment and noise reduction, and transmitting only the fusion image, thereby reducing data bandwidth.
This approach enhances image quality and sensitivity under poor visibility conditions while minimizing motion blur and data transmission requirements, enabling accurate object detection and ADAS functionality.
Smart Images

Figure EP2025063891_04122025_PF_FP_ABST
Abstract
Description
[0001] Description
[0002] Detection with improved image quality for identifying obstacles in the vicinity of a vehicle.
[0003] The present invention relates to a method, in particular a computer-implemented method, for detecting an object in the environment of a moving vehicle using image data from a vehicle camera system with at least one camera, a camera system for carrying out the method according to the invention, as well as a computer program and a computer program product.
[0004] Today's advanced driver-assistance systems (ADAS) offer drivers a wide range of functions. These functions can be used to support the driver while maintaining control of the vehicle. Depending on the level of automation, fully automated driving is also possible. Examples of such ADAS functions include various methods for detecting objects or obstacles on the road, methods for detecting lane markings and / or keeping the vehicle in its lane, methods for detecting rain on the windshield, and methods for assisting with or performing parking maneuvers. These and other functionalities are often based, at least in part, on images captured by cameras mounted on the vehicle.
[0005] For many common ADAS functions, three-dimensional perception of the vehicle's surroundings is a necessary prerequisite. Sensor systems used for this task can include, for example, a camera system with at least one camera to capture the vehicle's environment. Camera-based sensors offer the advantage of high lateral resolution. In addition to monocular cameras, multi-camera systems, such as stereo camera systems, are frequently used. One problem with camera-based ADAS systems is handling images captured in poor lighting or visibility conditions, especially at night. The quality of the captured image data decreases significantly in such cases, with the consequence that information can no longer be extracted from the image data, or at least not reliably.This in turn leads to a limitation of the ADAS functions based on the image data received from the camera system.
[0006] To improve image quality in poor visibility conditions, such as bad weather or at night, camera systems with active illumination are used, among other methods. However, such solutions are often complex, both in terms of design and operation, and are also associated with comparatively high costs, high energy consumption, and an increased likelihood of technical failures. Another alternative is the use of cameras that operate in wavelength ranges outside the visible spectrum, such as…
[0007] Thermal imaging cameras or infrared cameras. These cameras often don't use CMOS technology for their image sensors, but rather, for example, an indium gallium arsenide-based sensor. Therefore, these cameras are comparatively expensive. Among CMOS-based image sensors, so-called SPAD sensors (single-photon avalanche diode) achieve particularly high sensitivity. These sensors produce binary images.
[0008] Furthermore, there are various approaches to optimizing different recording parameters. In a typical vehicle camera system with multiple cameras positioned at different locations relative to the vehicle, image noise in poor visibility conditions is one of the biggest problems. This problem can generally be addressed with a longer exposure time. A longer exposure time leads to an improved signal-to-noise ratio. However, long exposure times do not necessarily result in a significantly higher signal-to-noise ratio.
[0009] Exposure cues carry the risk of motion blur, especially when the vehicle is in motion. Short exposure times, on the other hand, are advantageous for moving objects and / or moving cameras and result in significantly less motion blur, but are detrimental in terms of the signal-to-noise ratio and the achievable image quality. Finding the best possible compromise between optimal exposure time and minimal motion blur can be challenging.
[0010] Another way to improve image quality in poor visibility conditions is through image processing, often computer-aided, or the use of special image acquisition techniques. One example is burst photography, in which a sequence of images is captured and combined. These methods often rely on explicit motion estimation and / or precise spatial alignment of the captured image areas relative to each other to avoid or minimize motion blur caused by image fusion. By capturing and combining a large number of images with short exposure times within a predefined period, the advantages of both long and short exposure times can be combined.However, transmitting a large number of images from one or more cameras to a computing unit may be problematic due to the finite bandwidth of commercially available cameras for data transmission.
[0011] The present invention is therefore based on the objective of providing a simple, efficient and cost-effective way to generate high-quality camera image data regardless of the prevailing lighting conditions.
[0012] This problem is solved by the method according to claim 1, the camera system according to claim 9, the computer program according to claim 14 and the computer program product according to claim 15.
[0013] With regard to the method, the problem underlying the invention is solved by a method, in particular a computer-implemented method, for improving the image quality of images from a camera system of a moving vehicle with at least one camera, wherein the method comprises the following process steps:
[0014] Capturing a sequence of images of the surroundings of the moving vehicle with a predefined exposure time,
[0015] Selection of a predefinable number of consecutive images from the sequence, and
[0016] Generating a fusion image with a virtual exposure time that corresponds to the sum of the predefinable exposure times of the predefinable number of consecutive images.
[0017] With regard to the configurable exposure time, it is preferred to specify or select the shortest possible exposure time, in particular exposure times of less than 10 ms or less than 5 ms. It is also preferred to select a high frame rate, in particular greater than 30 fps or 60 fps.
[0018] The fusion image exhibits a significantly longer virtual exposure time, for example in the range of 20-30 ms. The virtual exposure time is not a conventional exposure time, meaning that a linear relationship with respect to the respective signal must be present: doubling the virtual exposure time therefore does not necessarily mean a doubling of the signal in the resulting image.
[0019] The fusion image thus combines the advantages of short and long exposure times, resulting in high light sensitivity with minimal motion blur. The virtual exposure time preferably corresponds to a high temporal fill factor. Within the scope of the present invention, this is understood to mean a suitable combination of a frame rate and the predefinable exposure time. A high temporal fill factor corresponds to high light sensitivity and can be achieved, for example, by a high frame rate with comparatively short predefinable exposure times. Preferably, the predefinable exposure time and the predefinable number of images are selected such that a high temporal fill factor is achievable. This leads to a higher light sensitivity of the camera system, particularly under poor visibility and / or lighting conditions.
[0020] The fusion image is advantageously used in a driver assistance system to implement an ADAS function, for example to detect at least one object in the vicinity of the moving vehicle.
[0021] An advantage of the method according to the invention is that only the fusion image, and not the entire predefinable number of image data in the sequence, needs to be transmitted. Rather, a fusion image only needs to be generated every n input images, where n corresponds to the predefinable number of images in the sequence.
[0022] This results in a significantly lower required data bandwidth for transmitting the images needed to implement various ADAS functions from the camera system to a central processing unit. The respective ADAS function is then performed using the fused images, which are characterized by a long virtual exposure time compared to the predefined exposure time, and consequently a high temporal fill factor, without any motion blur. This significantly improves image quality and light sensitivity, which is particularly advantageous when using the camera system in poor visibility conditions, such as at night.
[0023] In contrast to conventional serial photography, in the present invention the vehicle is in motion, so that both the movement of the vehicle and of moving objects that may be included in the images can be taken into account.
[0024] The generation of the fusion image can be carried out in many different ways, all of which fall within the scope of the present invention. For example, the individual images used to generate a fusion image can be combined, in whole or in part, for example by averaging or summing them.
[0025] According to an advantageous embodiment, generating the fusion image with the virtual exposure time includes spatial alignment and integration of a predefined number of consecutive images or of sub-areas of a predefined number of consecutive images into the fusion image. The spatial alignment is preferably performed with respect to both lateral and longitudinal movements. The spatial alignment and / or integration can be performed uniformly for the entire image, or it can be performed specifically and differently for different image regions (regions of interest).
[0026] According to another preferred embodiment, the specified number of consecutive images is provided as input to a trained neural network, which is designed to generate the fusion image based on this input. The neural network is thus designed to combine a large number of images to create the respective, improved fusion image.
[0027] The neural network is preferably a convolutional neural network.
[0028] During a training phase, the neural network can be trained, for example, using training data in the form of images for which the respective source images are enhanced with image noise and / or motion blur. The neural network is then trained to reconstruct the source images from the training data.
[0029] With regard to the neural network, it is advantageous if the network is designed to replace image pixels or image areas with high noise levels with learned, low-noise image pixels or image areas. This involves a motion-compensated accumulation and / or substitution of patches.
[0030] It is also advantageous to create a fusion image corresponding to a future point in time. In other words, a fusion image is created that, relative to the last captured or selected image in the sequence, represents a future point in time.
[0031] For this purpose, the network can be instructed during a training phase to calculate the fusion image according to the future time. A reinforcement learning process can be used in this context.
[0032] The calculated future fusion image can then be compared, for example, with the image taken by the camera system at the future time or with the corresponding fusion image.
[0033] In this context, it is advantageous to calculate a difference between the future fusion image and the current fusion image, or to generate a corresponding difference image. The fusion image is the fusion image that corresponds to the time of the calculated future fusion image. Based on the difference and / or the difference image, sudden events, such as fast-moving objects like a person suddenly jumping out from behind a stationary car, can be detected.
[0034] According to an advantageous embodiment of the method according to the invention, the most recently acquired images of the sequence are selected to generate the fusion image. This ensures that the current environmental situation of the vehicle's surroundings is always considered, i.e., that the current location of objects in the vehicle's environment is captured. A further embodiment involves using the most recently acquired image as the basis for generating the fusion image. It would actually be logical to use an image as the basis that has a capture time midway between the start and end capture times of the first and last selected images, with respect to the predefined number of images. Instead, however, the image with the latest capture time is used as the basis. In this way, it can be achieved that the fusion image is based on the current or...The system determines the most up-to-date status regarding the vehicle's movement and / or any moving objects in its vicinity. This is beneficial for the accuracy of downstream ADAS functions and for reducing system latency.
[0035] A fusion image generated by the method according to the invention is preferably used for a computer vision application or for a driver assistance system. One or more fusion images and / or the difference image determined according to the invention can be used.
[0036] In particular, the fusion image can be used to implement an ADAS function of the driver assistance system, for example, object detection and / or recognition. Furthermore, the fusion images generated according to the invention can be advantageously used to implement autonomous vehicle operation. Autonomous vehicles must reliably and promptly detect objects located in front of the vehicle, regardless of prevailing visibility conditions.
[0037] Furthermore, fusion images can be used to reliably detect small objects such as tires or tire parts and other objects such as lost cargo, regardless of the prevailing visibility conditions.
[0038] Computer vision involves the processing and analysis of images captured by cameras in a variety of ways to understand their content or to extract various, particularly geometric, pieces of information. The problem underlying the invention is further solved by a camera system for a vehicle comprising at least one camera, typically with a lens and an image sensor, a storage unit for storing images captured by the camera, and a processing unit, wherein the camera system is configured to carry out the inventive method according to one of the described embodiments. The use of such a camera advantageously reduces the data stream to a central processing unit of the driver assistance system, since only fusion images or further images determined from fusion images need to be transmitted at any given time.
[0039] Regarding the camera system, it is advantageous if the storage unit is a ring buffer. The ring buffer can conveniently store the images necessary to generate a fusion image, and the earliest image captured at any given time can be replaced by a newer, more recent image.
[0040] It is also advantageous if the storage unit is a recursive storage unit.
[0041] Another design of the camera system involves the computing unit being a single-chip system or an image processor.
[0042] Finally, a further embodiment of the camera system according to the invention includes the use of a CMOS image sensor, a SPAD image sensor, or a quantum-based image sensor, in particular a JOT-based image sensor. A JOT-based image sensor is described, for example, in "Jot devices and the Quanta Image Sensor" by J Ma. et al., published in the 2014 IEEE International Electron Devices Meeting, doi: 10.1109 / IEDM.2014.7047021. The problem underlying the invention is also solved by a driver assistance system comprising a camera system according to the invention based on at least one of the described embodiments.
[0043] Furthermore, the problem underlying the invention is solved by a computer program with instructions which, when the computer program is executed by a computer, cause the computer to execute the inventive method according to at least one of the described embodiments, and by a computer program product on which the inventive computer program is stored.
[0044] The invention and its advantageous embodiments are described in more detail with reference to the following figures. These show:
[0045] Fig. 1 is a schematic drawing of a camera system according to the invention;
[0046] Fig. 2 shows a flowchart relating to the method according to the invention;
[0047] Fig. 3 shows two possible embodiments for the camera system according to the invention;
[0048] Fig. 4 shows the generation of a fusion image using a neural network;
[0049] Fig. 5 shows a possible training method for the neural network;
[0050] Fig. 6 shows the replacement of noisy image areas or pixels with learned low-noise image areas or pixels;
[0051] Fig. 7 shows the generation of a future fusion image.
[0052] In the figures, identical elements are always designated with the same reference numeral. Fig. 1 schematically shows a camera system 1 according to the invention, comprising a camera 2, a processing unit 3, and a storage unit 4. The camera system 1 is configured to generate a fusion image Fl according to the present method, as will be explained in more detail in connection with the following figures. The fusion image Fl is generated in the processing unit 3 of the camera system 1; that is, only the fusion image Fl is transmitted at any given time, for example, to a central processing unit 5 of a driver assistance system (not shown here). The transmission between the camera system 1 and the central processing unit 5 can be wireless or wired.
[0053] Within the scope of the present invention, both the movement of the vehicle and the movements of various objects in the vehicle's environment, which are recorded by the camera system 1, are considered. Especially under poor visibility conditions, longer exposure times T are typically required, which leads to undesirable motion blur in the recorded images I. Therefore, to improve the image quality of images I from the camera system 1 of the moving vehicle, a sequence S of images I is recorded at different times tm — ti and with a predefinable exposure time T, according to the invention.
[0054] From this sequence, a predefinable number of consecutive images l(tn)... I(ti) are selected, preferably the most recently acquired images I. Based on the selected images l(t- n )... I(ti) then generates a fusion image Fl(to). The fusion image Fl(to) has a virtual exposure time T.V which of the sum of the predefinable exposure times T of the predefinable number of images l(t- n )... I(ti) corresponds. It is therefore possible to create a fusion image Fl(to) with a long exposure time T. V, a high temporal fill factor and thus high light sensitivity, but without the occurrence of motion blur. Fig. 3 shows two advantageous embodiments for the camera system 1 according to the invention. In Fig. 3a, the camera system 1 has at least one camera 2, which has a lens 2a and an image sensor 2b. Images I captured by the camera 2 are stored in the storage unit 3, which is configured here as a ring buffer. A predefinable number A of images are transferred to the processing unit 4 to generate a fusion image Fl, which in turn can be further transmitted via the interface 6, for example to the central processing unit 5, which is not shown again here. In contrast to the embodiment in Fig. 3a, the storage unit 3 in Fig. 3b is configured as a recursive storage unit.
[0055] Numerous methods are possible for generating the fusion image Fl from the images I, all of which fall within the scope of the present invention. Without limiting the generality of the invention, the following description refers to the generation of a fusion image Fl using a neural network NN, which can, for example, be designed as a convolutional neural network (CNN). Such a neural network NN is shown in Fig. 4.
[0056] The neural network NN is given the predefinable number A of images I of the sequence S, here l(t- n The value l(ti) is provided as input. The neural network NN then calculates the fusion image l(to) as output based on this input. It is therefore designed to determine a fusion image Fl based on the predefined number A of images I as input variables.
[0057] Such a neural network (NN) can be trained in different ways. One possible training procedure is illustrated in Fig. 5. In a first step, suitable training data is generated by displaying a sequence S of images I, here l(t- n ) - l(to) is provided with a predefined noise. The result is the noisy sequence SN, here with lN(t- n) - lN(to). Each noisy image IN corresponds to a corresponding original image I. In a second step, the images IN of the noisy sequence SN are provided to the neural network NN as input. The neural network NN then uses this input to generate a fusion image Fl(to). In a third step, the generated fusion images Fl(to) are compared with corresponding images I of the non-noise sequence S. For example, the images I used for comparison can be correlated with respect to their acquisition time t. Based on this comparison, parameters of the neural network NN can then be optimized.
[0058] According to one embodiment of the method according to the invention, the neural network NN can further be configured to replace image pixels or image areas with high noise levels with learned, low-noise image pixels or image areas, as illustrated in Fig. 6. In this way, a further improvement in image quality can be achieved.
[0059] It is also possible to use the neural network NN to create a future fusion image Fl(t+i), Fl(t+ m ) to generate, as illustrated in Fig. 7. To accomplish this task, the neural network NN can be trained in a corresponding training phase to determine a fusion image Fl(to) based on the respective input (comparable to the training phase described in connection with Fig. 6). The neural network NN also calculates further fusion images Fl(t+i), Fl(t+ m), which, relative to the current time t, correspond to future times t+i, t+m. The determined future fusion images Fl(t+i), Fl(t+ m These images can then be compared with temporally correlated images of sequence S, and parameters of the neural network NN can be optimized based on this comparison. A corresponding training phase can, for example, be designed as reinforcement learning.
Claims
Patent claims 1. Method, in particular a computer-implemented method, for improving the image quality of images (I) of a camera system (1) of a moving vehicle with at least one camera (2), wherein the method comprises the following process steps: Recording a sequence (S) of images (I) of the surroundings of the moving vehicle with a predefinable exposure time (T), Selection of a predefinable number (A) of consecutive images (l(tn), ... I (ti )) of the sequence(S), and Generating a fusion image (Fl) with a virtual exposure time (T) V ), which corresponds to the sum of the predefinable exposure times (T) of the predefinable number (A) of successive images (I).
2. The method of claim 1, wherein the generation of the fusion image (Fl) with the virtual exposure time (T) V) a spatial orientation and an integration of the predefinable number of consecutive images (l(t- n ), ... I(ti)) or of subsets of the predefinable number of consecutive images ( I (t- n ), ... I (ti )) to the fusion image (Fl) includes.
3. Method according to claim 1, wherein the predefinable number of consecutive images (l(t- n ), ... I(ti)) is provided as input to a trained neural network (NN), which neural network (NN) is designed to generate the fusion image (Fl) based on the input.
4. Method according to claim 3, wherein the neural network (NN) is configured to replace image pixels or image areas with high noise with learned low-noise image pixels or image areas.
5. Method according to claim 3 or 4, where a fusion image (Fl(t+i), Fl(t+ m )) is generated according to a future point in time (t+i , t+m).
6. Method according to at least one of the preceding claims, wherein the most recently recorded images of the sequence (S) are selected to generate a fusion image (Fl).
7. Method according to at least one of the preceding claims, wherein the last recorded image serves as the basis for generating the fusion image (Fl).
8. Use of the fusion image (Fl) for a computer vision application or a driver assistance system.
9. Camera system (1) for a vehicle comprising at least one camera (2), a storage unit (3) for storing images (I) taken by means of the camera (2) and a computing unit (4), wherein the camera system (1) is configured to perform the method according to at least one of the preceding claims.
10. Camera system (1) according to claim 9, wherein the storage unit (3) is a ring buffer.
11. Camera system (1) according to claim 9 or 10, wherein the storage unit (3) is a recursive storage unit.
12. Camera system (1) according to one of claims 9-11, wherein the computing unit (4) is a single-chip system or an image processor.
13. Camera system (1) according to one of claims 9-12, wherein the image sensor is a CMOS image sensor, a SPAD image sensor or a quantum-based image sensor, in particular a JOT-based image sensor.
14. Computer program with instructions which, when executed by a computer, cause the computer to execute the method according to any one of claims 1-7.
15. Computer program product on which the computer program according to claim 14 is stored.