Method for representing an additional graphic information in an image of an environmental region of a motor vehicle, electronic vehicle guidance system for a motor vehicle, motor vehicle, and computer program product
Semantic segmentation classifies image pixels to allow overlays only on drivable surfaces, ensuring obstacles remain visible, addressing the issue of obscured obstacles in vehicle environments.
Patent Information
- Application Number
- PCT/EP2025/070828
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-07-31
- Filing Date
- 2025-07-21
- Publication Date
- 2026-02-05
AI Technical Summary
Existing methods for representing additional graphic information in vehicle environments often fail to ensure that obstacles are clearly visible, as the information overlays can obscure them, especially when intensity is increased for visibility, leading to potential safety hazards.
Implementing semantic segmentation algorithms to classify image pixels into drivable and non-drivable surfaces, allowing overlays only on drivable surfaces and adapting or interrupting overlays based on pixel classification to maintain obstacle visibility.
Ensures that obstacles are not obscured by additional graphic information, enhancing safety and user friendliness by maintaining clear visibility of critical objects in the vehicle environment.
Smart Images

Figure EP2025070828_05022026_PF_FP_ABST
Abstract
Description
[0001] Method for representing an additional graphic information in an image of an environmental region of a motor vehicle, electronic vehicle guidance system for a motor vehicle, motor vehicle, and computer program product
[0002] The invention relates to a method for representing an additional graphic information in an image of an environmental region of a motor vehicle, wherein the image contains image contents that are represented by pixels on a display surface of a display device of the motor vehicle. According to the method described here, the additional graphic information is represented as an overlay on top of the image contents of the image. Accordingly, the additional graphic information may be an augmented image content. Further aspects of the invention relate to an electronic vehicle guidance system for a motor vehicle, a motor vehicle with such an electronic vehicle guidance system, a computer program, as well as a computer-readable storage medium with such a computer program.
[0003] The representation of overlaid image contents or overlays is well-known from the technical field of motor vehicles. They are for example used in the technical field of parking assistance systems to overlay on top of an image of an environmental region of a motor vehicle additional graphic information, which for example represent a driving path, which the vehicle will drive along with the current angle of the steering wheel. When changing the angle of the steering wheel, the driving path is correspondingly adapted by pivoting for instance from the left to the right. A driver of the motor vehicle may thus anticipatorily recognize, where his motor vehicle is going to drive, when he keeps the current angle of the steering wheel. If there is an obstacle in the region of the driving path, the driver intuitively knows that he has to change the angle of the steering wheel, if he wants to avoid a collision with the obstacle, for instance another vehicle.
[0004] Here, the problem may arise that the overlaid additional graphic information, for instance the driving path, frequently does not sufficiently visually stand out against other image contents of the image of the environmental region. In this case the additional information may be difficult to recognize for the driver. To solve this problem, WO 2023 / 007220 A1 envisages to recognize the visual similarity between the additional graphic information and a background of the image and to adapt the properties of the additional graphic information, for example the color and / or the intensity, accordingly. Nevertheless, it may happen that the additional graphic information is represented on top of objects or obstacles in the environmental region and thereby covers the same. Especially when for example the intensity of the representation of the additional graphic information is increased to make it easier to distinguish from the background, an overlaid or covered obstacle may no longer be recognizable for the driver in the image at all. The driver may then no longer reliably recognize whether or not he can safely drive in a certain region in the environmental region with his motor vehicle.
[0005] An objective of the present invention therefore consists in improving the representation of overlaid additional graphic information in a motor vehicle, in particular with regard to the visibility of obstacles in the environmental region of the motor vehicle.
[0006] The objective is solved by the subject matters of the independent claims. Advantageous further developments of the invention are described by the dependent claims, the description, and the figures.
[0007] The invention relates to a method for representing an additional graphic information in an image of an environmental region of a motor vehicle. The environmental region therein may comprise an area surrounding the motor vehicle, preferably an area of few meters around the motor vehicle. The environmental region may also comprise a region behind the motor vehicle, wherein this relates in particular to a region that may be captured by a rear view camera of the motor vehicle, that is a camera arranged in the rear part of the motor vehicle. The image contains image contents describing or representing the environmental region. The represented image contents may for example include static objects, in particular a road surface, road markings, curb stone edges, plants, or the like. The represented image contents may alternatively or additionally include movable or dynamic objects, in particular other moving vehicles, pedestrians, or the like. The image may be represented as real image of at least one on-board camera on a display surface of a display device of the motor vehicle. The image may also be created as virtual view from a plurality of real images of a multitude of on-board cameras. Therein, the virtual view may be calculated as a view from any virtual camera perspective.
[0008] The image contents, static or dynamic, therein are represented by pixels on the display surface of the display device of a motor vehicle. In other words, the image is a pixel-based representation of the environmental region, which comprises various image contents. In other words, the image contents are represented by pixels, wherein each pixel has characteristic pixel values, which describe the represented image content. A pixel value is for example a color value and / or an intensity value of the respective pixel.
[0009] The additional graphic information may be understood as additional image content, that is an image content that is not originally contained in the image, but rather may be added retrospectively in the sense of an augmentation to the image. This may be the driving path described in the above. The additional graphic information, however, may also include any conceivable additional information, for example an information relating to a building in the environmental region, which is represented in the image, or the like. Therein, the additional graphic information is represented as overlay on top of the image contents of the image, for instance in the manner of an AR image content (AR - Augmented Reality). Preferably, the additional graphic information moves along with the overlaid image contents, when the motor vehicle moves in the environmental region and / or when an object in the environmental region, which is added by the additional information, moves relative to the motor vehicle.
[0010] According to the invention the image contents of the image are classified on a pixel level or pixelwise, wherein each pixel of the image is assigned to a predetermined object class depending on the image content represented by the pixel and / or dpending on a color value assigned to the pixel. This pixelwise classification is known as semantic segmentation from the field of computer vision, that is from the field of computer-assisted visual perception. The semantic segmentation is performed by a semantic segmentation algorithm, that is by an algorithm for automatic visual perception.
[0011] Algorithms for automatic visual perception, which may also be referred to as computer vision algorithms, algorithms for machine vision, or machine vision algorithms, may be regarded as computer algorithms for automatically performing a visual perception task. A visual perception task, which is also referred to as computer vision task, may for example be understood as a task for extracting visual information from image data. In particular, the visual perception task in some cases may in principle be performed by a human capable of visually perceiving an image corresponding to the image data. In the present context, visual perception tasks, however, are also automatically performed, without the assistance through a human being required.
[0012] A computer vision algorithm may for instance comprise an image processing algorithm or an algorithm for image analysis, which is or was trained by machine learning and for example may be based on an artificial neural network, in particular a convolutional neural network. The computer vision algorithm may for instance comprise an object detection algorithm, an obstacle detection algorithm, an object tracking algorithm, a classification algorithm, a semantic segmentation algorithm, and / or a depth estimation algorithm.
[0013] Corresponding algorithms may in analogy also be executed on other input data than images that are visually perceivable by a human. For example, also point clouds or images of infrared cameras, lidar systems et cetera may be evaluated by correspondingly adapted computer algorithms. Strictly speaking, the corresponding algorithms are not algorithms for visual perception, since the corresponding sensors may work in the regions that are not visually, that is for the human eye, perceivable, for instance in the infrared region. For this reason such algorithms within the scope of the present invention are referred to as algorithms for automatic perception. Algorithms for automatic perception thus include algorithms for automatic visual perception, whilst not being restricted to the same with regard to a human perception. Consequently, an algorithm for automatic perception in this sense may also contain a computer algorithm for automatic performance of a perception task, which for example is or was trained by machine learning and in particular may be based on an artificial neural network. Such generalized algorithms for automatic perception may also include object detection algorithms, object tracking algorithms, classification algorithms and / or segmentation algorithms, for example semantic segmentation algorithms.
[0014] In case an artificial neural network for implementing an algorithm is employed for automatic visual perception, a frequently employed architecture is that of a convolutional neural network, CNN. In particular a 2D-CNN may be applied to corresponding 2D camera images. Also for other algorithms for automatic perception CNNs may be employed. For example, 3D-CNNs, 2D-CNNs, or 1 D-CNNs may be applied to point clouds, depending on the spatial dimensions of the point cloud and the details of the processing.
[0015] The result or the output of an algorithm for automatic perception is dependent on the specific underlying perception task. For instance, the output of an object algorithm may contain one or several bounding boxes, which define a spatial position and optionally an orientation of one or several corresponding objects in the environment and / or corresponding object classes for the one or several objects. An output of a semantic segmentation algorithm, which is applied to a camera image, may contain an object class on pixel level for each pixel of the camera image. In analogy, an objective of a semantic segmentation algorithm, which is applied to a point cloud, may contain a corresponding point level class for each of the points. The classes on pixel level or on point level may for example define an object type, to which the respective pixel or point pertains.
[0016] If the image is a distorted representation, for example a perspective representation of the environmental region from the perspective of a virtual camera, such as it is known from surround view systems in the field of automobiles, the semantic segmentation is preferably performed prior to a rectification of the image. The segmentation for example in an image taken by a fisheye camera is performed prior to the so-called de-fishing of the image.
[0017] After the semantic segmentation an assigning of a pixel value to a respective pixel in dependence on the object class, to which the respective pixel was assigned, is effected. The pixel value may be a virtual depth value, which is then stored in a depth buffer. In a simple case, two object classes may be provided, into which all pixels of the image may be classified, wherein one of the object classes comprises a drivable surface and the other one of the object classes comprises all other objects in the environmental region. A surface is defined as drivable by the motor vehicle, if it has mechanical properties allowing for a motor vehicle to drive on it. In other words, the surface should be robust enough to bear a motor vehicle that has a weight of several hundred kilograms. Alternatively or additionally, the surface that is drivable by the motor vehicle comprises regions, on which a motor vehicle is allowed to drive.
[0018] In this exemplary case, to all pixels that are classified into the object class “drivable surface” the pixel value “1” may be assigned. To all other pixels of the image the pixel value “0” may be assigned. Thus, a binary pixel value map of the environmental region may be generated. Since the pixel values are virtual depth values, this pixel value map may be referred to as depth map.
[0019] A next method step comprises the determining of the pixel values of the pixels within a current image region of the image to be overlaid by the additional graphic information. In other words, it is determined in which image region of the image the additional graphic information should currently be represented. This image region may migrate or shift in the case of a movement of the motor vehicle in the environmental region and / or a movement of the image contents with regard to the motor vehicle. Such shifting may lead to pixels with other pixel values than the originally present ones migrating into and / or out of the image region to be overlaid. Therefore, the method step to be described here is preferably repeated according to a predetermined update instruction at predetermined temporal and / or spatial intervals. To stick with the above-described example, this means that it is determined whether in the image region to be overlaid pixels with the values “0” and / or “1” are present. This may be effected by scanning the pixels in the image region to be overlaid.
[0020] Finally, an overlaying of the additional graphic information is effected only on top of such pixels within the image region to be currently overlaid, the pixel values of which lie within a predetermined range of pixel values, which consequently are assigned to certain object classes. For the above example this may mean that the additional graphic information is only overlaid on top of pixels, the pixel value of which is “1” for “drivable surface”. On top of all other pixels, that is for example those with the pixel value “0”, the overlaying may be omitted or be effected in an adapted form.
[0021] By the invention the advantage arises that an overlaying of additional graphic information on top of determined image regions is avoided and is admitted only on top of other image regions. Thus, it may be avoided that a driver may fail to perceive obstacles in the environmental region or perceive them only incompletely because in the image which is provided to him by the display surface in the motor vehicle they are covered by additional graphic information.
[0022] The invention also comprises embodiments which entail additional advantages.
[0023] An embodiment envisages that for each pixel within the image region of the image that is currently to be overlaid by the additional graphic information the pixel value is retrieved from the depth buffer and compared with the predetermined range of pixel values. In other words, for each pixel within the region to be overlaid by the additional graphic information it is checked whether the pixel value assigned to it lies in the predetermined value range or not. For this purpose the pixel values may simply be compared with the pixel value range. For the comparison the pixel values of the pixels are preferably retrieved in the sequence from the depth buffer, in which the pixels are to be overlaid by the additional graphic information. This means it is gradually checked for each pixel of the overlay or the additional graphic information to be represented whether the pixel value of the pixel of the image, which is to be overlaid by the pixel of the overlay, lies within the predetermined range of pixel values.
[0024] For the semantic segmentation preferably at least one first object class with an assigned first range of pixel values and a second object class with an assigned second range of pixel values are defined as object classes, wherein the first object class comprises surface drivable by the motor vehicle, that is for example the object class “road”, and wherein the second object class comprises surface not drivable by the motor vehicle, that is for instance the object class “everything except road”.
[0025] By this binary classification into object classes, very quickly and with relatively little computing effort it may be decided whether an overlay pixel may be overlaid on top of a respective image pixel. Preferably, the segmentation algorithm recognizes the road pixels by their color value since it may be assumed that the road pavement everywhere has more or less the same color value of the color grey or black. Thus, a particularly easy and fast pixelwise classification of the image contents into the object classes is facilitated.
[0026] An advantageous further development envisages that the additional graphic information is only overlaid on top of such pixels within the currently to be overlaid image region, the pixels of which lie in the first range of pixel values and in particular correspond to a predetermined pixel value from the first range of pixel values. In other words, in the example described here, the overlay pixels are only overlaid on top of the road pixels, not though on top of pixels of the object class “everything except road”.
[0027] A further embodiment envisages that the additional graphic information on top of such pixels within the currently to be overlaid image region, the pixels of which lie outside the first range of pixel values, in particular within the second range of pixel values, is not overlaid or only overlaid in an adapted manner. Preferably, the adapting of the additional graphic information comprises a change of color and / or a change of intensity of the additional graphic information starting from a respective initial color value and / or starting from a respective initial intensity value. If during checking the pixel values of the pixels to be overlaid it is found that a pixel value lies within the second range of pixel values, this may be captured as discontinuity. Upon reaching this discontinuity the representation of the additional graphic information may be stopped. In other words the representation of the additional graphic information may be cut off and / or interrupted at this point. If an object, for example another vehicle, enters the currently to be overlaid image region, the object in the representation of the image with the additional graphic information interrupts the representation of the additional graphic information. For the driver it may then appear as if the other vehicle moved in front of or on top of the additional graphic information. Thus, the object, in the example the other vehicle, can be immediately recognized by the driver. Therein, it is irrelevant whether the object moves as dynamic object into the image region since it drives or moves into the field of vision of the on-board camera(s), or whether the object is static and enters the image region through a movement of the motor vehicle, because by the movement of the motor vehicle also the field of vision of the onboard camera(s) shifts.
[0028] A further embodiment envisages that for each pixel of the image the pixel values stored in the depth buffer during a ride of the motor vehicle are updated at predetermined temporal and / or spatial intervals, wherein the overlay of the image with the additional graphic information is performed in dependence on the most current pixel values in each case. This means the possibility that over time a change of the spatial distribution of the pixel values in the image occurs is accommodated. In other words, in the course of time the pixel value map changes. Therefore, according to the embodiment described here, the pixel value distribution within the image is continuously updated.
[0029] As has already been mentioned, the motor vehicle may comprise a multitude of real onboard cameras or camera sensors, which may pertain to an environmental sensor system of the motor vehicle. According to a further embodiment by the multitude of real cameras of the motor vehicle real images of the environmental region are taken and from these real images the image is generated. In other words, the image is calculated from the different images taken by the on-board cameras. Therein, the image is represented from a perspective of a virtual camera arranged in the environmental region. The position of the virtual camera therein may for example be above the motor vehicle so that the image is represented as a top view from a bird’s perspective. The image contents of the image therein are represented according to a perspective of the virtual camera by projection of the pixels representing the image contents onto a three-dimensional mesh describing the perspective of the virtual camera. Depending on the perspective, the image contents may be represented distorted, for example if the perspective is a bowl view. The distortion may take place before or after the semantic segmentation. This means that, to start with, the semantic segmentation, that is the pixelwise classification of the image contents, is performed. Thereafter, the image may be rectified. This order, however, may also be inverted.
[0030] A further development envisages that during the projection of the pixels representing the image content onto the three-dimensional mesh describing the perspective of the virtual camera, the pixel values of the pixels are stored in the depth buffer. In other words, the projection and the creation of the described pixel value map or depth map may be effected in a single working step, which saves computing time and computing capacity. Preferably, the additional graphic information is equally adapted in dependence on the perspective of the virtual camera, in particular distorted. This may be effected by applying the projection equation, which underlies the described projection of the pixels, to the pixels of the additional graphic information. By the adaptation of the additional graphic information, that is by adaptation of the overlay, to the perspective of the virtual camera, from which the image is represented, it may be ensured that also in the case of distorted image contents invariably only those pixels are overlaid by the additional graphic information, the pixel values of which fulfill the required prerequisites, that is which lie within the predetermined range of pixel values.
[0031] Preferably the additional graphic information comprises a driving path and / or a direction arrow and / or a virtual parking space marking.
[0032] A further aspect of the invention relates to an electronic vehicle guidance system for a motor vehicle, comprising an environmental sensor system with at least one environmental sensor, in particular a camera sensor or a camera, a storage device comprising a depth buffer, a display device comprising a display surface, which is configured to display pixel-based image contents with overlaid additional graphic information, and a computing device comprising at least one computing unit. The electronic vehicle guidance system is configured to perform a method according to one of the described embodiments.
[0033] An electronic vehicle guidance system may be understood as an electronic system, which is configured to guide a vehicle in a fully automated or a fully autonomous manner and, in particular, without an intervention or control by a driver or user of the vehicle being necessary. The vehicle carries out all required functions, such as steering maneuvers, deceleration maneuvers and / or acceleration maneuvers as well as monitoring and recording the road traffic and corresponding reactions automatically. In particular, the electronic vehicle guidance system may implement a fully automatic or fully autonomous driving mode according to level 5 of the SAE J3016 classification. An electronic vehicle guidance system may also be implemented as an advanced driver assistance system, ADAS, assisting a driver for partially automatic or partially autonomous driving. In particular, the electronic vehicle guidance system may implement a partly automatic or partly autonomous driving mode according to levels 1 to 4 of the SAE J3016 classification. Here and in the following, “SAE J3016” refers to the respective standard in the version of dated April 2021. Guiding the vehicle at least in part automatically may therefore comprise guiding the vehicle according to a fully automatic or fully autonomous driving mode according to level 5 of the SAE J3016 classification. Guiding the vehicle at least in part automatically may also comprise guiding the vehicle according to a partly automatic or partly autonomous driving mode according to levels 1 to 4 of the SAE J3016 classification.
[0034] For this purposes, at least one control signal, for instance one or several actuators of the motor vehicle may be provided, amongst them for example one or several braking actuators and / or one or several steering actuators and / or one or several drive motors of the motor vehicle. The one or the several actuators may influence a longitudinal and / or transverse steering of the motor vehicle in order to guide the motor vehicle at least partly automatically.
[0035] Assistance information may be output via an output device of the motor vehicle, for example a display and / or an audio output system and / or a haptic output system.
[0036] In the present disclosure, a computing unit may for example be understood as a data processing device with processing circuitry. A computing unit can therefore perform computing operations in order to process data. The computing operations may also include indexed accesses to a data structure, for example a look-up table, LUT.
[0037] In particular, a computing unit may include one or more computers, one or more microcontrollers, and / or one or more integrated circuits, for example, one or more application-specific integrated circuits, ASIC, one or more field-programmable gate arrays, FPGA, and / or one or more systems on a chip, SoC. The computing unit may also include one or more processors, for example one or more microprocessors, one or more central processing units, CPU, one or more graphics processing units, GPU, and / or one or more signal processors, in particular one or more digital signal processors, DSP. The computing unit may also include a physical or a virtual cluster of computers or other of said units.
[0038] A computing unit may also comprise one or more hardware and / or software interfaces and / or one or more memory units. Therein, a memory unit may be implemented as a volatile data memory, for example a dynamic random access memory, DRAM, or a static random access memory, SRAM, or as a non-volatile data memory, for example a readonly memory, ROM, a programmable read-only memory, PROM, an erasable programmable read-only memory, EPROM, an electrically erasable programmable readonly memory, EEPROM, a flash memory or flash EEPROM, a ferroelectric random access memory, FRAM, a magnetoresistive random access memory, MRAM, or a phase-change random access memory, PCRAM.
[0039] A further aspect of the invention relates to a motor vehicle comprising an electronic vehicle guidance system. The motor vehicle according to the invention is preferably designed as a motor vehicle, in particular as passenger car or truck, or as passenger bus or motorcycle.
[0040] A further aspect of the invention relates to a computer program, comprising instructions, which, when they are executed by a computing unit, in particular by a computing unit of a computing device of an electronic vehicle guidance system, cause the same to perform a method according to one of the described embodiments.
[0041] A further aspect of the invention relates to a computer-readable storage medium, on which such a computer program is stored. This means that as a further solution, the invention also comprises a computer-readable storage medium, comprising program code, which, when executed by a computer or a computer network causes the same, to perform an embodiment of the method according to the invention. The storage medium may at least partially be provided as a non-volatile data storage (for example as a flash memory and / or as SSD - solid state drive) and / or at least partially as a volatile data storage (for example as a RAM - random access memory). The storage medium may be arranged in the computer or computer network. The storage medium may, however, also be operated for example as so-called Appstore Server and / or Cloud Server in the internet. By the computer or computer network a processor circuit with for example at least one microprocessor may be provided. The program code may be provided as binary code and / or as assembler code and / or as source code of a programming language (for example C) and / or as program script (for example Python).
[0042] Jointly, the computer-readable storage medium and the computer program may also be referred to as computer program product.
[0043] All advantages entailing in connection with the described embodiments of the method according to the invention also apply to the further aspects of the invention, and vice versa. The invention also includes further developments of all aspects according to the invention, which comprise features, as they have already been described in connection with the further developments of the method according to the invention and vice versa. For this reason the corresponding further developments of the further aspects according to the invention are here not described once again.
[0044] The invention also relates to the combinations of the features of the described embodiments. The invention also comprises realizations, each of which comprise a combination of the features of several of the described embodiments, unless the embodiments were described as mutually exclusive.
[0045] With the indications „top“, „bottom“, „front“, „rear“, horizontal", „vertical“, „depth direction", „width direction", height direction" et cetera the positions and orientations given in the case of use according to the intended purpose and arrangement according to the intended purpose of the virtual camera or the real camera or the motor vehicle are indicated.
[0046] Further features of the invention are apparent from the claims, the figures and the figure description. The features and feature combinations mentioned above in the description as well as the features and feature combinations mentioned below in the description of figures and / or shown in the figures may be comprised by the invention not only in the respective combination stated, but also in other combinations, without going beyond the scope of the invention. Thus, also embodiments of the invention are to be regarded as comprised and disclosed, which are not explicitly shown and explained in the figures, but derive by separated feature combinations from the explained embodiments and can be generated from them. Also embodiments and feature combinations are to be regarded as disclosed, which thus do not comprise all features of an originally formulated independent claim. Moreover, implementations and featured combinations are to be regarded as disclosed, in particular by the above-explained explanations, which go beyond the feature combinations set out in the back-references of the claims or deviate therefrom.
[0047] Embodiments of the invention are explained in further detail in the following by reference to schematic drawings. These show in:
[0048] Fig. 1 a schematic representation of a motor vehicle comprising an electronic vehicle guidance system according to an embodiment of the invention;
[0049] Fig. 2 a schematic representation of a method according to an embodiment of the invention; Fig. 3 a schematic representation of a method according to a further embodiment of the invention; and
[0050] Fig. 4 a schematic representation of a method according to a further embodiment of the invention.
[0051] The embodiments set out in the following are preferred embodiments of the invention. In the embodiments the described components of the embodiments each represent individual features of the invention, which are to be considered independently of one another and which also independently of one another further develop the invention. Therefore, the disclosure should also comprise other feature combinations of the embodiments than the ones represented. Moreover, the described embodiments are also capable of being supplemented by further ones of the already described features.
[0052] In the figures the same reference signs each designate elements having the same function.
[0053] In Fig. 1 schematically a top view of a motor vehicle 10 comprising an electronic vehicle guidance system 12 according to an embodiment of the invention is shown. The electronic vehicle guidance system 12 in the embodiment comprises a computing device 14 and a display device 16 with a display surface (not shown). Further, the electronic vehicle guidance system 12 in the embodiment comprises a first real camera 18, a second real camera 20, a third real camera 22, and a fourth real camera 24. According to the embodiment the first real camera 18 is arranged at a front 26 of the motor vehicle 10, the second real camera 20 is arranged on a right side 28 of the motor vehicle 10, the third real camera 22 is arranged at a rear 30 of the motor vehicle 10 and the fourth real camera 24 is arranged on a left side 32 of the motor vehicle 10. The arrangement of the real cameras 18, 20, 22, 24, however, is possible in manifold ways, preferably, however, in such a way that the motor vehicle 10 may be captured at least partially.
[0054] The real cameras 18, 20, 22, 24 in particular comprise a wide capture range, which for example may be larger than 180°. The wide capture range may for example be provided by a fisheye lens of a lens unit of the real camera 18, 20, 22, 24. For instance the electronic vehicle guidance system 12 may provide a surround view system, SVS, or an electronic rear view mirror or may be configured as a further driver assistance system of the motor vehicle 10, in which the environmental region 34 is captured at least partially. The real cameras 18, 20, 22, 24 may be CMOS cameras (complementary metal-oxide semi-conductor cameras) or CCD cameras (charge coupled device cameras) or also other image capturing devices, which may provide a frame of the environmental region 34 and / or the motor vehicle 10. The real cameras 18, 20, 22, 24 are in particular video cameras, which continuously provide an image sequence of frames. The computing device 14 may then process the image sequence of frames for example in real time. The computing unit 14 may be arranged for example within the respective real camera 18, 20, 22, 24 or within the display unit 16. The computing device 14 may, however, also be arranged outside the respective camera 18, 20, 22, 24 or the display device 16 in a random different position within the motor vehicle 10 and thus be configured as unit that is separate from the real camera 18, 20, 22, 24 and from the display device 16.
[0055] The display device 16 may for example be configured with a liquid crystal display, LCD. The display device 16 may be arranged in manifold ways in the motor vehicle 10, preferably, however, in such a way that a user of the motor vehicle 10 may direct a view devoid of obstacles at the display surface of the display device 16. The display unit 16 preferably comprises a touch-sensitive display surface that may be operated by touch command through the user.
[0056] By the real cameras 18, 20, 22, 24 a multitude of real images is taken. The real images show the environmental region 34 at least partially from the perspective of the respective real camera 18, 20, 22, 24. Preferably, the real images are taken in such a way that they are at least partially overlapping.
[0057] Additionally, the motor vehicle 10 comprises an environmental sensor system comprising a plurality of environmental sensors 36, at least some of which may be sensors of the so- called low speed maneuvering system (abbreviated: LSMS) of the motor vehicle 10. The environmental sensors 36 are preferably configured for distance measuring and may for example comprise radar, lidar, and / or ultrasonic sensors. In the embodiment of Fig. 1 , exemplarily and for reasons of clarity, only four environmental sensors 36 are shown. It goes without saying that there may be a multiple of environmental sensors 36 on and in the motor vehicle 10, which may be arranged on the motor vehicle 10 randomly and as needed.
[0058] In Fig. 1 , moreover, a boundary 38 of the environmental region 34 is indicated by a dashed line. It goes without saying that the environmental region 34 may also have different shapes than the one shown here. In particular, the boundary 38 of the environmental region 34 may also be dynamically adapted to a respective driving situation and be modified.
[0059] Fig. 2 shows a schematic representation of a method diagram, wherein here an exemplary embodiment of the method according to the invention is shown. The embodiment shown here is based on the exemplary embodiment of at least one of the vehicle cameras 18, 20, 22, 24 as so-called fisheye camera.
[0060] Fisheye cameras are counted amongst the non-gnomonic or non-rectilinear cameras. Such a camera may be understood as a camera with a non-gnomonic or non-rectilinear lens unit. A non-gnomonic or non-rectilinear lens unit may be understood as a lens unit, that is one or several lenses, having a non-gnomonic, that is non-rectilinear or curvilinear, mapping function. In particular, fisheye cameras represent non-gnomonic or non- rectilinear cameras.
[0061] The mapping function of the lens unit may be understood as a function r(0) mapping an angle 0 from the central axis of the radial distortion of the lens unit to a radial shift r out of the image center. The function depends parametrically on the focal length f of the lens unit.
[0062] For example, a gnomonic or rectilinear lens unit comprises a gnomonic mapping function, in particular r(0) = f tan(0). In other words, a gnomonic or rectilinear lens unit maps straight lines in the real world to straight lines in the image, at least up to lens imperfections.
[0063] A non-gnomonic, non-rectilinear or curvilinear lens unit in general does not map straight lines to straight lines. In particular, the mapping function of a non-gnomonic or non- rectilinear camera may be stereographic, equidistant, equisolid angle or orthographic. Mapping functions of non-rectilinear lens units may also be, at least approximately, be given by polynomial functions.
[0064] In a step S2.1 a camera image is received from a fisheye camera. The camera image may for example be transferred for further processing to the computing unit 14 of the electronic vehicle guidance system 12. There, the camera image may be prepared in a step S2.2 for the semantic segmentation. The semantic segmentation itself may then be performed by reference to a segmentation algorithm in a step S2.3. The result of the segmentation according to the embodiment described here may be a classification of the pixels of the camera image into only two object classes, namely “road” with a pixel value „1“ and ..everything except road" with a pixel value „0“. In other words, „1“ for the object class „road“ and „0“ for the object class ..everything except road" may be assigned as pixel values. This binary pixel value distribution may be effected in a step S2.4. In a step S2.5 on the basis of the pixel value distribution a virtual depth map of the image contents may be created. In a step S2.6 the depth map may be projected onto a 3D mesh, for example to rectify the fisheye image. The pixel value distribution may be stored in a depth buffer, which may be updated in a step S2.7. In a step S2.8, which preferably runs in parallel, the additional graphic information may be prepared for display. In a step S2.9 for each pixel of the image it may be checked whether its pixel value allows for an overlay with a pixel of the additional graphic information. For this purpose, in a step S2.10 in particular the question may be answered whether the pixel to be overlaid according to its pixel value belongs to the object class „road“, or not. If this question can be answered in the positive („+“), because the result of the checking of the pixel value is that the pixel value of the pixel to be overlaid is „1“, the pixel of the additional graphic information may be displayed there (step S2.11). If the result of the checking of the image value is, however, that the pixel value of the pixel to be overlaid the pixel of the additional graphic information is not shown there (step S2.12). Rather, in step S2.12 it may be decided to skip this pixel of the additional graphic information in the rendering. If the method is continued with step S2.11 , thereafter step S2.13 may be effected, namely the representation of the image content with the overlaid additional graphic information.
[0065] Fig. 3 shows a schematic representation of a method according to a further embodiment of the invention. Here it is in particular exemplarily shown how it can be proceeded if a discontinuity in the pixel value map is found. In other words, the method here shown exemplarily deals with the approach, in case during updating the pixel value map it is determined that a pixel of the image does not or no longer comprise a pixel value allowing for an overlay with the additional graphic information. This is described again in the following in an exemplary way by reference to the example of the binary pixel value distribution with “1” for object class “road” (overlay allowed) and “0” for object class “everything except road”, which has already been set out several times.
[0066] In a step S3.1 the method of the rendering, that is of the rendering with overlay or overlaid additional graphic information, respectively, is started. In a step S3.2 for a pixel of the image to be currently checked its pixel value is read from the depth buffer. In a step S3.3 it is checked whether the pixel value lies within the predetermined value range for „road“ or ..everything except road”. If the result is „road“, that is pixel value „1“, the pixel of the additional graphic information to be overlaid is indicated at the corresponding pixel of the image (step S3.4), before it is continued in step S3.5 and then in step S3.2 with the checking of the next pixel of the image. If the result of the checking, however, is “everything except road”, that is pixel value „0“, according to the embodiment described here the additional graphic information at this spot of the image is not overlaid, but rather interrupted or cut off (step S3.6). Here it appears to the viewer as if the image content, to which the pixel value „0“ is assigned, shifted in front of the additional graphic information. In the following step S3.7 the method is stopped. Alternatively and not shown here, the method could also continue with the checking of a subsequent pixel in step S3.2.
[0067] Fig. 4 making reference to the components shown and described in connection with the previous figures shows a schematic representation of a method according to a further embodiment of the invention. Steps S4.1 to S4.4 in the example described here are performed in analogy to steps S3.1 to S3.4. If it is found in step S4.3 that the pixel value of the pixel of the image to be currently checked has the value „0“, that is everything except road", in contrast to step S3.6, in step S4.5 the additional graphic information is not cut off, but rather overlaid with an increased transparency in comparison with an initial value so that the object in the environmental region, which is assigned to the object class everything except road", through the overlay with the additional graphic information is still visible to the viewer. Alternatively or additionally, here a warning color with increased transparency in comparison with an initial color may here be chosen. In the example shown, both after step S4.4 and after the alternative step S4.5, it is continued with a step S4.6, in which a checking of the subsequent pixel is performed. This means that after step S4.6 the method described here continues again with step S4.2.
[0068] In modern surround view visualization systems for assisting drivers frequently a bowlshaped 3D mesh is employed as uniform projection surface for the images of several cameras, which usually are arranged in four positions around the vehicle. This provides the driver with an integrated view of the environment of his vehicle from a bird’s eye perspective and improves the spatial awareness.
[0069] With the introduction of dynamic graphic overlays in this virtual 3D-world, however, one decisive challenge by the covering of the critical objects results. Dynamic graphic overlays, such as for example projected driving paths or highlighted obstacles, may inadvertently overlay important objects in the environment, in particular on top of adjacent vehicles. In view of the 3D nature of the visualization these overlays may cover parts of the projected camera image on the mesh. As a result amongst other things this may mislead the driver if important objects, such as proximate vehicles, pedestrians or obstacles are covered by these graphic overlays. Thus, for example a trajectory path drawn in the visualization on top of a proximate car may convey the false impression of an unobstructed driving path. Moreover, a challenge to the user-friendliness and the interpretation of the contents of the display exists. The primary objective of the surround view system, however, is to improve the understanding of the driver of its environment. Overlaid additional graphic information or overlays that can be represented without any difference on top of every image background, however, may interfere with the visualization such that it is counter-intuitive and difficult for the driver to interpret.
[0070] In view of this problem the present invention provides an innovative method which seamlessly integrates the semantic segmentation into the graphic representation of overlaid image contents. The method can predetermine in an intelligent way where graphic elements appear in a 3D environment, and can ensure that critical objects remain clearly visible in the camera transmissions so that both the safety as well as the user friendliness of the visualization system are improved.
[0071] As part of the depth-controlled overlay rendering a graphic rendering module, equipped with the depth texture of the image and the necessary overlays, can start with the rendering process. Using the depth texture or depth map as reference, the depth check during the rendering is activated. Thereby, it is ensured that overlay regions are avoided, which are marked with a „closer“ depth value, such as for example locations of other vehicles, and that critical objects remain visible and are not covered by the overlays.
[0072] The capturing of the depth information is effected on pixel level. On the basis of the information on the camera feeds a “depth map” is created - a representation of the scene, in which things that are closer to the ego vehicle than others, such as for instance other vehicles, are clearly marked. If it is time to add the overlays to the 3D view, they are not simply blindly overlaid on top of all image contents. Rather, they are represented in reference to the depth map. In other words, the overlays use the depth map as orientation aid. If the depth map says: “Hey, something is in the close proximity here, for example a car”, the overlay knows that it should not cover this region. It is as if the overlay were given a haptic sense. It can “feel” where things are located in the 3D space and can move around them. The result is a 3D view, in which the overlays seamlessly move around the vehicles and other “close” objects so that the same are invariably visible and are not covered. It is as if the overlays perceive the environment and adapt in real time to them.
[0073] The updating of the depth texture or the depth buffer is effected with the aid of bit mask textures. To this end, after the semantic segmentation a binary bit mask texture may be generated, in which the pixels that were assigned to the identified vehicles (or other important objects), are set to a determined value (for example 1 or 255). All other pixels representing uncritical regions are set to an alternative value (for example 0). A frame buffer object (FBO) is initialized in OpenGL. The bit mask texture representing the depth values derived from the segmentation is attached to this FBO as texture attachment. This texture serves as guideline for the updating of the depth buffer values. In the mapping of the camera feeds to the bowl-shaped 3D mesh the associated shaders make reference to the bit mask texture. For each rendered fragment or pixel the shader checks the value in the bit mask texture. If the fragment or the pixel is assigned to a vehicle (or a marked object in the bit mask), the value of the depth buffer for this fragment is set to indicate that it is located “closer” to the viewer. Fragments that do not correspond to the non-vehicle regions obtain a depth value, which indicates that they are “more distant”. After this process, the depth value is filled with values corresponding to the vehicle regions that are “closer” and other regions that are “more distant”. This depth buffer is then used during the overlay rendering process to ensure that overlays are not drawn on top of regions that are marked as “closer” (such as vehicles).
[0074] The invention integrates advanced semantic segmentation techniques with 3D graphic rendering in order to dynamically and intelligently control the placing of overlays in a surround visualization system, which preferably employs a bowl-shaped 3D mesh for the projection of camera images. The semantic segmentation is preferably effected in real time, wherein each camera image is exposed to a semantic analysis, which classifies various image contents in the image. Therein different elements in the image are classified and in particular objects, such as vehicles, pedestrians, and other critical obstacles identified.
[0075] Based on the segmentation results a depth map is created, which assigns a „closer“ depth value to critical objects, such as for example vehicles. This depth map is then applied to the bowl-shaped 3D mesh. When rendering dynamic overlays the system makes reference to the depth map. The overlays take the depth values into account and ensure that they are not drawn on top of the “closer” regions of vehicles or other identified critical objects.
[0076] For instance, a dynamic operation is effected in real time: In contrast to static or predefined overlay systems, the invention may work in real time and may adapt to the constantly changing environment of the motor vehicle. The system can be integrated into the various camera configurations and may be expanded without any problem to recognize and prioritize besides vehicles also other critical objects, for example pedestrians or certain road signs. By employment of advanced semantic models, which are optimized for the real time operation, and of graphic hardware the system provides a smooth performance without any noteworthy computing effort.
[0077] A concrete application example of the invention may provide the employment as parking aid on an overcrowded parking lot. In this connection a driver of a car tries to find a suitable parking spot on a particularly overcrowded parking lot, without colliding with other parked cars, shopping trolleys, or pedestrians.
[0078] According to the described example, the vehicle of the driver is equipped with a series of cameras providing a 360 degree view around the car. When the driver approaches the parking spot, the cameras transmit live images to a computing device, for example a data processing unit of the vehicle. The data processing unit employs the semantic segmentation in order to analyze the camera images and to distinguish between drivable ground, parked vehicles, pedestrians and other objects. Each object type or each object class in the system is assigned an unambiguous designator or pixel value.
[0079] On the basis of the semantic segmentation the system generates an overlay. This overlay may contain lines and indicators, which suggest the optimum steering path for safe parking of the car. The overlay is conceived in such a way that it only appears on the ground surfaces, wherein the object identifiers or pixel values are used to avoid that other vehicles or pedestrians are covered.
[0080] The driver may see the real-time feed via a display in the dashboard or a heads-up display (HUD) in the vehicle interior space, where the overlay appears as clear guideline on the ground and guides it safely into the parking spot. Owing to the intelligent rendering of the overlay, which does not cover any recognized objects, obstacles remain well visible Following the visual indications of the overlay, the driver maneuvers the car successfully into the parking spot, without risking a collision. The overlay can dynamically adapt to the movements of the vehicle and provides updated notifications if the driver needs to correct his course.
[0081] This practical example shows how the invention improves the safety and the comfort of the driver in the case of a frequent but complex driving task. It illustrates the capability of the system to adapt in real time to a changing environment, whilst maintaining visual clarity and avoidance of obstacles.
[0082] According to an exemplary embodiment of the invention a segmented raw image of a fisheye camera may be used, wherein the individual pixels of the raw image preferably are assigned either a black color value (0) or a white color value (255). The black pixels pertain to a road or to a surface that is drivable by a car, respectively. The white pixels pertain to everything except road (remainder). If from the color image of the fisheye camera a bowl view is generated, the pixels of the color image are projected according to a projection equation onto the three-dimensional shape of the bowl view. In other words, the image contents of the color image are transferred by projection of the image contents of pixels representing the color image to a three-dimensional mesh describing the bowl shape into the bowl shape. The same projection equation may be used to transfer the pixel values (that is the color values 0 and 255) into the bowl shape. This means that thereby at the same time information as to whether the image content is classified as drivable or not drivable is transferred into the bowl shape.
[0083] Usually, the bowl shape or bowl view, however, is generated only by projection of the texture of the fisheye camera image, wherein only the color information or color values from the fisheye camera image are used. Therefore, in the known case additional graphic information or overlays or further augmentations cannot know anything about the assignment of certain image regions to the object class “everything except road”, on top of which they should not be overlaid, since the same should remain visible for the driver. This means that artifacts, such as the drawing of overlays on top of a wall or other vehicles, may be produced.
[0084] According to the described embodiment, which may also be applied to views of virtual cameras, the depth buffer (2D buffer), is updated in order to retrieve the depth values together with the color from the fisheye camera raw images. The updating of the depth buffer preferably is effected together with the above-described projection of the segmented fisheye image. Together with the updating of the color buffer the raw fisheye image segmented in a binary manner and the fish eye projection for updating the depth buffer is employed. The virtual camera view now has the correctly updated depth information available. Then the overlay with the additional graphic information (the overlay rendering) is effected, wherein prior to the rendering the depth value of each pixel of the image is checked as to whether the pixel may be overlaid.
[0085] The invention may be applied to each random view besides the bowl view since the depth information is updated preferably by the same projection equation which also updates the color information.
[0086] This means that the depth buffer is updated in order to render a binary depth map, which differentiates between the object class “road” and all other objects on the basis of the semantic segmentation. Binary depth value assignment means that in the output of the semantic segmentation each pixel is assigned a binary depth value. Therein, the value “1 ” may be employed for ground regions or drivable ground surfaces, which are suitable for the overlay rendering. The value „0“ may be used for all other objects, such as vehicles, pedestrians, or the sky, where no overlay should be generated.
[0087] The depth buffer projection comprises the mapping of the segmented fisheye image to the bowl-shaped 3D mesh. During this projection the binary depth value for each pixel is written into the depth buffer. When writing the depth buffer, a shader may be employed, which writes the binary depth values into the depth buffer when the segmented fisheye image is projected onto the 3D mesh. This shader ensures that ground pixels or road pixels obtain a depth value of „1“ and all the other ones a value of „0“.
[0088] By the binary depth buffer the overlay rendering preferably is effected only on pixels, for which the depth value is „1“ (ground or road). The depth test for overlays is only successful if the depth buffer value is „1“ and overlays are only rendered on ground regions or road regions, respectively.
[0089] The creating of the overlays comprises the generating of the overlay content, which may contain navigation paths, direction arrows, or parking guidelines, which are to appear on the ground level. This means that the invention provides for a conditioned rendering on the basis of depth values of image contents. Therein, the depth test in the rendering phase ensures that each overlay pixel retrieves the depth buffer and is only drawn if the corresponding depth buffer value is “1”. If the depth buffer value is “0”, which indicates a non-ground pixel, the overlay is not rendered. On the whole, the examples show how the display of overlaid or overlaying additional graphic information may be improved.
Claims
Claims1 . Method for representing an additional graphic information in an image of an environmental region (34) of a motor vehicle (10), wherein the image contains image contents that are represented by pixels on a display surface of a display device (16) of the motor vehicle (10), wherein the additional graphic information is represented as an overlay on top of the image contents of the image, the method comprising the steps of- pixelwise classification of the image contents of the image, wherein each pixel of the image in dependence on the image content represented by the pixel and / or in dependence on a color value assigned to the pixel is assigned to a predetermined object class,- assigning a pixel value to a respective pixel in dependence on the object class, to which the respective pixel was assigned, and storing the pixel values as depth values in a depth buffer,- determining the pixel values of the pixels within an image region of the image to be currently overlaid by the additional graphic information, and- overlaying the additional graphic information only on top of such pixels within the image region to be currently overlaid, the pixel values of which lie within a predetermined range of pixel values.
2. Method of claim 1 , wherein for each pixel within the image region of the image to be currently overlaid by the additional graphic information the pixel value is retrieved from the depth buffer and is compared with the predetermined range of pixel values.
3. Method according to any one of the preceding claims, wherein as object classes at least one first object class with an assigned first range of pixel values and a second object class with an assigned second range of pixel values are defined, wherein the first object class comprises surface drivable by the motor vehicle (10) and wherein the second object class comprises surface not drivable by the motor vehicle (10).
4. Method according to claim 3, wherein the additional graphic information is overlaid only on top of such pixels within the image region to be currently overlaid, the pixel values of which lie in the first range of pixel values and in particular correspond to a predetermined pixel value from the first range of pixel values.
5. Method according to any one of the preceding claims, wherein the additional graphic information is not overlaid or overlaid only in an adapted manner on top of such pixels within the image region to be currently overlaid, the pixel values of which lie outside the first range of pixel values, in particular within the second range of pixel values.
6. Method according to claim 5, wherein the adapting of the additional graphic information comprises a color change and / or an intensity change of the additional graphic information starting from a respective initial color value and / or starting from a respective initial intensity value.
7. Method according to any one of the preceding claims, wherein for each pixel of the image the pixel values stored in the depth buffer during a ride of the motor vehicle (10) are updated at predetermined intervals, wherein the overlaying of the image with the additional graphic information is performed in dependence on the most current pixel values in each case.
8. Method according to any one of the preceding claims, wherein by a plurality of real cameras (18, 20, 22, 24) of the motor vehicle (10) real images of the environmental region (34) are taken and from these real images the image is generated, wherein the image is represented from a perspective of a virtual camera arranged in the environmental region (34), wherein the image contents according to the perspective of the virtual camera are represented by projecting the pixels representing the image contents onto a 3D mesh describing the perspective of the virtual camera.
9. Method according to claim 8, wherein during projecting, the pixel values of the pixels are stored in the depth buffer.
10. Method according to any one of claims 8 or 9, wherein the additional graphic information is adapted in dependence on the perspective of the virtual camera.11 . Method according to any one of the preceding claims, wherein the additional graphic information comprises a driving path and / or a direction arrow and / or a virtual parking space marking.
12. Electronic vehicle guidance system (12) for a motor vehicle (10), comprising- an environmental sensor system with at least one environmental sensor (36), in particular a camera (18, 20, 22, 24),- a memory device with a depth buffer,- a display device (16) with a display surface, which is configured to display pixelbased image contents with overlaid additional graphic information, and- a computing device (14) with at least one computing unit, wherein the electronic vehicle guidance system (12) is configured to perform a method according to any one of claims 1 to 11 .
13. Motor vehicle (10) with an electronic vehicle guidance system (12) according to claim 12.
14. Computer program, comprising instructions, which, when they are executed by a computing unit, in particular by a computing unit of a computing device (14) of an electronic vehicle guidance system (12) according to claim 12, cause these to perform a method according to any one of claims 1 to 11 .
15. Computer-readable storage medium, on which a computer program according to claim 14 is stored.
Citation Information
Patent Citations
Overlay adaptation for visual discrimination
WO2023007220A1
Image stitching with dynamic SEAM placement based on ego-vehicle state for surround view visualization
WO2023192752A1