Methods for the situational output of visual cues

The adaptive method aligns vehicle display systems with the driver's view using environmental sensors and interior cameras to selectively display cues, addressing excessive information in current systems and improving road safety.

DE102024000847B4Active Publication Date: 2026-02-19MERCEDES BENZ GROUP AG
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
DE102024000847
Authority / Receiving Office
DE · DE
Patent Type
Patents
Current Assignee / Owner
Filing Date
2024-03-14
Publication Date
2026-02-19
Estimated Expiration
2044-03-14

AI Technical Summary

Technical Problem

Current vehicle display systems provide excessive and distracting information, leading to potential loss or neglect of critical information in complex driving conditions, particularly in poor visibility, thus compromising road safety.

Method used

An adaptive method that utilizes environmental sensors and interior cameras to create a representation of the vehicle's surroundings, aligning it with the driver's field of view, and selectively displays visual cues only where necessary to minimize distractions.

Benefits of technology

Enhances road safety by ensuring critical driving information is visible and minimizing distractions by tailoring visual cues to the driver's perception, even in adverse conditions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000000_0000_ABST
    Figure 00000000_0000_ABST
Patent Text Reader

Abstract

Method for the situational output of optical cues to support the driving task of a person (6) driving a vehicle (1) on a display device (7), wherein a position of the vehicle (1) and a representation of the environment related to the current position of the vehicle (1), determined by means of sensors and / or maps, wherein an image of the person (6) driving the vehicle (1) is captured by means of an interior camera (2), from which a corneal area with corneal reflection images is extracted, after which the extracted image area is transformed as a driver image model and the representation of the environment and projected onto a common projection surface in a common coordinate system, wherein the driver image model and the representation of the environment are reduced to the extent of the smaller of the two models, after which a registration of the representation of the environment and the driver image model takes place.wherein differences between the representation of the environment and the driver image model are determined, and according to which an adjustment of optical cues superimposed on the display device (7) of the environment is made depending on the number of differences, wherein an adjustment of a degree of automation is made depending on the number of differences.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] The invention relates to a method for the situational output of optical cues to support the driving task of a person driving a vehicle on a display device.

[0002] German patent application DE 10 2021 213 332 A1 discloses a method in which data relating to a vehicle's driving situation are recorded. Based on this data, a visual cue appropriate to the respective situation is then displayed on a screen. The visual cue is thus adapted to the driving situation and serves to support the driving task of a person operating a vehicle in precisely that situation. It could, for example, be a warning generated based on the recorded environmental conditions around the vehicle.

[0003] The problem with such methods is that in current series production applications, they visualize a multitude of elements via the display device, such as a head-up display. This includes relatively trivial information like speed, route guidance instructions, and the like. However, it extends to the augmentation of lane data for assisted lateral control systems, such as lane keeping assist, or, preferably, contact-analog augmentation of visible or, for example, weather-related obscured objects in the vehicle's surroundings. In many respects, the needs of the driver are not used as the basis for the visualization; instead, the data available in the vehicle is visualized, or, in the case of the aforementioned state of the art, is generated situationally based on situations detected by the environmental sensors.All in all, this provides a driver with a multitude of information, which is often difficult to overlook and, in complex driving situations, such as driving in poor visibility conditions like heavy rain, does not serve road safety but rather leads to considerable distraction and disruption. This can ultimately result in truly important information being lost in the sheer volume of visual information, or in information being completely ignored in certain situations because it is perceived as disruptive.

[0004] The Wikipedia entry under the link “https: / / en.wikipedia.org / w / index.php?title=Image.registration&oldid=1192049344” reveals a method for registering images.

[0005] Wikipedia reveals under the link “https: / / de.wikipedia.org / w / index.php?title=Eye-Tracking&oldid=240397037” Methods for gaze direction detection.

[0006] DE 10 2009 020 300 A1 discloses a method in which a representation is generated that includes a virtual road surface with corresponding road boundaries and other objects from the vehicle's environment. The other objects are arranged at positions where the geometric relationships between the respective object and the road boundaries correspond to the real geometric relationships from the perspective of the vehicle's driver.

[0007] From DE 10 2009 020 328 A1, it is known to capture a vehicle's surroundings using both laser sensors and a camera. The positions and movements of objects within the captured environment are determined from the data of the laser sensors and the camera. By automatically comparing the captured objects detected by the two types of sensors, it can be determined which objects in the vehicle's vicinity are visible or invisible to the driver under given light and visibility conditions. Objects invisible to the driver are projected into the driver's field of vision using a head-up display or a 3D display.

[0008] The object of the present invention is to provide an improved method for the situational output of optical cues, which actually supports a person driving the vehicle, adaptively to the extent that this appears useful or necessary.

[0009] According to the invention, this problem is solved by a method with the features in claim 1, and in particular with the features in the characterizing part of claim 1. Advantageous embodiments and further developments of the method according to the invention also result from the dependent claims.

[0010] The method according to the invention provides an approach in which the output of optical cues is adaptively adjusted to the view of the person driving the vehicle, in order to support the driving task situationally and not generate unnecessary distraction.

[0011] The method focuses on providing support in situations where cues in the vehicle's surroundings, such as road markings and lateral lane boundaries, are difficult to perceive, for example, due to heavy rain and standing water on the road. To achieve this, the method first creates a representation of the environment using sensors, preferably distance-measuring sensors like lidar, radar, and / or maps. This is accomplished by identifying environmental features from one or, preferably, several available data sources. These sources can include active environmental sensors on the vehicle, particularly those with distance-measuring capabilities (e.g., LiDAR, radar, laser, ultrasound, time-of-flight, or similar), an infrared camera, or even a higher-level fusion of data from various environmental sensors.The aforementioned sensors are designed to operate in darkness and / or poor visibility conditions such as fog, rain, etc. – in other words, to detect objects invisible to the human eye and create a representation of the environment. This representation can be derived from a data source such as a high-precision map, preferably received from a backend server, swarm maps, and / or infrastructure such as surveillance cameras. The vehicle's self-localization to determine its current position is achieved, for example, via GPS sensors and / or lane attributes determined by sensors, which include, in particular, distance to the edge of the road, signs, significant objects, etc. The objects identified from the map data are then used to represent the environment relative to the vehicle's own position.

[0012] For determining the representation of the environment, it is particularly advantageous to utilize and merge all available data sources so that the information complements each other, for example, in a lane model. Based on this representation of the environment, a higher-level coordinate system can then be used, into which this representation is transformed. Preferably, the representation of the environment comprises the environmental features determined from the aforementioned data sources that lie in the vehicle's direction of travel or within a possible line of sight of the driver. In other words, the representation of the environment includes areas that are preferably within the driver's field of vision while driving.

[0013] In the next step, an image of the person driving the vehicle is captured using an interior camera. From this image, a corneal area containing corneal reflections from the driver's eyes is extracted. This extracted image, also known as the corneal image, is subsequently referred to as the driver image model. This driver image model allows us to determine what the driver can see. The driver image model and the representation of the environment are then transformed and projected onto a common projection surface in a shared coordinate system. The larger of the two models is selected as the area to which the larger model is reduced or cropped. This process is also known as cropping.The driver image model and the representation of the environment are now available in the same coordinate system and of the same size and can be superimposed.

[0014] The driver image model and the representation of the environment are now compared by registering the images against each other. Registration algorithms use predefined metrics to determine whether the registration is optimal and whether the representation of the environment and the driver image model can be approximated sufficiently to identify identical content and any missing content in either model. This registration process allows for the identification of differences between the representation of the environment and the driver image model. Depending on the number of differences, the overlaid environmental cues displayed on the screen are then adjusted. The cues are displayed in a contact-like manner, ensuring that the driver perceives the corresponding cues at precisely the same location where they would be visible in reality.

[0015] The system can adaptively determine which information the driver can perceive and which they cannot. If the driver cannot perceive all the information necessary for safe driving, the system can then display visual cues, such as lane markings. These cues are individually tailored to the driver's visual field, ensuring they are only displayed when needed. This avoids or minimizes unnecessary distractions, thus minimizing the driver's exposure to distractions. This ensures that even under adverse conditions like heavy rain, the driver can clearly see the relevant lane information.This allows the driver to travel safely and reliably in their designated lane, which enhances road safety. Overall, this creates an adaptive, situation-dependent display system that ideally supports the driver and contributes to increasing road safety.

[0016] The method can be used largely independently of the scenarios in the vehicle's environment. It is unaffected by wide and rapid saccades, as it only uses reflections on the cornea. It is largely independent of weather and / or lighting conditions. And it does not require externally provided depth reference values. In the case of using high-resolution maps to determine the representation of the environment, it is also independent of the availability of environmental sensors or the quality of the data they provide.

[0017] Such a procedure is already easy to implement today, as no additional hardware beyond the current standard is required.

[0018] According to the invention, the degree of automation is adjusted depending on the number of differences. The vehicle can therefore be driven with assistance or semi-automatically by the person driving it. Based on the number of differences between the representation of the environment and the driver image model based on the perspective of the person driving the vehicle, it can ultimately be determined whether an increase or decrease in the degree of automation, which is typically communicated to the person driving the vehicle, is advisable.

[0019] According to a particularly advantageous embodiment of the inventive method, no visual cues are displayed when the number of differences falls below a predetermined threshold. In this case, the extracted driver image model of the reflections on the cornea is used to determine that the driver sees all relevant information, thus preventing the display of additional information to avoid distracting the driver. However, if the number of differences rises above this threshold, the necessary additional information is displayed situationally and adaptively, as explained above.

[0020] A head-up display (HUD) can be used as a display device, for example, which is ideally suited for overlaying visual cues with the surroundings. Alternatively, the display device can be a screen to show a real or animated image of the surroundings.

[0021] During registration, the differences can be determined according to a particularly advantageous embodiment of the inventive method using a so-called Mutual Information Coefficient or an index of structural similarity, the so-called SSIM (Structural Similarity Index Measure). The SSIM is typically used to measure similarities between two images. It is a so-called index with a complete reference, meaning that the measurement or estimation of image quality is based on an uncompressed or noise-free original image. The Mutual Information Coefficient, which is also generally known, serves essentially to capture the relationship between any two variables. It can therefore determine similarities between different quantities, in this case between the driver image model and the representation of the environment.In principle, however, other methods for measuring similarity would also be conceivable.

[0022] A very advantageous further development of the inventive method uses a spherical segment as a common projection surface for the driver image model and the representation of the environment, thus ultimately transforming the representation of the environment into the surface on which the driver image model, i.e. the extracted image area of ​​the corneal reflections of a person driving the vehicle, is already present.

[0023] As mentioned above, the method according to the invention can be designed such that only those sections which the person driving the vehicle cannot see are overlaid with the visual cues. In addition to the lane markings already mentioned, these visual cues can also include, for example, signs, characteristic points, and markings on the road surface such as directional arrows, speed limits, or the like.

[0024] In the case of the heavy rain scenario described above, this means, for example, that some of the lane markings are perceptible to the driver, while others are not. The inventive method, in this embodiment, makes it possible to display precisely these parts on the display device and, in particular, to superimpose them on the surroundings in a contact-like manner. The situational output of visual cues therefore does not obstruct the driver's view of the surroundings they can perceive, but rather supplements their view only in those areas where perception is hindered by environmental factors such as standing water on the road.Ideally, the entire section of the road relevant to the person, including all necessary lane markings, can now be visible to them, regardless of whether they see them themselves or whether these have been projected into the image perceived by the person via the display device using the inventive method.

[0025] The invention further relates to a vehicle with an interior camera, a display device, a positioning device for determining its position, and a computing unit arranged in or connected to the vehicle, which is configured to carry out the inventive method according to one of the embodiments described above. Preferably, this vehicle may additionally have environmental sensors with at least one or, in particular, several environmental sensors such as cameras, LiDAR sensors, or the like. In principle, however, the representation of the environment, as described above, could also be derived solely from suitable map data if the vehicle's position is known.

[0026] Further advantageous embodiments of the method according to the invention and of a vehicle equipped for carrying out the method according to the invention will become apparent from the exemplary embodiment, which is described in more detail below with reference to the figures.

[0027] This shows: Fig. 1 a schematically indicated vehicle according to the invention; Fig. 2 a flowchart of the method according to the invention; Fig. 3. A presentation of the calculation of suitable metrics to derive the similarity of image content after registration, with three exemplary examples; and Fig. 4 An exemplary illustration of poor visibility due to heavy rain and the image information displayed as a result.

[0028] In the presentation of the Fig. Figure 1 shows a vehicle designated 1. It is intended to be suitable for carrying out the method according to the invention and has an interior camera designated 2. The vehicle also has a positioning device for determining its position, which is indicated here as a GPS receiver 3. The vehicle 1 also contains a computing unit 4 for carrying out the method described later, which may also have a communication connection (not shown here) to a cloud server or another type of external server to which computing operations can be outsourced, either wholly or partially. The vehicle 1 also shows optional environmental sensors 5, here in the form of a camera. These environmental sensors 5 may also include other sensors (not shown here), such as a LiDAR sensor or the like.The data from these sensors, or preferably a fusion of the data from all environmental sensors of this environmental sensor system 5, can then optionally be used in the procedure described below.

[0029] In the presentation of the Fig. In addition, vehicle 1 is indicated as having a person 6 who is to drive the vehicle 1. To provide this person 6, who is driving the vehicle 1, with situational and adaptive visual cues to support their driving task as needed, the vehicle 1 also has a display device 7, which is shown in the representation of the Fig. 1 is designed as a projection device for a head-up display, which projects the relevant information onto the windshield and thus into the field of vision of person 6.

[0030] To adaptively and situationally provide the appropriate support information and display it only when needed, the computing unit 4 is configured to implement the following procedure. This procedure is represented by its individual steps S1-S9 in the diagram below. Fig. 2 is schematically indicated in a flowchart.

[0031] The process step S1 consists of detecting and deriving a representation of the environment, for example by means of the environmental sensors 5, if these are available, and / or alternatively or preferably additionally by means of data from a high-resolution map, on the basis of which the representation of the environment or parts of the representation of the environment is derived based on the position of the vehicle 1, which is available through the positioning device 3.

[0032] As can be seen from the representation of the Fig. As can be seen from the two arrows indicated on the left, this process step S1 is followed by process step S2. For this step, the representation of the environment is transferred into a higher-level coordinate system, in particular, it is transformed into a coordinate system of vehicle 1.

[0033] In the subsequent process step S3, an image of person 6 driving vehicle 1 is captured by the interior camera 2. In process step S4, a corneal region of person 6 driving vehicle 1, containing corneal reflection images, is extracted from this captured image area. In process step S5, a corresponding transformation is then performed to project the extracted image region, which represents a driver image model, and the representation of the environment onto a common projection surface in a common coordinate system. The common projection surface can, in particular, be a spherical surface, so that ultimately the representation of the environment determined in process step S1 is projected onto the driver image model according to S4. This makes it possible to view what person 6 can see.

[0034] In process step S6, the first step of the actual image processing takes place in the form of a so-called cropping, in which the transformed representation of the environment from process step S5 is reduced to the basis of the driver image model. The typically larger representation of the environment is thus adjusted in size to the driver image model and ultimately to the area seen by person 6 driving the vehicle. In process step S7, the driver image model and the transformed and cropped representation of the environment are registered. In process step S8, suitable metrics are then calculated, for example, the Mutual Information Coefficient (sometimes also called the Mutual Coefficient Index), to derive the similarity of the image content of the representation of the environment on the one hand and the driver image model on the other after registration.The Mutual Information Coefficient describes how well one image can be approximated based on the signal intensity of another image. If the image contents are identical, the index is high. However, if the image contents are completely different, its metric tends towards zero.

[0035] In the presentation of the Fig. Figure 3 illustrates this with an example. The diagram shows the relevant metrics, such as the Mutual Information Coefficient (M), plotted on the top axis. The x-axis to the right represents the degree of visibility of the surroundings to the driver, for example, due to rain or heavy rain.

[0036] Within the scope of the presentation options using black and white line graphics, the following is displayed below the diagram: Fig. Figure 3 attempts to illustrate some of these scenarios with examples. The image on the far right shows a section from the perspective of person 6, who is driving vehicle 1, as would be expected under excellent visibility conditions. The road boundaries, road markings, and an example of a vehicle ahead are clearly visible here. In the middle image, visibility is already limited, for example, due to light rain or similar conditions. Visibility for person 6 is significantly more limited in the image on the far right. Fig. 3, in which, for example, due to heavy rain, only individual areas are still indistinctly recognizable and, for example, road markings, an edge boundary of the roadway or the like are no longer recognizable at all due to standing water in this area.

[0037] The sequence of process steps S1 to S8 is now repeated as shown by the arrow labeled A on the left of the flowchart. As long as the calculated metrics remain above the threshold shown in the example diagram... Fig. If the metrics calculated in step S8 are below the limit value x shown in the diagram, it can be assumed that the surroundings are sufficiently recognizable to person 6, so no intervention is necessary. If the metrics calculated in step S8 are below this limit value x, which applies here to the five metrics shown on the right, then, according to arrow B in the diagram, the following action is taken. Fig. 2. The process is switched to step S9 and optical cues are displayed on the display device 7 or the head-up display of the vehicle 1. The optical cues are preferably superimposed on the optical representation or an object of the optical representation in a contact-analogous manner.

[0038] In this scenario, the augmentation of lane data is therefore linked to the lowest possible coefficient or a high registration error, as this indicates that the content diverges from each other, meaning that the lane elements are not present in the driver image model. This has the advantage that the augmentation content is adaptively linked to the needs of person 6, who is driving vehicle 1.

[0039] In the presentation of the Fig. This is shown below in section 4. In the Fig. 4a) The scenario shown above is from the illustration in Fig. 3 shown at the bottom right. For example, due to heavy rain, the visibility of the individual lanes is extremely limited, so the contact-like display of visual cues regarding the individual lanes, which has been chosen here as an example, would be helpful for person 6. In the representation of the Fig. 4b) Below, the same scenario is shown again, whereby the corresponding lane markings from the representation of the environment are displayed in the head-up display 7 in such a way that they overlay the real markings and edges, which are not or only very poorly visible, thus enabling safe driving.

[0040] In another example not shown, S1 is analogously derived from the Fig. 2. A cyclist, at least partially obscured in fog, is detected using radar sensors as a representation of the surroundings. In another example, a traffic sign, not recognizable due to weather conditions (i.e., rain, fog, or darkness), is determined from map data as a representation of the surroundings according to step S1. After completing steps S1 to S8, a visual cue is output according to step S9, depending on the similarity between the driver image model and the representation of the surroundings. Preferably, the visual cue for the representation of the surroundings (i.e., the cyclist or the traffic sign) comprises a contact-analog superimposed virtual representation of a cyclist or a traffic sign on the head-up display, so that the driver sees the virtual representation at the position on the windshield where the cyclist or traffic sign, invisible in the fog, would be visible.

[0041] In addition to the basic possibility of displaying visual cues as soon as the calculated metrics fall below the threshold x, as shown here, it would of course also be possible, in principle, to identify, based on a more detailed specification of the corneal reflections image of person 6, those areas where the necessary information is still recognizable for person 6, and those areas where this is not the case. In this case, it might then be sufficient to display visual cues only where person 6 cannot recognize sufficient information. For example, in the case of a greater accumulation of fluid on the right lateral edge of the cornea, the following could be displayed: Fig.In the example shown for the four streets, only the lane markings are displayed, whereas they would not be displayed for the dashed center line, which is still sufficiently recognizable for the person. Information that the person can perceive would therefore not be displayed, so as not to distract them from the traffic situation by overlaying the real information they perceive with the augmented information.

Claims

[1] Method for the situational output of optical cues to support the driving task of a person (6) driving a vehicle (1) on a display device (7), wherein a position of the vehicle (1) and a representation of the environment related to the current position of the vehicle (1), determined by means of sensors and / or maps, wherein an image of the person (6) driving the vehicle (1) is captured by means of an interior camera (2), from which a corneal area with corneal reflection images is extracted, after which the extracted image area is transformed as a driver image model and the representation of the environment and projected onto a common projection surface in a common coordinate system, wherein the driver image model and the representation of the environment are reduced to the extent of the smaller of the two models, after which a registration of the representation of the environment and the driver image model takes place.wherein differences between the representation of the environment and the driver image model are determined, and depending on the number of differences, an adjustment is made to the optical cues superimposed on the display device (7) of the environment, and depending on the number of differences, an adjustment of the degree of automation is made. [2] Method according to claim 1, characterized by , that no visual indicators are displayed if the number of differences falls below a predefined threshold (x). [3] Method according to claim 1 or 2, characterized by , that a head-up display is used as the display device (7). [4] Method according to claim 1 or 2, characterized by , that a screen is used as a display device (7) to display a real or animated image of the environment. [5] Method according to any one of claims 1 to 4, characterized bythat the differences are described using a Mutual Information Coefficient (M) or an index of structural similarity. [6] Method according to any one of claims 1 to 5, characterized by , that a spherical segment surface is used as the common projection surface. [7] Method according to any one of claims 1 to 6, characterized by , that the visual indicators are designed in such a way that they do not overlap, or at least do not partially overlap, visible lane markings, signs or characteristic points. [8] Method according to any one of claims 1 to 7, characterized by , that radar data, backend data, swarm maps, map data, longitudinal markers, marginal markers and / or characteristic points are used in determining the representation of the environment. [9] Vehicle (1) comprising an interior camera (2), a display device (7), a position determination device (3) for determining its position, and a computing unit (4) arranged in or connected to the vehicle (1), which is configured to perform the method according to any one of claims 1 to 8.

Citation Information

Patent Citations

  • Method for displaying partially automatically determined surrounding of vehicle for user of vehicle from outlook of passenger, involves detecting automatically surrounding of vehicle with one or multiple object recognition devices

    DE102009020300A1

  • Method for displaying objects from the surroundings of a vehicle with varying degrees of visibility on the display of a display device

    DE102009020328A1

  • Method, computer program and device for controlling an augmented reality display device

    DE102021213332A1