Method for outputting an augmented reality display in a vehicle and vehicle
The method classifies environmental objects in vehicles to mask distracting and highlight relevant ones, using AI-generated content and user preferences, enhancing safety and comfort by reducing distractions.
Patent Information
- Application Number
- DE102024002745
- Authority / Receiving Office
- DE · DE
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2024-08-24
- Publication Date
- 2025-08-21
- Estimated Expiration
- 2044-08-24
AI Technical Summary
Existing augmented reality systems in vehicles often distract drivers with irrelevant or distracting environmental objects, reducing driving safety and comfort.
A method that classifies environmental objects into 'relevant', 'deflecting', and 'sedative' categories, masking deflecting objects and highlighting relevant ones using augmented reality displays, enhanced by AI-generated content and user preferences.
Enhances driving safety and comfort by reducing distractions, allowing drivers to focus on critical objects and improving maneuver anticipation.
Smart Images

Figure 00000000_0000_ABST
Abstract
Description
[0001] The invention relates to a method for outputting an augmented reality display in a vehicle according to the type defined in more detail in the preamble of claim 1 and to a vehicle according to the type defined in more detail in the preamble of claim 9.
[0002] While driving, the driver of a vehicle must devote their attention to relevant aspects of the current driving situation. This particularly applies to observing the path ahead and the behavior of nearby road users. When driving a vehicle, there is always a risk of distraction, which can lead to accidents. For example, hectic traffic situations can arise in which road users drive chaotically around each other or pedestrians or cyclists suddenly cross the road. Great care must be taken, particularly in construction zones, to avoid driving past a temporary exit or, due to the narrow roadway, risking hitting neighboring vehicles or a structural barrier. Sensory overload can also occur, for example, if there are many flashing lights or advertising boards in the driver's field of vision.In addition, surrounding advertising can be perceived as distracting or intrusive by vehicle occupants. Furthermore, novice drivers have comparatively less experience with driving, so they must gradually learn which aspects are best focused on in each driving situation.
[0003] Various driver assistance systems are available to support the driver in controlling their own vehicle. Environmental sensors such as cameras, LIDAR, ultrasonic sensors, and / or radar sensors are used to detect static and dynamic objects in the vehicle's vicinity and provide appropriate information to the driver.
[0004] This makes it possible, for example, to provide a blind spot assistant, an emergency braking assistant or the like.
[0005] Machine vision also enables a particularly differentiated assessment of the driving situation. For example, objects detected in camera images can be classified, making it possible to distinguish between traffic signs, traffic lights, cars, trucks, buildings, road markings, and the like.
[0006] Furthermore, the application of augmented reality (AR) in vehicles is well known. For example, vehicles ahead can be highlighted with a marker on a head-up display (HUD) in a similar way to contact.
[0007] Systems and methods for an augmented reality application in a vehicle are also known from US 2021 / 0318135 A1. A navigation route is determined from the vehicle's current location to a destination. Various features of this navigation route are determined, and a map representation of the navigation route is enhanced by displaying additional information. This enhanced map representation is displayed on the vehicle's windshield. Additional information can include, for example, weather information, mobile phone signal strength, references to points of interest along the navigation route, and the like. An image database is stored in a data storage device, with images stored in the image database being projected onto the windshield using the augmented reality projector.Depending on the sensor data generated by the vehicle's sensors, the display of said images from the image database can be adapted to the respective traffic situation with the help of artificial intelligence. Furthermore, DE 10 2023 003 826 A1 describes a method for displaying hidden road users and a vehicle, and DE 10 2022 1 13 342 A1 concerns a selective head-up display for a vehicle with active gaze guidance for the driver in the event of danger.
[0008] The present invention is based on the object of providing an improved method for outputting an augmented reality display in a vehicle, with the aid of which the driving safety and comfort of the person driving the vehicle can be further increased compared to the prior art.
[0009] According to the invention, this object is achieved by a method for outputting an augmented reality representation having the features of claim 1. Advantageous embodiments and further developments as well as a vehicle in which the method according to the invention is carried out emerge from the dependent claims.
[0010] A generic method for outputting an augmented reality representation in a vehicle, wherein the vehicle detects its surroundings using environmental sensors, a computing unit evaluates sensor data generated by the environmental sensors, wherein the computing unit detects and classifies environmental objects present in the environment, and wherein the augmented reality representation is output via a display device in the vehicle, wherein environmental objects detected by the computing unit are masked in a contact-analogous manner, is further developed according to the invention in that the computing unit assigns environmental objects to one of the following classes: relevant, distracting or calming; wherein environmental objects classified as distracting are hidden from the augmented reality representation.The method according to the invention thus provides for masking out irrelevant and even potentially distracting objects from the augmented reality display in order to direct the driver's focus to relevant surrounding objects when viewing the augmented reality display. This supports the driver in steering their vehicle, which can ultimately increase road safety. Masking out irrelevant objects is preferable to highlighting relevant objects, as this results in fewer objects being visible overall in the display, which allows for faster cognitive processing by the user's brain. This means that the driver is less distracted by irrelevant aspects of the traffic situation and perceives relevant aspects with increased concentration.This increases the probability that the driver will be able to assess the driving maneuvers of other road users early and correctly, which allows them to control their own vehicle more reliably.
[0011] The vehicle can, in particular, comprise one or more cameras, LiDARs, ultrasonic sensors, and / or radar sensors as environmental sensors. By analyzing the sensor data, the computing unit can detect and classify corresponding environmental objects. With the help of radar sensors, ultrasonic sensors, LiDARs, and stereo cameras, it is possible to generate depth information and thus locate environmental objects relative to the vehicle. In addition, characteristic dimensions of the environmental object can be measured for a particular object class. For example, a truck or a tractor is distinguished by characteristic dimensions and a characteristic silhouette compared to, for example, a cyclist, an e-scooter rider, or the like. Particularly preferably, the vehicle comprises at least one environmental camera as a sensor.With the help of proven methods of machine vision, known as "computer vision" (CV), this also enables the visual detection and classification of environmental objects. For each sensor modality, characteristic measurable properties of the respective sensor type can be stored, allowing the classification and assignment of the respective environmental objects to the three classes: "relevant," "distracting," and "calming." For example, a calming sky can be recognized by the large area covered in a camera image and its blue color. Particularly preferably, a corresponding machine learning model of machine vision can be trained to assign the respective recognized image content according to the categories. For example, the machine learning model can be trained using supervised learning.For this purpose, pairs of images and text are fed into the machine learning model, with the text specifying which image content can be recognized in each image. These are so-called labeled datasets.
[0012] Relevant objects can include, for example, roads, roadways, road users, traffic signs, traffic lights, lane markings, and the like. Distracting objects can include, for example, flashing lights, advertising, video screens, billboards, buildings or building facades, road users, and the like. The following environmental objects, in particular, can be classified as calming: objects found in nature such as lakes, trees, mountains, viewpoints, landmarks, the sky, the sea, and the like. In particular, those environmental objects that are related to the execution of the driving task are added to the "relevant" class. In particular, objects that can lead to sensory overload are assigned to the "distracting" class. In particular, objects that contribute to the relaxation of the person driving the vehicle are added to the "calming" class.Environmental objects can also be classified into two or even all three of these categories depending on the situation, for example a road user. Road users can be relevant to traffic, for example a pedestrian or cyclist crossing the vehicle's own path. However, road users can also be distracting, such as a lightly dressed person walking past on the sidewalk. Based on measured relative distances, the trajectory of the road user can be predicted, which allows for corresponding classification into one of the respective categories. If it is detected that the person is crossing the road, they are classified as relevant. However, if it is detected that they are walking along the sidewalk, they are classified as distracting.Visual characteristics can also be taken into account, such as the clothing worn by the person, so that, for example, people dressed in striking clothing are classified as distracting and people dressed “normally” are ignored.
[0013] There are various ways to display the augmented reality display in the vehicle and how to mask out surrounding objects classified as distracting. This will be discussed in more detail below.
[0014] For example, an advantageous development of the method according to the invention provides that, in order to hide environmental objects to be masked, a dummy icon is displayed in a contact-analogous area of the display device to a respective environmental object, the contrast is reduced, the brightness is reduced, the gray value is increased and / or the environmental object is virtually cut out. For example, the environmental object to be masked can be replaced or overlaid with a dummy icon. The dummy icon is displayed in the contact-analogous area to the environmental object to overlay the environmental object. This can be any conceivable shape, pictogram, image, or the like. A rendered animation can also be used as a dummy icon. In the simplest case, the corresponding contact-analogous area of the display device is filled with a geometric object such as a polygon, e.g.a square, a rectangle, a circle, an oval, a star or the like. The corresponding polygon can be monochrome and opaque, for example black. Such a dummy icon can be compared to a black censorship bar for censoring a person's eyes in a photo. However, the distracting environmental object can also be overlaid with an image or a rendered animation. If the object to be masked is a vehicle such as a particularly luxurious sports car, a rarely found classic car or the like, the object can be overlaid with an image of a "boring" vehicle such as a small or mid-size car. This reduces the risk that the person driving the vehicle will look away from what is happening and instead look at the "special" vehicle.
[0015] Reducing the contrast and / or brightness also reduces the visibility of the contact-analog area on the display, thus also reducing the visibility of the distracting ambient object. Increasing the gray value of the display in this area also reduces the visibility of the ambient object to be masked.
[0016] Particularly advantageously, the surrounding object to be masked can also be partially or completely cut out of the augmented reality representation for hiding. For example, it is known to cut out image content from an image using suitable algorithms, particularly those using artificial intelligence. The cutout area is then filled with the background visible in the image. For example, if a person or vehicle in front of a house wall needs to be removed from the image area, the corresponding person or vehicle is cut out, and the resulting empty area is filled with the texture and color of the house wall.
[0017] Thanks to the classification into the categories "relevant," "distracting," and "calming," relevant objects are prevented from being removed from the augmented reality display, so that even if distracting environmental objects are hidden, an increase in the risk of accidents is not to be expected. If, for example, a hidden person suddenly runs across the road, they can be quickly highlighted in the augmented reality display as a relevant environmental object, thus changing the class. This will be discussed in more detail later. In particular, the brightness of the entire augmented reality display can be reduced, with the exception of the environmental object to be highlighted.
[0018] According to a further advantageous embodiment of the method according to the invention, it is further provided that a head-up display, AR glasses worn by a vehicle occupant, or a vehicle-integrated display device is used as the display device. Multiple vehicle occupants can also perceive a respective augmented reality display using multiple corresponding display devices. For example, the person driving the vehicle can perceive the surroundings and the augmented reality display via a HUD. The projection surface of the HUD is preferably integrated into the windshield of the vehicle. The person driving the vehicle or any vehicle occupant could also wear corresponding AR glasses, such as a so-called head-mounted display (HMD).
[0019] Such AR glasses or augmented reality glasses comprise means through which the environment and corresponding additional virtual content can be perceived simultaneously. For example, such AR glasses can comprise a transparent display surface for each eye, through which the environment can be viewed, and onto which respective additional content can be displayed. It is also conceivable for the AR glasses to include a camera with which the environment is filmed. The corresponding camera feed is then displayed together with the additional virtual content on an opaque display surface. The augmented reality representation can also be projected onto the retinas of the vehicle occupants using AR glasses.
[0020] The augmented reality display can also be presented via a vehicle-integrated display device. This could be, for example, a display used as an instrument cluster, the head unit display, a dedicated passenger display, or the like. The vehicle's surroundings are captured using one or more of the vehicle's surround-view cameras. The camera feed generated by this or these cameras is then output to the corresponding display device, with the surrounding objects masked accordingly.
[0021] A further advantageous embodiment of the method according to the invention further provides that environmental objects classified as relevant and / or calming are highlighted in the augmented reality display. This means that not only can irrelevant or distracting environmental objects be hidden, but relevant and / or calming environmental objects can also be highlighted, improving their perceptibility for the driver or vehicle occupants. This allows the driver to specifically direct their focus. For example, a road user traveling ahead or crossing the road can be displayed more clearly, allowing the driver to recognize the road user early on. This makes it easier for the driver to adapt their driving behavior in a timely manner and thus control the vehicle appropriately.For example, the visibility of lane markings can also be increased, making it easier for the driver to steer the vehicle while maintaining lane stability, particularly in poor visibility conditions.
[0022] To highlight environmental objects to be masked, a dummy icon is preferably displayed, a bounding box is displayed, the contrast is increased, the brightness is increased, the environmental object is virtually illuminated and / or the environmental object is enlarged in an area of the display device that is contact-analogous to a respective environmental object. Various modalities for improving the perceptibility of respective environmental objects are thus possible. Here, too, the dummy icon can be formed by any geometric surface such as a circle, ellipse or polygon, or even a pictogram, image or rendered animation. For example, the object to be highlighted can be a small car or mid-size car. The object is then overlaid with the image of a sports car, a convertible, a classic car or the like, which draws the gaze of the person driving the vehicle.This allows the driver to perceive the corresponding surrounding object more reliably. Various options are available for determining which vehicle model from which vehicle manufacturer should be used as the dummy icon. For example, the driver's personal preferences can be taken into account. It would also be conceivable to have an advertising partnership with a specific vehicle manufacturer, so that vehicles produced by that vehicle manufacturer are preferentially used as dummy icons.
[0023] To highlight relevant surrounding objects, they can also be provided with a bounding box. The bounding box can enclose the surrounding object completely or intermittently. The bounding box can be of a suitable thickness and can be illuminated in any color and brightness. The brighter the color and the brighter the bounding box, the more clearly the driver is alerted to the correspondingly masked, or in this case, marked, surrounding object.
[0024] To increase perceptibility, the contrast and / or brightness of the corresponding contact-analogous area on the display device can also be increased.
[0025] In addition, the object can be virtually illuminated. This means that the surrounding object is displayed on the display as if it were actually illuminated by a light. Proven image manipulation algorithms can be used for this purpose. Additionally or alternatively, the surrounding object can also be enlarged on the display. This increases the area of the surrounding object, increasing the likelihood that it will be perceived by the driver.
[0026] The method according to the invention further provides that a generative pre-trained transformer, processing driving scenario data generated by the vehicle, generates an input request for generating image content for a generative AI model for generating image content. The AI model, processing the input request, artificially generates image content. The artificially generated image content is output on the display device in a region of the display device that is contact-analogous to a respective environmental object. The method according to the invention thus provides for the cascade connection of two artificial intelligence-based models for the artificial generation of image content.Using the generative pre-trained transformer, particularly in the form of a so-called large language model (LLM), appropriate input prompts can be generated for a downstream generative AI model to generate image content. It is also conceivable for a downstream generative AI model to generate videos, i.e., image sequences with a fixed or variable number of frames per second. This ensures that consecutively generated images are consistent and do not, for example, change brightness, illumination, disappearance or appearance, or change the size or location of objects.
[0027] The current traffic situation or driving scenario is described using the driving scenario data. The image content generated by the AI model is then used to mask corresponding surrounding objects. Such image content can be used both to hide and to highlight surrounding objects. For example, artificially generated image content can be used as a dummy icon. For example, the aforementioned image of a sports car, mid-size vehicle, or compact car can be generated using artificial intelligence.
[0028] Generative pre-trained transformers and generative AI models for generating image content are very powerful. This means that a wide variety of results can be generated given appropriate input data. For example, the driving scenario can be extensively paraphrased using the generative pre-trained transformer, enabling a high degree of variation for image generation. By adjusting random parameters such as a so-called "seed," the variety of variations can be increased even further.
[0029] It is also conceivable that the generative pre-trained transformer controls the representation of the augmented reality representation, for example by determining which method should be used to mask an environmental object, such as virtual cutting out, overlaying with a dummy icon, reducing the contrast, or the like.
[0030] According to a further advantageous embodiment of the method according to the invention, it is further provided that at least one of the following information is used as driving scenario data: - sensor data generated by the environmental sensors; - a textual description of the driving scenario generated by an image-to-text model, wherein the textual description is generated by the image-to-text model by processing a camera image generated by an environmental camera of the vehicle; and / or - a text input or voice input from a vehicle occupant.
[0031] Thus, various options for generating the driving scenario data are also possible. These options can also be combined with one another as desired. For example, sensor data generated by the environmental sensors can be fed directly as input data to the generative pre-trained transformer. This allows the generative pre-trained transformer to analyze corresponding camera images, LiDAR data, ultrasonic sensor data, and / or radar sensor data itself. With the help of the image-to-text model, a textual description of the driving scenario can also be fed to the generative pre-trained transformer. The image-to-text model can advantageously be a generative AI model itself. Users can also enter text or voice input to describe the driving scenario.Text input can be entered, for example, on a touch-sensitive display in the vehicle or via a mobile device linked to the vehicle, such as a smartphone, tablet computer, laptop, or the like. Microphones can also be installed in the vehicle interior to allow the recording of voice commands. The microphones integrated into mobile devices such as smartphones, wearables, or the like can also be used for this purpose. The user or vehicle occupant can not only describe the driving situation but also transmit their personal preferences to the generative, pre-trained transformer. For example, the vehicle occupant can communicate that they prefer a certain vehicle model from a certain vehicle manufacturer.Objects detected in the environment in the form of small cars or mid-size vehicles can then be overlaid by a corresponding dummy icon, which represents the vehicle favored by the vehicle occupant.
[0032] A further advantageous embodiment of the method according to the invention further provides that the vehicle uses a gaze direction detection device to determine the gaze direction of the person driving the vehicle, and the computing unit carries out the contact-analog masking of the surrounding objects in the augmented reality display depending on the gaze direction. An interior camera arranged in the vehicle interior, for example, can be used as the gaze direction detection device. Such an interior camera can be aimed at the eye area of the person driving the vehicle and track their gaze direction. Based on the known installation position of the interior camera, it is thus possible to determine which spatial region of the vehicle's surroundings the person driving the vehicle is looking into. Taking into account camera images generated by surrounding cameras, it is then also possible to determine which surrounding objects are being viewed by the person driving the vehicle.This can be used to determine which object classes the driver particularly frequently or rarely views, or which properties a particular object in the environment has. This can, in turn, be used to determine the personal interests of the driver and whether or not the driver is paying attention to the respective traffic-relevant objects. This makes it possible, for example, to uncover the corresponding lack of skill of a novice driver. This information can be used specifically to adapt the display content to be shown in the augmented reality representation. If, for example, the driver particularly frequently views sports cars, other vehicles can be displayed with preference than such a sports car. It would also be possible to specifically hide advertisements detected in the environment.Taking into account the preferences of the person driving the vehicle, advertisements for products that the person driving the vehicle might like cannot be censored or removed.
[0033] According to a further advantageous embodiment of the method according to the invention, it is further provided that the computing unit executes the intensity of the contact-analog masking of the surrounding objects in the augmented reality display depending on the driver state, the driving situation, and / or a manual user input. The intensity describes how many of the surrounding objects in the augmented reality display are to be masked or the frequency with which surrounding objects are masked. This includes both masking distracting surrounding objects and highlighting relevant or calming surrounding objects.
[0034] The driver's state can be recorded using proven methods. For example, the driver's eyes can be recorded. The blinking frequency or the reaction time with which the pupil dilates or constricts in response to changes in brightness can be used as a measure to assess fatigue or alertness. The driver's entire face or countenance can also be recorded with the interior camera, making it possible to detect, for example, when the driver yawns. Data can also be retrieved from third-party devices, such as wearables worn by the driver, such as fitness trackers. If the driver is in a state of comparatively high concentration, it is possible to allow the detection of a larger number of distracting environmental objects, thus reducing the number of distracting environmental objects that need to be masked out.In addition, the number of surrounding objects to be highlighted can be reduced, as it can be assumed that the driver will perceive these objects sufficiently early anyway. Likewise, the coverage of surrounding objects with artificially generated image content can be increased, which potentially distracts the driver more, but does not increase the risk of an accident due to the driver's increased concentration.
[0035] Additionally or alternatively, the driving situation can also be taken into account to control the intensity of the display. Various parameters can be taken into account to describe the driving situation, such as whether the vehicle is currently in manual, semi-automated or even autonomous operating mode, traffic density, weather conditions, and the like. In manual operating mode with a reflective road surface due to rain and a low sun in heavy traffic, for example, a tense driving situation exists. In contrast, autonomous driving on a motorway on a cloudy day with light traffic represents a relaxed driving situation. In this way, distracting surrounding objects can be hidden in the tense driving situation and relevant surrounding objects can be highlighted, and vice versa in the relaxed driving situation.
[0036] In addition, the intensity of the display can be controlled via manual user input. For example, the user or vehicle occupant can change the intensity using a control device such as a physical slider or a virtual rotary control displayed as a graphical control element on a graphical user interface.
[0037] To determine which environmental objects should be shown or hidden first, the respective environmental objects can be additionally enhanced with an evaluation metric during the environmental object classification step. Environmental objects that then exhibit a correspondingly high or low value of the evaluation metric are then preferentially shown or hidden. For example, sports cars detected in camera images using the aforementioned machine vision algorithms can be assigned a high evaluation metric, whereas mopeds are assigned a low evaluation metric. The evaluation metric can be set particularly preferentially for a particular environmental object, taking into account the personal preferences of the respective vehicle occupant.
[0038] In a generic vehicle comprising a computing unit, environmental sensors, and a display device configured to display an augmented reality representation, the computing unit, the environmental sensors, and the display device are configured according to the invention to carry out a method described above. The computing unit accordingly has at least read access to a computer-readable storage medium storing machine-interpretable instructions which, when executed by a processor of the computing unit, cause the processor to carry out the method described above. The vehicle can be any road vehicle such as a car, truck, van, bus, or the like. In general, it would also be conceivable for it to be a rail vehicle, watercraft, or aircraft.
[0039] Further advantageous embodiments of the method according to the invention for outputting the augmented reality representation in a vehicle also emerge from the exemplary embodiments which are described in more detail below with reference to the figures.
[0040] Showing: Fig. 1 a schematic representation of a traffic scene from the perspective of a person driving a vehicle; Fig. 2 a schematic representation of the display content of an augmented reality display of the vehicle according to the invention shown in Fig. 1 according to a first embodiment; Fig. 3 a schematic representation of the display content of the augmented reality representation of the scene from Fig. 1 according to a second embodiment; Fig. 4 a schematic representation of the display content of the augmented reality representation of the scene from Fig. 1 according to a third embodiment; and Fig. 5 a schematic representation of a data flow graph for generating artificially generated image content.
[0041] Fig. Figure 1 shows a representation of an everyday traffic scene in a city from the perspective of a person driving a vehicle. This is a non-enhanced or non-augmented representation. Various environmental objects 1 can be seen, such as traffic lights, lane markings, vehicles, traffic signs, buildings, the sky, an advertising poster, trees, and the like.
[0042] To ensure road safety, the driver should steer the vehicle in such a way as to avoid collisions with any surrounding objects 1. However, depending on the external conditions and their experience, the driver may be more or less distracted, which increases the likelihood of accidents. To prevent this, the driver's focus is specifically influenced by an augmented reality (AR) image displayed on a display device in the vehicle. This augmented reality (AR) image is generated using a method according to the invention and is described in the Fig. 2 to 4 shown in different embodiments.
[0043] According to the invention, the vehicle detects and classifies environmental objects 1 using environmental sensors. The environmental objects 1 are then assigned to one of the following classes: "relevant," "distracting," and "calming." Environmental objects 1 classified as distracting are hidden from the augmented reality display (AR).
[0044] This is in Fig. 2, where the advertising poster is overlaid by a dummy icon 2. The dummy icon 2 is the Fig. The embodiment shown in Figure 2 is a black rectangular box. This prevents the driver from being distracted and / or disturbed by the advertising. The method according to the invention thus makes it possible to implement a type of "ad blocker" for advertising in the surrounding area of a vehicle.
[0045] It is also possible to highlight additional environmental objects 1 classified as relevant and / or calming in the augmented reality display (AR). This is Fig. 3. For example, the color and brightness of the areas occupied by the sky and treetops on the display device can be increased. The intense blue and green colors can reassure the driver. In addition, surrounding objects 1 relevant to the driving situation can be highlighted, such as the Fig. 3 shows the traffic lights and lane markings. The left traffic light in the figure is highlighted by a bounding box 3, while the right traffic light is enlarged.
[0046] Fig. Figure 4 shows a further embodiment in which corresponding surrounding objects 1 are covered by an artificially generated image content. Such an artificially generated image content can, for example, serve to form said dummy icons 2. How Fig. As Figure 4 shows, the small car parked on the left side of the road can thus be replaced with a high-quality, luxurious sports car. Advertising, particularly advertising tailored to the driver, could also be placed on the building's facade. Particularly preferred is the consideration of the driver's preferences when creating the artificially generated image content. These preferences can be proactively queried by the vehicle occupant or entered by the driver themselves, derived from the driver's observed behavior, and the like.
[0047] Fig. 5 shows a data flow graph for generating said artificially generated image content BILD.
[0048] The vehicle according to the invention detects its surroundings using environmental sensors 503, such as ultrasonic sensors, radar sensors, LiDAR, and the like. Corresponding sensor data is processed via an ADAS system 504. A plausibility check 505 of the detected environmental objects 1 can be performed, for example, through sensor fusion.
[0049] In addition, the vehicle records its surroundings with at least one surrounding camera 502. The respective camera images are in Fig. 5 symbolized by the abbreviation "IMG" for "Image". Using a semantic segmentation model 506, areas irrelevant to the driving task are marked in the corresponding camera images. Furthermore, the respective traffic scene can be described using an image-to-text model I2T. The image-to-text model I2T paraphrases what is shown in the respective camera image and thus generates a textual description of the driving scene 501. A respective text is Fig. 5 symbolized by the abbreviation “TXT”.
[0050] In addition, the person driving the vehicle can make a voice input 507, make a text input 508 via a touch-sensitive display device in the vehicle, and / or make a corresponding text input via a mobile device 509. A corresponding textual description of the driving scene and / or the personal preferences 510 is fed, together with the textual description of the driving scene 501, to a generative pre-trained transformer GPT. Using the generative pre-trained transformer GPT, an input prompt PROMPT is generated for a downstream generative AI model MIGM for generating artificial image content. The generative AI model MIGM for generating artificially generated image content BILD can also be referred to as a "multimodal image generation model." Such a multimodal generative AI model can generate results taking into account different types of input data, here taking into account images and text.The pre-trained generative transformer GPT can pass the self-processed input data in unmanipulated or manipulated form to the generative AI model MIGM. For the functionality of such AI models, see MUMU: BOOTSTRAPPING MULTIMODAL IMAGE GENERATION FROM TEXT-TO-IMAGE DATA, Berman, William et al., Researcher and Residence, Sutter Hill Ventures, arXiv: 2406.18790v1 [cs.CV] 26 Jun 2024, https: / / arxiv.org / pdf / 2406.18790.
[0051] The artificially generated image content IMAGE created in this way is then transferred to the corresponding display device for output in the augmented reality representation AR.
[0052] Optionally, the state of the vehicle can also be recorded in a step 511, for example, by analyzing whether the vehicle is controlled manually, semi-automatically, fully automated, highly automated, or even autonomously. For example, with increasing levels of automation, the augmented reality (AR) representation can become increasingly complex.
[0053] Furthermore, in an optional step 512, the state of the person driving the vehicle can be recorded, for example, through gaze tracking, fatigue detection, emotion recognition, and / or by analyzing the driver's speech, as well as a visual analysis. Data from wearables such as fitness trackers or the like can also be used to assess the driver's state. Such wearables can be coupled to the vehicle in the usual way, in particular via a Bluetooth connection, and transmit vital data relevant for assessing the driver's state. If the person driving the vehicle is in a state of high alertness, the display content of the augmented reality representation can be made more complex, and vice versa for a distracted and / or tired driver.
Claims
[1] Method for outputting an augmented reality (AR) representation in a vehicle, wherein the vehicle detects its surroundings with an environmental sensor system (503), and a computing unit evaluates sensor data generated by the environmental sensor system (503), wherein the computing unit detects and classifies environmental objects (1) present in the environment, wherein environmental objects (1) detected by the computing unit are masked in a contact-analog manner, wherein the computing unit assigns environmental objects (1) to one of the following classes: relevant, distracting or calming, and wherein the augmented reality representation (AR) is output via a display device in the vehicle, wherein environmental objects (1) classified as distracting are hidden from the augmented reality representation (AR), characterized by , that a generative pre-trained transformer (GPT) generates an input request (PROMPT) for generating an image content (IMAGE) for a generative Kl model (MIGM) for generating image content by processing driving scenario data generated by the vehicle, the Kl model (MIGM) artificially generates an image content (IMAGE) by processing the input request (PROMPT), and the artificially generated image content (IMAGE) is output on the display device in a contact-analogous area of the display device to a respective environmental object (1). [2] Method according to claim 1, characterized by in that, in order to hide environmental objects (1) to be masked, a dummy icon (2) is displayed in a contact-analogous area of the display device to a respective environmental object (1), the contrast is reduced, the brightness is reduced, the gray value is increased and / or the environmental object (1) is virtually cut out. [3] Method according to claim 1 or 2, characterized by that a head-up display, AR glasses worn by a vehicle occupant or a vehicle-integrated display device is used as the display device. [4] Method according to one of claims 1 to 3, characterized by that environmental objects (1) classified as relevant and / or calming are highlighted in the augmented reality (AR) display. [5] Method according to claim 4, characterized by in order to highlight environmental objects (1) to be masked, a dummy icon (2) is displayed in a region of the display device which is contact-analogous to a respective environmental object (1), a boundary frame (3) is displayed, the contrast is increased, the brightness is increased, the environmental object (1) is virtually illuminated and / or the environmental object (1) is enlarged. [6] Method according to claim 1, characterized bythat at least one of the following information is used as driving scenario data: - sensor data generated by the environmental sensor system (503); - a textual description (501) of the driving scenario generated by an image-to-text model (I2T), wherein the textual description (501) is generated by the image-to-text model (I2T) by processing a camera image generated by an environmental camera (502) of the vehicle; and / or - a text input or voice input from a vehicle occupant. [7] Method according to one of claims 1 to 6, characterized by that the vehicle uses a viewing direction detection device to determine a viewing direction of the person driving the vehicle and the computing unit carries out the contact-analog masking of the surrounding objects (1) in the augmented reality display (AR) depending on the viewing direction. [8] Method according to one of claims 1 to 7, characterized bythat the computing unit carries out the intensity of the contact-analog masking of the surrounding objects (1) in the augmented reality display (AR) depending on a driver state, the driving situation and / or a manual user input. [9] Vehicle comprising a computing unit, an environmental sensor system (503) and a display device configured to display an augmented reality (AR) display, characterized by that the computing unit, the environmental sensor system (503) and the display device are configured to carry out a method according to one of claims 1 to 8.
Citation Information
Patent Citations
Selective head-up display for a vehicle with active gaze guidance for the driver in case of danger
DE102022113342A1
Methods for displaying hidden road users and vehicles
DE102023003826A1