Procedure for interaction with a road user and vehicle

By using environmental sensors and an interior camera to replicate real driver behavior through a display device, enhanced by machine learning, the method improves interaction reliability between vehicles and road users, especially in autonomous scenarios.

DE102023004208B4Active Publication Date: 2025-06-18MERCEDES BENZ GROUP AG
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
DE102023004208
Authority / Receiving Office
DE · DE
Patent Type
Patents
Current Assignee / Owner
Filing Date
2023-10-19
Publication Date
2025-06-18
Estimated Expiration
2043-10-19

AI Technical Summary

Technical Problem

Existing methods for interacting with road users, particularly in autonomous vehicles, lack reliability as they often result in misinterpretation due to limited understanding of the vehicle's intentions, especially under adverse conditions or without a human driver.

Method used

A vehicle uses environmental sensors and an interior camera to monitor surroundings, processing data with a computing unit to control a display device showing a representation of the driver's gestures, facial expressions, and gaze direction, based on real driver behavior recorded by the interior camera, enhanced by a machine learning model trained on various traffic scenarios.

Benefits of technology

This method significantly increases the reliability of interaction by ensuring road users correctly interpret the vehicle's intentions, even in adverse conditions, through a dynamic and context-aware representation of the driver, enabling bidirectional communication.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000000_0000_ABST
    Figure 00000000_0000_ABST
Patent Text Reader

Abstract

The invention relates to a method for interacting with a road user (1), wherein a vehicle (2) monitors its surroundings with at least one environmental sensor, sensor data generated by the environmental sensor is evaluated by an in-vehicle computing unit to detect an interaction situation, and the computing unit controls a display device (3) of the vehicle (2) directed towards the surroundings to display a representation (4) of the person driving the vehicle, wherein the person driving the vehicle interacts with the road user (1) in the representation (4) by performing a specific gesture, facial expression, and / or changing the direction of view. The method according to the invention is characterized in that the representation (4) of the person driving the vehicle is based on camera images that are or were recorded by an interior camera that records the person driving the vehicle.
Need to check novelty before this filing date? Find Prior Art

Description

The invention relates to a method for interaction with a road user of the type defined in more detail in the preamble of claim 1 and to a vehicle of the type defined in more detail in the preamble of claim 11.In road traffic, situations are often encountered that require an interaction of a vehicle-guiding person with another road user, such as, for example, the driver of another car, a cyclist, a pedestrian or the like.For example, two vehicles travel from opposite directions towards a bottleneck which can be simultaneously passed by only one vehicle. Typically, the driver of the vehicle will give priority to the oncoming vehicle on the side of the bottleneck. It is usual to subsequently consider this by a certain gesture, such as, for example, a head kicking, a sinking or the lifting of the fingers from the steering wheel.In another traffic situation, a pedestrian, scooter driver or cyclist would like to cross a road. The vehicle-guiding person of a vehicle travelling on the road would like to stop in order to let the road user pass over the road. In this case, the vehicle-guiding person could like to stop the intention and let the road user pass, likewise by means of a corresponding gesture, for example a wink movement in the direction transverse to the road. In addition, the targeted establishment of an eye contact between the vehicle-guiding person and the road user can signal that the vehicle-guiding person has also perceived the road user.However, there are situations in which the recognition of corresponding gestures, mimicries and the like is only possible with difficulty or not at all for the further road user. Due to the weather, vision through the windshield of a vehicle may be obstructed, for example, when moisture has accumulated on the windshield and reflections occur due to a deep sun. Even in adverse lighting conditions such as in twilight or at night, the further road user cannot look into the vehicle interior if no interior lighting should be switched on. The road user can therefore not perceive the vehicle-guiding person, and thus also corresponding gestures or mimicries.In the future, autonomous driving vehicles will also increasingly participate in the traffic situation. Such vehicles may be controlled by a computer system and therefore do not require a vehicle driver. Separate means are thus to be provided which enable an interaction with the road user in the usual way. Although it would be conceivable to actuate the lighting devices of the vehicle, such as, for example, the short illumination of the high beam, this could be ambiguous for the further road user due to the situation.Methods and means for interacting with road users in the context of automated or autonomous driving are known, for example, from U.S. Pat. No. 2020 / 0 223 352 A1. The publication describes an autonomous vehicle having a display device oriented toward the front of the vehicle, for example integrated into the windshield. An avatar may be displayed on the display device, which avatar is intended to represent a non-present vehicle-guiding person of the autonomous vehicle. The vehicle is able to detect its environment with the aid of sensors and thus to identify the presence of further road users. If the traffic situation requires an interaction with the road user, the avatar can carry out a specific reaction, such as showing a specific gesture or mimic. The reaction of the avatar is derived from the control behavior of the autonomous vehicle. Supplementary information can be represented in the form of a text message, for example. The appearance of the avatar can be adapted to the preferences of a vehicle occupant. Since the avatar reacts as a function of the control commands of the autonomous vehicle, only a greatly restricted interaction with the road user is possible due to the limited scope of generally executable driving maneuvers. This increases the risk that a response that is inappropriate in the respective traffic situation is output, which response can be incorrectly interpreted accordingly by the road user.DE 10 2014 224 484 A1 discloses a method for adapting the external appearance of a motor vehicle. In a first step, an automatic detection of occupant information relating to a state of a vehicle occupant, a number of the vehicle occupants and / or a seat position of a vehicle occupant takes place, wherein in a second step of the method, an outer element of the motor vehicle influencing the external appearance of the motor vehicle is controlled as a function of the occupant information.US 2023 / 0 400 913 A1 discloses methods for an interaction between an autonomous vehicle and an external observer. Virtual models from the driver can be generated and displayed to the external observer. The virtual models may facilitate interaction between the external observer and the autonomous vehicle through gestures or other visual cues.US 2022 / 0 153 207 A1 describes a method for communication between a vehicle and a human traffic participant. For this purpose, an interaction with the road user is recorded on the basis of information recorded with a vehicle sensor. The interaction with the road user takes place by means of a representation of human gestures generated by a human-machine interface.The object of the present invention is to specify an improved method for interaction with a road user, which is distinguished by an increased reliability of the interaction between a vehicle-guiding person or his vehicle and a further road user. This means that the further road user, compared to the prior art, understands the extension of a message output by the vehicle-guiding person or the vehicle more frequently correctly.According to the invention, this object is achieved by a method for interaction with a road user having the features of claim 1. Advantageous embodiments and refinements and a vehicle for carrying out the method are evident from the claims dependent thereon.A method of the generic type for interaction with a road user, wherein a vehicle monitors its environment with at least one environment sensor, sensor data generated by the environment sensor are evaluated by a vehicle-internal computing unit for detecting an interaction situation, and the computing unit controls a display device of the vehicle directed into the environment for the representation of a representation of the vehicle-guiding person, wherein the vehicle-guiding person interacts in the representation with the road user by performing a specific gesture, mimic and / or changing the viewing direction, is further developed according to the invention in that the representation of the vehicle-guiding person is based on camera images which are recorded by means of an interior camera detecting the vehicle-guiding person.In contrast to the solution known from the prior art, the behavior of the vehicle-guiding person in the representation shown on the display device is not based on the control behavior of an autonomous vehicle, but rather on the actually carried out reaction of a real vehicle-guiding person in road traffic. Such reactions are or have been recorded by means of the interior camera in a wide variety of traffic situations. Thus, the response is based on the real behavior of a human, which increases the probability that the response output in the representation is easier to interpret and thus understand for the road user in the respective traffic situation. This correspondingly increases the reliability with which the traffic participant understands the message underlying the interaction or reaction.The vehicle can be any road vehicle, such as a car, truck, transporter, bus or the like. The vehicle may have a wide variety of environmental sensors, such as, in particular, one or more environmental cameras, ultrasonic sensor systems, radar sensor systems and / or laser scanners such as a LiDAR. Sensor data generated by the environmental sensors are processed by the computing unit. This allows the computing unit to recognize a corresponding interaction situation. For this purpose, the computing unit can also evaluate and take into account sensor data that are generated by other sensors of the vehicle, such as a wheel rotational speed sensor, a steering angle sensor, an acceleration sensor and the like.By evaluating the sensor data generated by the environmental sensors, the vehicle or the computing unit can capture and analyze its environment. This makes it possible to recognize static and dynamic objects and ultimately thus also the further road users. Some sensor systems, such as LiDARe, allow depth information to be generated and the environment to be measured therewith. Thus, the dimensions of the road user can be determined and, based thereon, it can be concluded what type of road user is. Thus, for example, pedestrians, cyclists and trucks have different dimensions. Evaluating camera images allows a particularly differentiated evaluation of the traffic situation, since using proven image recognition algorithms, objects in the environment of the vehicle can not only be recognized but can also be classified. This enables a particularly reliable determination of the type of road user. In addition, the context of the traffic situation can be detected and evaluated, for example taking into account traffic signal systems detected in the scene, traffic signs, roadway markings, gestures performed by the road user, mimetics and the like. Thus, not only can messages be communicated to the road user by the representation displayed on the display device, but also messages still returned by the road user can be recognized and interpreted.Taking into account the other vehicle sensor data, such as odometry data, it can then be detected, for example, that the vehicle-guiding person is braking, which is an indication that the vehicle-guiding person wishes to leave the road user via the road.A fleet of vehicles of a vehicle manufacturer can be operated for a period of time in order to carry out the method according to the invention, in order to identify gradually typical interaction situations in the everyday life between vehicles and road users and to record the interaction behavior which has occurred in this case. This information can be distributed accordingly to the computing units of the vehicles, so that the respective computing units are enabled to reliably identify interaction situations in the most varied traffic situations.The vehicle may have a wide variety of outwardly directed display devices. Such a display device can be integrated into the exterior of the vehicle at any point, for example connected to an outer skin element such as the engine hood, a fender, a bumper, the grille or the like. The display device can also be integrated into the windshield or another pane of the vehicle or be arranged behind this pane from the point of view of the road user.Particularly advantageously, a plurality of display devices are provided at different locations on the vehicle, so that the representation of the vehicle-guiding person can be perceived from different whereabouts by the road user.According to the invention, the method provides that a camera image stream generated by the interior camera is transmitted live to the display device in the interaction situation. This increases the reliability of the interaction between the road user and an actually present vehicle-guiding person of the vehicle in a situation in which the vehicle-guiding person could not be correctly recognized by the road user without providing the display device through a view through the windshield of the vehicle. As mentioned at the outset, this could be the case, for example, in darkness and in the absence of interior lighting or in the case of bright reflections on the windshield. However, by operating the display device, the road user can view the vehicle-guiding person live and thus understand the gestures, mimicries and / or changes in viewing direction carried out by the vehicle-guiding person. Particularly preferably, the display device emits light itself, which also allows reliable observation of the vehicle-guiding person in the event of darkness.In particular, the camera images of the camera image stream can be processed by the computing unit before output on the display device. For example, the contrast, the brightness and the like can be changed, in particular increased, zoomed into a specific section of the camera image, and the like. In addition, the interior camera can be formed by an infrared camera, which allows the vehicle interior and thus also the vehicle-guiding person to be recognized in the event of darkness. In this case, the vehicle interior can additionally be actively illuminated by an infrared light source during the recording of the camera image stream.According to a further advantageous embodiment of the method according to the invention, the computing unit feeds at least some of the sensor data as input data to a machine learning model which, as output data, in particular during an automated driving operation, determines the representation of the vehicle-guiding person to be displayed on the display device, wherein the machine learning model has been trained in an initial training phase, wherein sensor data of the environment generated as input data in the training phase and camera images of the interior camera correlating in time therewith are fed to the machine learning model in the training phase. This also enables a representation of the representation of a virtual vehicle-guiding person if no vehicle-guiding person is actually present at all and the vehicle drives automatically or autonomously. The use of a machine learning model for ascertaining the representation thus allows the virtual representation of a vehicle-guiding person in the context of automated or autonomous driving. The machine learning model determines directly or indirectly the response to be output in the representation to be displayed on the display device, and is able to determine particularly suitable responses for the respective traffic situation on the basis of the training. This means that the vehicle is able to display in the respective traffic situation exactly reactions of a "vehicle-guiding person" on the display device of the type that a real vehicle-guiding person would actually also carry out in the traffic situation. This further improves the interaction behavior between the vehicle and the road user for autonomous vehicles.In particular, the machine learning model is independently capable of identifying corresponding interaction situations. This capability is provided to the machine learning model in addition to the capability of determining appropriate responses through the training itself. Thus, the machine learning model associates the traffic situation evaluated by analyzing the sensor data with the behavior of the real vehicle driver that is currently observed by analyzing the interior camera images. If corresponding gestures, mimicries and the like are observed, an interaction situation is present.In the training phase, the camera images of the interior camera (indirect determination of the reaction) and / or information derived therefrom (direct determination of the reaction) are used as ground truth. Training of the machine learning model is completed once the machine learning model is able to predict appropriate responses for a "vehicle-guiding person" in a certain number of traffic situations. For this purpose, the vehicle manufacturer can define an absolute or relative frequency for reactions determined by the machine learning model, which must match the reaction carried out by a real person. In test series, it is then checked with what frequency the machine learning model in the respective test situation has actually determined such a "real" reaction.In this case, the machine learning model preferably predicts, as output data, the image content to be displayed on the display device in the form of a camera image stream. This is the indirect determination of the reaction. Thus, the machine learning model directly determines the display content to be displayed on the display device. For this purpose, the machine learning model is advantageously designed as a convolutional LSTM network or by a division model. As a loss function for training, a pixel-by-pixel MSE loss may be used. For this purpose, the camera images of the interior camera are used directly as ground truth in the training phase.An alternative approach, on the other hand, provides thatthe machine learning model itself determines the mimic, gesture and / or viewing direction of the vehicle-guiding person from the respective camera images in the training phase or accesses the mimic, gesture and / or viewing direction determined from the camera images using an algorithm and links these to the traffic situation described by the sensor data; andthe vehicle-guiding person is represented on the display device by an avatar to which a gesture, mimic and / or viewing direction matching the current traffic situation by the machine learning model as a function of the processed sensor data is impressed.The machine learning model is itself configured to determine the user behavior represented by mimic, gesture and / or viewing direction; alternatively, this can also be done by classical algorithms accessed by the machine learning model.This is a prediction of the response or of the user behavior determined from historical data. The machine learning model determines, as a function of the traffic situation, the user behavior to be impressed on an avatar to be displayed on the display device, said user behavior being represented by mimic, gesture and / or viewing direction. The user behavior to be impressed on the avatar is effected, for example, using a model of the generative artificial intelligence.For this purpose, information derived from the camera images is supplied to the machine learning model in the training phase as ground truth. Here, too, the machine learning model can advantageously be formed by an artificial neural network. It can likewise be a convolutional LSTM or a diffusion model.For recognizing the mimic, gesture and viewing direction of the vehicle-guiding person, conventional methods can be used. Such methods are usually based on the recognition and tracking of characteristic features within the face of the detected person. The position of these characteristic features can be described here by scalars, such as X and Y coordinates, for example. Preferably, an infrared camera, optionally with active illumination, is used to track the viewing direction of the vehicle-guiding person. The viewing direction can likewise be expressed as a scalar, for example by two angle specifications. The detection of the body pose of a human based on artificial intelligence is also known from the prior art. Body poses can also be described as scalars.Although the representation of the representation of the vehicle-guiding person as avatar is accompanied by an increased processing effort of corresponding data, it allows a more flexible and more extensive adaptation of the display content to be displayed.A further advantageous embodiment of the method according to the invention further provides that the computing unit follows the viewing direction of the avatar relative to the vehicle during the interaction with the road user to the current location of the road user. As already mentioned, the vehicle is able to locate the road user relative to the vehicle by evaluating the sensor data. The installation position of the display device on the vehicle is known, so that the computing unit can calculate the viewing direction of the avatar such that the avatar views the road user. This illustrates to the road user that the vehicle has detected the road user in the respective traffic situation by means of his sensor system. This gives the road user a certainty that the vehicle will avoid an accident with the road user, in particular in the case of autonomous control.According to a further advantageous embodiment of the method according to the invention, the appearance of the avatar is adapted by the computing unit. In the simplest case, the avatar is merely a 2D image. For corresponding gestures, different images are displayed on the display device. Preferably, however, it is an animated avatar, in particular in the form of a three-dimensional rendered representation. The appearance of the avatar may be changed. This can be done either automatically by the computing unit or else manually by a vehicle occupant. For this purpose, the vehicle occupant can input corresponding commands via an in-vehicle human-machine interface, such as a touch-sensitive display device. For example, the sex of avatar, the skin color, the style, the face shape, the clothes, and the like can be changed. Also, the abstraction degree of the avatar may be changed. Thus, the avatar can either be represented in a styled and simplified manner, for example as a combi figure, or else in a realistic manner, in particular using high-resolution photorealistic textures. In particular, the vehicle occupant, for example in the form of the vehicle-guiding person, can be captured by means of the interior camera and the appearance of the vehicle occupant can be transferred to the avatar.The computing unit can automatically adapt the appearance of the avatar, for example, as a function of the current weather. For example, if the sun is especially looking, the avatar may wear a sun hat and a pair of sunglasses. If, on the other hand, the avatar could wear a rain jacket with a raised hood. This can increase the acceptance for interaction with the avatar for the road user.A further advantageous embodiment of the method according to the invention further provides that the vehicle outputs an acoustic message for interaction with the road user by means of external loudspeakers, in particular a message recorded by the vehicle-guiding person by means of a microphone. The recorded message can also be output live parallel to the camera image stream. If the camera image stream generated by the interior camera is thus transmitted live to the display device as display content, the road user can conduct a conversation with the vehicle occupant in a way. Thus, the vehicle-guiding person can not only give the road user the intention to control the vehicle via gestures, mimicses and viewing directions, but also additionally by voice. For example, the car driver could say "I let you pass. They can cross the road.". The vehicle can have acoustic detection means such as external microphones, so that a response of the road user in the vehicle interior can also be output via loudspeakers. This particularly preferably also allows bidirectional communication.By contrast, prerecorded or computer-generated voice messages can also be output to the road user via the external loudspeakers of the vehicle. This is suitable in particular for acoustic communication with the road user in an autonomously controlled vehicle.According to a further advantageous embodiment of the method according to the invention, the mouth of the avatar is moved during the output of the acoustic message, in particular synchronously with words contained in the acoustic message. This further improves the interaction with the road user. In particular, if the avatar is displayed on the display device in a realistic representation, the interaction sequence of the road user with the vehicle is thereby carried out particularly naturally. This can further improve the acceptance for the road user for interaction with the vehicle.A further advantageous embodiment of the method according to the invention further provides that the computing unit detects a reaction of the road user by analyzing the sensor data and then adjusts the representation of the representation of the vehicle-guiding person on the display device and / or outputs an acoustic response message via external loudspeakers of the vehicle. The response of the road user can likewise be formed by a gesture, mimic, viewing direction and / or spoken language. The road user and the vehicle-guiding person or the virtual vehicle-guiding person can thus interact. For example, the avatar may perform a wink movement to illustrate to the road user that he can safely pass the road. In this case, the road user could lift the hand and pitch to think about. The avatar's display can then be adjusted to the effect that the avatar closes the eyes and smudging them. This further improves the interaction between the road user and the vehicle.The road user could also call in the direction of the vehicle: "I cannot recognize him correctly", whereupon the computing unit adjusts display parameters of the representation output on the display device. For example, the brightness or the contrast can be increased and / or switched to an infrared mode in a live camera image of the interior camera.For the determination of acoustic response messages to be output to the road user by the vehicle, prefabricated response messages for the respective traffic situation can be read out from a data memory. Response messages can, however, also be generated anew for the respective traffic situation. Methods of natural language processing, also referred to as natural language processing (NLP), or also generative AI models, for example large language models, also referred to as large language model (LLM), can be used for this purpose.In a vehicle comprising at least one environment sensor, an interior camera, a computing unit and at least one display device directed into the environment, according to the invention the environment sensor, the interior camera, the computing unit and the display device are configured to execute a method described above.The vehicle preferably has at least two display devices oriented differently with respect to the environment, wherein the computing unit is configured to locate the road user in the environment while processing the sensor data and to actuate at least that display device for displaying the representation of the vehicle-guiding person which faces the road user. This increases the reliability that the road user can actually also visually capture the representation of the vehicle-guiding person.It is also conceivable to display the representation of the vehicle-guiding person on a plurality of display devices simultaneously. The representation can preferably be adapted to the respective orientation of the display devices. If, for example, a display device is integrated into the windshield of the vehicle and a display device is integrated into the rear window of the vehicle, the representation of the vehicle-guiding person on the windshield can show the vehicle-guiding person frontally and from behind on the rear window. This ensures a particularly natural interaction between road user and vehicle.Further advantageous embodiments of the method according to the invention for interaction with a road user and the vehicle also result from the exemplary embodiments which are described in more detail below with reference to the figures.The following are shown: FIG. 1 shows a schematic top view of a vehicle according to the invention in a first traffic situation; and FIG. 2 shows a schematic top view of a vehicle according to the invention in a second traffic situation.Frequently, the vehicle-guiding person of a vehicle 2 shown in FIG. 1 interacts with a road user 1 by gestures, mimicses and / or changing the viewing direction. This requires, on the one hand, the presence of a corresponding vehicle-guiding person and the ability of the road user 1 to be able to see the vehicle-guiding person within the vehicle 2 as well. However, an autonomously controlled vehicle lacks the vehicle-guiding person. In addition, a vehicle-guiding person actually present cannot be detected from the outside, for example at night in the absence of vehicle interior lighting.In order nevertheless to enable an interaction between the vehicle 2 or the vehicle-guiding person and the road user 1, a method according to the invention for interaction with the road user 1 is carried out by the vehicle 2.FIG. 1 shows a first traffic situation in which the vehicle 2 is approaching a road user 1 in the form of a pedestrian on a road 6, wherein the pedestrian wishes to cross the road 6. The embodiment shown in FIG. 1 is a manually controlled vehicle 2.The vehicle-guiding person has recognized the road user 1 and wishes to leave it over the road 6. The vehicle-guiding person brakes the vehicle 2 and performs a corresponding wink gesture. Since it is night and therefore dark, the road user 1 cannot normally recognize the vehicle-guiding person.However, the vehicle 2 has a display device 3 directed into the environment, here integrated into the windshield. The vehicle 2 detects its environment with the aid of surrounding sensors, not shown in more detail, and is thereby able to detect and evaluate the current traffic situation. Corresponding sensor data are evaluated by a computing unit, not shown in more detail. The computing unit also serves for controlling the display device 3. The vehicle 2 or the computing unit recognizes an interaction situation in the current traffic situation, which requires the activation of the display device 3 and the presentation of a representation 4 of the vehicle-guiding person, see FIG. 2.In the exemplary embodiment shown in FIG. 1, the person driving the vehicle 2 is detected using a passenger compartment camera, not shown in detail, and the camera image stream generated by the passenger compartment camera is output live on the display device 3. Thus, the road user 1 can still recognize the vehicle-guiding person and perceive the winking movement. The interior camera is advantageously designed as an infrared camera and the vehicle interior is actively illuminated with an infrared light source during the camera image activation. Thereupon, the road user 1 can start moving and cross the road 6.FIG. 2 shows a further exemplary embodiment in which the vehicle 2 and a road user 1 in the form of an oncoming vehicle wish to pass by a bottleneck 7. In this case, the road user 1 should wait and the vehicle 2 should first pass through the bottleneck 7, since the corresponding obstacle 8 is in the form of a parked car on the roadway side of the road user 1.The vehicle 2 is shown at two different times. At the position P 1, the vehicle 2 is approaching the obstacle 8. At the position P 2, the vehicle 2 has just passed the obstacle 8. Typically, the vehicle-guiding person of the vehicle 2 would be thought of at the position P 2 at the road user 1 by a corresponding gesture.FIG. 2 also shows the representation 4 shown on the display device 3. Thus, no real car-guiding person is present. Instead of the vehicle-guiding person, an avatar 5 is displayed on the display device 3, which avatar lifts the hand when passing the obstacle 8. This conveys the assistance to the road user 1.In this exemplary embodiment, the computing unit of the vehicle 2 executes a machine learning model which reads in the sensor data of the surrounding sensors generated by the vehicle 2 as an input variable. As an output variable, the machine learning model supplies either the image content to be displayed on the display device 3 directly or a corresponding gesture, mimic and / or viewing direction to be impressed on the avatar 5 in the representation 4.In a training phase preceding the deployment phase, the machine learning model receives as input data the sensor data generated by the surrounding sensors and camera images generated by the interior camera, showing a person actually present and controlling the vehicle 2. As a result, the machine learning model gradually learns, on the one hand, to recognize respective interaction situations and, on the other hand, to generate information for generating the representation 4. For this purpose, as ground truth in the training phase of the machine learning model, either the camera images of the interior camera itself or information derived therefrom in the form of the gesture, mimic and / or viewing direction actually performed by the vehicle-guiding person are used.

Claims

Method for interaction with a road user (1), wherein a vehicle (2) monitors its environment with at least one environment sensor, sensor data generated by the environment sensor are evaluated by a vehicle-internal computing unit for detecting an interaction situation, and the computing unit controls a display device (3) of the vehicle (2) directed into the environment for displaying a representation (4) of the vehicle-guiding person, wherein the vehicle-guiding person interacts with the road user (1) in the representation (4) by carrying out a specific gesture, mimic and / or changing the viewing direction, wherein the representation (4) of the vehicle-guiding person is based on camera images which are recorded by means of an interior camera detecting the vehicle-guiding person, characterized in that a camera image stream generated by the interior camera is transmitted live to the display device (3) in the interaction situation.Method according to Claim 1, characterized in that the arithmetic unit feeds at least some of the sensor data as input data to a machine learning model which, as output data, determines the representation (4) of the vehicle-guiding person to be displayed on the display device (3), wherein the machine learning model has been trained in an initial training phase, wherein sensor data generated as input data in the training phase and camera images correlated with respect thereto in time are fed to the machine learning model in the training phase of the interior camera.Method according to Claim 2, characterized in that the machine learning model predicts, as output data, the image content to be displayed on the display device (3) in the form of a camera image stream.Method according to Claim 2, characterized in that - the machine learning model in the training phase determines the mimic, gesture and / or viewing direction of the vehicle-guiding person in the respective camera images or accesses the mimic, gesture and / or viewing direction determined from the camera images using an algorithm and links this with the traffic situation described by the sensor data; and - the vehicle-guiding person is represented on the display device (3) by an avatar (5), to which a gesture, mimic and / or viewing direction matching the current traffic situation by the machine learning model as a function of the processed sensor data is impressed by the computing unit.Method according to Claim 4, characterized in that the arithmetic unit tracks the viewing direction of the avatar (5) during the interaction with the road user (1) to the current location of the road user (1) relative to the vehicle (2).Method according to claim 4 or 5, characterised in that the appearance of the avatar (5) is adapted by the computing unit.Method according to one of Claims 1 to 6, characterized in that the vehicle (2) outputs, by means of external loudspeakers, an acoustic message for interaction with the road user (1), in particular a message recorded by the person carrying the vehicle by means of a microphone.Method according to claim 7 and one of claims 4 to 6, characterised in that the mouth of the avatar (5) is moved during the output of the acoustic message, in particular synchronously with words contained in the acoustic message.Method according to one of Claims 1 to 8, characterized in that the arithmetic unit detects a reaction of the road user (1) by analyzing the sensor data and then adapts the representation of the representation (4) of the vehicle-guiding person on the display device (3) and / or outputs an acoustic response message via external loudspeakers of the vehicle (2).Vehicle (2) comprising at least one environment sensor, an interior camera, a computing unit and at least one display device (3) directed into the environment, characterized in that the environment sensor, the interior camera, the computing unit and the display device (3) are configured to carry out a method according to one of Claims 1 to 10.Vehicle (2) according to Claim 10, characterized byat least two display devices (3) which are oriented differently with respect to the environment, wherein the arithmetic unit is configured, while processing the sensor data, to locate the road user (1) in the environment and to actuate at least that display device (3) which faces the road user (1) for the purpose of presenting the representation (4) of the vehicle-guiding person.

Citation Information

Patent Citations

  • Method for customizing the exterior appearance of a motor vehicle and motor vehicle with an customizable exterior appearance

    DE102014224484A1

  • System and method for providing automated digital assistant in self-driving vehicles

    US20200223352A1

  • Communication between a vehicle and a human road user

    US20220153207A1

  • Virtual models for communications between user devices and external observers

    US20230400913A1