Method for controlling an audiovisual communication interface

A neural network-based method for optimizing audiovisual communication in vehicles determines suitable times for activities based on driver state, minimizing distractions and enhancing safety by scheduling or postponing activities.

DE102024133096B3Active Publication Date: 2026-02-12DR ING H C F PORSCHE AG
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
DE102024133096
Authority / Receiving Office
DE · DE
Patent Type
Patents
Current Assignee / Owner
Filing Date
2024-11-12
Publication Date
2026-02-12
Estimated Expiration
2044-11-12

AI Technical Summary

Technical Problem

Existing audiovisual communication interfaces in vehicles do not consider the driver's state, leading to potential distractions and safety risks by delivering information at inopportune moments.

Method used

A method using a neural network trained with annotated context and vehicle data to determine suitable times for communication activities based on the driver's cognitive capacity, allowing activities to be scheduled or postponed to minimize distractions.

Benefits of technology

Reduces driver distractions and improves road safety by ensuring communication activities occur at optimal times, reducing the risk of accidents and eliminating the need for additional sensors.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000000_0000_ABST
    Figure 00000000_0000_ABST
Patent Text Reader

Abstract

The invention enables an improvement in road safety during journeys with a motor vehicle by means of a method (1000) for controlling an audiovisual communication interface in a driver-controlled motor vehicle, wherein, with the aid of a neural network trained with training data, a time at which an activity, in particular a communication activity, of the communication interface takes place is determined based on context data and / or vehicle data, and / or a time for the activity is evaluated, wherein the determined time and / or the evaluation of the time is related to an available cognitive capacity of a driver of the motor vehicle, wherein the training data comprises annotated context data and / or vehicle data from several participants, and wherein the annotations of the training data are based on assessments, in particular by the participants, of real situations.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] The invention relates to a method for controlling an audiovisual communication interface. State of the art

[0002] US20070063854A1 describes a system for adaptively estimating a driver's workload based on machine learning.

[0003] US20190375426A1 describes an adaptive driver assistance system based on machine learning.

[0004] DE102023004039A1 describes a system for detecting driver workload. This system uses sensors in the vehicle that collect data on driver behavior, vehicle handling, and other relevant characteristics, such as steering wheel angle, vehicle speed, acceleration, and lane change data.

[0005] GB2504583A describes a method to determine driver load from observed vehicle, driver and environment data for the adaptation of a human-machine interface.

[0006] DE 11 2018 000 968 B4 discloses an image display system that displays information about a behavioral assessment of a vehicle, in particular about future positions of the vehicle and surrounding vehicles. Disclosure of the invention

[0007] Modern vehicles continuously communicate data to their drivers using audiovisual communication interfaces. These include, for example, visual elements such as pop-up messages, dynamic user interfaces, or warning lights, and auditory cues such as beeps or voice assistant output.

[0008] Currently, this data is typically passed on to the driver immediately, without considering the driver's state at that time. This can be suboptimal, as the data may be made available to drivers at inconvenient times when they need to concentrate on the primary driving task. For example, communication interface activities, such as issuing messages and / or initiating interaction with the driver, can cause interruptions and distractions between the communication interface and the driver at an inopportune moment, thus posing a safety risk.

[0009] The previously mentioned, known methods of detecting driver workload using sensors or similar devices can only be used in appropriately equipped vehicles. Consequently, there are sometimes considerable difficulties in continuously providing the required data.

[0010] The object of the present invention is therefore to offer a cost-effective method by which the road safety of a motor vehicle can be improved during activities of its audiovisual communication interface.

[0011] The problem is solved by a method for controlling an audiovisual communication interface in a driver-operated motor vehicle, wherein a neural network trained with training data determines, based on context data and / or vehicle data, a time at which an activity, in particular a communication activity, of the communication interface takes place, and / or a time for the activity is evaluated, wherein the determined time and / or the evaluation of the time is related to an available cognitive capacity of a driver of the motor vehicle, wherein the training data includes annotated context data and / or vehicle data from several participants, and wherein the annotations of the training data are based on assessments, in particular by the participants, of real-world situations.

[0012] Accordingly, one aspect of the invention is to determine a suitable time at which the activity can be performed so that the driver is only minimally distracted from the traffic situation. Such a suitable time could be, for example, when the driver is waiting at a red light and is not otherwise occupied. The activity can then be scheduled for this suitable time.

[0013] Alternatively or additionally, it is provided that a point in time, for example the current time, is assessed to determine whether it is suitable for carrying out the activity. If the time is unsuitable, the activity can be postponed. If the time is suitable, the activity can be carried out.

[0014] This helps prevent driver overload and distractions, thus avoiding driver errors. The risk of accidents can be reduced, and overall road safety can be improved.

[0015] Another idea is to use a neural network to evaluate the context data and / or the vehicle data. This would allow even complex, for example non-linear, relationships to be taken into account.

[0016] The neural network can also be designed to learn autonomously. In particular, the neural network can be designed for continuous improvement, for example, through further supervised learning. Network weights of the neural network can be continuously optimized in a central system. The optimized network weights can then be regularly updated and integrated into the neural networks of the respective vehicles. This allows for continuous improvement in the planning and execution of activities.

[0017] Another consideration is that training data is used as the starting point for the neural network, particularly its network weights. This training data is based on annotations derived from assessments of real-world situations. This allows for the linking of real-world situations with participants' subjective assessments. Furthermore, this enables the neural network's behavior, defined as network weights and / or network structure, to be tailored to the driver's actual perceived workload. By incorporating these assessments, a direct measure for annotating real-world situations can be used. It eliminates the need to estimate driver workload using additional sensor data for training purposes. As a result, the training data can provide particularly accurate predictions of a driver's subjective workload.It is to be expected that subjective assessments of the stress situation correlate particularly strongly with potential errors in behavior, for example due to distractions or interruptions. This also allows for the consideration of interindividual differences, such as those arising from varying levels of driving experience, which may manifest themselves in different degrees of automation of processes.

[0018] Furthermore, additional external sensors can be dispensed with, making the process particularly cost-effective overall.

[0019] The training data can include annotated datasets with contextual and vehicle data from multiple participants. The training data may have been recorded during real-world driving experiences.

[0020] The training data can be annotated by regularly asking participants, for example every 30 to 120 seconds, whether the respective times are favorable or unfavorable times to perform an activity, such as issuing a message.

[0021] For example, contextual data describing the landscape to which the driver is exposed can be particularly interesting for assessing whether a moment is suitable or not.

[0022] The contextual data can include, for example, weather data, traffic data, land use data, time of day data and / or road characteristics.

[0023] This data is often already available in the vehicle at high temporal resolution and / or can be easily retrieved from the internet, for example, and therefore does not require any additional sensor equipment.

[0024] The vehicle data can include speed data, in particular acceleration data, especially data from an inertial measurement unit (IMU), steering angle data, brake and / or accelerator pedal angle data, satellite navigation data such as GPS data, and / or occupancy information. This data is also often already available in the vehicle at high temporal resolution.

[0025] Furthermore, such contextual and vehicle data offer the advantage that, unlike techniques that rely on video data, they can derive the suitability of moments in a privacy-friendly manner.

[0026] The vehicle data can be read via a CAN bus, especially one internal to the vehicle, thus avoiding the need for additional device installations and saving costs.

[0027] The context data can be retrieved from a remote computer system, for example by calling a web-based API and especially via the Internet.

[0028] It is conceivable, in particular, that the determination and / or assessment of the timing depends on the type of activity. For example, if a message is to be issued, the urgency of the message can be taken into account. Messages concerning chat messages, emails, press releases, or similar communications received by the driver, for instance, might be assigned a low urgency rating. Messages relating to immediate traffic or the vehicle's technical aspects, such as messages about vehicles traveling the wrong way or malfunctions in the vehicle's propulsion system, could, however, be classified as highly urgent. Messages classified as highly urgent can be issued immediately, particularly regardless of other criteria for assessing the suitability of the timing.

[0029] Messages classified as non-urgent may be postponed until a time deemed suitable for such non-urgent messages.

[0030] The activity could include an interaction with, for example, a voice assistant. It is conceivable to automatically trigger an interaction via context-dependent interruptions, particularly when significant changes in the driver's environment, especially with regard to contextual data, make adjustments necessary or likely. This method would allow the voice assistant to initiate such activities "intelligently" at an appropriate time.

[0031] Once we have combined this data, we can use it as input for the trained model and infer the probability of a suitable moment in real time. The entire pipeline can be executed efficiently and in real time on an embedded system on edge devices such as a car. Based on the calculated probabilities, interaction loops can be deferred or downstream systems triggered.

[0032] The activity can also include switching between the user interfaces of the communication interface. With an increasing number of vehicle functions, audiovisual communication interfaces, especially user interfaces, tend to become more complex. One way to enable "shortcuts" to the functions with which users can interact is through dynamically changing user interfaces that adapt clickable surfaces based on contextual data and user profiles. This method allows switches between different user interfaces to occur at appropriate times, ensuring the driver is not distracted from the road and can also devote sufficient time and attention to the communication interface to familiarize themselves with the changed user interface.

[0033] It is also conceivable that the activity involves the output of a message via the communication interface. The message can be visual, for example, in the form of a pop-up. Alternatively or additionally, the message can also be audible. It is also conceivable that the message is transmitted to the driver haptically. For example, if the driver unintentionally leaves a lane, the steering wheel and / or the seat of the vehicle could vibrate to warn them of the impending crossing of the shoulder.

[0034] The timing for starting an activity, for example, starting an interaction with the voice assistant, can be determined in particular by taking into account the urgency of the activity.

[0035] A neural network can be used as a network topology for edge devices. In particular, it can be designed to run the process offline, for example on a vehicle's onboard computer.

[0036] An offline version allows the determination and / or evaluation of timing to be independent of external influences, such as the availability of an internet connection. This can be particularly important if highly critical activities, such as warnings about technical malfunctions of the vehicle or similar issues, or warnings of impending collisions, are also to be possible via the audiovisual communication interface.

[0037] The audiovisual interface can be part of the on-board computer. Alternatively, the on-board computer can be part of the audiovisual interface. The audiovisual interface can also include a navigation system, a radio, a video player, web access software (e.g., an internet browser), a text-to-speech unit, and / or a speech recognition unit.

[0038] Further features and advantages of the invention will become apparent from the following detailed description of an embodiment of the invention with reference to the figures of the drawing, which show details essential to the invention, as well as from the claims.

[0039] The training data is based on real, annotated situations. It is conceivable that the training data could be augmented with simulated data. For example, it is possible to account for delays or a lack of timeliness, particularly of contextual data. For instance, additional weather data could be generated that is time-shifted to a certain extent and / or modified in its content. For example, in an additional training dataset, the amount and / or probability of precipitation could be altered compared to an original value to account for potential inaccuracies in weather data. Furthermore, additional training data can be generated through modification to enable particularly stable determinations and / or evaluations of times.

[0040] The individual features can be implemented individually or in any combination in various versions of the invention. The schematic drawing illustrates exemplary embodiments of the invention, which are explained in more detail in the following description. Brief description of the drawings The only character ( Fig. Figure 1) shows a method 1000 for controlling an audiovisual communication interface. Embodiments of the invention

[0041] In particular, it shows Fig. 1, such as how activities of the audiovisual communication interface can be controlled during a driver's journey in a motor vehicle.

[0042] This requires that a neural network trained with training data is available. Furthermore, it is assumed that training data for the neural network has already been prepared and stored within the network. As explained in more detail above, the training data comprises annotated contextual data and / or vehicle data from multiple participants. The training data has been compiled as described above. Specifically, participants drove vehicles while sensor data was logged via the vehicle's CAN bus, corresponding to preselected vehicle data such as speed, steering angle, braking, position from the vehicle's satellite navigation system, accelerator pedal angle, acceleration, and occupancy data, for example, based on seatbelt contacts.In addition, contextual data, in particular weather data retrieved via the Internet, traffic data, time of day data, especially whether it is daytime, twilight or nighttime driving, as well as data on road properties, especially also retrieved via the Internet, such as data on road surface dryness, road surface slip properties, especially whether there is a risk of slipping, and traffic volume along the routes driven, are collected during the journeys.

[0043] In subsequent sessions, the participants annotated the collected data.

[0044] It is conceivable that participants could, for example, select at 30-second intervals by pressing a key whether the respective time is considered suitable or unsuitable for an activity of the audiovisual communication interface.

[0045] One variant of the procedure may provide for offering more than two selection options.

[0046] For example, it may be intended that participants assess the suitability for different types of activities, such as issuing a message versus initiating a voice-based interaction between the driver and the audiovisual communication interface.

[0047] In order to offer participants further support during the annotations, it is conceivable to also record video footage of the vehicle interior and / or the surroundings of the vehicle during the journeys and to present this to the participants during the annotations in a timely manner together with the other data.

[0048] During a journey in a motor vehicle, 1010 contextual and vehicle data can be collected in a first step according to procedure 1000. The types of contextual and vehicle data should correspond to those used to generate the training data.

[0049] If the vehicle's audiovisual communication interface is to perform an activity, for example output a message, the time is evaluated according to a second step 1020.

[0050] The timing is evaluated using a previously trained neural network. This network is fed with the currently available contextual and vehicle data. If the neural network determines that the current time, and consequently the currently recorded contextual and vehicle data, are suitable for message output, the message is sent. For example, a voice output is triggered, which audibly delivers the message. Visual output is also possible, for example, on a display of the communication interface.

[0051] If the timing is deemed unsuitable, the audiovisual communication interface waits according to step 1030 until a suitable time arises.

[0052] During the waiting period, data can be continuously collected in accordance with step 1010 and evaluated as described above for step 1020.

[0053] It is conceivable that this waiting period could be interrupted by high-urgency messages. Such messages could, for example, be issued without further delay. The waiting period for the preceding, less urgent message could then continue with continuous data collection and evaluation of the respective time.

[0054] Thus, during step 1030, activities that originally occurred at an inappropriate time can be postponed to a more appropriate time. Reference symbol list 1000 procedures 1010 steps 1020 steps 1030 steps

Claims

[1] Method (1000) for controlling an audiovisual communication interface in a driver-operated motor vehicle, wherein, using a neural network trained with training data, a time point at which an activity, in particular a communication activity, of the communication interface takes place is determined based on context data and / or vehicle data, and / or a time point for the activity is evaluated, where the specific point in time and / or the assessment of the point in time is related to an available cognitive capacity of a driver of the motor vehicle, wherein the training data includes annotated context data and / or vehicle data from multiple participants, and wherein the annotations of the training data are based on assessments, in particular by the participants, of real-life situations. [2] Method (1000) according to the preceding claim, characterized bythat the contextual data includes weather data, traffic data, land use data, time of day data and / or road characteristics. [3] Method (1000) according to any one of the preceding claims, characterized by that the vehicle data includes speed data, steering angle data, brake and / or accelerator pedal angle data, IMU data, satellite navigation data and / or an occupancy. [4] Method (1000) according to any one of the preceding claims, characterized by that the vehicle data is read via a CAN bus. [5] Method (1000) according to any one of the preceding claims, characterized by that the context data is retrieved from a remote computer system. [6] Method (1000) according to any one of the preceding claims, characterized by that the determination and / or evaluation of the timing depends on a type of activity. [7] Method (1000) according to any one of the preceding claims, characterized by that the activity involves an interaction with a voice assistant. [8] Method (1000) according to any one of the preceding claims, characterized by that the activity includes switching a user interface of the communication interface. [9] Method (1000) according to any one of the preceding claims, characterized by , that the activity includes the output of a message through the communication interface.

Citation Information

Patent Citations

  • Image display system, image display method and recording medium

    DE112018000968B4

  • Methods and systems for generating adaptive instructions

    US20190375426A1