Method and device for outputting an audio file in a vehicle

By determining the occupant's focus of attention using gaze, gestures, and verbal cues, the method and device generate personalized audio files to enhance the driving experience and emotional bond with the vehicle.

DE102025000731B3Active Publication Date: 2026-06-11MERCEDES BENZ GROUP AG
View PDF 6 Cites 0 Cited by

Patent Information

Authority / Receiving Office
DE · DE
Patent Type
Patents
Current Assignee / Owner
MERCEDES BENZ GROUP AG
Filing Date
2025-02-27
Publication Date
2026-06-11

AI Technical Summary

Technical Problem

Existing vehicle audio systems fail to personalize the audio experience based on the occupant's interaction and focus of attention, missing opportunities to enhance the driving experience and emotional bond with the vehicle.

Method used

A method and device that determine the occupant's focus of attention using gaze direction, gestures, and verbal utterances, generating and overlaying audio files based on this information to create a personalized soundscape that enhances the driving experience and well-being.

Benefits of technology

The method and device enable a personalized audio experience by generating situationally appropriate audio files, enhancing the driving experience and emotional bond with the vehicle through interactive occupant participation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000000_0000_ABST
    Figure 00000000_0000_ABST
Patent Text Reader

Abstract

The invention relates to a method for outputting an audio file (A) in a vehicle (1), wherein, according to the invention, the audio file (A) is generated and / or selected depending on a determined focus of attention relating to the vehicle's interior and / or exterior environment of at least one occupant (4) of the vehicle (1) and is output in the vehicle (1), wherein the generated and / or selected audio file (A) is partially superimposed on a currently output audio file, and the generated and / or selected audio file (A) and the currently output audio file are output together as a soundscape (K). The invention further relates to a device for carrying out the method.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] The invention relates to a method and a device for outputting an audio file in a vehicle.

[0002] German patent application DE 10 2023 002 174 B3 discloses a method for calibrating a vehicle-integrated binaural 3D audio system. The 3D audio system emits sound via two loudspeakers to simulate a virtual sound source for a user at a location different from the loudspeakers.The following calibration steps are performed by a control unit: a) placing the virtual sound source inside or outside a vehicle containing the 3D audio system; b) emitting sound via the virtual sound source; c) detecting a location indication provided by the user, where the location indication describes the user's presumed location of the virtual sound source; d) determining a location difference between the actual location and the location described by the location indication of the virtual sound source; e) modifying, depending on the location difference, a head transfer function used by the control unit to adapt the audio signal emitted via the two loudspeakers to the location of the virtual sound source; and f) repeating steps a) to d) until the location difference falls below a defined threshold.This involves using vehicle-integrated loudspeakers and the location indication is captured using vehicle interior sensors.

[0003] US Patent 11,687,155 B2 describes a method for eye-tracking in a vehicle. The method includes the acquisition of sensor data based on the gaze tracking of a driver. The method also includes the identification of sensor data associated with the vehicle's external environment within the acquired sensor data. Furthermore, the method includes filtering out sensor data associated with the external environment from the acquired sensor data to provide internal vehicle sensor data. The method also includes controlling a vehicle interface based on the internal vehicle sensor data.

[0004] DE 10 2013 225 736 A1 describes a help module that provides a user with help content by specifying an application for which help should be displayed. DE 10 2024 109 178 A1 and DE 10 2019 129 392 A1 describe methods for gesture recognition.

[0005] The invention is based on the objective of providing a method and a device for outputting an audio file in a vehicle.

[0006] The problem is solved according to the invention by a method which has the features specified in claim 1 and by a device which has the features specified in claim 10.

[0007] Advantageous embodiments of the invention are the subject of the dependent claims.

[0008] According to the invention, a method for outputting an audio file in a vehicle provides that the audio file is generated and / or selected and output in the vehicle depending on a determined focus of attention relating to a vehicle interior environment and / or a vehicle exterior environment of at least one occupant of the vehicle.

[0009] By applying this method, it is possible to enhance the experience, in particular the driving experience, and the well-being of at least one occupant of the vehicle through digital means available, especially on the vehicle side.

[0010] Furthermore, the process can create or even strengthen an emotional bond between at least one occupant and the vehicle.

[0011] Furthermore, the procedure offers at least one occupant another option for personalizing the vehicle.

[0012] In one implementation, the focus of attention of at least one occupant is determined based on a recorded gaze direction and / or a recorded gesture and / or a recorded verbal utterance. In particular, by combining a recorded gaze direction, a recorded gesture, and a recorded verbal utterance, the focus of attention of at least one occupant of the vehicle can be determined with considerable accuracy, and based on this, the audio file can be generated in a targeted manner and / or selected from an available database and played back in the vehicle.

[0013] In one implementation, a pointing trajectory of at least one occupant is recognized based on a captured gesture, so that it can be determined what, for example, which environmental object, the at least one occupant is likely pointing at, i.e., what their focus of attention is directed towards.

[0014] One implementation of the method involves taking into account changes in the detected pointing trajectory over time and / or changes in the vehicle's direction of travel. This means that a change in a gesture by at least one occupant, from which the pointing trajectory was detected, is considered, provided that the change is based on the vehicle continuing to travel. For example, a tree that at least one occupant is pointing at is initially located to the right in front of the vehicle, and a short time later the tree is located to the right behind the vehicle, and the at least one occupant is still pointing at the tree, so the pointing trajectory has changed.

[0015] In one possible implementation, a probability related to the recognized pointing trajectory is determined prompt-based using a language vision model and / or conditional probabilities and / or other machine learning methods. This allows for a fairly accurate determination of the probability of what at least one occupant of the vehicle is pointing to, both inside and especially outside the vehicle—that is, what their attention is focused on—in order to generate and / or select an audio file, particularly one that is situationally appropriate, and play it back in the vehicle.

[0016] In a further implementation, the focus of attention of at least one occupant in relation to the vehicle's external environment is determined based on the detected gaze direction and / or the recognized pointing trajectory and on detected environmental signals from the vehicle's sensors and / or on a detected verbal utterance from the occupant. In particular, the focus of attention of at least one occupant is determined by combining all detectable signals, so that it can be identified as accurately as possible where the attention of at least one occupant is directed, and accordingly an audio file is generated and / or selected and played back in the vehicle.

[0017] In another possible implementation, the focus of attention of at least one occupant in relation to the vehicle interior is determined based on the occupant's gaze direction and / or haptic signals detected by a touch-sensitive display unit and / or a spoken utterance from the occupant. Thus, the focus of attention of at least one occupant can be determined with a high degree of accuracy, and an audio file can be generated and / or selected and subsequently played back accordingly.

[0018] According to the invention, the generated and / or selected audio file is partially overlaid on a currently output audio file. For example, if a song is playing and the vehicle is currently driving along a road by the sea, and at least one occupant has their attention focused on the sea, the song will be partially overlaid with the sound of the sea.

[0019] According to the invention, the generated and / or selected audio file and the currently played audio file are combined and displayed as a soundscape in the vehicle, thereby enhancing the experience, particularly the driving experience, and the well-being of at least one occupant of the vehicle. In particular, this soundscape comprises previously played music and nature sounds, which can significantly enhance the experience and well-being of the at least one occupant in the vehicle.

[0020] The invention further relates to a device for outputting an audio file in a vehicle. According to the invention, the device comprises - a vehicle camera for capturing image signals to determine the direction of gaze and / or gesture of at least one occupant, - a microphone for recording audio signals to determine a spoken utterance of at least one occupant, - a control unit that is or can be linked to the vehicle camera and microphone for determining the focus of attention of at least one occupant relating to the vehicle interior and / or the vehicle exterior, - a sound generator that is or can be linked to the control unit for generating and / or selecting a focus-dependent audio file and - a loudspeaker device that is or can be coupled to the sound generator for outputting the generated and / or selected audio file in the vehicle.

[0021] The device makes it possible to generate and / or select an audio file that focuses the attention of at least one occupant and to output it in the vehicle, so that an experience, in particular a driving experience, as well as the well-being of at least one occupant, especially during the vehicle's operation, can be enhanced.

[0022] Exemplary embodiments of the invention are explained in more detail below with reference to drawings.

[0023] This shows: Fig. 1 schematically depicts components of a device for performing a method for outputting an audio file in a vehicle and Fig. 2 schematically a control unit with a sound generator for generating and / or selecting an audio file.

[0024] Corresponding parts are marked with the same reference symbols in all figures.

[0025] Fig. Figure 1 shows, by way of example and in a highly simplified form, components of a device for carrying out a method for outputting an audio file A in a vehicle 1 and in Fig. 2 is a control unit 2 with a sound generator 3 for generating and / or selecting such an audio file A.

[0026] It is generally known that so-called sound signatures, similar to an exterior design, particularly of a vehicle 1, serve as a distinguishing feature for products, especially branded products. This also means that the sound signature is created and defined during the development phase of a branded product, and, for example, the engine noise of a vehicle 1 can be adjusted based on driving behavior. Interactive elements for interacting with a driver as a passenger 4 and / or other occupants of the vehicle 1 are essentially not considered, regardless of driving behavior with regard to steering wheel and / or pedal use.

[0027] The method described below makes it possible to enhance an individual immersive experience, particularly a driving experience, during the operation of vehicle 1 using an audio file A, thereby increasing brand value, for example, for a vehicle manufacturer. This can strengthen a personal connection between the occupant 4 and the vehicle 1 and increase the personalization of the vehicle 1.

[0028] In particular, the method represents a technical solution for a personalized audio experience, especially a sound experience, based on interactive occupant participation.

[0029] As in Fig. Figure 1 shows that music M is being played in a vehicle 1 while it is in motion, for example, by tuning in to a radio station and / or using a streaming service. This could be conventional music M or generated music M being played in vehicle 1.

[0030] For example, the vehicle 1 has a binaural audio system which emits sound by means of a loudspeaker device 5 with at least two loudspeakers, in particular in a vehicle interior environment, in order to simulate a virtual sound source for an occupant 4 at a location different from the loudspeakers, as disclosed in the prior art according to DE 10 2023 002 174 B3.

[0031] Vehicle 1 has a vehicle camera 6 which continuously records image signals S1 during the vehicle 1's operation, based on which the control unit 2 determines the direction of view of the occupant 4 and a gesture G of the occupant 4.

[0032] In addition, the vehicle 1 has a microphone 7 which continuously records audio signals S2 during the vehicle 1's operation, on the basis of which the control unit 2 determines a linguistic utterance of the occupant 4 and / or occupants 4 and, by means of an analysis, recognizes the linguistic content of the utterance.

[0033] Furthermore, the operation of a touch-sensitive display unit 8 of the vehicle 1 is recorded using haptic signals S3, and in particular, it is determined which display content was tapped.

[0034] Furthermore, the vehicle 1 has an environmental sensor system 9, which comprises a number of detection units arranged in and / or on the vehicle 1. These units may be camera-based, lidar-based, and / or ultrasound-based. During vehicle 1 operation, the environmental sensor system 9 continuously detects environmental signals S4, which are used to detect the vehicle 1's surroundings and objects located within them.

[0035] The captured image signals S1 and / or audio signals S2 and / or haptic signals S3 and / or environmental signals S4 are fed to the control unit 2 for evaluation and processing, and the focus of attention of the occupant 4 is determined. In particular, it is determined whether the focus of attention relates to the vehicle interior and / or the vehicle exterior.

[0036] Based on signals S1 to S4, the system determines where occupant 4 is looking, i.e., what their attention is focused on. Based on the detected signals S1 to S4, in particular the determined direction of gaze and / or the recognized gesture G, it is also recognized if occupant 4 wants to draw the attention of another occupant 4 to something, for example, in the vehicle's exterior surroundings, or especially if they want to point something out.

[0037] For example, it is possible to determine that the focus of inmate 4's attention is directed towards a meadow and / or a wooded area and / or a lake and / or a schoolyard with children and / or a pet shop and / or a music shop and / or a display board.

[0038] In one embodiment of the procedure, a pointing trajectory Z is recognized based on the gesture G made by occupant 4, in particular if occupant 4 points to something in the vehicle's external environment, for example to draw the attention of another occupant 4.

[0039] To determine the pointing trajectory Z of a gesture by occupant 4, in particular a hand movement, a trajectory of a hand and / or finger gesture projected into the distance is used to determine what occupant 4 is pointing at. Additionally, for this purpose, for example, the audio signals S2 relating to a spoken utterance by occupant 4 and the environmental signals S4 detected by the environmental sensors 9 are taken into account.

[0040] Furthermore, a change in the pointing trajectory Z with respect to time is considered, especially since the pointing trajectory Z changes depending on the direction of travel of the vehicle 1, from which, for example, a distance to a focused object can be optimally estimated.

[0041] To determine the probability of the detected pointing trajectory Z, specifically the direction the occupant 4 is pointing, a prompt-based probability calculation for the pointing trajectory Z can be performed using a language vision model and an output of a section of the image captured by the environmental sensors 9. For example, a prompt for this would be: "Extract the section of the image in the environmental signals S4 that is most likely meant, based on the input data." The input data includes the captured signals S1 to S4 from the vehicle camera 6, the microphone 7, the display unit 8, the environmental sensors 9, and a current situation within the vehicle, such as music M currently playing, which is determined based on captured situation signals S5.

[0042] Alternatively or additionally, conditional probabilities are modeled, for example using Bayes' theorem and / or the Markov model: P(environmental data section)=P(S1,S2,S3,S4,S5) with: P stands for probability.

[0043] One implementation of the method includes weighting based on the detected signals S1 to S5. Various weighting options are suitable, one of which, described below, refers to prompt-based probability using the language vision model.

[0044] It is possible that signals S1 to S5 representing input data may be missing. For example, audio signal S2, captured by microphone 7, is missing because occupant 4 did not speak. Therefore, the other captured signals S1, S3, S4, and S5 are automatically weighted more heavily than if all signals S1 to S5 were present.

[0045] The prompt can then be repeated at short intervals with the currently available signals S1 to S5, according to a history of previous prompts. In this way, for example, short-term changes regarding gesture G can be taken into account, which may have resulted from the vehicle 1 continuing to drive and moving. For example, a meadow that occupant 4 pointed to was to the right in front of vehicle 1 two seconds ago; now the meadow is directly to the right of vehicle 1.

[0046] Subsequently, a training set for the language-vision model can be provided for further training. This set might demonstrate, for example, that an eye movement of occupant 4 should be weighted less than a gesture G, particularly a hand-pointing movement, since occupant 4 keeps his gaze directed towards the roadway. Such weighting can therefore be implicitly trained from a training set of the language-vision model, especially as a fine-tuning exercise.

[0047] Alternatively or additionally, a variant for the development time of the Language-Vision model can also be specified, in particular which

[0048] Signals S1 to S5 should be given particularly high weighting, for example when using conditional probabilities, an influence of a viewing direction in relation to a condition of the pointing trajectory Z.

[0049] Based on a determined focus of attention of inmate 4, an audio file A is generated and / or selected using a database DB in which a large number of so-called sound profiles are stored as audio files A.

[0050] For example, if vehicle 1 drives past a meadow and occupant 4 focuses their attention on the meadow, whereby the focus of attention is detected based on the determined gaze direction of occupant 4 and / or based on a pointing trajectory Z determined based on a gesture G, a passing bee and / or another insect and / or a sound of a molehill under construction and / or an exaggerated sound relating to plant growth is generated as an audio file A and / or selected using the database DB and output using the loudspeaker device 5 in vehicle 1.

[0051] In particular, the audio file A is partially superimposed on the currently playing music M as a soundscape K in vehicle 1.

[0052] Several options can be used to determine whether audio file A is spontaneously generated using sound generator 3 or selected from the database DB.

[0053] For example, the selection of audio file A in relation to the database DB is based on a predefined link, for example in relation to a recognized object that inmate 4 is pointing to.

[0054] Alternatively or additionally, the selection is based on feedback from at least one inmate 4 as a control variable with regard to a match between the recognized focus of attention and an audio file A that was played in the past.

[0055] Alternatively or additionally, inmate 4 gives direct feedback on which audio file A should be output when a specific focus of inmate 4's attention is detected, whereby this audio file A may first need to be created.

[0056] Alternatively or additionally, if different selection options for audio files A are available, inmate 4 can select one or more audio files A, which will then be output when the inmate 4's attention is focused accordingly.

[0057] As described above, the music M is superimposed on the generated and / or selected audio file A as a soundscape K and output via the loudspeaker device 5 in the vehicle 1.

[0058] An amplification of the soundscape K can be achieved by means of a feedback channel, in which a previous output of the soundscape K is used as an input signal for the further generation of an audio file A.

[0059] Alternatively or additionally, the soundscape K can be enhanced by automatically recognizing an emotional reaction of the occupant 4, for example by evaluating image signals S1 from the vehicle camera 6.

[0060] Inmate 4 can also actively interact with sound generator 3 to enhance the soundscape K, particularly with regard to receiving feedback, and can evaluate the focus-related generated and / or selected and output audio file A via a user interface using text and / or voice input. For example, it is possible to favor a file, with both positive and negative evaluations.

[0061] The soundscape K can therefore be amplified, in particular, when an interaction, for example a gesture G, of an occupant 4 occurs on the soundscape K, especially based on captured image signals S1 from the vehicle camera 6 and / or captured audio signals S2 from the microphone 7. Based on the captured image signals S1, for example a positive reaction of the occupant 4 can be detected, for example via the pupils.

[0062] In a further embodiment of the procedure, particularly when only one occupant 4, especially the driver, is in the vehicle 1, the audio file A is selected according to the detected signals S1 to S5, whereby a head movement towards an object and / or a head and / or body position are also taken into account when determining the focus of attention of occupant 4. For example, a suggestion is displayed, i.e., visually output, with regard to the selected audio file A. If occupant 4 does not reject this suggestion, the audio file A is output via the loudspeaker device 5 in the vehicle 1.

[0063] If there are multiple occupants 4 in the vehicle 1, a distinction is made as to whether the focus of attention of one occupant 4 or multiple occupants 4 is determined with regard to the generation and / or selection of audio file A. A setting can be configured to consider the focus of attention of all occupants 4 in the vehicle 1, for example, based on a distinction between driver / front passenger / rear passenger / age. This setting can be configured by the device, a driver, or the owner of the vehicle 1. In a vehicle 1 with a driver, it can be configured so that the determination of the focus of attention for the generation and / or selection of audio file A is not based on the driver of the vehicle 1.

[0064] Further development of the procedure provides for a distinction to be made according to the relevance of the recorded signals S1 to S5 according to the preferences of the occupant 4, for example with regard to a time of day etc., in particular preset or learned.

[0065] It is also conceivable that occupants 4 could be weighted, for example, depending on their seat position in vehicle 1 and / or their preferences. It can also be taken into account that the type of trip influences the weighting. For example, on a holiday trip, all occupants 4 and / or the preferences of all occupants 4 would be considered, whereas on a business trip, only the driver and front passenger and / or their preferences regarding audio file A and the associated soundscape K would be considered. Reference symbol list 1 vehicle 2 Control unit 3 Sound generator 4 occupants 5 speaker device 6 Vehicle camera 7 microphone 8 Display unit 9 Environmental sensors An audio file DB database G gesture K Soundscape M Music S1 image signal S2 audio signal S3 haptic signal S4 ambient signal S5 Situation Signal Z-axis trajectory

Claims

Method for outputting an audio file (A) in a vehicle (1), wherein the audio file (A) is generated and / or selected and output in the vehicle (1) depending on a determined focus of attention of at least one occupant (4) of the vehicle (1) relating to a vehicle interior environment and / or a vehicle exterior environment, characterized in that - the generated and / or selected audio file (A) is partially superimposed on a currently output audio file, and - the generated and / or selected audio file (A) and the currently output audio file are output together as a soundscape (K). Method according to claim 1, characterized in that the focus of attention of the at least one occupant (4) is determined on the basis of a detected gaze direction and / or on the basis of a detected gesture (G) and / or on the basis of a verbal utterance of the at least one occupant (4). Method according to claim 1 or 2, characterized in that a pointing trajectory (Z) of the at least one occupant (4) is recognized on the basis of a detected gesture (G). Method according to claim 3, characterized in that a change in the detected pointing trajectory (Z) with respect to time and / or a change in the direction of travel of the vehicle (1) is taken into account. Method according to claim 3 or 4, characterized in that a probability with respect to the recognized pointing trajectory (Z) is determined prompt-based by means of a language vision model and / or by means of conditional probabilities and / or further machine learning methods. Method according to one of claims 2 to 5, characterized in that a focus of attention of at least one occupant (4) with regard to the vehicle's external environment is determined on the basis of the detected gaze direction and / or on the basis of the detected pointing trajectory (Z) and on the basis of detected environmental signals (S4) of an environmental sensor system (9) of the vehicle (1) and / or on the basis of a detected verbal utterance of the occupant (4). Method according to one of claims 2 to 6, characterized in that a focus of attention of at least one occupant (4) with regard to the vehicle interior environment is determined on the basis of the detected gaze direction of the occupant (4) and / or on the basis of detected haptic signals (S3) of a touch-sensitive display unit (8) and / or on the basis of a verbal utterance of the occupant (4). Device for outputting an audio file (A) in a vehicle (1) according to the method of any one of claims 1 to 7, characterized by: - ​​a vehicle camera (6) for capturing image signals (S1) to determine a gaze direction and / or a gesture (G) of the at least one occupant (4), - a microphone (7) for capturing audio signals (S2) to determine a spoken utterance of the at least one occupant (4), - a control unit (2) that is or can be coupled to the vehicle camera (6) and the microphone (7) for determining a focus of attention of the at least one occupant (4) relating to the vehicle interior and / or the vehicle exterior.- a sound generator (3) that is or can be coupled to the control unit (2) for generating and / or selecting a focus-dependent audio file (A) and - a loudspeaker device (5) that is or can be coupled to the sound generator (3) for outputting the generated and / or selected audio file (A) in the vehicle (1).

Citation Information

Patent Citations

  • User-specific help

    DE102013225736A1

  • Graphical user interface, means of transport and method for operating a graphical user interface for a means of transport

    DE102019129392A1

  • Method for calibrating a vehicle-integrated binaural 3D audio system and vehicle

    DE102023002174B3

  • PINCH DETECTION AND REJECTION

    DE102024109178A1

  • Method for vehicle eye tracking system

    US11687155B2