Method for acoustic interaction between a vehicle and a user and information technology system

The method uses AI to analyze music tracks and generate personalized sound responses for vehicles, addressing the limitations of existing systems by enhancing user immersion and safety through adaptable acoustic interactions.

DE102024134314B4Active Publication Date: 2026-05-28MERCEDES BENZ GROUP AG +1
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
DE102024134314
Authority / Receiving Office
DE · DE
Patent Type
Patents
Current Assignee / Owner
Filing Date
2024-11-21
Publication Date
2026-05-28
Estimated Expiration
2044-11-21

AI Technical Summary

Technical Problem

Existing vehicle sound systems fail to adequately personalize soundscapes to individual customer preferences and are limited by reliance on specific sound tracks and driving situation parameters, lacking adaptability and recognition of musical characteristics.

Method used

A method utilizing artificial intelligence to analyze music tracks, break them into snippets, assign metadata, and generate sound responses based on predefined blueprints, allowing personalized acoustic interactions with the vehicle through acoustic and visual feedback.

Benefits of technology

Enhances user immersion and emotional connection to the vehicle by providing personalized sound responses tailored to individual preferences, improving driving safety and experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000000_0000_ABST
    Figure 00000000_0000_ABST
Patent Text Reader

Abstract

The invention relates to a method for acoustic interaction between a vehicle and a user, wherein the vehicle, in response to an action (1) performed by the user, outputs a sound response (2) via acoustic output means, wherein the sound response (2) is artificially generated from the musical patterns inherent in a piece of music (3), which are identified by an analysis of the piece of music (3) using artificial intelligence. The method according to the invention is characterized in that - the piece of music (3) is analyzed by a processing cascade (4), comprising at least one interacting digital signal processor algorithm and a suitable pre-trained machine learning model, which divides the piece of music (3) into music snippets (5) corresponding to individual temporal sections of the piece of music (3), and assigns metadata (6) to the music snippets (5), describing music-characterizing properties of the respective music snippet (5); - the music snippets (5) are stored in a snippet database (7); - for various actions (1) that can be performed by the user, at least one sound response blueprint (8) is provided, containing a metadata specification; - the sound response blueprints (8) are applied to the music snippets (5) stored in the snippet database (7), whereby for each sound response (2) at least one music scheme is generated from at least one music snippet (5) depending on the metadata specification in order to generate the sound response (2) for the respective action (1); and - a corresponding sound response (2) is emitted when performing a respective action (1).
Need to check novelty before this filing date? Find Prior Art

Description

[0001] The invention relates to a method for acoustic interaction between a vehicle and a user according to the type defined in more detail in the preamble of claim 1, and to an information technology system set up for carrying out the method.

[0002] A user and the vehicle they are controlling interact in a variety of ways. For example, when the user makes steering movements, the vehicle moves into a circular path, causing the user to experience lateral acceleration. This lateral acceleration force, or apparent force, provides a kind of haptic feedback for the steering movement. If the user presses the accelerator pedal, the vehicle accelerates, which may be accompanied by a revving of the engine. This provides the user with acoustic feedback. Acceleration and braking also change the vehicle's speed, causing the surroundings to pass the user faster or slower, which can be interpreted as visual feedback. Thus, actions performed by the user result in a reaction.This interplay of action and reaction leads to a deep immersion of the user in the driving experience. Such immersion is beneficial for driving safety, as it reduces the user's susceptibility to external distractions. Furthermore, it fosters a particularly strong emotional connection between the user and their vehicle, which can enhance the overall user experience.

[0003] It is known that users are given the option to personalize their vehicle or parts of the human-machine interface used to interact with the vehicle. For example, the Mercedes-Benz Sound Experience and the "Sound Worlds" used in this context are a case in point. Users can select different Sound Worlds from a predefined catalog, with each Sound World defining a unique sound profile for the vehicle. A Sound World comprises a set of sounds that are emitted by the vehicle through its speakers as a sound response to various events. The individual Sound Worlds are based on acoustic concepts such as "reduced," "soft," "synthetic," "spacey," "ambient," "rock," "80s," and the like. Using such predefined sound databases, an initial level of personalization is possible.However, the Sound Worlds are relatively broad gradations of different music or sound concepts, meaning that specific customer preferences can only be inadequately met. Due to the sheer number of different music genres and songs, it is currently impossible to fulfill every individual customer request for a suitable soundscape in the vehicle.

[0004] Since battery-electric vehicles typically produce less operating noise than vehicles with combustion engines, it is common practice to artificially generate driving sounds for battery-electric vehicles and play them back through speakers. The company Glydsphere offers a customizable speaker for playing driving sounds from battery-electric vehicles. Using an app, users can select different sound profiles for the speaker. Additionally, songs can be processed using granular synthesis to generate new sounds for the speaker. Further information can be found on the manufacturer's website: https: / / glydsphere.com.However, a disadvantage is that, due to the distortion resulting from granular synthesis, the original sound file used to generate the artificial driving sounds is hardly recognizable.

[0005] The generation and output of adaptive soundscapes depending on the driving situation is known, for example, from US 2023 / 0410774 A1 or US 2022 / 0019402 A1. Various vehicle parameters can be used to adjust the soundscape to be output in the vehicle interior, such as steering angle, speed, acceleration, the vehicle's location (e.g., in the form of a GPS position), and the like.

[0006] A vehicle capable of applying effects to the sound tracks contained in a song, depending on vehicle parameters, before outputting them in the vehicle interior, is also known from German patent application 10 2023 004 347.8, which was not yet published at the time of filing this application.

[0007] Furthermore, German patent application 10 2024 119 744.7, which was also unpublished at the time of filing this application, discloses a method for outputting an adaptive soundscape in a vehicle interior. The soundscape in the vehicle interior is adjusted depending on the driving situation. Sound tracks in a song are identified, and noises are generated based on these tracks. These noises can be enhanced with acoustic effects, such as reverberation. Depending on the driving situation, more and more different noises are output in the vehicle interior. The soundscape can thus be built up and reduced layer by layer. The addition or reduction of corresponding noise layers occurs depending on driving-related events taking place during the journey.

[0008] A disadvantage of the method described above is that the artificially generated sounds are based solely on the individual sound tracks of a particular song. Each sound track can be assigned to a single instrument, such as drums, bass, electric guitar, or vocals. This limits the possibilities for creating artificial soundscapes. Furthermore, the adjustment of the soundscapes is solely dependent on the driving situation, i.e., changing vehicle and driving parameters such as acceleration, speed, and the like.

[0009] Furthermore, DE 10 2021 110 268 A1 discloses a method and system for the scene-synchronous selection and playback of audio sequences for a motor vehicle. The system serves to output music that matches the scenery the vehicle is driving through. A map containing a route is pre-divided into areas of different scenario types. Audio sequences from various audio files are parameterized and assigned to the different scenario types depending on the respective parameters. The system then determines which areas of the map the planned route of the vehicle passes through and plays an audio sequence appropriate to the respective scenario type in the vehicle as it drives through that area.

[0010] The present invention is based on the objective of providing an improved method for acoustic interaction between a vehicle and a user, which enables safer operation of the vehicle.

[0011] According to the invention, this problem is solved by a method for acoustic interaction between a vehicle and a user with the features of claim 1. Advantageous embodiments and further developments as well as an information technology system for carrying out the method are set forth in the dependent claims.

[0012] A generic method for acoustic interaction between a vehicle and a user, wherein the vehicle, in response to an action performed by the user, emits a sound response via acoustic output means, wherein the sound response is artificially generated from the musical schemes inherent in a piece of music, which are identified by an analysis of the piece of music using artificial intelligence, provides that - the piece of music is analyzed by a processing cascade, comprising at least one interacting digital signal processor algorithm and a suitable pre-trained machine learning model, which breaks down the piece of music into musical snippets, corresponding to individual temporal sections of the piece of music, and assigns metadata to the musical snippets, describing musically characterizing properties of the respective musical snippet; - the music snippets are stored in a snippet database; wherein according to the invention - for various actions that can be performed by the user, at least one sound response blueprint is provided, containing a metadata specification; - the sound response blueprints are applied to the music snippets stored in the snippet database, whereby for each sound response, depending on the metadata specification, at least one music scheme is generated from at least one music snippet in order to create the sound response for the respective action; and - a corresponding sound response is emitted when performing a respective action; and - the processing cascade determines the music snippets in such a way that the respective music snippets can be assigned to the formal parts of the piece of music or to a subsection of the respective formal part of the piece of music.

[0013] Using the method according to the invention, the acoustic response behavior of the vehicle can be adapted to individual customer preferences, i.e., personalized. This contributes to a particularly deep immersion and emotional connection between the user and the vehicle or the driving experience. Due to the increased immersion and the associated concentration on the driving experience, road safety can be improved.

[0014] The sound responses emitted by the vehicle in response to specific actions are based on the corresponding piece of music. This music track is provided to the processing cascade, and thus to the appropriately pre-trained machine learning model, in the form of computer-readable data. A piece of music can therefore be provided as an audio file, using any format such as MP3, WAV, WMA, AAC, OGG, FLAC, and the like. Music tracks can be provided to the processing cascade in a variety of ways. For example, the processing cascade can read music tracks from a storage medium, receive them via the internet, and so on. At least one machine learning model, through its pre-training, is able to analyze the music track based on the musical patterns present within it.Such musical schemes depend in particular on the instruments used in the song, the tempo, the chords played, the mode, the key, the pitch, the volume, the time signature, the timbre, the harmony, the rhythm, and the corresponding vocal parts of different song sections. A musical scheme thus describes sections of a song that are characterized by the same musical features. Using pattern recognition, a task that can be solved particularly reliably and efficiently by artificial intelligence, especially in the form of machine learning models, specific passages in musical pieces can be located. The processing cascade identifies the respective musical snippets and assigns them the aforementioned metadata. This metadata describes the respective musical features of the musical snippet.In a broader sense, not only a complete piece of music or song can be fed into the processing cascade, but also individual musical stems, also referred to as "stem".

[0015] The user particularly prefers to select the music track to be fed into the processing cascade. The user can input the music track directly, for example, by storing the track as an audio file on a private storage medium. Alternatively, the user can simply specify the artist and title, after which at least one machine learning model in the processing cascade retrieves the corresponding track from a music database, such as an online one.

[0016] The music snippets, enriched with metadata, are then stored in the aforementioned snippet database. These snippets are available for further use. Depending on the sound response blueprints, the aforementioned sound responses are generated from the music snippets. The sound response blueprints serve to generate sound responses that each follow a specific, predefined pattern. Specific sound response blueprints are used for different actions, ensuring that even when different pieces of music are used to generate the sound responses, the sound responses associated with each action can be clearly distinguished and identified by the user. This enables recognition of the sound responses and the linking of sound responses to actions.

[0017] Each sound response blueprint contains metadata specifications that allow suitable music snippets to be extracted from the snippet database to generate the respective sound response for the underlying action. These music snippets can be output individually or in combination, unchanged, to create a specific musical scheme and thus a specific sound response. In other words, individual sections of the music track can be quoted. It is also possible to add acoustic effects to individual sections of a music snippet, or even to entire music snippets, which will be discussed in more detail later, or even to generate entirely new music snippets.

[0018] For example, unlocking the vehicle can be an action. The sound response blueprint associated with unlocking the vehicle might specify, for instance, the use of a musical scheme characterized by a particular tempo and rhythm. The appropriate music snippets are selected from the snippet database, and the sound response is generated based on them. The sound response is then emitted by the vehicle when the action is performed.

[0019] A corresponding acoustic response from the vehicle can be enhanced by responses output via other transmission channels. For example, a visual response can be provided by the vehicle. The vehicle might include various interior lights, such as ambient lighting. This ambient lighting could, for instance, consist of a colored LED strip. Depending on the music snippet being played, or the underlying piece of music, different light patterns can be displayed via the ambient lighting. For example, the brightness, dynamics, color, or other characteristics of the light pattern can be altered based on the music's features, particularly its genre. For instance, if the music is pop with male vocals, a flashing blue light could be displayed.However, if it is country music with female vocals, a pink running light can be displayed, for example.

[0020] The various sound response blueprints can be customized. For example, developers can modify or create new sound response blueprints. This allows them to cater to different tastes in various markets. Furthermore, different stakeholders can provide sound response blueprints, such as the vehicle manufacturer, third parties from the music industry, the user themselves, and so on. Ideally, users should have several sound response blueprints available for the same action. This allows them to further personalize their sound experience, ultimately enhancing the user experience and immersion in the driving experience. By selecting the appropriate sound response blueprint, the user essentially chooses different music snippets to generate a specific sound response for each action.

[0021] As previously described, the processing cascade is designed to determine the music snippets in such a way that each snippet can be assigned to the structural sections of the piece of music, specifically one of the following: intro, verse, chorus, bridge, and outro. This ensures that the different structural sections of a piece of music are particularly distinct from one another. Thus, distinct music snippets can be generated that, due to their origin from the same piece of music, nevertheless exhibit a certain similarity. The most common structural sections in songs are intro, verse, chorus, bridge, and outro, which can be used accordingly to assign the respective music snippets. Each music snippet is always assigned to only one of the aforementioned structural sections. However, a music snippet can also correspond to only a subsection of a particular structural section.

[0022] Preferably, musical snippets from a section preceding the chorus are used to create a sound response to an action performed before the journey begins, and snippets from a section following the chorus are used to create a sound response to an action performed after the journey ends. This establishes a temporal correspondence between the progression of the music and vehicle use, further enhancing the user's immersion in the driving experience. The beginning of the music is thus mentally linked to the start of vehicle use, and the end of the music to the end of vehicle use. Musical snippets from the chorus are particularly well-suited for shaping the vehicle's soundscape during the journey, as the chorus is a passage that is repeated frequently within the music.Therefore, the refrain is particularly suitable for shaping the acoustic backdrop during the journey, as the journey typically occupies the largest portion of the vehicle's use.

[0023] A further advantageous embodiment of the method according to the invention provides that new music snippets are artificially generated based on the music snippets stored in the snippet database, wherein the newly generated music snippets share at least one musical characteristic with those music snippets on which they are based. This allows for an even more diverse generation of sound responses. The respective newly generated music snippets can be created using a synthesizer or artificial intelligence, in particular generative artificial intelligence. For this purpose, properties of music snippets extracted from the musical piece are adopted. For example, a music snippet can include a specific tempo, a specific chord progression, or the like, which is transferred to a new music snippet.For example, a different instrument or several instruments can be used for the new music snippet.

[0024] According to a further advantageous embodiment of the method according to the invention, at least one sound response blueprint is structured in multiple layers, with a separate sound track being generated for each blueprint layer based on the music snippets selected according to the metadata specification. Each blueprint layer is thus assigned its own channel. The layered structure allows for simple and reliable manipulation of the respective content of the music snippets. Individual instruments or a background sound track can be assigned to the individual blueprint layers. Individual metadata specifications can be included for each blueprint layer. Different music snippets can also be assigned to the individual blueprint layers or sound tracks, so that several music snippets can be combined particularly easily to generate a sound response.Multiple blueprint layers can also be based on a single music snippet. In particular, different effects are applied to the music snippets according to the different blueprint layers.

[0025] A further advantageous embodiment of the method according to the invention provides that at least one music snippet is imprinted with an acoustic effect specified by a sound response blueprint, in particular an effect dependent on a driving parameter or vehicle parameter. Thus, the sound response blueprints can also contain specifications for the effects to be imprinted on the respective music snippets. These effects can be dependent on driving parameters or vehicle parameters, such as the accelerator pedal position, the engine speed, the vehicle's speed, and the like.

[0026] Thus, the music snippets stored in the snippet database and extracted from the music track can be used not only to generate so-called event sounds, but also to create the sound effects for the vehicle's driving noises. For example, the pitch of such a sound can increase with rising engine speed.

[0027] According to a further advantageous embodiment of the method according to the invention, it is further provided that the piece of music is assigned to a genre and genre-specific sound response blueprints are used. Since genres can differ significantly from one another, for example in terms of tempo, the use of genre-specific sound response blueprints allows the respective genre-specific musical characteristics to be better captured. Thus, the sound response blueprints represent a kind of template or fingerprint for generating a sound characteristic of a particular action. This characteristic sound then changes from genre to genre, which makes it possible to transmit the different musical sensibilities of various genres to the vehicle.

[0028] A further advantageous embodiment of the method according to the invention further provides that at least one of the following actions triggers a reaction from the vehicle: - the user approaching or moving away from the vehicle; - unlocking or locking the vehicle; - the opening or closing of a cover in the outer skin of the vehicle, in particular a door, a window, a trunk lid, a fuel filler cap or a loading lid; - the user getting in or out of the vehicle; - activating or deactivating the ignition or the vehicle's electronics; and / or - connecting or disconnecting the vehicle from a charging station; and / or the vehicle proactively issues a response, in particular: - when a warning message appears; and / or - when performing a kickdown.

[0029] This means the vehicle can react not only as a consequence of user action, but can also proactively produce its own responses. This allows the vehicle to emit characteristic sounds for a wide variety of events. For each event, the user can choose the sound based on the music being played. The characteristic sound signature is maintained using predefined sound response blueprints. This results in a consistent and uniform acoustic behavior across the manufacturer's fleet. The manufacturer's corporate design can thus be maintained acoustically. Furthermore, individually customizable sound responses can create a striking visual effect, further enhancing the user's emotional connection to the vehicle.

[0030] According to a further advantageous embodiment of the method according to the invention, it is further provided that at least one loudspeaker emitting sound into the vehicle interior is used to output a sound response to an action performed by the user, at least partially, inside the vehicle; and / or at least one loudspeaker emitting sound to the vehicle's surroundings is used to output a sound response to an action performed by the user, at least partially, outside the vehicle. The vehicle can thus output sound responses not only inside the vehicle but also externally. For this purpose, the vehicle can be equipped with appropriate loudspeakers. In particular, stereo output or even surround sound output, for example using the so-called Dolby Atmos sound format, is thus possible.By directing the sound response to the user's location, it is ensured that the user can perceive it optimally. For example, when the user approaches the vehicle or unlocks the doors, the sound response can be emitted into the vehicle's surroundings. When the user enters the vehicle, the sound can be emitted simultaneously both inside and outside the vehicle. If the user starts the vehicle's electronics or closes the doors after entering, the sound response can be prioritized for the vehicle's interior.

[0031] A generic information technology system, comprising a vehicle with acoustic output means, a computing unit, and a communication unit, as well as an external central computing unit, is further developed according to the invention in that the computing unit is configured to receive a specification for a piece of music via a user interface and to generate sound responses for actions from the piece of music according to a method described above, and to store these in a data memory of the vehicle's computing unit via the communication unit, and the vehicle is configured to output the sound responses stored in the data memory via the acoustic output means when a corresponding action is detected. The computing unit and the vehicle, respectively, are connected to the vehicle's computer unit.The vehicle's processing unit has at least read access to a computer-readable storage medium, each containing machine-interpretable instructions. When executed by a processor of the processing unit, these instructions cause the processor to perform the respective process steps of the method according to the invention. The vehicle's processing unit can detect the occurrence of a particular action, or an event triggering the sound response, by retrieving information via the vehicle's data bus. Thus, the processing unit can access information from control units of vehicle subsystems, such as a system for monitoring and controlling the door locking status.

[0032] The central computing unit can communicate with the vehicle, particularly via the internet. For this purpose, the vehicle can establish a wireless internet connection via mobile network using the communication unit, for example, a telecommunications unit.

[0033] The central computer can provide the user interface in a variety of ways. For example, the user can log in to a web portal using appropriate login credentials such as a username and password. The web portal can be accessed from the vehicle, via an application running on a mobile device such as a smartphone, via a desktop computer, and so on. This allows the user to communicate their music preferences. In particular, the user specifies the music track to be selected. However, the central computer can also automatically select the music track based on the user's musical preferences. Specifically, the underlying music track can be changed automatically, for example, after a set interval, such as once a week, so that using the vehicle remains engaging and exciting even over extended periods.This can further improve the user's emotional connection to the vehicle and thus also the immersion.

[0034] The sound responses stored in the vehicle's data storage can also be exchanged between vehicles. Specifically, corresponding copies of the sound responses are kept in a section of a data storage system on the central computer, assigned to the user's profile. If the user then uses a different vehicle, the respective sound responses can be transferred to the other vehicle, so that the familiar acoustic behavior is also available in this new vehicle. This reduces the time it takes for the user to become acquainted with the new vehicle.

[0035] Further advantageous embodiments of the inventive method for acoustic interaction between a vehicle and a user also result from the exemplary embodiments, which are described in more detail below with reference to the figures.

[0036] This shows: Fig. 1 a schematic flowchart of a method according to the invention for acoustic interaction between a vehicle and a user; Fig. 2 a schematic representation of the assignment of sound reactions to formal parts of a piece of music; Fig. 3 a schematic representation of a graphical user interface; Fig. 4 a highly schematic flowchart of the method according to the invention; Fig. 5 a detailed representation of an analysis step carried out within the framework of the method according to the invention in accordance with a first embodiment; Fig. 6 a detailed representation of an analysis step carried out within the framework of the method according to the invention in accordance with a second embodiment; Fig. 7 a schematic detailed representation of a snippet database; Fig. 8 a schematic representation of the process of generating sound responses; and Fig. 9 a schematic representation of the generation of a sound response based on a multi-layered sound response blueprint.

[0037] Fig. Figure 1 illustrates, in a highly schematic way, the process of a method according to the invention for acoustic interaction between a vehicle and a user. The user manually selects a piece of music 3 via a corresponding user interface, for example, provided via the infotainment system of their vehicle. The piece of music 3 is then fed to a processing cascade 4 executed on a computing device, such as a server. This cascade consists of at least one interacting digital signal processor algorithm and a suitably pre-trained machine learning model. The at least one machine learning model of the processing cascade 4 is trained to analyze the piece of music 3 and divide it into music snippets 5. The music snippets 5 correspond to temporal segments of the piece of music 3, each exhibiting specific musical characteristics.Based on these music snippets 5, sound reactions 2 are generated, which are output via the vehicle's acoustic output devices when the user performs actions 1 and optionally also when the vehicle itself performs them. Such a sound reaction 2 can also be referred to as an event sound.

[0038] How Fig. As shown in Figure 2, each music snippet 5 is assigned to the individual sections 9 of the music piece 3. Thus, the sound response 2 to be output for each action 1 is based on a corresponding section 9. Music snippets 5 can, for example, be assigned to the sections 9: intro, verse, chorus, bridge, and outro.

[0039] The user is given extensive means to personalize the respective sound reactions. 2. For example, Fig. 3 A possible display of a graphical user interface, by means of which the user can further influence the sound reactions 2 assigned to the respective actions 1. Using a corresponding play button 12, the user can preview the sound reaction 2 set for the respective action 1. By pressing an edit button 13, the user accesses a submenu for adjusting the respective sound reaction 2 or an underlying sound reaction blueprint 8 shown in the following figures. After editing a respective sound reaction 2, it can be uploaded to a data storage system of the central computer unit via an export button 14, in order to be made available for later use, including in other vehicles.

[0040] In an editing submenu (not shown in detail), the user can, for example, manually select a music snippet 5 that best represents music track 3, trim it, and join it into a seamless loop. Various acoustic effects can be applied, particularly to different sound tracks or channels of music snippet 5. The intensity of individual effects can be adjusted using a slider or rotary control. Furthermore, additional layers for sound tracks can be added and modified as needed.

[0041] Fig. Figure 4 shows a rough outline of the processing chain. In step 401, the music track 3 is selected and fed into the machine learning model of processing cascade 4. Processing cascade 4 is preferably executed on a central computing facility. Such a server or server cluster can be configured as a high-performance computing cluster and thus have comparatively powerful hardware components, ensuring the efficient and rapid processing of large amounts of data.

[0042] In step 402, the machine learning model analyzes the music piece 3. Based on its training, the machine learning model delivers individual music snippets 5, each with associated metadata 6. The music snippets 5 and the metadata 6 are stored in a snippet database 7. Subsequently, in step 403, sound processing is performed to generate the aforementioned sound responses 2 from the music snippets 5, according to the previously mentioned and described parameters. Fig. The 9 sound response blueprints shown in section 8. Furthermore, in step 404, music snippets 5 can be artificially generated from existing music snippets 5. This can also be referred to as sound synthesis. Finally, in step 405, the generated results are output, i.e., the respective sound responses 2 are transferred for application in a vehicle.

[0043] Fig. Figure 5 shows the analysis in detail according to a first embodiment. Step 401, in which the music track 3 is provided, is shown. In step 501, the genre of the music track 3 is determined. In step 502, source separation can take place. The processing pipeline now splits into a left and a right branch. According to the left branch, the aforementioned sound reactions 2, or event sounds, are generated. According to the right branch, sound reactions 2 are generated that are adapted depending on driving or vehicle parameters during vehicle use in order to represent driving noises.

[0044] Step 503 involves beat detection. Step 504 identifies key moments by recognizing auditory characteristics in the music piece 3. Step 505 performs pattern recognition, specifically based on a self-similarity matrix. Step 506 analyzes the chords contained in the music piece 3. Step 507 delivers the respective timestamps at which music snippets 5 begin and end in the music piece 3, as well as the metadata 6 associated with these music snippets 5.

[0045] Step 508 involves beatmatching. Step 509 determines the key of music piece 3. Step 510 transfers the results to the snippet database 7.

[0046] Step 511 involves onset detection. Step 512 identifies the key of piece 3. Step 513 performs pattern recognition using the aforementioned self-similarity matrix. Step 514 detects chords within piece 3. Step 515 identifies tonic segments from the most prominent parts of piece 3 or from the entire stem. Step 516 performs a spectral analysis, and step 517 a harmonic analysis.

[0047] Alternatively, the in Fig. The analysis process shown in Figure 6 is executed. Thus, in step 601, the aforementioned beat detection takes place, in step 602, the source separation, and in step 603, the key signature is detected. Here too, the subsequent process is divided into a left and right branch.

[0048] In step 604, key moments are identified based on energy and so-called spectral novelty curves.

[0049] Step 605 involves transcribing the singing. Step 606 involves comparing sentence similarity. Step 607 detects vocal activity. Steps 605, 606, and 607 are executed iteratively, preferably starting with step 607.

[0050] In step 608, chords are recognized according to a chord pattern detection.

[0051] In step 609, the aforementioned beatmatching takes place. In step 610, based on the insights gained so far from the analysis of the music track 3, iconic moments in the underlying song are identified and tailored as needed. In step 611, the music snippets 5 can be sorted according to relevance. In step 612, a similarity-based comparison of key music snippets can be performed based on pattern recognition. In step 613, each music snippet 5 can be marked based on an energy evaluation, also known as "labeling."

[0052] Here too, the left branch is used to generate event sounds.

[0053] The right branch shows the procedure for generating sounds to accompany the driving action.

[0054] Step 614 involves calculating a chromagram. Step 615 is optional and allows for the separation of harmonic and percussive events. Step 616 determines the energy of pitches, depending on the key.

[0055] Step 617 involves chord detection. Steps 618 to 620 are also optional. Step 618 can detect dissonant intervals in chords or the key signature. Step 619 can determine the degree of dissonance based on a harmonic-to-noise ratio. Step 620 can detect non-harmonic content. Step 621 determines the frequency center and bandwidth.

[0056] Step 622 involves counting the peak values ​​of the novelty curves and estimating their amplitude. Finally, step 623 analyzes the dynamics and energy, taking into account statistical parameters such as the root mean square (RMS) and peak factors.

[0057] In step 624, a value for harmony and tonality is assigned. In step 625, a value for variability and intensity can be assigned.

[0058] In step 626, appropriate music snippets 5 are then selected to generate dynamic vehicle sounds.

[0059] Fig. Figure 7 shows the snippet database 7 in greater detail. The snippet database 7 is divided into a left path for processing music snippets 5 and a right path for generating new music snippets 5.

[0060] In step 701, music snippets 5 are selected. In step 702, function-dependent sound response blueprints 8 are applied to these music snippets 5. In step 703, an effect modulation supporting the respective function is applied. In step 704, an effect mix adjustment is made. In step 705, the envelope is adjusted. This is also referred to as "Envelope Adjustment".

[0061] In step 706, the aforementioned music snippets 5 are selected for conditioning. In step 707, the psychoacoustic properties of each music snippet 5 are determined. In step 708, a continuous background melody is extracted from the aforementioned music snippets 5. In steps 709, 710, and 711, an Aura Bass Layer, an Aura Melody Layer, and an Aura Noise Layer are generated based on this. In step 712, the individual Aura Layers are adjusted using layer-specific effects with the help of sound response blueprints 8.

[0062] For a specific action 1, several music snippets 5 or sound reaction blueprints 8 can preferably be provided for the user to choose from, so that the user can personalize the sound reaction 2 to be generated even more specifically according to their preferences.

[0063] Fig. Figure 8 shows the process for generating the aforementioned sound reactions 2 from the music snippets 5 again, according to an alternative. In step 801, the content of the music snippets 5 is adjusted so that the tempo and key correspond to a specification.

[0064] The specification can be derived from the underlying tenor of the respective music-characterizing properties of the music piece 3. In step 802, the respective contents of the music snippets 5 are assigned to individual blueprint layers 10 of the sound response blueprints 8 in a function-dependent manner - see also Fig. 9. In step 803, a function is used to support effect modulation. In step 804, five effects applied to the music snippet are mixed and adjusted. In step 805, the surrounding textures are adjusted. In step 806, the results are converted to a target format and exported.

[0065] Fig.Figure 9 shows the generation of the sound response 2 based on a multi-layered sound response blueprint 8. In the illustrated embodiment, action 1 is a welcome signal, for example, when the user approaches the vehicle or unlocks it. The upper path shows the music snippets 5 determined by the processing cascade 4 based on the music track 3 used, which were stored in the snippet database 7. Examples of content for drums, electric guitar, vocals, and piano are shown. The individual music snippets 5 do not necessarily have to be instrument- or channel-specific.

[0066] The sound response blueprint 8 contains several blueprint layers 10, each comprising its own sound track 11 or containing a specification for generating a sound track 11 based on the music snippets 5. This could, for example, include a first and second bass accent as well as a background melody, referred to here as noise swells. The design of the content of the blueprint layers 10 is adapted in particular to the tempo and key of the music piece 3 and / or is genre-specific.

[0067] A mixer 15 assigns the content of the respective music snippets 5 to the respective sound tracks 11, optionally supplemented by acoustic effects. The result is the sound response 2. This can then be stored in a data memory of a computing unit in the vehicle. As soon as the vehicle detects that the aforementioned action 1 is being executed, the sound response 2 is output via acoustic output devices, in particular internal and / or external speakers.

Claims

[1] Method for acoustic interaction between a vehicle and a user, wherein the vehicle, in response to an action (1) performed by the user, outputs a sound response (2) via acoustic output means, wherein the sound response (2) is artificially generated from the musical schemes inherent in a piece of music (3), which are identified by an analysis of the piece of music (3) using artificial intelligence, wherein - the piece of music (3) is analyzed by a processing cascade (4), comprising at least one interacting digital signal processor algorithm and a suitable pre-trained machine learning model, which divides the piece of music (3) into music snippets (5) corresponding to individual temporal sections of the piece of music (3), and assigns metadata (6) to the music snippets (5), describing music-characterizing properties of the respective music snippet (5); - the music snippets (5) are stored in a snippet database (7); characterized by , that - for various actions (1) that can be performed by the user, at least one sound response blueprint (8) is provided, containing a metadata specification; - the sound response blueprints (8) are applied to the music snippets (5) stored in the snippet database (7), whereby for each sound response (2) at least one music scheme is generated from at least one music snippet (5) depending on the metadata specification in order to generate the sound response (2) for the respective action (1); and - a corresponding sound response (2) is emitted when performing a respective action (1); and - the processing cascade (4) determines the music snippets (5) in such a way that the respective music snippets (5) can be assigned to the form parts (9) of the piece of music (3) or to a subsection of the respective form part (9) of the piece of music (3). [2] Method according to claim 1, characterized by , that the respective music snippets (5) can be assigned to one of the following formal parts (9): Intro, Verse, Chorus, Bridge and Outro. [3] Method according to claim 2, characterized by , that musical snippets (5) originating from a section (9) preceding the refrain are used to generate a sound reaction (2) to an action (1) performed before the start of the journey, and musical snippets (5) originating from a section (9) following the refrain are used to generate a sound reaction (2) to an action (1) performed after the end of the journey. [4] Method according to any one of claims 1 to 3, characterized by, that based on the music snippets (5) stored in the snippet database (7) new music snippets (5) are artificially generated, whereby newly generated music snippets (5) have at least one music-characterizing property in common with those music snippets (5) on which they are based. [5] Method according to any one of claims 1 to 4, characterized by , that at least one sound response blueprint (8) is multi-layered, with a separate sound track (11) being generated for each blueprint layer (10) based on the music snippets (5) selected according to the metadata specification. [6] Method according to any one of claims 1 to 5, characterized by , that at least one music snippet (5) is given an acoustic effect specified by a respective sound reaction blueprint (8), in particular an effect dependent on a driving parameter or vehicle parameter. [7] Method according to any one of claims 1 to 6, characterized by, that the piece of music (3) is assigned to a genre and genre-specific sound response blueprints (8) are used. [8] Method according to any one of claims 1 to 7, characterized by , that at least one of the following actions (1) triggers a response from the vehicle: - the user approaching or moving away from the vehicle; - unlocking or locking the vehicle; - the opening or closing of a cover in the outer skin of the vehicle, in particular a door, a window, a trunk lid, a fuel filler cap or a loading lid; - the user getting in or out of the vehicle; - activating or deactivating the ignition or the vehicle's electronics; and / or - connecting or disconnecting the vehicle from a charging station; and / or the vehicle proactively issues a response, in particular: - when a warning message appears; and / or - when performing a kickdown. [9] Method according to any one of claims 1 to 8, characterized by , that at least one loudspeaker emitting sound into the vehicle interior is used to produce a sound response (2) to an action (1) performed by the user at least partially inside the vehicle interior; and / or at least one loudspeaker emitting sound into the surroundings of the vehicle is used to produce a sound response (2) to an action (1) performed by the user at least partially outside the vehicle. [10] Information technology system comprising a vehicle with acoustic output devices, a computing unit and a communication unit as well as an external central computing unit, characterized by, that the computing device is configured to receive a specification for a piece of music (3) via a user interface and to generate sound reactions (2) for actions (1) from the piece of music (3) according to a method according to one of claims 1 to 9 and to store these via the communication unit in a data storage of the computing unit of the vehicle and that the vehicle is configured to output the sound reactions (2) stored in the data storage via the acoustic output means when a corresponding action (1) is detected.

Citation Information

Patent Citations

  • Method for outputting an adaptive soundscape in a vehicle interior and vehicle

    DE102024119744B3

  • System to create motion adaptive audio experiences for a vehicle

    US20220019402A1

  • Dynamic sounds from automotive inputs

    US20230410774A1

  • Method and system for scene-synchronous selection and playback of audio sequences for a motor vehicle

    DE102021110268A1

  • DE102023004347A1