Method and device for the simultaneous acoustic output of a music signal and a speech information signal

By adapting language information signals to match music signal parameters, the system addresses the issue of uncoordinated interactions in infotainment systems, enhancing user experience and intelligibility during music playback.

DE102023210929A1Pending Publication Date: 2025-05-08ROBERT BOSCH GMBH
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
DE102023210929
Authority / Receiving Office
DE · DE
Patent Type
Applications
Current Assignee / Owner
Filing Date
2023-11-03
Publication Date
2025-05-08

AI Technical Summary

Technical Problem

Existing infotainment systems in vehicles often experience uncoordinated interactions between voice announcements and music playback, leading to annoying or irritating experiences for drivers and impaired intelligibility of language announcements.

Method used

The system adapts the output of language information signals based on at least one parameter of the music signal, such as speed, rhythm, or style, to synchronize and enhance the intelligibility and pleasantness of the audio experience.

Benefits of technology

This approach improves user experience by ensuring that language information is more intelligible and pleasant during music playback, reducing annoyance and enhancing the overall effectiveness of the infotainment system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000000_0000_ABST
    Figure 00000000_0000_ABST
Patent Text Reader

Abstract

A method for outputting a music signal and a speech information signal, wherein the speech information signal is available for output during the output of the music signal, where at least one music parameter of the music signal is determined, wherein at least one speech signal parameter of the speech information signal is adjusted depending on at least one of the music parameters, where the music signal on the one hand and the speech information signal on the other hand are output simultaneously with at least one speech signal parameter set, an audio signal processing device, set up to carry out this procedure, a vehicle infotainment system comprising such an audio signal processing device and configured for the simultaneous playback of the music signal on the one hand and the speech information signal on the other hand with the set at least one speech signal parameter via at least one loudspeaker, wherein the at least one loudspeaker is included by or connected to the vehicle infotainment system, and a program code which is designed to execute the procedure according to the procedure described above when processed by a processor.
Need to check novelty before this filing date? Find Prior Art

Description

State of the art

[0001] Today's motor vehicles typically feature combined or integrated driver entertainment and information systems, or infotainment systems for short. These systems are typically configured to play music, for example, from a radio program or a data storage device, and also to assist the driver with voice announcements, for example, from a vehicle navigation system. Voice announcements that are time-critical, i.e., intended to effectively support or inform the driver at a specific time, often coincide with the ongoing playback of a music piece. An uncoordinated juxtaposition of the voice announcements and the music piece playback is often perceived as annoying or irritating, and can also impair the intelligibility of the voice announcements. Advantages of the invention

[0002] The invention with the features of the independent patent claims has the advantage of improving the user experience (UX) when outputting voice information during ongoing music playback. This is essentially achieved by detecting at least one parameter of the music signal and adjusting the output of the voice information depending on this at least one music signal parameter, or "music parameter" for short.

[0003] For this purpose, a method is proposed for outputting a music signal and a speech information signal, wherein the speech information signal is available for output during the output of the music signal, wherein at least one music parameter of the music signal is determined, wherein at least one speech signal parameter of the speech information signal is adapted as a function of at least one of the music parameters, wherein the music signal on the one hand and the speech information signal with the set at least one speech signal parameter on the other hand are output simultaneously.

[0004] Advantageously, a speed parameter of the music signal is recorded and the speech information signal is adapted accordingly to the speed of the music signal. For this purpose, an advantageous embodiment of the invention provides that the at least one music parameter represents a number of bars per unit of time or BPM and / or that the at least one music parameter represents a rhythm and / or that in the case of a music signal containing a vocal component, the at least one music parameter represents a number of syllables per unit of time and / or a syllable rhythm of the vocal component of the music signal.

[0005] Advantageously, it can also be provided that the at least one or one further recorded music parameter represents a music style of the music signal.

[0006] Advantageously, it can further be provided that a speech signal parameter adapted as a function of the detected music parameter(s) is a speed of the speech signal, in particular a sequence of syllables or words and / or a rhythm of the speech signal.

[0007] Furthermore, it can advantageously be provided that, as speech signal parameters, a frequency and / or a pitch and / or a timbre of the speech signal is adapted depending on the recorded music parameter(s).

[0008] All of these measures outlined contribute, partly depending on the type and design of the music signal, to improving the intelligibility of the voice information output of the assistance system for the user, in particular the driver, and to making it appear more pleasant.

[0009] The invention further relates to an audio signal processing device which is designed to carry out the method according to one of the preceding claims.

[0010] The invention is further directed to a vehicle infotainment system comprising an audio signal processing device as described above, configured for the simultaneous reproduction of the music signal on the one hand and the speech information signal with the set at least one speech signal parameter on the other hand via at least one loudspeaker, wherein the at least one loudspeaker is comprised by the vehicle infotainment system or is connected to the vehicle infotainment system.

[0011] Finally, according to the invention, there is a program code with program code means, i.e. instruction codes, which is designed to carry out the method as described above when processed by a processor. Drawings

[0012] Embodiments of the invention are illustrated in the figures and explained in more detail below. The same reference numerals in the figures denote the same or equivalent elements. Fig. 1 shows a block diagram of a device according to the invention for the simultaneous acoustic output of a music signal and a voice information signal using the example of an infotainment system which is designed for installation and operation in a motor vehicle, in particular a lorry (truck) or a passenger car (car). Fig. 2 shows a flow chart of the method according to the invention Description of the embodiments

[0013] The Fig. The infotainment system 1 shown in Figure 1 comprises a music signal source 110, which provides a music signal 111 intended for playback via loudspeakers in the vehicle interior or in the passenger compartment. The music signal source 110 can, for example, be a radio receiver that extracts music signals from a received radio signal. Alternatively or additionally, the music signal source can also be a playback device for a storage medium on which music pieces are stored. Furthermore, the music signal source can also be formed by an internet access point, via which an internet radio service or an internet music storage device is accessed. The list of possible embodiments of the music signal source is to be understood as exemplary and not exhaustive.

[0014] The music signal 111 from the music signal source 110 is fed to an analysis module 120. The analysis module is configured to determine at least one, i.e., one or more, parameters of the supplied music signal 111. The determined parameters of the music signal are referred to as music parameters. A first determined music parameter can, for example, be a speed parameter of the music signal, in particular, for example, the number of bars per minute or, in the case of modern music, BPM (beats per minute).

[0015] In a preferred embodiment, alternatively or additionally, further music parameters are determined, How -a rhythm, -a musical style -a number of syllables per unit of time and / or a syllable rhythm of a vocal part of the music signal, -a frequency and / or a pitch and / or a timbre of a singing voice of the music signal.

[0016] The determined music parameter(s) are available in the form of a music parameter signal at the output of the analysis module 120 for forwarding and further processing by a shaping module 140.

[0017] The infotainment system 1 further comprises an assistance information source 130, which, typically on an occasion-specific basis, provides assistance information 131 in the form of text modules. This can, for example, be a vehicle navigation component 130 that outputs text modules 131 intended to be converted into voice navigation instructions using a known TTS (text-to-speech) method. Alternatively or additionally, the assistance information source 130 can also be configured to provide, for example, warnings 131 such as warnings regarding critical traffic situations ahead or, for example, in the case of a fully or partially autonomous or automatically controlled vehicle, a takeover request 131 for the driver to assume the driving task or other assistance information. This additional assistance information is also provided in text form for subsequent TTS conversion.

[0018] To convert the assistance information(s) 131 provided by the assistance information source 130 into one or more speech information signals, these are fed to a shaping module 140. The shaping module 140 has a known TTS (Text-to-Speech) unit, which is designed to synthesize text modules or text-based components of an assistance information 131 supplied to it into an acoustically reproducible speech information signal 141.

[0019] For this purpose, the shaping module 140 receives not only the assistance information(s) 131 from the assistance information source 130, but also the music parameter signal 121 provided by the analysis module 120. The music parameter signal 121 represents at least one, preferably a majority or all of the determined music parameters of the music signal 111 currently provided by the music signal source 110.

[0020] The shaping module 140 is now designed such that the determined music parameter or preferably the determined music parameters are incorporated into the text-to-speech conversion of the assistance information(s) 131. The text-to-speech conversion is controlled taking into account at least one or preferably a plurality of the determined music parameters in such a way that -the speed of the speech information signal is adapted to the speed of the music, that is, for example, a beat rate, i.e. beats per minute, or BPM of the music and / or a speed of a vocal part of the music and / or -the rhythm of the speech information signal is adapted to the beat of the music and / or -the frequency or pitch of the speech information signal is adapted to that of a singing voice of the music signal or a musical style.

[0021] In a further embodiment, it can be further provided that, depending on the recorded music parameters, previously stored reference voices are used in the context of the text-to-speech conversion of the assistance information(s), such as a female or male voice, soft or hard voice tone or similar.

[0022] In a further embodiment, it is preferably further provided that the output of the synthesized speech information signal is synchronized in time with the music signal, taking into account the associated delay and urgency of the output, so that, for example, a voice guidance instruction as an example of a speech information signal begins synchronously with a bar of the piece of music or the beginning of a sentence or a verse of a vocal part of the piece of music.

[0023] The essentially unchanged music signal 111 and the synthesized speech information signal 141 adapted to the music signal 111 are combined in a multiplexer 150, i.e., superimposed, and output as a common audio signal 151 via an output unit 160 in the vehicle interior that has at least one loudspeaker. The at least one loudspeaker arranged in the vehicle interior can be part of the infotainment system 1 or connected to it.

[0024] In an advantageous embodiment, the described infotainment system 1 is designed, of course with the exception of the loudspeaker, as a control unit with a processor that executes a control program with program code means, i.e., control commands, to implement the method according to the invention. The control program is stored in a program memory of the control unit. The control program can be permanently stored in the infotainment system or loaded into the program memory from a storage medium as needed.

[0025] In Fig. 2 outlines the procedure of the method according to the invention.

[0026] In step 61, the music signal 111 is received from the music signal source 110.

[0027] In step 62, the music signal 111 is analyzed in the analysis module 120. Specifically, at least one music parameter characterizing the music signal 111 is determined, for example, a speed parameter such as beats per minute or BPM and / or a rhythm of the piece of music and / or, in the case of a piece of music with a vocal part, words or syllables of the vocal part per minute or per bar. In a preferred embodiment, alternatively or additionally, further music parameters are determined, such as a musical style, such as classical, pop, or rock music, and / or a frequency and / or a pitch and / or a timbre of a vocal part of the music signal.

[0028] In one possible embodiment, it may also be provided, in particular, that the music signal is first demultiplexed, i.e., broken down into individual components, such as rhythm, melody, and, in the case of a vocal part, vocal elements. Software that enables such decomposition and subsequent extraction or determination of the music parameters is known per se.

[0029] The determined music parameter(s) are made available as music parameter signal 121.

[0030] In step 63, the information 131 of the assistance component 130, for example, text modules for a route guidance instruction from a vehicle navigation system, is read in. These text modules are typically stored in a database and consist of complete words or parts of words. Text modules can, for example, be a distance indication in meters and the unit of length "meter" for specifying a distance to the point to which the route guidance instruction refers, a direction indication, e.g., "right" or "left" or "straight ahead," and an instruction such as "turn" or "continue driving." This list is, of course, not exhaustive.

[0031] In step 64, the information 131 received from the assistance component 130 is combined into a speech information signal 141 using the text-to-speech method in the shaping module 140. The speech information signal 141 is synthesized taking into account the music parameter signal 121 or the music parameter(s) represented by the music parameter signal 121. The music parameters are used to generate speech signal parameters, which are used to adjust or adapt the speech information signal 141. For example, the speed of the speech information signal 141 is adjusted depending on the music signal parameter(s). Furthermore, the pitch or frequency of the voice, i.e., pitch or frequency, or tone of voice, i.e., female or male, soft or hard, calm or aggressive, and the like, can advantageously be adjusted depending on the determined music parameters.For example, in the case of quiet classical music 111, the speech information signal 141 can be configured with a calm, female voice. Preferably, the volume of the speech information signal 141 is also adapted in a suitable manner to the volume of the music signal, so that, on the one hand, good intelligibility of the speech information signal 141 is achieved, but, at the same time, the playback is not perceived as disturbing.

[0032] In step 65, the unaltered music signal 111 and the synthesized speech information signal 141 are superimposed on each other in the multiplexer 150.

[0033] In step 66, the resulting overlay signal 151 is output via the playback module 160 in the vehicle interior of the motor vehicle for the entertainment and information of the vehicle driver.

Claims

[1] Method for outputting a music signal (111) and a speech information signal (141), wherein the speech information signal is available for output during the output of the music signal, wherein at least one music parameter of the music signal (111) is determined (62), wherein at least one speech signal parameter of the speech information signal (141) is adapted (64) in dependence on at least one of the music parameters, wherein the music signal (111) on the one hand and the speech information signal (141) with the set at least one speech signal parameter on the other hand are output simultaneously (66). [2] Method according to claim 1, characterized by that the at least one music parameter represents a number of beats per unit of time or BPM. [3] Method according to claim 1 or 2, characterized by that at least one musical parameter represents a rhythm. [4] Method according to one of the preceding claims, characterized by that the at least one music parameter represents a number of syllables per unit of time and / or a syllable rhythm of a vocal portion of the music signal (111). [5] Method according to one of the preceding claims, characterized by that the at least one music parameter represents a music style of the music signal (111). [6] Method according to one of the preceding claims, characterized by that the at least one speech signal parameter represents a speed of the speech information signal, in particular a sequence of syllables or words and / or a rhythm of the speech information signal (141). [7] Method according to one of the preceding claims, characterized by that the at least one speech signal parameter represents a frequency and / or a pitch and / or a timbre of the speech signal. [8] Audio signal processing device arranged to carry out the method according to one of the preceding claims. [9] Vehicle infotainment system (1) comprising an audio signal processing device according to claim 8 and configured for simultaneous reproduction of the music signal (111) on the one hand and, on the other hand, of the speech information signal (141) with the set at least one speech signal parameter via at least one loudspeaker, wherein the at least one loudspeaker is comprised by the vehicle infotainment system (1) or is connected to the vehicle infotainment system. [10] Program code which is designed to execute the method according to one of claims 1 to 5 when processed by a processor.

Citation Information

Patent Citations

  • Method for audible display of system information in vehicle, involves assigning system information to audible music attribute in vehicle inner chamber and audible display of the system information is carried out depending on attribute

    DE102007008815A1

  • Method, apparatus and computer program for operating an audio system in a vehicle

    DE102019102941A1

  • Controlling playback of audio data

    EP4156181A1

  • JP002001125599A