Vehicle and method for determining characteristic lip movement patterns
Patent Information
- Application Number
- EP2024755234
- Authority / Receiving Office
- EP · EP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-09-23
- Filing Date
- 2024-08-09
- Publication Date
- 2025-08-13
Smart Images

Figure EP2024072556_27032025_PF_FP_ABST
Abstract
Description
[0001] Vehicle and method for determining characteristic lip movement patterns
[0002] The invention relates to a vehicle of the type defined in more detail in the preamble of claim 1 and to a method for determining characteristic lip movement patterns.
[0003] A vehicle's infotainment system can play media content such as movies, series, music videos, audiobooks, podcasts, songs, and the like. The music played in the vehicle can inspire passengers to sing or hum along. This can be viewed both positively and negatively by individual vehicle occupants.
[0004] For example, if a passenger sings off-key, meaning they're not on the right pitch or aren't singing in sync with the lyrics, this can easily be perceived as distracting. In the worst case, this can lead to an argument, which can significantly distract the driver from their driving duties.
[0005] However, singing together in particular can also have an uplifting effect and thus make using the vehicle more fun.
[0006] Using so-called audio decomposition, voices can be separated from accompanying music or background noise. For this, see, for example, Driedger, J., & Müller, M. (2015, April). Extracting singing voice from music recordings by cascading audio decomposition techniques. In 2015 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) (pp. 126-130). IEEE.
[0007] Furthermore, methods for computer-assisted evaluation of karaoke singing performances are known, allowing computers to evaluate singing performance in the same way as humans would. For further information, see Tsai, WH, & Lee, HC (2011). Automatic evaluation of karaoke singing based on pitch, volume, and rhythm features. IEEE transactions on audio, speech, and language processing, 20(4), 1233-1243 and Mayor, O., Bonada, J., & Loscos, A. (2009, February). Performance analysis and scoring of the singing voice. In Proc. 35th AES Inti. Conf., London, UK (pp. 1-7).
[0008] From US 2011 / 0251841 A1, the coordination and mixing of vocals recorded by singers who are geographically distant from one another is also known. This document describes an app that can be run on a mobile device such as a smartphone to provide a karaoke function. Songs can be listened to via the smartphone, with the lyrics and pitch information in particular being shown on the smartphone display. The user of the app can then sing along, with the user's singing being recorded using the smartphone's integrated microphone. Data generated via the smartphone can be transmitted to a central server. Several people who are geographically distant from one another can sing the same song. Recordings made by the different people can be mixed together by the server. The app orThe server also offers the functionality to perform automatic pitch correction, correcting the sound of off-key passages. Vocal performance can also be evaluated, and the best singers can be ranked. Information can be exchanged between the respective devices, including in MIDI format. The app user must independently demonstrate the intention to perform karaoke. The corresponding app operation requires a high level of cognitive effort and is therefore completely unsuitable for use by a person driving a vehicle.
[0009] In addition, US 2019 / 0147841 A1 discloses means for displaying a karaoke interface. This document also describes the use of a mobile device such as a smartphone to participate in a collaborative karaoke activity. Individual users can film themselves with the front camera of their smartphone and record videos or images of their karaoke performance and share them with each other. Camera-based facial recognition can be performed to identify characteristic facial features and use these as reference points for imprinting a virtual mask. For example, a user's face can be replaced with the face of a famous personality, a mascot, or a cartoon character. The singing quality of the individual karaoke performances can also be determined and evaluated. This task can be performed either by a computer system or by the users of a corresponding app.The vocal quality, as assessed by the computer model or the users, can also be weighted. Successful or popular users can be ranked at the top.
[0010] Furthermore, DE 102015 014652 B4 discloses a motor vehicle and a method for operating the motor vehicle, in which the lyrics of a piece of music are output. While karaoke is performed by the vehicle occupants in the vehicle, the vehicle is operated at least semi-autonomously. A soundtrack is recorded in the vehicle interior, and this determines how "well" the vehicle occupant sings along to the song.
[0011] Furthermore, CN 1 13270082 A discloses a vehicle with a karaoke function. The vocals contained in a song are filtered out of an audio signal using voice suppression methods, and the audio signal is output in the vehicle. Furthermore, the acoustics in the vehicle interior are recorded to create an ambient audio signal. Echoes are filtered out of the ambient audio signal, making the singing along of a vehicle occupant more perceptible.
[0012] Furthermore, DE 102018214 976 A1 discloses a method for controlling a multimedia device and a computer program. In this method, a vehicle occupant's reaction to the playback of a multimedia piece in a vehicle is detected by sensors, and the subsequent playback of multimedia pieces is planned using a machine learning method.
[0013] Furthermore, DE 102022 107293 A1 discloses an assistance system and an assistance method for a vehicle. In this system, a vehicle occupant's reaction to the playback of a multimedia piece in a vehicle is detected by sensors, and a vehicle function is activated depending on the reaction.
[0014] Furthermore, CN 1 08 346427 A discloses a method for voice recognition, as well as a device and storage medium suitable for this purpose. This involves comparing acoustically recorded speech with speech recorded by lip reading. The present invention is based on the object of providing an improved vehicle characterized by improved user interaction, particularly while driving.
[0015] According to the invention, this object is achieved by a vehicle having the features of claim 1. Advantageous embodiments and further developments as well as a method for determining characteristic lip movement patterns emerge from the dependent claims.
[0016] A generic vehicle comprising media playback means for outputting songs in the vehicle interior, comprises song determination means for detecting and recording the singing along of at least one vehicle occupant during the output of a song;
[0017] Vocal analysis means for detecting deviations between the vehicle occupant’s vocals and the vocals contained in the song; and
[0018] Vocal indication means which are designed to indicate to the vehicle occupant at least the deviations between the singing along and the singing and to record the singing along, to compensate for deviations in the singing along by means of an automatic pitch correction and to output the compensated singing along via the media playback means.
[0019] According to the invention, the singing indication means are further configured to train a machine learning model with the singing behavior of a specific vehicle occupant, so that the machine learning model learns to predict deviations between the singing and the singing along for defined songs when the respective song is played in the vehicle interior before the vehicle occupant sings along to a respective song passage, and to determine pitch adjustments that can be used for automatic pitch correction by means of the trained machine learning model and to apply these.
[0020] With the aid of the vocal determination means, vocal analysis means, and vocal indication means, the vehicle according to the invention is thus capable of reacting appropriately and in a situation-specific manner to a particular situation in which a vehicle occupant sings along while the vehicle is in use. By displaying the discrepancies between the singing along and the singing, the respective vehicle occupant is informed by another source that they may be singing out of tune. If this is already being criticized by other vehicle occupants, but the singing vehicle occupant does not believe the other vehicle occupants, this can prevent an emerging argument because the vehicle evaluates the singing along from a neutral standpoint. Furthermore, the off-key singing of a vehicle occupant can also be corrected automatically by applying the aforementioned automatic pitch correction, also known as auto-tune.This is preferably done in a subtle manner so that the vehicle occupant in question is not aware of it. This also helps prevent arguments from arising due to the singing occupant's poor singing skills. This ultimately improves road safety, as the driver is not distracted by the potential argument in the vehicle interior.
[0021] However, individual vehicle occupants can also enjoy singing along. By displaying the discrepancies between the singing and the actual singing, the respective vehicle occupant can practice and thus improve their singing skills. By correcting off-key singing using Auto-Tune, even poor singers can enjoy singing along. This improves the mood of the vehicle occupants, which ultimately improves the vehicle's user experience.
[0022] The interaction between the vehicle occupant(s) and the vehicle is carried out in a subtle manner, so that the respective vehicle occupants, especially the driver, are not subjected to excessive cognitive strain. In particular, the vocal detection devices can proactively detect singing along in the vehicle interior without having to be activated by a vehicle occupant. Accordingly, the discrepancy between the singing along and the actual singing is automatically detected, and the vocal detection devices are activated accordingly.
[0023] In the simplest case, only the person driving the vehicle is in the vehicle and is driving from a starting point to a destination. Music is played via the media playback devices, for example the vehicle's infotainment system. For example, the vehicle's radio is used as the audio source, a sound file stored on a storage medium such as a hard drive, SSD, SD card or USB stick is read, or music is obtained from a music streaming service. Randomly or intentionally, a song particularly favored by the person driving the vehicle is played. Suddenly, the person driving the vehicle starts singing along, even though this was not their intention. Without the person driving the vehicle noticing, the singing is recorded, the deviation from the singing is determined, the singing is balanced out using automatic pitch correction, and the balanced singing is output accordingly via the media playback devices.This way, the driver doesn't notice that they're singing off-key, which could otherwise frustrate them and impair their ability to drive the vehicle safely. This can be reliably prevented with the help of the vehicle according to the invention.
[0024] There could also be more passengers in the vehicle, such as two, three, four, or even more. The singing can also be recorded and evaluated individually for each passenger. For this purpose, the vehicle is equipped with a corresponding number of vocal identification, vocal analysis, and vocal indication devices.
[0025] In another scenario, the singing of a specific vehicle occupant could be evaluated, and the deviation of the singing from the singing in the vehicle could be displayed. Any display device in the vehicle could be used for this purpose, such as the instrument cluster display, the head unit display, a dedicated passenger display, a display integrated into a headrest, and the like. Having the other vehicle occupants follow the singing of the vehicle occupant can be entertaining and thus enhance the driving experience.
[0026] The media playback means comprise at least acoustic output means such as sound transducers, in particular in the form of loudspeakers, and a corresponding processing unit for controlling the acoustic output means. The processing unit also requires a source of songs, for example in the form of the aforementioned radio, in the form of a sound file read from a computer-readable storage medium, and / or in the form of an internet stream. The media playback means or the processing unit can also analyze the audio content to be output via the acoustic output means and identify underlying musical characteristics. However, these could also be explicitly present in a corresponding sound file, for example, determined by the respective distributor of the song.This can be, for example, the lyrics, the semantic content of the lyrics, the harmony, the melody and / or the rhythm of at least a passage of the song or preferably the entire song.
[0027] Harmony, melody, and rhythm represent three fundamental systems for classifying music. Harmony describes the structure of pitches, i.e., the simultaneous sounding of several individual sounds. Melody describes the temporal arrangement or sequence of notes. Rhythm indicates the dynamics of music and can be further divided into bar, meter, and tempo.
[0028] Concrete embodiments of the song determination means are mentioned below.
[0029] With the help of the vocal identification tools, vocal characteristics corresponding to the corresponding musical characteristics can be identified, which allows a comparison using the vocal analysis tools.
[0030] The vocal analysis tools compare the vocal characteristics with the musical characteristics and can thus identify discrepancies between the singing along and the actual singing. In particular, the vocal analysis tools examine whether the singing vehicle occupant is hitting the correct pitch at the correct time and / or singing the correct lyrics. In the simplest case, the vocal analysis tools are implemented by a program routine running on a processing unit. This can be the same processing unit that also serves to implement the media playback tools. In general, the media playback tools, the vocal determination tools, the vocal analysis tools, and / or the vocal indication tools can utilize resources provided by one or more different processing units.
[0031] Singers identified by the vehicle can also be entered into a leaderboard, allowing their respective singing performances to be tracked. A leaderboard can be maintained locally within a specific vehicle, allowing passengers to compete against each other in a singing competition within the vehicle. A corresponding leaderboard can also be maintained across the vehicles in a fleet, for example, managed by a central computing device external to the vehicle, such as a cloud server. This allows users of different vehicles to compete against each other in a singing competition.For this purpose, the respective vehicle users can log in to their vehicle with an individual user ID, for example in the form of a user name and a secret password, which links the respective singing services with the user ID and stores them accordingly in the central computer system external to the vehicle.
[0032] The vehicle can have communication means such as a telecommunications unit. A respective computing unit of the vehicle can use the telecommunications unit to establish a wireless communication connection to the vehicle-external central computing device, for example, via mobile radio and / or Wi-Fi. This allows information to be exchanged with the internet via mobile radio or via a Wi-Fi hotspot and then transmitted to the central computing device connected to the internet.
[0033] As already described above, the vocal indication means are further configured to train a machine learning model with the vocal behavior of a specific vehicle occupant. This allows the machine learning model to predict deviations between the vocals and the vocals-along for defined songs when the respective song is played in the vehicle interior before the vehicle occupant sings along to a particular passage of the song. The machine learning model then uses the trained machine learning model to determine and apply usable pitch adjustments for automatic pitch correction. This increases the reliability of determining balanced vocals. While traditional Auto-Tune methods are capable of performing automatic pitch correction almost in real time, this still involves a certain latency.As a result, the balanced sing-along sound played in the vehicle interior does not completely synchronize with the singing. However, with the help of the appropriately trained machine learning model, this disadvantage can be overcome. The corresponding machine learning model gradually learns the singing behavior of the respective vehicle occupant and can thus be enabled to predict song passages that are sung off-key. This not only determines the general off-key singing itself, but also the extent to which notes are not pitched correctly. Accordingly, the machine learning model is able to determine a correction suggestion as to the extent to which pitch correction is necessary. This makes it possible to correct the pitch while singing along, so that the balanced sing-along sound can be played simultaneously with the singing in the vehicle interior.This further improves user interaction with the vehicle.
[0034] In particular, temporally convolutional neural networks can be used in conjunction with short-term memories such as GRUs, also known as gated recurrent units, or LSTMs, also known as long short-term memories.
[0035] An advantageous development of the vehicle according to the invention further provides that the vocal determination means comprise at least one microphone for detecting the acoustics in the vehicle interior and a computing unit for processing acoustic signals generated by the microphone. The computing unit is configured to recognize vocal characteristics in the acoustic signals, in particular in the form of the lyrics, the semantic content of the lyrics, the harmony, the melody, and / or the rhythm of at least one passage of the song. The computing unit is also preferably configured to receive the song content from the media playback and subtract it from the acoustic signals. This enables acoustic detection of the singing along in the vehicle interior.
[0036] Preferably, several microphones distributed throughout the vehicle interior are used to capture singing along. This allows for even more reliable detection of singing along. By analyzing the time difference with which the singing along is captured by the respective microphones and / or taking into account the varying amplitude of the respective sound signal, the singing vehicle occupant can also be located within the vehicle interior. This makes it possible to determine which vehicle occupant is currently singing.
[0037] A song is played in the vehicle interior via the media playback device. The song overlays the vocals of the vehicle occupant(s). This can make it difficult to capture the vocals acoustically. To counteract this, techniques known from so-called audio decomposition can be used, such as harmonic-percussive-residual decomposition, melody-residual decomposition, transient-residual decomposition, robust principle component analysis and / or machine learning methods, such as discriminators in the context of a generative adversarial network. This makes it possible to separate different sound sources from an audio signal. In this way, the vocals can be differentiated from the vocals contained in the song and the vocals of different vehicle occupants can be distinguished from one another. This makes it possible to hear the vocals of the vehicle occupant(s) even more clearly and thus extract them more reliably.In particular, the computing unit obtains the song content, i.e. the sounds to be output via the vehicle's loudspeakers, directly from the media playback means, so that the computing unit is reliably able to filter the song content from the acoustic signal picked up by the microphone(s).
[0038] However, with a purely acoustic-based analysis for detecting singing along, there is always a residual risk that the singing along in the vehicle interior will not be correctly identified. According to a further advantageous embodiment of the vehicle according to the invention, the singing determination means thus comprise at least one camera that captures at least the face of at least one vehicle occupant and a computing unit for processing camera images generated by the camera, wherein the computing unit is configured to recognize lip movements of the vehicle occupant in the camera images, to compare these with lip movement patterns characteristic of singing, and to recognize the singing by the vehicle occupant if the camera images contain at least one such characteristic lip movement pattern. Thus, the singing of a vehicle occupant can be determined not only acoustically but also visually.
[0039] This can be used particularly advantageously to automatically initiate the corresponding analysis of the singing along and display of the discrepancies between the singing along and the singing, or to determine a balanced singing along and output it via the media playback means. This preferably eliminates any acoustic detection of the vehicle interior, which improves data protection and thus the privacy of the vehicle occupants. However, as soon as lip movements characteristic of singing are identified, the acoustic vehicle interior detection can be activated and the singing analysis means and singing indication means can be controlled accordingly. This further reduces the level of distraction for the driver, since no manual control actions are required.This further improves user interaction with the vehicle and also increases road safety by reducing the level of distraction for the driver.
[0040] The lip movement patterns characteristic of singing can be stored in a database stored on a respective computing unit. Characteristic lip movement patterns can, for example, have been identified by a service provider or the vehicle manufacturer in previous measurement campaigns. Lip movements can be identified in the camera images using machine vision. Artificial intelligence-based methods, such as machine learning models, particularly in the form of artificial neural networks, can also be used for this purpose.
[0041] Preferably, the computing unit is further configured to recognize singing characteristics in the recorded lip movements by comparing them with the characteristic lip movement patterns, in particular in the form of the lyrics, the semantic content of the lyrics, the harmony, the melody, and / or the rhythm of at least one passage of the song. Thus, the visual detection of vehicle occupants not only generally allows the recognition of singing along, but also allows the singing characteristics to be determined. This can support or replace the acoustic detection of the singing characteristics. The respective singing characteristics can be linked to the characteristic lip movement patterns. For this purpose, appropriate measurement campaigns can be carried out to establish a connection between singing characteristics and characteristic lip movement patterns.
[0042] The computing unit can thus be enabled to read lips, i.e. to identify the spoken or sung words by purely visually observing the lip movements of a vehicle occupant. These words can be converted into a text string and read into a speech recognition model. This makes it possible to determine not only the lyrics themselves, but also their semantic content. A further advantageous embodiment of the vehicle according to the invention further provides that the vocal determination means are further configured to control the media playback means upon detecting at least one vehicle occupant singing along, in order to cause a karaoke version of the song played in the vehicle interior to be played, wherein the karaoke version comprises a vocal track with reduced volume or is entirely free of vocal tracks.This makes it easier for the vehicle's occupants to sing along, as the vocals are less audible or even completely absent, allowing the vehicle's occupants to better concentrate on their own singing. For this purpose, appropriate karaoke versions can be stored in a computer unit in the vehicle or accessed online via a suitable service provider.
[0043] A further advantageous embodiment of the vehicle according to the invention further provides that the vocal indication means are further configured to control the media playback means to reduce the volume of the vocal track contained in the output song and / or to increase the volume of the balanced vocals. The computing unit can also be configured to reduce the volume of only the vocal track contained in the song, instead of a karaoke version of the respective song played in the vehicle interior, and instead to play back the balanced vocals determined by the vocal indication means at a louder level. This makes it easier for the vehicle occupant(s) to perceive their own vocals. Since this is balanced vocals, this significantly increases the well-being of the respective vehicle occupants.In this way, the vehicle occupants perceive that they are hitting the notes of the singing, which ultimately convinces them of their own singing abilities.
[0044] A further advantageous embodiment of the vehicle according to the invention further provides that the singing analysis means are further configured, upon detection of a temporal discrepancy between the singing and the singing along, to output at least one visual and / or at least one haptic cue in the vehicle interior at the precise time at which the singing requires the vehicle occupant's intervention. This can improve the vehicle occupant's ability to sing in rhythm with the song. This can prevent the respective vehicle occupant from singing a particular passage of the song too early or too late. For the temporal indication of when which lyrics must be sung, methods and means are already known from the prior art, such as the temporal changing of a color gradient of the displayed lyrics, wherein the boundary of the color gradient precisely coincides with the passage of the song to be sung.According to the invention, however, in-vehicle devices can be used to output the visual and / or haptic cues. For example, a light located in the interior of the vehicle can be switched on precisely when the vehicle occupant's intervention is required. So-called ambient lighting could also be used for this purpose. Additionally or alternatively, haptic cues can be output, for example, via the actuators of a massage seat or by imparting vibrations to vehicle elements touched by the vehicle occupant, such as the steering wheel or a gearshift.
[0045] In particular, the output of haptic cues is particularly subtle, as only the respective vehicle occupant notices the haptic cue. This makes it easier for them to sing along in time with the song without the other vehicle occupants noticing the assistance.
[0046] According to a further advantageous embodiment of the vehicle according to the invention, the vocal analysis means are further configured to objectively evaluate the deviation between the vocals and the singing along using predefined evaluation metrics and / or to subjectively evaluate them using a machine learning model, wherein the machine learning model was trained with the subjective evaluations of the singing along of a large number of test listeners in measurement campaigns for selected songs. A weighting can be carried out between an objective evaluation and a subjective evaluation. For the objective evaluation of the deviation of the singing along to the singing, for example, differences on a Musical Instrument Digital Interface (MIDI) scale can be determined to determine a correct pitch. Errors in the relative volume can be determined using short-term log-energy.Hidden Markov models can be used to detect deviations in the sung rhythm. Expression classification can also be performed, whereby the most likely desired expression is assigned to a sequence. Using this additional dimension, it is possible to evaluate even vocal elements that deviate from the target at a higher level of abstraction.
[0047] According to the invention, a method for determining characteristic lip movement patterns comprises the following method steps:
[0048] - transmitting, by a computing unit provided in a vehicle as described above, an association of lip movements and singing features to a computing device external to the vehicle, wherein the computing unit receives music features from the media reproduction means and extracts the singing features to be associated with the lip movements from the music features;
[0049] - aggregating the assignments obtained from a plurality of different computing units in a match database by the computing device;
[0050] - Examining the aggregated assignments for characteristic patterns in order to identify lip movement patterns characteristic of certain singing features, by the computing device; and
[0051] - Providing the lip movement patterns thus identified for retrieval by the respective computing units.
[0052] Using the method according to the invention, it is possible to identify new characteristic lip movement patterns and distribute them to the respective vehicle processing units for application. This allows for even more reliable detection of singing through visual detection of the vehicle occupants and, in addition, to determine the respective singing characteristics. To examine the aggregated assignments for characteristic patterns, similar lip movement patterns can be assigned to groups. Proven grouping algorithms, also known as clustering algorithms, such as the k-means algorithm, can be used for this purpose.
[0053] As already mentioned above, the musical features such as lyrics, semantic content of the lyrics, harmony, melody, and / or rhythm of at least one song passage can be contained in a corresponding sound file or determined by the processing unit through analysis of the sound file. These musical features are then assigned to corresponding vocal features, which in turn allows a corresponding assignment to the lip movements. Further advantageous embodiments of the vehicle according to the invention also emerge from the exemplary embodiment, which is described in more detail below with reference to the figure.
[0054] Figure 1 shows a schematic representation of a vehicle according to the invention.
[0055] Figure 1 shows a vehicle 1 according to the invention, which is designed to detect singing along 6 of at least one vehicle occupant 4 to a song played in the vehicle interior and to react thereto in a situation-specific manner.
[0056] For this purpose, the vehicle 1 comprises media playback means 2 for outputting songs. The media playback means 2 comprise at least acoustic output means, for example in the form of loudspeakers 13, as well as a computing unit 9 for controlling the loudspeakers 13. The computing unit 9 can use different song sources, such as a radio (not shown in detail), a computer-readable storage medium to which the computing unit 9 has access, and / or an internet stream. A mobile device (not shown in detail), such as a smartphone, can also be coupled to the computing unit 9, for example wirelessly via Bluetooth, NFC, or Wi-Fi, or wired, for example via a USB cable or Ethernet cable, and a sound file stored on a computer-readable storage medium of the mobile device can be read out.
[0057] The vehicle 1 further comprises vocal determination means 3 for detecting and recording the singing along 6 of at least one vehicle occupant 4 during the song playback. For this purpose, purely acoustic monitoring of the vehicle interior using one or more microphones 8 is possible, as is alternatively or additionally visual recording of the vehicle interior using at least one camera 10, by means of which at least the face of the at least one vehicle occupant 4 can be recorded. This makes it possible to detect lip movements of the vehicle occupant 4. Sensor data generated by the microphone(s) 8 or the camera(s) 10 can also be evaluated by the computing unit 9. The singing along 6 of the vehicle occupant 4 can be particularly easily recognized in corresponding acoustic signals.However, there is a risk that the singing along 6 will be drowned out, for example by the songs played in the vehicle interior or other noise, and thus cannot be recorded correctly. Lip movement patterns detected in the corresponding camera images can be compared by the computing unit 9 with known lip movement patterns that occur during singing, whereby the singing along of the vehicle occupant 4 can be identified. The characteristic lip movement patterns can also contain indications of singing characteristics, such as the lyrics as well as the harmony, the melody and / or the rhythm of the song. In this way, the computing unit 9 can read lips by processing the camera images and thereby determine the spoken lyrics. Through a semantic analysis of the lyrics, for example with the help of natural language recognition models, the semantic content of the lyrics can also be determined.
[0058] The vocal analysis means 5 can be implemented by program routines executed on the computing unit 9. The vocal analysis means 5 detects deviations between the vocals 6 of the vehicle occupant 4 and the vocals contained in the song. Furthermore, the vehicle 1 has vocal indication means 7, which are configured to indicate to the vehicle occupant 4 at least the deviations between the vocals 6 and the vocals, to record the vocals 6, to compensate for deviations in the vocals 6 through automatic pitch correction, and to output the compensated vocals via the media playback means 2. Corresponding program routines for providing corresponding method steps can also be stored on the computing unit 9 and executed there accordingly.Furthermore, at least one display device 14 is part of the singing indication means 7 in order to be able to display the corresponding deviations between the singing 6 and the singing.
[0059] With the aid of a method according to the invention, new lip movement patterns characteristic of singing can also be identified. For example, the computing unit 9 determines music features corresponding to the respective singing features from information obtained from the media reproduction means 2 and assigns these to the correspondingly identified lip movements. These assignments are transmitted from the computing unit 9 to a computing device 11 external to the vehicle, which maintains a correspondence database 12 for storing these assignments. For this purpose, the vehicle 1 can have a telecommunications unit 15. The vehicle-external central computing device 11 examines the assignments of music or singing features to lip movements for the presence of characteristic patterns. In this way, similar characteristic lip movement patterns for the same singing features can be grouped. Lip movement patterns with a high degree of similarity are assigned to the same music or singing features.Singing characteristics are assigned and can thus be made available for retrieval by the respective computing units 9 in order to be used there to determine singing characteristics by analyzing the lip movements.
Claims
Patent claims 1. Vehicle (1), comprising media playback means (2) for outputting songs in the vehicle interior, comprising Vocal determination means (3) for detecting and recording the singing along of at least one vehicle occupant (4) during the output of a song; vocal analysis means (5) for detecting deviations between the singing along (6) of the vehicle occupant (4) and the singing contained in the song; and vocal indication means (7) which are configured to display at least the deviations between the singing along (6) and the singing to the vehicle occupant (4) and to record the singing along (6), to compensate for deviations in the singing along (6) by automatic pitch correction, and to output the compensated singing along via the media playback means (2), characterized in that the vocal indication means (7) are further configured to train a machine learning model with the singing behavior of a specific vehicle occupant (4).so that the machine learning model learns to predict deviations between the singing and the singing along (6) for defined songs when the respective song is played in the vehicle interior before the vehicle occupant (4) sings along to a respective song passage, and to determine and apply pitch adjustments that can be used for automatic pitch correction using the trained machine learning model.
2. Vehicle (1) according to claim 1, characterized in that the vocal determination means (3) comprise at least one microphone (8) detecting the acoustics in the vehicle interior and a computing unit (9) for processing acoustic signals generated by the microphone (8), wherein the computing unit (9) is is configured to recognize singing features in the acoustic signals, in particular in the form of the lyrics, the semantic content of the lyrics, the harmony, the melody and / or the rhythm of at least one passage of the song, and wherein the computing unit (9) is further preferably configured to receive the song content from the media reproduction means (2) and to subtract it from the acoustic signals.
3. Vehicle (1) according to claim 1 or 2, characterized in that the singing determination means (3) comprise at least one camera (10) capturing at least the face of at least one vehicle occupant (4) and a computing unit (9) for processing camera images generated by the camera (10), wherein the computing unit (9) is configured to detect lip movements of the vehicle occupant (4) in the camera images, to compare these with lip movement patterns characteristic of singing and to detect the performance of singing by the vehicle occupant (4) if the camera images contain at least one such characteristic lip movement pattern.
4. Vehicle (1) according to claim 3, characterized in that the computing unit (9) is further configured to recognize singing features in the detected lip movements by comparing them with the characteristic lip movement patterns, in particular in the form of the lyrics, the semantic content of the lyrics, the harmony, the melody and / or the rhythm of at least one passage of the song.
5. Vehicle (1) according to one of claims 1 to 4, characterized in that the vocal determination means (3) are further configured to control the media playback means (2) upon detection of singing along by at least one vehicle occupant (4) in order to cause the output of a karaoke version of the song output in the vehicle interior, wherein the karaoke version comprises a vocal track with reduced volume or is completely free of vocal tracks.
6. Vehicle (1) according to one of claims 1 to 5, characterized in that the vocal indication means (7) are further configured to control the media playback means (2) to reduce the volume of the vocal track contained in the output song and / or to increase the volume of the balanced vocals.
7. Vehicle (1) according to one of claims 1 to 6, characterized in that the singing analysis means (5) are further configured, upon detection of a temporal deviation between the singing and the singing along (6), to cause the output of at least one visual and / or at least one haptic indication in the vehicle interior at the precise time at which the singing requires the intervention of the vehicle occupant (4).
8. Vehicle (1) according to one of claims 1 to 7, characterized in that the vocal analysis means (5) are further configured to objectively evaluate the deviation between the vocals and the singing along (6) using predefined evaluation metrics and / or to subjectively evaluate it using a machine learning model, wherein the machine learning model was trained with the subjective evaluations of the singing along (6) of a large number of test listeners in measurement campaigns for selected songs.
9. Method for determining characteristic lip movement patterns, characterized by the following method steps: - transmitting, by a computing unit (9) provided in a vehicle according to one of claims 4 or 5 to 8 with reference back to claim 4, an association of lip movements and singing features to a vehicle-external computing device (11), wherein the computing unit (9) receives music features from the media reproduction means (2) and extracts the singing features to be associated with the lip movements from the music features; - aggregating the assignments obtained from a plurality of different computing units (9) in a match database (12) by the computing device (11); - examining the aggregated assignments for characteristic patterns in order to identify lip movement patterns characteristic of certain singing features, by the computing device (11); and - Providing the lip movement patterns thus identified for retrieval by the respective computing units (9).