Method for controlling vehicle functions

By installing microphones and cameras in vehicles to detect and analyze acoustic signals and lip movements, identifying singing and musical characteristics, and automatically controlling vehicle functions, the problem of passengers singing distracting the driver is solved, improving driving safety and passenger experience.

CN121816614APending Publication Date: 2026-04-07MERCEDES BENZ GRP
View PDF 5 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-09-05
Publication Date
2026-04-07

AI Technical Summary

Technical Problem

Singing while driving can distract the driver and reduce road safety, and current technology struggles to effectively identify and address this situation.

Method used

By installing microphones and cameras in the vehicle, in-vehicle noise and lip movements are detected. The acoustic signals and lip images are analyzed using a computing unit to identify singing characteristics and musical features, and the vehicle functions are automatically controlled to respond to the occupants' singing behavior.

Benefits of technology

It enables automated recognition and control of singing in the car, and can stop or encourage passengers to sing when necessary, thus improving driving safety and passenger enjoyment.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121816614A_ABST
    Figure CN121816614A_ABST
Patent Text Reader

Abstract

The invention relates to a method for controlling vehicle functions, in which noise in the interior of a vehicle is detected by means of at least one microphone (1) and / or at least the face of at least one vehicle occupant (3) is detected by means of at least one camera (2), wherein a computing unit (4) inside the vehicle processes sensor data generated by the microphone (1) and / or the camera (2) in order to control vehicle functions in accordance with information obtained from the processing of the sensor data, and wherein the computing unit (4) checks a camera image generated by the camera (2) for lip movements carried out by a vehicle occupant (3). The method according to the invention is characterized in that the computing unit (4) checks the acoustic signal generated by the microphone (1) for a singing sound (5) of at least one vehicle occupant (3) and / or compares the identified lip movement with a known lip movement pattern generated during singing, and when the computing unit (4) detects singing of the at least one vehicle occupant (3), determines the singing sound (5) of the at least one vehicle occupant (3). The computing unit (4) actuates the vehicle function.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a method for controlling vehicle functions of the type defined in more detail in the preamble of claim 1. Background Technology

[0002] In-vehicle entertainment systems can reproduce media content such as audiobooks, podcasts, songs, TV series, or movies. Audio sources for songs can include, for example, radio receivers, optical data carriers such as CDs, flash memory such as USB sticks, smartphones connected via Bluetooth, or internet streaming. Playing songs in a vehicle can encourage occupants to sing together. Singing together can bring enjoyment to occupants, especially as a group. However, singing together can also be distracting to other occupants. In particular, it can distract the driver, reducing road safety. Therefore, methods and devices that allow for appropriate responses to situations where occupants sing while driving are desired.

[0003] A vehicle recognition system for detecting the speech activity of vehicle occupants is known from DE 10 2013 222 645 A1. This system is capable of detecting camera-based lip movements of vehicle occupants while they are speaking. If the system determines that speech activity is increased within the vehicle's interior space, this can be interpreted as a reduction in the driver's attention. To encourage the driver to refocus their attention on driving, signals can be output within the vehicle. To further confirm whether someone is indeed speaking within the vehicle's interior space, at least one microphone can be used to monitor the interior space.

[0004] For example, US 2018 / 0268812 A1 discloses the use of personal assistants in connection with the detection of a person's lip movements using a camera. Voice assistants, for example, can be implemented on smartphones and initiate tasks such as creating calendar entries, asking for weather forecasts, and making calls via voice input. If the corresponding voice input is made in a noisy background, the voice assistant may partially or completely fail to hear the underlying voice commands. Furthermore, individual words may be misunderstood. This disclosed text describes the detection of the corresponding user's facial region using a camera for detecting lip movements. Then, to analyze the underlying voice commands, only those moments when lip movements are simultaneously detected on the user's side during voice input are considered. Here, the corresponding camera images can be evaluated using artificial intelligence. Not only can the lip movements themselves be identified, but the lips themselves can also be read, i.e., the words spoken by the user can be deduced from characteristic lip movement patterns.

[0005] Furthermore, EP 4 163 913 A1 discloses a vehicle integration method for implementing control via voice commands. This disclosure describes a vehicle-integrated voice assistant that can be operated collaboratively by multiple vehicle occupants. Relevant to some vehicle functions is identifying which vehicle occupant issues the corresponding voice command, such as changing the volume of a seat-specific speaker, setting a seat-specific air conditioning unit, or preventing children in the vehicle from issuing control commands to the autonomously controlled vehicle. For this purpose, the vehicle occupant who monitors the vehicle interior space using a camera and who moves their lips during voice input recording is identified as the initiator of the voice command.

[0006] In addition, US 2012 / 01310196 A1 discloses a method for recognizing a user’s emotions based on facial signals, voice signals, or physical signals, and corresponding sensing devices for detecting these emotions.

[0007] Furthermore, DE 10 2018 214 976 A1 discloses a method, a computer program, and an apparatus for controlling a multimedia device. Multimedia works are output to a vehicle via the multimedia device, and the reactions of vehicle occupants to these reactions are detected using sensors. For example, vehicle occupants singing or dancing along to music can be acoustically identified by detecting corresponding lip movements using a microphone or a camera. This is used to control the subsequent temporal output of the multimedia works.

[0008] Furthermore, DE 10 2016 204 183 A1 discloses a method for selecting music using gesture control and voice control. In this method, the user performs user activities, which are recorded, analyzed, and associated with music titles provided in a database. If the respective user activities are re-identified, the musical piece assigned to the user activity is reproduced from the database. The user activity may, for example, consist of singing. Summary of the Invention

[0009] The purpose of this invention is to provide an improved method for controlling vehicle functions, which improves the user interaction between vehicle occupants and the vehicle.

[0010] According to the invention, this objective is achieved by a method for controlling vehicle functions having the features of claim 1. Advantageous designs and improvements are given in the dependent claims.

[0011] The method of the type described for controlling vehicle functions includes detecting noise in the vehicle interior space using at least one microphone and / or detecting the face of at least one vehicle occupant using at least one camera, wherein a computing unit examines camera images generated by the camera for lip movements performed by the vehicle occupant, and wherein a computing unit inside the vehicle processes sensor data generated by the microphone and / or camera to control vehicle functions based on information obtained from the processing of the sensor data. Specifically, the computing unit examines acoustic signals generated by the microphone for singing by at least one vehicle occupant, and / or compares identified lip movements with known lip movement patterns generated during singing, and the computing unit controls vehicle functions when it detects singing by at least one vehicle occupant. Furthermore, the computing unit is configured to collect and compare musical features and vocal features, and if at least one feature matches, determine that the vehicle occupant is singing along with the music. The musical features and vocal features include the song's text, semantic content, and / or the song's harmony, melody, and / or rhythm. The computing unit extracts musical features from acoustic signals and / or by analyzing media content reproduced in the vehicle, and extracts vocal features from acoustic signals and / or by comparing identified lip movements with lip movement patterns. According to the invention, the computing unit determines musical features that trigger or prevent singing by the vehicle occupant and stores these musical features in the form of vocal triggering features and vocal stopping features in a trigger database on the computing unit and / or a computing device external to the vehicle.

[0012] The method according to the invention allows for the automated control of vehicle functions based on the recognition of human singing within the vehicle's interior space. Therefore, there are distinctly different scenarios where singing within a vehicle can be evaluated not only positively but also negatively.

[0013] For example, if a vehicle occupant sings while the vehicle is in motion, the singing can be evaluated as negative because it may distract the driver. If the computing unit identifies that at least one vehicle occupant, particularly the driver, is singing while performing a driving task, the computing unit can take corresponding measures to stop the occupant from singing.

[0014] If, for example, a driver spontaneously sings, visual and / or acoustic warnings may be triggered in the vehicle to prompt the driver to stop singing. Media content, such as songs, may also be played via the vehicle's entertainment system. The driver may then sing along. If the computing unit detects this singing, it can control the entertainment system and pause or stop the song. Music may also enter the vehicle's interior from outside, particularly when windows are open. If the computing unit detects this singing during this driving situation, it can operate the power windows to close the open windows and thus stop the music from entering the vehicle's interior. In vehicles with relatively poor sound insulation, music may also enter the interior when windows are closed, allowing the computing unit to control the vehicle's entertainment system to play a noise-canceling mode (also known as active noise cancellation (ANC)) or to mask the music with other noise.

[0015] Detecting singing within the vehicle's interior space is feasible in various ways and methods. Therefore, the computing unit can evaluate only the acoustic signals generated by the microphone, analyze only the camera images used to detect lip movements, or fuse both sensor modalities. Preferably, the computing unit evaluates the acoustic signals from multiple microphones distributed at different locations within the vehicle's interior space. Furthermore, this makes it possible to locate the sound source within the vehicle's interior space, taking into account the propagation time difference and / or volume difference of the corresponding sound signals. Thus, just as in analyzing which vehicle occupant is moving their lips, it is also possible to extract which vehicle occupant is singing. If singing is detected together within the vehicle's interior space, the audio signal reproduced via the vehicle's entertainment system can be analyzed and subtracted from the acoustic signal generated by the microphone. This removes interfering noise from the acoustic signal, allowing the computing unit to hear the vehicle occupant's voice more clearly. The computing unit then analyzes the acoustic signal and identifies the singing by recognizing the presence of signal components corresponding to the singing (such as specific pitches and their time-varying curves) within the acoustic signal. Furthermore, speech analysis can be performed to determine the words sung by the vehicle occupant. These words can then be compared with known song text to detect the singing of a particular song.

[0016] By using one or more cameras that detect the interior space of a vehicle, the faces of vehicle occupants can be detected. This makes it possible to detect lip movements. The computing unit compares these lip movements with known lip movement patterns. Typical lip movement patterns generated during singing can be stored in a database in the computing unit's data storage. This makes it possible to read the lips, thereby identifying the lyrics being sung based solely on the visual detection of the vehicle occupants, making it possible to determine the corresponding song text. Furthermore, based on characteristic lip movement patterns, similar to acoustic monitoring of the vehicle interior space, singing with specific pitch, rhythm, etc., can be identified. Therefore, the singing of a specific song can be identified by the computing unit not only based on the song text but also on the melody, etc.

[0017] There are also situations where singing inside the vehicle can be considered positive. For example, this could happen when a vehicle occupant is in a good mood and wants to listen to a specific song. In this case, the occupant can sing the song, which is then recognized by the computing unit, which in turn controls the vehicle's entertainment system to reproduce the sung song. If the vehicle itself does not have storage media containing the song, a streaming service provider can be used to obtain the song via the internet.

[0018] As described above, the configuration here involves the computing unit collecting and comparing musical features and vocal features, and determining that a vehicle occupant is singing along with the music if at least one feature matches. The musical features and vocal features include the song's text, semantic content, and / or the song's harmony, melody, and / or rhythm. The computing unit extracts musical features from acoustic signals and / or by analyzing media content reproduced in the vehicle, and extracts vocal features from acoustic signals and / or by comparing identified lip movements with lip movement patterns. By collecting musical features and vocal features and comparing them, the computing unit is thus able to detect singing along in a particularly reliable manner and method. In other words, the computing unit compares the music reproduced in the vehicle with the vocal sounds detected in the vehicle.

[0019] As mentioned earlier, a song can be clearly identified based on its song text, the semantic content of the song text, but also based on the music itself, namely harmony, melody and rhythm.

[0020] Therefore, harmony, melody, and rhythm represent three basic systems of order used to classify music. Here, harmony describes the structure of pitch, that is, the simultaneous sounding of multiple single notes. Melody describes the temporal arrangement or sequence of tones. Rhythm describes the dynamics of music and can be further divided into meter, tempo, and tempo.

[0021] The computing unit can determine all these features not only for music reproduced in the vehicle but also for detected vocals. It is particularly feasible to reliably analyze songs by analyzing the media content reproduced in the vehicle. Therefore, the computing unit can, for example, read the corresponding MP3 file and evaluate the information contained therein. Thus, songs flowing live into the vehicle can be processed by the computing unit. Acoustic signals reproduced via the vehicle's speakers can also be analyzed by the computing unit. Therefore, the corresponding speaker input signals can be analyzed by the computing unit, or corresponding noise in the vehicle's interior space can be detected using a microphone, and the corresponding acoustic signals generated using the microphone can be evaluated. Furthermore, corresponding vocal features can be collected by comparing identified lip movements with lip movement patterns. Particularly preferably, the various analysis types can be combined to determine the corresponding music and vocal features in a particularly reliable manner and method.

[0022] To determine if occupants are singing together, for a given song, at least one musical feature and at least one vocal feature must be specifically consistent in time. The more consistent the features are over the duration of the reproduced song, the more likely the vehicle occupants are to sing the song together. For example, an occupant might sing off-key, meaning they haven't sung the correct pitch. Different musical and vocal features can be compared simultaneously for the same part of a song. For example, for the same song position, there might be harmonic discrepancies, but the semantic content might be the same. Therefore, even if a vehicle occupant sings off-key, the computing unit can still determine the consistency between the song and the vocals. Conversely, vehicle occupants might be unfamiliar with the song text and instead simply hum along or use placeholder words like "la la la." If the vehicle occupant sings the correct pitch, the computing unit can still determine the consistency between the reproduced song and the vocals within the vehicle's interior space.

[0023] As already described, the computing unit determines musical features that trigger or prevent singing by vehicle occupants and stores these musical features in a trigger database on the computing unit and / or a computing device external to the vehicle, in the form of singing trigger features and singing prevention features. Therefore, it is possible to determine, for a wide variety of users or vehicle occupants, entire songs or only portions of songs that explicitly motivate the corresponding vehicle occupants to sing or prevent singing.

[0024] In its simplest form, singing trigger features can be identified by marking the corresponding song positions where vehicle occupants begin singing. Conversely, song positions where vehicle occupants pause or do not begin singing at all can be associated with singing prevention features. Alternatively, artificial intelligence can be used to analyze musical features and filter out corresponding singing trigger or singing prevention features. Therefore, certain text passages, harmonies, melodies, or rhythms may encourage or prevent vehicle occupants from singing.

[0025] The corresponding trigger database can be stored not only in the computing unit inside the vehicle but also in a computing device outside the vehicle. It is also conceivable to update and expand the corresponding trigger database in the vehicle's computing unit and correspondingly aggregate the newly acquired information in the computing device outside the vehicle. The computing device can then distribute the aggregated information to the various computing units in the fleet of vehicles, so that the information acquired in a specific vehicle can also be used in other vehicles.

[0026] Here, the trigger can be person-specific, thus saving individual triggers for different people. Individuals can be explicitly identified in known ways and methods, such as facial recognition based on camera images generated using interior space cameras, particularly when using machine vision, with the aid of personalized vehicle keys and / or with the aid of personal PIN codes, etc.

[0027] According to an advantageous design of the method according to the invention, the computing unit uses artificial intelligence to detect singing. All commonly used AI methods can be used for this purpose, particularly machine learning methods, with artificial neural networks being especially preferred. This makes it possible for the computing unit to detect singing more reliably. Here, the corresponding machine learning model can be trained by the vehicle manufacturer before implementation or operation on the corresponding computing unit. Corresponding measurement activities can then be performed. For example, test subjects can sing a fixed number of different songs together, thereby extracting characteristic lip movement patterns for each song. Here, the corresponding machine learning model can be trained to such an extent that not only the song itself, but also individual harmonies, melodies, and / or rhythms can be assigned to corresponding lip movement patterns, which can then be transferred to other songs.

[0028] Furthermore, another advantageous embodiment of the method according to the invention is specified herein as follows: the computing unit examines camera images and / or sensor data generated by sensors for generating depth information from the vehicle interior space in order to detect a dancing motion pattern characteristic of at least one vehicle occupant, and determines that singing is taking place in the vehicle interior space as soon as the computing unit identifies dancing motion. Typically, vehicle occupants not only sing but also dance simultaneously. Since vehicle occupants in the vehicle interior space are usually seated in their respective vehicle seats, preferably wearing seatbelts, "dancing" herein refers to rhythmic reciprocating movements of various body parts. Therefore, actions such as nodding, snapping fingers, shaking the head, and swaying the body from side to side can be interpreted as dancing. Therefore, detecting dancing is another indicator for determining singing in the vehicle interior space. Thus, the computing unit can also detect singing in the vehicle interior space more reliably.

[0029] To this end, the limbs and body parts of vehicle occupants can be identified based on visual features and their movement tracked in camera images. Vehicles can also have motion sensors that detect the interior space, such as radar sensors, laser scanners, and radio sensors, which allow the generation of depth information. Therefore, motion can be detected by tracking the distance to a measured object over time.

[0030] Computer-supported viewing methods, also known as "computer vision," can be used to identify the body posture or changes in body posture of vehicle occupants. This can be supported, for example, by using deep neural networks for "posture estimation." Furthermore, pressure sensors within the vehicle, such as those integrated into the vehicle seats, can be used to identify movement patterns characteristic of dance.

[0031] According to another advantageous design of the method according to the invention, the computing unit sends consistent pairs consisting of musical features and vocal features to a computing device outside the vehicle. The computing device merges the pairs obtained by multiple different computing units into a consistency database. By processing the pairs merged in the consistency database, the computing device groups lip movements assigned to specific musical features and provides these lip movements for recall by the computing unit, wherein the corresponding groupings of the computing unit can be used to find consistency between musical features and vocal features. The computing device outside the vehicle can be, for example, a cloud server. The cloud server can be run by the vehicle manufacturer, for example. Thus, the computing units of the fleet vehicles can detect the corresponding user behavior of particularly large groups of people and transmit it to the computing device for central analysis. This makes it possible to extract characteristic lip movement patterns for individual musical features, as well as for individual song parts or also for the entire song. This allows it to be determined that when singing the same song part, people perform similar lip movements. If these characteristic lip movement patterns are then re-detected in the vehicle, this is a reliable indicator that the corresponding song part is being sung. This makes it possible to more reliably detect singing of songs based on lip movement analysis.

[0032] Communication between the vehicle-integrated computing unit and external computing devices can be achieved in proven ways and methods. Therefore, the computing devices can be accessed via the Internet. A telecommunications unit can be installed in the vehicle, allowing the computing unit to connect to the Internet via mobile radio, Wi-Fi, etc.

[0033] Furthermore, another advantageous design of the method according to the invention specifies that an external computing device uses an unsupervised machine learning algorithm to group lip movements. Unsupervised machine learning enables the reliable discovery of patterns in the data, allowing for the extraction of particularly accurate groups of lip movements. For example, the so-called k-means algorithm, known for its high reliability and simplicity, can be used for this purpose.

[0034] Furthermore, another advantageous embodiment of the method according to the invention specifies that, when at least two vehicle occupants are in the vehicle, the computing unit detects the emotional state of the vehicle occupants by processing sensor data and correlates the corresponding emotional state with detected or absent singing. This can be used to provide additional functionality. It is feasible to detect and evaluate the emotions of vehicle occupants in drastically different ways and methods, for example, as described in US 2012 / 0130196 A1. Therefore, life sensors can be installed in the vehicle or life sensors carried or borne by the vehicle occupants can be used to detect vital parameters, such as those integrated into mobile terminal devices, like smartphones or wearable devices like smartwatches. The specific emotional state of the vehicle occupant can then be inferred from the time-varying curves of the vital parameters. In particular, changes in emotional state can be determined, for example, when the user's pulse rate, skin conductivity, and / or skin temperature increase or decrease.

[0035] Here, a connection is established between the emotions of the second person and the singing of the first person. Therefore, the singing of the first person has a positive or negative impact on the emotions of the second person. Individuals can clearly distinguish this, as already described.

[0036] Speech analysis can also be used to detect and analyze conversations taking place inside the vehicle. The computing unit can then determine whether the vehicle occupants are arguing or happy to continue singing together. For example, if one vehicle occupant wants to sing but this interferes with at least one other vehicle occupant, an argument may occur. Using computer-supported linguistics, also known as natural language processing (NLP), the computing unit can thus identify "singing" as a conversation topic.

[0037] It is also possible to evaluate the singing of the song by other passengers in other vehicles as a positive emotion.

[0038] Another advantageous design of the method according to the invention further specifies that the calculation unit: - If an improvement in the emotional state of the vehicle occupants is detected when the singing begins and / or a deterioration in the emotional state of the vehicle occupants is detected when the singing ends, automatically trigger the reproduction of another song in the vehicle, particularly a song that includes at least one singing trigger feature; and / or - If the vehicle occupants' emotional state deteriorates when the singing begins and / or the vehicle occupants' emotional state improves when the singing ends, the playback of the song currently playing in the vehicle is automatically stopped or the playback of a song in the vehicle including at least one singing blocking feature is initiated.

[0039] Therefore, the computing unit controls the reproduction of media content, especially songs, in the vehicle's interior space according to the corresponding circumstances, in order to either further enhance the enjoyment of the vehicle occupants or prevent the driver's mood from deteriorating or becoming distracted.

[0040] If a vehicle occupant begins singing and the mood of another vehicle occupant improves, or if a vehicle occupant stops singing and the mood of another vehicle occupant subsequently declines, then the singing within the vehicle's interior is considered a positive indicator. To maintain a positive atmosphere, the computing unit can automatically play additional songs to encourage further singing and a more pleasant mood among the occupants. Specifically, singing triggering features associated with the respective vehicle occupant are identified, and songs that precisely include these features are selected for reproduction. These could be, for example, songs with specific text containing a particular theme or specific song sections with highly preferred harmonies, melodies, and / or rhythms.

[0041] Conversely, if the singing inside the vehicle is evaluated as negative—specifically determined by the calculation unit if at least one vehicle occupant experiences a decline in emotional state when the singing begins and / or a corresponding increase in emotional state when the singing ends—then the calculation unit can prevent the playback of songs via the vehicle's entertainment system, or songs that the occupants are unlikely to sing along to. To this end, the calculation unit selects songs that possess particularly strong singing-blocking characteristics for the relevant vehicle occupants. This thus avoids arguments and improves road traffic safety by reducing driver distraction.

[0042] Furthermore, another advantageous embodiment of the method according to the invention specifies that, for a song to be reproduced in the vehicle, the computing unit determines the respective vocal triggering features and vocal blocking features contained in the song by accessing a trigger database before its reproduction. Typically, the corresponding vocal triggering features or vocal blocking features for some song segments or the entire song can be collected at any point in time and stored with the corresponding song, so that the corresponding features can be recalled at any time. However, songs for which vocal triggering features and vocal blocking features have not yet been determined can also be output in the vehicle interior space. The computing unit can perform this correspondingly when the corresponding song is first reproduced in the vehicle interior space. Then, as described above, the corresponding extracted vocal triggering features and vocal blocking features can be stored so that they do not need to be extracted again. This task can also be undertaken by a computing device outside the vehicle for at least some song segments or for the entire song. The corresponding vocal triggering features and vocal blocking features can then be correspondingly communicated from the computing device outside the vehicle to the computing unit in the fleet vehicles.

[0043] Because individuals have different musical tastes, the singing trigger characteristics and singing blocking characteristics will differ for different individuals. Therefore, the present invention also includes the detection of singing trigger characteristics and singing blocking characteristics that are specific to a particular type or individual. For this purpose, the corresponding characteristics can be associated with musical taste or user type. Vehicle occupants can also be explicitly identified, for example by registering with a clear username and password at the vehicle's entertainment system, identifying biometrics, or identifying carried identification features, such as a clear Bluetooth address or MAC address of a mobile terminal device. The respective method steps according to the invention described above can then be tailored to the specific vehicle occupant.

[0044] In the corresponding trigger database, the song text, as well as the harmony, melody, and / or rhythm, can be identified and tagged accordingly. For example, for someone who likes heavy metal music, these can be identified as vocal trigger features, while for someone who likes pop music or hit songs, these may be identified as vocal blocking features, and vice versa. Attached Figure Description

[0045] Further advantageous designs of the method for controlling vehicle functions according to the invention also arise from the embodiments described in more detail below with reference to the accompanying drawings.

[0046] in: Figure 1 A schematic diagram illustrating the implementation of a method for controlling vehicle functions in a vehicle according to the present invention is shown; and Figure 2 This diagram illustrates music, singing, and the emotional states of vehicle occupants identified by a computing unit inside the vehicle. Detailed Implementation

[0047] Singing while driving can be pleasant, distracting, or even dangerous due to the driver's distraction. By means of the method according to the invention, it is possible to [redacted - likely referring to a specific method or procedure]. Figure 1 The vehicle shown in the image has a function to detect and automatically control the vehicle by using the sound of singing 5 inside the vehicle's interior space.

[0048] To this end, the vehicle 6 according to the invention includes at least one microphone 1 for detecting background noise in the vehicle's interior space and at least one camera 2 aimed at the corresponding face of at least one vehicle occupant 3. Sensor data generated by the microphone 1 and camera 2 is evaluated by a computing unit 4 inside the vehicle. Here, the computing unit 4 is able to identify not only music but also singing voice 5 from the corresponding acoustic signals generated by one or more microphones 1. Furthermore, the computing unit 4 is also able to evaluate camera images generated by the camera 2 and thereby identify the lip movements of the vehicle occupant 3. The computing unit 4 compares these lip movements with known lip movement patterns to correspondingly identify the singing voice 5. Here, not only can the corresponding song text and its semantic content be identified, but also features representing the music, such as harmony, melody, and / or rhythm, can be identified.

[0049] Here, computing unit 4 identifies in Figure 2 The musical feature 10 and singing feature 11 shown are compared with each other, and when at least one musical feature 10 and at least one singing feature 11 match, the singing together with the song played in the vehicle 6 is identified. Furthermore, the vehicle 6 may have life sensors 13, and the computing unit 4 can collect life parameters of the vehicle occupants 3 by means of these life sensors. By analyzing these life parameters, the computing unit 4 can extract the corresponding life parameters of the vehicle occupants 3. Figure 2 Emotional state 12 is shown here. Figure 2 Emotional state 12 is represented as the change of emotion e over time t. This also allows for the detection and evaluation of the semantic content of conversations within vehicle 6, thereby identifying, for example, arguments or agreements between vehicle occupants 3. Another vehicle occupant 3 singing along to the song 5 sung by the first vehicle occupant 3 can also be designed as a positive emotional state 12.

[0050] Furthermore, vehicle 6 may have a communication device 14, such as a telecommunications unit, by means of which vehicle 6 or computing unit 4 can establish a wireless communication connection with computing unit 7 outside the vehicle. A consistency database 8 is maintained on computing unit 7 for merging pairs consisting of music features 10 and vocal features 11 associated with multiple different computing units 4. Computing unit 7 can process the pairs merged in consistency database 8 and thus can group lip movements assigned to a specific music feature 10 and provide those lip movements for invocation by the corresponding computing unit 4. This is feasible because each corresponding vocal feature 11 is assigned to a lip movement itself. These corresponding groups can then be used by computing unit 4 to find consistency between music features 10 and vocal features 11 by analyzing the corresponding lip movements or lip movement patterns.

[0051] This leads to the exchange of characteristic lip movement patterns newly discovered during use in the corresponding vehicle 6 of convoy 15.

[0052] Here, the computing unit 4 extracts the musical features 10 that trigger singing by the vehicle occupant 3 and stores these musical features as singing trigger features 10.1 in the computing unit 4 and / or the trigger database 9 on the computing device 7 outside the vehicle. Furthermore, the computing unit 4 determines the musical features 10 that prevent singing by the vehicle occupant 3. These musical features 10 are also stored in the trigger database 9 as singing prevention features 10.2. The songs or song segments that trigger or prevent singing by the vehicle occupant 3 can be identified by means of the singing trigger features 10.1 and singing prevention features 10.2. Furthermore, the corresponding singing trigger features 10.1 and singing prevention features 10.2 can be assigned to various semantic content and / or musical characteristics of the song, such as harmony, melody, and / or rhythm. Particularly preferably, individual singing trigger features 10.1 and singing prevention features 10.2 are determined for different types of people or preferred musical directions.

[0053] Figure 2 The process for this is explained. Figure 2 The upper region shows the musical features 10 when the song is reproduced in vehicle 6, represented by the CD shown to the left of the corresponding box. Here, the song text is shown in the upper row of the box and the musical notes in the lower row. The singing features 11, identified by the calculation unit 4, are shown in the lower box. The calculation unit 4 compares the singing features 11 with the musical features 10 and thereby determines whether the corresponding vehicle occupants 3 are singing together. The corresponding time window (where the corresponding vehicle occupants 3 sing together) is shown in... Figure 2 This is indicated by the following method: the corresponding song text portion or note is only displayed in the singing feature 11 at a certain section. Therefore, in this case, vehicle occupants 3 only sing the "Oh, Christmas tree, oh, Christmas tree" section together. Specifically, when the corresponding vehicle occupant 3 begins to sing, the calculation unit 4 assigns the singing trigger feature 10.1 to the corresponding music feature 10. Figure 2 In the illustrated embodiment, for example, song text motivates vehicle occupants 3 to sing together. If a corresponding vehicle occupant 3 stops singing, the calculation unit 4 preferably assigns the singing-stopping feature 10.2 to the corresponding music feature 10. For example, in Figure 2 In the embodiment shown, the melody of the corresponding song segment is not liked by the vehicle occupant 3, so the vehicle occupant stops singing 5.

[0054] Below the box showing the singing characteristics 11 of vehicle occupant 3, the emotional state 12 of another vehicle occupant 3 is shown. The emotional state 12 of the other vehicle occupant 3 improves at the beginning of singing 5 and worsens after singing 5 ends. The calculation unit 4 can thus deduce that the other vehicle occupant 3 likes the singing 5.

[0055] A particularly preferred embodiment of the method according to the invention specifies that, in this case, the calculation unit 4 causes additional songs to be played in the vehicle 6, particularly songs that include a particularly large number of song-triggering features 10.1 and a particularly small number of song-blocking features 10.2 for each vehicle occupant 3. Conversely, if the song 5 triggers dislike for other vehicle occupants 3, it is possible to alternatively block the playback of additional songs or at least play songs that include a particularly large number of song-blocking features 10.2 and / or a particularly small number of song-triggering features 10.1, particularly excluding any song-triggering features 10.1.

[0056] The method according to the invention improves the user interaction between vehicle occupant 3 and the vehicle. Therefore, corresponding vehicle functions are automatically activated, deactivated, or controlled based on the detection of the singing sound 5. In particular, it improves driving safety because, in corresponding situations, stopping the singing sound 5 can reduce the driver's distraction.

[0057] To execute the method according to the invention, a computer-readable storage medium is implemented in the computing unit 4, on which a corresponding computer program product is stored. The implementation unit of the computing unit 4 executes the computer program product to provide the method steps according to the invention. Therefore, such a computer-readable storage medium and computer program product are also part of the invention. The method according to the invention is executed in a vehicle 6 according to the invention, which is also part of the invention.

Claims

1. A method for controlling vehicle functions, wherein noise in the interior space of the vehicle is detected by means of at least one microphone (1) and / or the face of at least one vehicle occupant (3) is detected by means of at least one camera (2), wherein a computing unit (4) examines a camera image generated by the camera (2) for lip movements performed by the vehicle occupant (3), wherein the computing unit (4) inside the vehicle processes sensor data generated by the microphone (1) and / or the camera (2) to control the vehicle functions based on information obtained from the processing of the sensor data, and wherein The computing unit (4) examines the acoustic signal generated by the microphone (1) for the singing (5) of at least one vehicle occupant (3), and / or compares the identified lip movements with known lip movement patterns generated during singing, and when the computing unit (4) detects the singing of at least one vehicle occupant (3), the computing unit (4) operates the vehicle function, and wherein The computing unit (4) collects and compares musical features (10) and vocal features (11) with each other, and determines that the vehicle occupant (3) sings along with the music if at least one feature is consistent, wherein the musical features and the vocal features (11) include the song text, semantic content and / or the harmony, melody and / or rhythm of the song, and wherein the computing unit (4) extracts the musical features (10) from the acoustic signal and / or by analyzing the media content reproduced in the vehicle (6) and extracts the vocal features (11) from the acoustic signal and / or the comparison between the identified lip movements and the lip movement patterns. Its features are, The computing unit (4) determines a music feature (10) that triggers or prevents singing by a vehicle occupant (3) and stores the music feature in the form of a singing trigger feature (10.1) and a singing prevent feature (10.2) in a trigger database (9) on the computing unit (4) and / or a computing device (7) outside the vehicle.

2. The method according to claim 1, Its features are, The computing unit (4) uses artificial intelligence to detect singing (5).

3. The method according to claim 1 or 2, Its features are, The computing unit (4) examines the camera images and / or sensor data generated by sensors that detect the interior space of the vehicle to generate depth information in order to detect the characteristic dancing motion pattern of at least one vehicle occupant (3), and determines that singing is taking place in the interior space of the vehicle as soon as the computing unit (4) identifies the dancing motion in the interior space of the vehicle.

4. The method according to any one of claims 1 to 3, Its features are, The computing unit (4) sends a consistent pair consisting of musical features (10) and vocal features (11) to a computing device (7) outside the vehicle. The computing device (7) merges the pairs obtained by multiple different computing units (4) in a consistency database (8). The computing device (7) processes the pairs merged in the consistency database (8), groups the lip movements assigned to a specific musical feature (10), and provides the lip movements for invocation by the computing unit (4). The corresponding grouping of the computing unit (4) can be used to find the consistency between musical features (10) and vocal features (11).

5. The method according to claim 4, Its features are, The computing device (7) uses an unsupervised machine learning algorithm to group lip movements.

6. The method according to any one of claims 1 to 5, Its features are, When at least two vehicle occupants (3) are in the vehicle (6), the computing unit (4) detects the emotional state (12) of the vehicle occupants (6) by processing sensor data and associates the corresponding emotional state (12) with the detected or absent singing (5).

7. The method according to claim 6, Its features are, The computing unit (4): - In the event that the emotional state (12) of the vehicle occupant (3) improves when the singing (5) begins and / or in the event that the emotional state (12) of the vehicle occupant (3) deteriorates when the singing (5) ends, the reproduction of another song in the vehicle (6), particularly a song including at least one singing trigger feature (10.1), is automatically triggered; and / or - If the emotional state (12) of the vehicle occupant (3) deteriorates when the singing (5) begins and / or the emotional state (12) of the vehicle occupant (3) improves when the singing (5) ends, the reproduction of the song currently playing in the vehicle (6) or the reproduction of a song including at least one singing blocking feature (10.2) in the vehicle (6) is automatically terminated.

8. The method according to claim 7, Its features are, The computing unit (4) determines, before reproducing a song in the vehicle (6), the respective song triggering feature (10.1) and song blocking feature (10.2) contained in the song by accessing the trigger database (9).

Citation Information

Patent Citations

  • Recognition system in a vehicle for recording the speech activity of a vehicle occupant

    DE102013222645A1

  • Method for Music Selection by Gesture and Speech Control

    DE102016204183A1

  • Method for controlling a multimedia device, as well as computer program and equipment for this purpose.

    DE102018214976A1

  • Mood Sensor

    US20120130196A1

  • Query endpointing based on lip detection

    US20180268812A1