VOLUME CONTROL METHOD
Patent Information
- Application Number
- DE602021033325
- Authority / Receiving Office
- DE · DE
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2020-06-02
- Filing Date
- 2021-05-10
- Publication Date
- 2025-07-02
- Estimated Expiration
- 2041-05-10
AI Technical Summary
Existing vehicle sound volume control systems do not adapt automatically to the intensity of passenger conversations, requiring manual adjustment and failing to consider conversation intensity for volume modulation.
A method utilizing machine learning to classify sound situations and visual communication intensity, combined with echo cancellation and personalized volume control based on passenger behavior, to dynamically adjust sound levels in response to conversations.
Automatically adjusts sound volume to match conversation intensity, enhancing user experience by minimizing manual intervention and providing personalized, zone-specific volume adjustments.
Description
[0001] The invention relates to the control of sound volume in an enclosed or partially enclosed space. It finds an advantageous application in the form of a method for controlling sound volume in a cabin, the cabin corresponding to an enclosed or partially enclosed space, and in particular in a cabin constituted by the passenger compartment of a motor vehicle.
[0002] It also relates to a computer program product whose instructions are capable of implementing such a method.
[0003] It also relates to a multimedia system implementing such a method.
[0004] It also relates to a motor vehicle comprising such a multimedia system.
[0005] Current vehicles have a Speed Dependent Volume Control (SDVC) feature. As the name suggests, this feature adjusts the volume level based on the vehicle's speed. Therefore, the faster the vehicle goes, the more noise there is in the cabin, and the more it is necessary to increase the multimedia volume to ensure that it remains audible for passengers. Thanks to SDVC, the user is supposed to have less need to manipulate the volume knob. However, today, if two passengers start a conversation (more or less sustained), the multimedia volume does not adapt.
[0006] We also know the document JP2019137167 concerning a method of controlling audio output based on the conversation established between passengers, which is determined by capturing image and sound, but no notion of intensity of the conversation is determined or taken into consideration to modulate the sound volume.
[0007] Similarly, document JP2012025270 is known concerning a method for regulating the volume of sound inside the passenger compartment based on elements captured by microphones and images of the occupants using a camera to differentiate sounds coming from inside the passenger compartment and those coming from outside, however this method does not determine or take into consideration the intensity of the conversation to modulate the sound volume. Document DE102016003401 is also known concerning a method for detecting conversation in a vehicle using a microphone and camera, with speech recognition, but this document does not allow the sound volume to be modulated based on the intensity of the conversation.
[0008] The document GOTTFRIED BEHLER ET AL: "Automatic Loudness Control", PROCEEDINGS OF THE 20TH INTERNATIONAL CONGRESS ON ACOUSTICS, August 23, 2010 discloses a method for automatic volume control in a vehicle, with echo cancellation allowing the signal emitted by the loudspeaker to be subtracted from the signal obtained by the microphone. This signal is used to control the playback level.
[0009] Finally, document US 2020 / 083856 A1 describes automatic volume adjustment in a vehicle based on a classification of the sound environment and context determined by cameras.
[0010] Existing methods therefore do not allow the multimedia volume to be automatically adapted according to the intensity of the passengers' conversation. One of the aims of the invention is to overcome at least some of the drawbacks of the prior art by providing a method for controlling the sound volume in the passenger compartment of a motor vehicle.
[0011] To this end, the invention provides a method for controlling sound volume according to claim 1.
[0012] Thanks to the invention, the sound level can automatically and dynamically adapt to users' conversations so as not to force them to modify the volume of the multimedia system themselves.
[0013] According to an advantageous feature, the third step uses a machine learning model, which allows pre-training for the recognition of given sound situation classes.
[0014] According to another advantageous feature, the cabin sound situation classification distinguishes at least three situations, namely the absence of communication in the cabin, speaking with a single voice in the cabin, multi-voice communication in the cabin, which makes it possible to classify common sound situations. In addition, the classification can also include the situation of singing correlated with the sound of the multimedia system, so as to distinguish this common situation, which is easily recognizable by comparing the sound signals.
[0015] Another advantageous feature is that multi-voice communication in the cabin is differentiated between light and sustained communication, which allows sound situations to be classified according to their intensity.
[0016] According to another advantageous feature, the fourth step uses a head pose estimator, which makes it possible to know the orientation of the users' gaze.
[0017] The advantage of the feature that the fourth step outputs an estimate of gaze frequency is that it allows the intensity of visual communication in the cabin to be determined.
[0018] According to another advantageous feature, the fifth step uses a mapping taking as input the classified sound situation and the determined visual communication intensity, which makes it possible to limit calculation times by combining robustness and efficiency.
[0019] According to another advantageous characteristic, the fifth step is personalized by previously identified passenger, for example by adapting the mapping to each passenger, to take into consideration the individual behavior of each passenger which allows not only to adapt the volume by zone according to the location by the camera of the passengers participating in the conversation and which allows at the same time each passenger to correct the volume individually if necessary, and thus allows to personalize the choices of the passengers and apply them to each speaker in a differentiated manner reinforced in particular by facial recognition.
[0020] According to another advantageous feature, the fifth stage includes a synthesis with SDVC functionality, which makes it possible to take into account both ambient noise and conversation intensities and in particular robust operation both in enclosed and partially enclosed spaces, with an open window for example.
[0021] The invention also relates to a computer program product comprising program code instructions recorded on a computer-readable medium, comprising instructions which, when the program is executed by the computer, cause the latter to implement the method according to the invention, which has advantages similar to those of the method.
[0022] The invention also relates to a multimedia system according to claim 8 which has advantages similar to those of the method.
[0023] The invention also relates to a motor vehicle comprising a multimedia system according to the invention, which has advantages similar to those of the method in a vehicle.
[0024] Other aims, characteristics and advantages of the invention will appear on reading the following description, given solely by way of non-limiting example, and made with reference to the appended drawings in which: [ Fig.1 ] there figure 1 which has already been mentioned, schematically illustrates a method of controlling sound volume in accordance with the invention, and, [ Fig. 2 ] there figure 2 illustrates a multimedia system in accordance with the invention. For clarity, identical or similar elements are identified by identical reference signs throughout the figures.
[0025] We have schematically represented on the figure 1 an embodiment of the method according to the invention. The method PROC for controlling the sound volume generated by a loudspeaker in a cabin, characterized in that it comprises: a first step E1 of acquiring the sound in the cabin, this first step E1 is carried out in particular by means of microphones which are connected to the multimedia system which implements the method. At the output of this first step is therefore obtained the raw sound Sb in the cabin, which is here for example the passenger compartment of a motor vehicle; a second step of filtering by cancellation in the sound acquired in the first step of the sound generated by the loudspeaker, this second step E2 makes it possible to remove from the sound captured by the microphones the sound generated by the loudspeakers. The operating mode is identical to known echo cancellation algorithms.This second step E2 therefore consumes as input the raw sound Sb coming from the first step E1 and information If, for example frequency, coming from the multimedia system and corresponding to the sound generated by the loudspeakers, to which the process then applies the transfer function of the loudspeaker so as to integrate the mechanical component of the loudspeaker with the purely frequency acoustic component. Alternatively, this information If preferentially corresponds to the excitation signals of the loudspeakers to which the process applies the frequency responses of the electronic amplification circuits, the mechanics of the loudspeakers and the acoustic paths between the loudspeakers and the microphones. This second filtering step E2 makes it possible to clean the raw sound captured by the microphones by removing that generated by the loudspeakers (music or voice when listening to a radio broadcast).At the output of this second step E2 is therefore obtained the net sound Sn in the cabin; a third step E3 of cabin sound situation classification from the sound filtered in the second step. This third step E3 consumes as input the net sound Sn and uses a machine learning model. The model is previously trained with a database representing different classes of predetermined sound situations, which are in particular: the absence of communication in the cabin, speaking with a single voice in the cabin, multi-voice communication in the cabin, multi-voice communication in the cabin being differentiated between light communication and sustained communication.The silence times and the balance of speaking time illustrate for example the distinction between mono and multi-voice communication as well as the intensity of the exchange in multi-voice communication, thus a type of tit for tat response with balance of speaking time corresponds to a real intense multi-voice communication, and not to a unidirectional mono-voice communication (telephone call...), users singing along to the music emitted by the IVI multimedia system.
[0026] The use of machine learning thus makes it possible to classify a communication according to its exchange intensity, which will subsequently allow a different volume attenuation. At the output of this third step E3, the sound situation class Css in the cabin is therefore obtained; a fourth step E4 of determining an intensity of visual communication in the cabin, by consuming as input visual information Iv coming from at least one camera. This visual information Iv makes it possible to evaluate the intensity of visual communication between the passengers through mutual glances. An algorithm of the pose estimator type (or head pose estimator in English) can be used. This algorithm generates as output an estimate of the frequency of the gaze (or frequency of eye-movement in English).The visual information Iv thus makes it possible to know whether the passengers in the cabin regularly look at each other based on the gaze frequency obtained, which makes it possible to obtain at the output of the fourth step E4 a visual communication intensity Icv; the fifth step E5 controls the sound volume generated by the loudspeaker based on the classified sound situation Css in the third step E3 and the visual communication intensity Icv determined in the fourth step E4. To do this, a map taking as input the classified sound situation Css and the determined visual communication intensity Icv is preferably used. Examples of values are given in the rest of the description in connection with the . figure 2 . In order to take into account the personalization of passengers at this stage, and if the number of loudspeakers and microphones in the passenger compartment allows it, the mapping can characterize the classified sound situations Css and the visual communication intensities Icv by seat, or for each passenger, thanks to the identification by row, and seat. Indeed, in the previous stages the means such as the microphones MIC or the camera CAM geographically identify the sounds or images in the passenger compartment, whether intrinsically by analysis or by their location, this data is therefore available, and the classified sound situations Css and the visual communication intensities Icv can thus from the start be by seat. This allows in particular to update the mapping according to each passenger, and to memorize it, for example if a single passenger voluntarily modifies the sound volume after the implementation of the process.The elements of the map could also be memorized by completely identified passenger, whether by facial recognition or other type of recognition if it is available in the vehicle (fingerprint for example), so as to create a database, thus the elements concerning him will be automatically updated in the map in relation to his location in the passenger compartment from the start of driving. In addition, at this stage, the synthesis with the SDVC functionality is also preferably executed, that is to say that the classic SDVC functionality is applied which provides a calculated sound volume and the volume modification determined according to the communication intensities Css, Icv, stored for example in the map, is applied to this value.
[0027] There figure 2illustrates an IVI (In-Vehicle Infotainment System) multimedia system equipped with an application processor that implements the method. This IVI system is here installed in the passenger compartment of a motor vehicle, it is connected to at least two MIC microphones located in the cabin, these microphones are preferably facing the passengers P1, P2, and connected to two HP loudspeakers located in the passenger compartment, for example in the dashboard and / or the headrests, and connected to a CAM camera located in the passenger compartment so as to perceive the head movements of the passengers P1, P2. Several cameras synchronized with each other can also be used and their images merged according to the size of the passenger compartment.
[0028] The MIC microphones record the raw sound Sb which is transmitted to the first module M1 for acquiring the sound in the passenger compartment, then processed by a second module M2 for filtering by cancellation in the raw sound Sb acquired by the first module of the sound Sg generated by the loudspeaker, which makes it possible to obtain at the output of the second module M2 the net sound Sn which corresponds to the sound emitted by the passengers P1, P2 including ambient noise (engine noise, rolling, air conditioning, open window, etc.). Here we focus on the sounds emitted by the passengers because the ambient noise is, moreover, already taken into consideration in the volume management operated by the classic SDVC functionality which is applied in parallel. The third module M3 for classifying the cabin noise situation consumes the net sound Sn filtered by the second module M2, and generates at output the noise situation class Css in the passenger compartment.At the same time, the fourth module M4 determines a visual communication intensity Icv in the passenger compartment from the visual information Iv coming from the CAM camera. Then, the fifth module M5 dynamically controls the sound volume generated by the HP speakers according to the sound situation classified by the third module M3 and the communication intensity determined by the fourth module M4. Preferably, for the needs of the SDVC functionality, this fifth module also consumes the speed information V of the vehicle so as to automatically control the multimedia volume according to the speed of the vehicle. The fifth module M5 then synthesizes the information received Css, Icv, V and then modifies the volume of the multimedia player LM (web radio, music, etc.).
[0029] For example : in the absence of communication in the passenger compartment, there is no change in the volume in relation to the information received on communication intensity Css, Icv, in the case of a single voice speaking in the passenger compartment, a reduction of approximately 3 dB is applied to the volume in relation to the information received on communication intensity Css, Icv, in the case of multi-voice communication in the cabin, if it is a light communication a reduction of approximately 6 dB is applied to the volume in relation to the information received on communication intensity Css, Icv, and if it is a sustained communication a reduction of approximately 12 dB is applied to the volume in relation to the information received on communication intensity Css, Icv. In the case of singing correlated with the sound of the multimedia system in the passenger compartment, an increase of approximately 3dB is applied to the volume, in connection with the information received on communication intensity, essentially Css sound classification.
[0030] Preferably, whether in a vehicle or a cabin, when a sound reduction action is applied, user behavior is analyzed. Two scenarios can arise: 1- either the user accepts the sound adjustment by not correcting it 2- or the user does not accept the proposed adjustment, and manually modifies the sound up or down
[0031] These two cases are analyzed to be taken into account in the future E5 step. For example, if the reduction is 12 dB, but the user corrects by increasing the volume by 5 dB, then the M5 module retains the user behavior, which it can also associate with an identified user, thanks to facial recognition by the CAM camera. When the case arises again, the algorithm will then decide to lower the volume not by 12 but by 7 dB.
[0032] Similarly, depending on the position of passengers P1, P2 in row 1 or row 2 of the vehicle and the number of speakers in the passenger compartment, the volume can be affected differently, for example if the passengers are in row 1, it is possible to reduce the front left speaker by 12dB, reduce the front right speaker by 9dB and reduce the rear speakers by 3dB depending on the position of the users involved in the conversation, which is detected in particular by means of the camera. If there are private acoustic zones in the passenger compartment, as described for example in patent FR3078931, the user experience will be greater by allowing attenuation depending on the zone.Indeed, there is then at least one loudspeaker per private acoustic zone, that is to say per passenger, which not only makes it possible to adapt the volume per zone according to the location by the CAM camera of the passengers participating in the conversation and which also allows each passenger to individually correct the volume if necessary, and thus makes it possible to personalize the choices of the passengers and apply them to each loudspeaker in a differentiated manner reinforced in particular by facial recognition.
[0033] Thanks to the invention, when vehicle passengers want to start a conversation, the system identifies the situation and slowly lowers the volume, for example over five to ten seconds so that the transition goes unnoticed, without the user having to manually change the multimedia sound and at the end of the conversation, the volume slowly returns to the initial level, especially if the speed has not increased.
Claims
1. Method for controlling the volume generated by a loudspeaker (HP) in a cabin containing passengers, characterized in that it comprises: - a first step (E1) of acquiring sound (Sb) in the cabin, - a second step (E2) of filtering by cancellation from the sound (Sb) acquired in the first step the sound (Sg) generated by the loudspeaker (HP), - a third step (E3) of classifying a cabin sound situation (Css) based on the sound (Sn) filtered in the second step (E2), - a fourth step (E4) of determining a visual communication intensity (Icv) in the cabin, said visual communication intensity (Icv) being representative of a frequency with which the passengers of the cabin look at one another, - a fifth step (E5) of controlling the volume generated by the loudspeaker (HP) depending on the sound situation (Css) classified in the third step (E3) and on the visual communication intensity (Icv) determined in the fourth step (E4).
2. Method for controlling volume according to the preceding claim, characterized in that the third step (E3) uses a machine-learning model.
3. Method for controlling volume according to any of the preceding claims, characterized in that the classification of the cabin sound situation distinguishes between at least three situations, namely absence of communication from the cabin, a single voice speaking in the cabin, and multi-voice communication in the cabin.
4. Method for controlling volume according to the preceding claim, characterized in that multi-voice communication in the cabin is differentiated depending on a loudness of exchange in the multi-voice communication.
5. Method for controlling volume according to any of the preceding claims, characterized in that the fourth step (E4) uses a head pose estimator.
6. Method for controlling volume according to any of the preceding claims, characterized in that the fourth step (E4) generates as output an estimation of the frequency of eye-movement.
7. Method for controlling volume according to any of the preceding claims, characterized in that the fifth step (E5) is personalized by previously identified passenger.
8. Multimedia system (IVI) equipped with an application processor, said system being installed in a cabin containing passengers and being connected to at least one loudspeaker (HP) located in the cabin, to at least one camera (CAM) located in the cabin, and to at least two microphones (MIC) located in the cabin, characterized in that it comprises: - a first module (M1) for acquiring sound in the cabin, - a second module (M2) for filtering by cancellation from the sound acquired by the first module the sound generated by the loudspeaker, - a third module (M3) for classifying a cabin sound situation (Css) based on the sound filtered by the second module, - a fourth module (M4) for determining a visual communication intensity (Icv) in the cabin, said visual communication intensity (Icv) being representative of a frequency with which the passengers of the cabin look at one another, - a fifth module (M5) for controlling the volume generated by the loudspeaker (HP) depending on the sound situation (Css) classified by the third module (M3) and on the communication intensity (Icv) determined by the fourth module (M4).
9. Motor vehicle, characterized in that it comprises a multimedia system (IVI) according to the preceding claim.