System and method for automatically adjusting vehicle multimedia volume

By acquiring in-car audio and image information, extracting features and generating control instructions, the volume of the in-car multimedia is automatically adjusted, solving the problem of existing technologies that cannot adjust the volume in non-dialogue scenarios, and improving the driving experience and safety.

CN116533909BActive Publication Date: 2025-09-05CHERY NEW ENERGY AUTOMOBILE TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310484164.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-04-27
Publication Date
2025-09-05
Estimated Expiration
2043-04-27

AI Technical Summary

Technical Problem

Existing technologies are unable to automatically adjust the volume of in-vehicle multimedia in non-conversation scenarios, especially when the driver is tired or the passengers are resting. The volume cannot be adjusted according to actual needs, resulting in a poor user experience.

Method used

By obtaining in-car audio information and facial image information, pre-processing and extracting audio and image features, generating final control instructions, and automatically adjusting the volume of the vehicle's multimedia, including using microphones and cameras to collect data, converting digital-to-analog converters to digital signals, using the librosa library for audio noise reduction, and the dlib library for facial feature detection.

Benefits of technology

It realizes automatic adjustment of the vehicle multimedia volume in non-dialogue scenarios, improves the comfort and intelligent experience of drivers and passengers, avoids frequent manual adjustment operations, and improves driving safety and riding comfort.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116533909B_ABST
    Figure CN116533909B_ABST
Patent Text Reader

Abstract

The system and method proposed in this invention automatically adjusts the volume of in-vehicle multimedia devices. This system obtains and preprocesses audio and image information from within the vehicle, extracts audio and image features, and then analyzes and generates final control instructions. The system then adjusts the volume of the in-vehicle multimedia device based on these final control instructions. While the in-vehicle multimedia device is playing audio, audio and image analysis is used to generate the final control instructions. This eliminates the need for manual adjustment by the driver or passengers, ensuring that the volume of the in-vehicle multimedia output is always automatically adjusted to the desired level. This enhances the intelligence of the in-vehicle multimedia device and improves the user experience for the driver and passengers.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This solution belongs to the field of in-vehicle intelligent cockpits, and specifically relates to a system and method for automatically adjusting the volume of in-vehicle multimedia. Background Art

[0002] With the advent of the intelligent and connected era of cars, the smart cockpit has become a key competitive area in the smart car market. Domestic automakers and emerging brands are investing heavily in this area, upgrading features and evolving functionality to enhance their brand competitiveness. Consumers are also increasingly interested in smart cars, especially younger consumers, who often view the smart cockpit as a "third living space" where they can listen to music, watch movies, and play games, enjoying the wonderful experience brought by smart cars. However, current in-car multimedia systems for listening to music and other such experiences often lack intelligent control, requiring manual adjustment, either through the large screen or voice control. Existing automatic volume adjustment solutions often simply adjust the current media volume based on vehicle speed.

[0003] Prior art discloses methods that collect sound intensity and adjust it according to pre-set policies to meet user expectations for in-car volume. Other prior art discloses methods that detect passengers' real-time mouth shapes to adjust the volume of in-car multimedia playback, allowing conversations to be lowered, improving user experience.

[0004] However, for some non-dialogue scenarios, such as when the driver is fatigued while driving or when the passengers are resting, automatic volume adjustment cannot be achieved. Summary of the Invention

[0005] In order to solve the above problems, the present invention proposes a system and method for automatically adjusting the volume of in-vehicle multimedia, which realizes the effect of automatically adjusting the volume of in-vehicle multimedia in non-dialogue scenarios, bringing a more comfortable and intelligent driving experience to the occupants of the vehicle.

[0006] In order to achieve the above object, the present invention adopts the following technical solutions:

[0007] The present invention proposes a method for automatically adjusting the volume of vehicle-mounted multimedia, comprising:

[0008] Obtain audio information and facial image information in the car;

[0009] Preprocessing the audio information to obtain preprocessed audio information, and preprocessing the image information to obtain preprocessed image information;

[0010] The signal-to-noise ratio is calculated by comparing the pre-processed audio information with the original audio information to obtain audio features, and the facial features of the pre-processed image information are extracted to obtain image features.

[0011] Generate final control instructions according to audio features and image features;

[0012] The volume of the in-vehicle multimedia is adjusted according to the final control instruction.

[0013] Preferably, the audio information is preprocessed to obtain preprocessed audio information, specifically including: the audio information includes multimedia sound and noise in the car, the audio information is short-time Fourier transformed to obtain transformed audio information, the transformed audio information is denoised by spectral subtraction to obtain denoised audio information, and the denoised audio information is inversely short-time Fourier transformed to obtain preprocessed audio information.

[0014] Preferably, the image information is preprocessed to obtain preprocessed image information, specifically including gray-scaling the collected image information to obtain grayscale image information, and then removing noise from the grayscale image information through median filtering and mean filtering to obtain the preprocessed image information.

[0015] Preferably, the audio characteristics are compared according to a preset signal-to-noise ratio range to obtain the in-vehicle noise situation, wherein the in-vehicle noise situation includes normal noise and excessive noise.

[0016] Preferably, facial features of preprocessed image information are extracted to obtain image features, specifically including detecting the coordinates of the eyes and mouth to obtain the eye opening degree, eye opening and closing frequency, mouth distance value and mouth opening and closing frequency of the face, and the image features include eye opening and closing degree features, eye opening and closing frequency features, mouth distance features, and mouth opening and closing frequency features.

[0017] Preferably, the image features are compared according to a preset image library to obtain state scenes corresponding to the image features, wherein the state scenes corresponding to the image features include a driver fatigue scene and a passenger rest scene;

[0018] The driver fatigue scenario is that the characteristic value of the driver's eye opening degree characteristic is higher than a preset first opening degree, the characteristic value of the eye opening frequency characteristic is lower than a first preset frequency, the characteristic value of the mouth distance characteristic is lower than a preset first distance, and the characteristic value of the mouth opening frequency characteristic is lower than the first preset frequency;

[0019] The image features of the passenger rest scene are that the characteristic value of the passenger's eye opening degree feature is lower than a preset second opening degree, the characteristic value of the eye opening and closing frequency feature is lower than the second preset frequency, the characteristic value of the mouth distance feature is lower than the preset second distance, and the characteristic value of the mouth opening and closing frequency feature is lower than the second preset frequency.

[0020] Preferably, the image features are compared according to a preset image library to obtain the state scene corresponding to the image features, a query is performed in the preset image library to obtain the face picture with the highest similarity to the image features, and the state scene corresponding to the face picture is used as the state scene corresponding to the image features.

[0021] Preferably, the first control instruction is generated according to the in-car noise situation corresponding to the audio feature, and the second control instruction is generated according to the state scene corresponding to the image feature;

[0022] When the noise condition inside the vehicle changes from normal to excessive noise, the result of the first control instruction is an increase instruction; when the noise condition inside the vehicle changes from excessive noise to normal, the result of the first control instruction is a decrease instruction;

[0023] When the driver is tired, the result of the second control command is an increase command; when the passenger is resting, the result of the second control command is a decrease command;

[0024] Perform an AND operation on the results of the first control instruction and the second control instruction to obtain the final control instruction. The decrease instruction is agreed to be , and the increase instruction is ;

[0025] When the result of the first control instruction is a decrease instruction and the result of the second control instruction is an increase instruction, the final control instruction obtained is to decrease the volume;

[0026] When the result of the first control instruction is a lowering instruction and the result of the second control instruction is a lowering instruction, the final control instruction obtained is to lower the volume;

[0027] When the result of the first control instruction is an increase instruction and the result of the second control instruction is a decrease instruction, the final control instruction obtained is to decrease the volume;

[0028] When the result of the first control instruction is an increase instruction and the result of the second control instruction is also an increase instruction, the final control instruction obtained is to increase the volume.

[0029] The present invention proposes a system for automatically adjusting the volume of vehicle-mounted multimedia, comprising:

[0030] The acquisition module is used to obtain audio information and facial image information in the car;

[0031] A preprocessing module, configured to preprocess the audio information to obtain preprocessed audio information, and preprocess the image information to obtain preprocessed image information;

[0032] The feature extraction module calculates the signal-to-noise ratio between the preprocessed audio and the original audio to obtain audio features; and extracts facial features from the preprocessed image information to obtain image features;

[0033] Analysis module, which generates final control instructions based on audio features and image features;

[0034] The vehicle-mounted multimedia module adjusts the volume of the vehicle-mounted multimedia according to the final control instruction.

[0035] Preferably, the acquisition module includes a microphone, a camera and a digital-to-analog converter, the microphone is used to obtain audio information in the car, the camera is used to obtain image information of the face, and the digital-to-analog converter converts the information obtained by the microphone and the camera into digital signals.

[0036] Compared with the prior art, the present invention has the following beneficial effects:

[0037] This invention proposes a method for automatically adjusting the volume of in-vehicle multimedia devices. By acquiring and preprocessing audio and image information within the vehicle, extracting audio and image features, and then analyzing and generating final control instructions, the system adjusts the volume of the in-vehicle multimedia device based on these final control instructions. While the in-vehicle multimedia device is playing audio, audio and image analysis is used to generate the final control instructions. This eliminates the need for manual adjustment by the driver or passengers, ensuring that the volume of the in-vehicle multimedia output is always automatically adjusted to the desired level. This enhances the intelligence of the in-vehicle multimedia device and improves the user experience for the driver and passengers.

[0038] The present invention proposes a system for automatically adjusting the volume of in-vehicle multimedia. The acquisition module acquires data, the preprocessing module preprocesses the data to facilitate subsequent use of the data, and the analysis module analyzes and judges the data. The system generates relevant control instructions based on the noise conditions in the vehicle and the status of the passengers and the driver. The system controls the volume through the in-vehicle multimedia module, automatically meeting the volume adjustment requirements of the in-vehicle multimedia and improving the riding experience. BRIEF DESCRIPTION OF THE DRAWINGS

[0039] Figure 1 The present invention discloses a flow chart of a method for automatically adjusting the volume of in-vehicle multimedia.

[0040] Figure 2 This is a module diagram of a system for automatically adjusting the volume of in-vehicle multimedia disclosed in the present invention. DETAILED DESCRIPTION

[0041] This section will describe in detail the specific embodiments of the present invention. The preferred embodiments of the present invention are shown in the accompanying drawings. The purpose of the accompanying drawings is to supplement the description of the text part of the specification with graphics, so that people can intuitively and vividly understand each technical feature and the overall technical solution of the present invention, but it should not be understood as a limitation on the scope of protection of the present invention.

[0042] In the description of the present invention, it should be understood that descriptions involving orientations, such as up, down, front, back, left, right, etc., indicating orientations or positional relationships, are based on the orientations or positional relationships shown in the accompanying drawings. They are only for the convenience of describing the present invention and simplifying the description, and do not indicate or imply that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation. Therefore, they cannot be understood as limitations on the present invention.

[0043] In the description of the present invention, "several" means one or more, "many" means more than two, "greater than," "less than," and "exceed" are understood to exclude the number itself, while "above," "below," and "within" are understood to include the number itself. The use of "first" and "second" in the description is solely for the purpose of distinguishing technical features and should not be construed as indicating or implying relative importance, implicitly specifying the number of the indicated technical features, or implicitly specifying the order of the indicated technical features.

[0044] In the description of the present invention, unless otherwise clearly defined, terms such as setting, installing, and connecting should be understood in a broad sense, and technicians in the relevant technical field can reasonably determine the specific meanings of the above terms in the present invention based on the specific content of the technical solution.

[0045] The present invention will be further described in detail below with reference to specific embodiments, which are intended to explain the present invention rather than to limit it.

[0046] The present invention proposes a method for automatically adjusting the volume of vehicle-mounted multimedia, comprising:

[0047] Obtain audio information and facial image information in the car;

[0048] Preprocessing the audio information to obtain preprocessed audio information, and preprocessing the image information to obtain preprocessed image information;

[0049] The signal-to-noise ratio is calculated by comparing the pre-processed audio information with the original audio information to obtain audio features, and the facial features of the pre-processed image information are extracted to obtain image features.

[0050] Generate final control instructions according to audio features and image features;

[0051] The volume of the in-vehicle multimedia is adjusted according to the final control instruction.

[0052] Obtain audio information and facial image information in the car;

[0053] Preprocessing the audio information to obtain preprocessed audio information, and preprocessing the image information to obtain preprocessed image information;

[0054] The signal-to-noise ratio is calculated based on the relationship between the pre-processed audio information and the original audio information to obtain audio features. The audio features are compared according to a preset signal-to-noise ratio range to obtain the in-vehicle noise conditions, which include normal noise and excessive noise. Facial features are extracted from the pre-processed image information to obtain image features. The image features are compared with a preset image library to obtain the state scenes corresponding to the image features, which include the driver fatigue scene and the passenger rest scene.

[0055] generating a first control instruction based on the vehicle interior noise condition corresponding to the audio feature, wherein when the vehicle interior noise condition changes from normal to excessive noise, the result of the first control instruction is an increase instruction, and when the vehicle interior noise condition changes from excessive noise to normal, the result of the first control instruction is a decrease instruction;

[0056] Generate a second control instruction based on the state scenario corresponding to the image feature. When the driver is tired, the result of the second control instruction is an increase instruction; when the passenger is resting, the result of the second control instruction is a decrease instruction;

[0057] Selecting the results of the first control instruction and the second control instruction according to the priority order of the decrease instruction over the increase instruction to obtain a final control instruction;

[0058] The volume of the in-vehicle multimedia is adjusted according to the final control instruction.

[0059] The system takes into account the current noise level inside the vehicle, whether passengers are resting or sleeping, and whether the driver is fatigued. It generates different control commands to lower or raise the volume based on these different scenarios. This prevents the driver from frequently adjusting the volume, allowing them to focus on driving, thereby improving driving safety. It also provides the appropriate volume level for the front passenger and rear passengers in different scenarios, enhancing travel comfort.

[0060] Adjust the volume of the in-car multimedia according to the control instructions. Adjust the volume of the currently playing video, music, etc.

[0061] In the specific embodiment of the present invention, please refer to Figure 1 , preprocessing the audio information to obtain preprocessed audio information, specifically including: the audio information includes multimedia sound and noise in the car, the audio information includes multimedia sound and noise in the car, performing short-time Fourier transform on the audio information to obtain transformed audio information, denoising the transformed audio information through spectral subtraction to obtain denoised audio information, and performing short-time Fourier inverse transform on the denoised audio information to obtain preprocessed audio information.

[0062] Use librosa library to perform short-time Fourier transform, then use spectral subtraction to remove noise, and use librosa library to perform short-time Fourier inverse transform on the denoised audio information to obtain preprocessed audio information;

[0063] librosa is a Python library for audio and music analysis. It provides functions such as stft for short-time Fourier transforms and isft for inverse short-time Fourier transforms. Spectral subtraction is a classic noise reduction algorithm with fast processing speed. It is widely used in speech noise reduction and can produce relatively pure audio information.

[0064] In the specific embodiment of the present invention, please refer to Figure 1 , preprocessing the image information to obtain preprocessed image information, specifically including gray-scaling the collected image information to obtain grayscale image information, and then removing the noise of the grayscale image information through median filtering and mean filtering to obtain preprocessed image information. Due to external environments such as noise and lighting or the equipment itself, the quality of the original image obtained is usually not very high, so it is necessary to enhance the original image to make the image clearer and highlight the image features to facilitate further image recognition and analysis. Image enhancement methods include spatial domain enhancement methods, and spatial domain enhancement methods include median filtering and mean filtering.

[0065] Median filtering is a nonlinear filtering technique that suppresses impulse noise and eliminates spike interference noise while effectively preserving the edges of the target image. Mean filtering is a linear filtering technique that offers the advantages of simplicity, efficiency, and ease of implementation. It can provide a rough description of object features, with selective masking being particularly effective.

[0066] In the specific embodiment of the present invention, please refer to Figure 1 , comparing the audio features according to a preset signal-to-noise ratio range to obtain the in-car noise situation, wherein the in-car noise situation includes normal noise and excessive noise.

[0067] Classifying the in-car noise conditions corresponding to the audio features can further refine the situations in actual applications, which is conducive to generating control instructions that better meet user needs and improve user comfort.

[0068] In the specific embodiment of the present invention, please refer to Figure 1 , detect the coordinates of the eyes and mouth to obtain the eye opening degree, eye opening and closing frequency, mouth distance value and mouth opening and closing frequency of the face, and the image features include eye opening and closing degree features, eye opening and closing frequency features, mouth distance features, and mouth opening and closing frequency features.

[0069] Use dlib to detect the coordinates of the eyes and mouth to obtain the eye opening degree, eye opening and closing frequency, mouth distance value and mouth opening and closing frequency of the face.

[0070] The image feature extraction process uses Python's dlib library to extract facial information from the image, mainly including the degree of eye openness, which is used to determine whether the driver is resting and whether the driver is tired based on the preset degree of openness and frequency; the mouth distance value is used to determine whether the driver is speaking based on the preset distance and frequency.

[0071] Facial information and key points are detected through Python's dlib library. The facial key point detection technology in the dlib library is mainly based on two algorithms: the HOG feature extraction algorithm and the support vector machine classifier. The HOG algorithm is a feature extraction algorithm used for target detection. It obtains the features of the image's corners, edges, and textures by calculating gradient information. It aims to extract target features with characteristics such as scale invariance, translation invariance, and rotation invariance, and fit a feature vector so that this feature vector can accurately describe the face in the current image. The support vector machine is an excellent classifier that has been verified by many practices and is widely used in many computer vision applications. Its main idea is to find a hyperplane to separate certain samples from a group of samples and maximize this group of intervals.

[0072] In the specific embodiment of the present invention, please refer to Figure 1 , according to the preset image library, the image features are compared to obtain the state scenes corresponding to the image features, the state scenes corresponding to the image features include the driver fatigue scene and the passenger rest scene,

[0073] The driver fatigue scenario is that the characteristic value of the driver's eye opening degree characteristic is higher than a preset first opening degree, the characteristic value of the eye opening frequency characteristic is lower than a first preset frequency, the characteristic value of the mouth distance characteristic is lower than a preset first distance, and the characteristic value of the mouth opening frequency characteristic is lower than the first preset frequency;

[0074] The image features of the passenger rest scene are that the characteristic value of the passenger's eye opening degree feature is lower than a preset second opening degree, the characteristic value of the eye opening and closing frequency feature is lower than the second preset frequency, the characteristic value of the mouth distance feature is lower than the preset second distance, and the characteristic value of the mouth opening and closing frequency feature is lower than the second preset frequency.

[0075] In the specific embodiment of the present invention, please refer to Figure 1 The method of comparing the image features according to the preset image library to obtain the state scene corresponding to the image features specifically includes searching in the preset image library to obtain the face picture with the highest similarity to the image features, and using the state scene corresponding to the face picture as the state scene corresponding to the image features.

[0076] In the specific embodiment of the present invention, please refer to Figure 1 , generating a first control instruction according to the in-car noise situation corresponding to the audio feature, and generating a second control instruction according to the state scene corresponding to the image feature;

[0077] When the noise condition inside the vehicle changes from normal to excessive noise, the result of the first control instruction is an increase instruction; when the noise condition inside the vehicle changes from excessive noise to normal, the result of the first control instruction is a decrease instruction;

[0078] When the driver is tired, the result of the second control command is an increase command; when the passenger is resting, the result of the second control command is a decrease command;

[0079] Perform an AND operation on the results of the first control instruction and the second control instruction to obtain the final control instruction, with the decrease instruction being 1 and the increase instruction being 0;

[0080] When the result of the first control instruction is a decrease instruction and the result of the second control instruction is an increase instruction, the final control instruction obtained is to decrease the volume;

[0081] When the result of the first control instruction is a lowering instruction and the result of the second control instruction is a lowering instruction, the final control instruction obtained is to lower the volume;

[0082] When the result of the first control instruction is an increase instruction and the result of the second control instruction is a decrease instruction, the final control instruction obtained is to decrease the volume;

[0083] When the result of the first control instruction is an increase instruction and the result of the second control instruction is also an increase instruction, the final control instruction obtained is to increase the volume.

[0084] The present invention proposes a system for automatically adjusting the volume of in-vehicle multimedia. Figure 2 ,include,

[0085] The acquisition module is used to obtain audio information and facial image information in the car;

[0086] A preprocessing module, configured to preprocess the audio information to obtain preprocessed audio information, and preprocess the image information to obtain preprocessed image information;

[0087] The feature extraction module calculates the signal-to-noise ratio between the preprocessed audio and the original audio to obtain audio features; and extracts facial features from the preprocessed image information to obtain image features;

[0088] Analysis module, which generates final control instructions based on audio features and image features;

[0089] The vehicle-mounted multimedia module adjusts the volume of the vehicle-mounted multimedia according to the final control instruction.

[0090] The acquisition module collects data, the preprocessing module pre-processes the data to facilitate subsequent use, the feature extraction module extracts features from the data, and the analysis module analyzes and judges the data. It generates relevant control instructions based on the noise level in the car and the status of the passengers. The volume is then controlled through the in-vehicle multimedia module. If the noise level in the car is high, the multimedia volume is appropriately increased. If someone is talking or the passenger is resting, the multimedia volume is appropriately lowered. If the driver is showing signs of fatigue, the volume is appropriately increased to provide a warning. This automatically meets the needs of in-vehicle multimedia volume adjustment, improving the driving experience.

[0091] In the specific embodiment of the present invention, please refer to Figure 2 The acquisition module includes a microphone, a camera and a digital-to-analog converter. The microphone is used to obtain audio information inside the car, the camera is used to obtain image information of the face, and the digital-to-analog converter converts the information obtained by the microphone and camera into digital signals.

[0092] The sound and image information inside the car is obtained through the microphone and camera. The obtained analog signal is converted into a digital signal by a digital-to-analog converter and then passed to the preprocessing module.

[0093] The above content is a further detailed description of the present invention in combination with specific preferred embodiments. It cannot be considered that the specific embodiments of the present invention are limited to these. For ordinary technicians in the technical field to which the present invention belongs, they can make several simple deductions or substitutions without departing from the concept of the present invention, which should be regarded as belonging to the scope of patent protection determined by the submitted claims of the present invention.

Claims

1. A method for automatically adjusting the volume of in-vehicle multimedia, characterized in that: include, Obtain audio information and facial image information in the car; Preprocessing the audio information to obtain preprocessed audio information, and preprocessing the image information to obtain preprocessed image information; The signal-to-noise ratio is calculated by comparing the pre-processed audio information with the original audio information to obtain audio features, and the facial features of the pre-processed image information are extracted to obtain image features. Generate final control instructions according to audio features and image features; adjusting the volume of the vehicle multimedia according to the final control instruction; Generate a first control instruction based on the in-car noise situation corresponding to the audio feature, and generate a second control instruction based on the state scene corresponding to the image feature; When the noise condition inside the vehicle changes from normal to excessive noise, the result of the first control instruction is an increase instruction; when the noise condition inside the vehicle changes from excessive noise to normal, the result of the first control instruction is a decrease instruction; When the driver is tired, the result of the second control command is an increase command; when the passenger is resting, the result of the second control command is a decrease command; Perform an AND operation on the results of the first control instruction and the second control instruction to obtain the final control instruction, with the decrease instruction being 1 and the increase instruction being 0; When the result of the first control instruction is a decrease instruction and the result of the second control instruction is an increase instruction, the final control instruction obtained is to decrease the volume; When the result of the first control instruction is a lowering instruction and the result of the second control instruction is a lowering instruction, the final control instruction obtained is to lower the volume; When the result of the first control instruction is an increase instruction and the result of the second control instruction is a decrease instruction, the final control instruction obtained is to decrease the volume; When the result of the first control instruction is an increase instruction and the result of the second control instruction is also an increase instruction, the final control instruction obtained is to increase the volume.

2. The method for automatically adjusting the volume of in-vehicle multimedia according to claim 1, characterized in that: The audio information is preprocessed to obtain preprocessed audio information, specifically including: the audio information includes multimedia sound and noise in the vehicle; the audio information is short-time Fourier transformed to obtain transformed audio information; the transformed audio information is denoised by spectral subtraction to obtain denoised audio information; and the denoised audio information is inversely short-time Fourier transformed to obtain preprocessed audio information.

3. The method for automatically adjusting the volume of in-vehicle multimedia according to claim 1, characterized in that: The image information is preprocessed to obtain preprocessed image information, specifically including gray-scaling the collected image information to obtain grayscale image information, and then removing noise from the grayscale image information through median filtering and mean filtering to obtain preprocessed image information.

4. The method for automatically adjusting the volume of in-vehicle multimedia according to claim 1, characterized in that: The audio features are compared according to a preset signal-to-noise ratio range to obtain the in-vehicle noise situation, where the in-vehicle noise situation includes normal noise and excessive noise.

5. The method for automatically adjusting the volume of in-vehicle multimedia according to claim 4, characterized in that: Extracting facial features from preprocessed image information to obtain image features, specifically including detecting the coordinates of the eyes and mouth to obtain the eye opening degree, eye opening and closing frequency, mouth distance value, and mouth opening and closing frequency of the face, wherein the image features include eye opening degree features, eye opening and closing frequency features, mouth distance features, and mouth opening and closing frequency features.

6. The method for automatically adjusting the volume of in-vehicle multimedia according to claim 5, characterized in that: Comparing image features according to a preset image library to obtain state scenes corresponding to the image features, wherein the state scenes corresponding to the image features include a driver fatigue scene and a passenger rest scene; The driver fatigue scenario is that the characteristic value of the driver's eye opening degree characteristic is higher than a preset first opening degree, the characteristic value of the eye opening frequency characteristic is lower than a first preset frequency, the characteristic value of the mouth distance characteristic is lower than a preset first distance, and the characteristic value of the mouth opening frequency characteristic is lower than the first preset frequency; The image features of the passenger rest scene are that the characteristic value of the passenger's eye opening degree feature is lower than a preset second opening degree, the characteristic value of the eye opening frequency feature is lower than the second preset frequency, the characteristic value of the mouth distance feature is lower than the preset second distance, and the characteristic value of the mouth opening frequency feature is lower than the second preset frequency.

7. The method for automatically adjusting the volume of in-vehicle multimedia according to claim 6, characterized in that: The image features are compared with the preset image library to obtain the state scene corresponding to the image features. A query is performed in the preset image library to obtain the face picture with the highest similarity to the image features, and the state scene corresponding to the face picture is used as the state scene corresponding to the image features.

8. A system for automatically adjusting the volume of in-vehicle multimedia, characterized in that: include, The acquisition module is used to obtain audio information and facial image information in the car; A preprocessing module, configured to preprocess the audio information to obtain preprocessed audio information, and preprocess the image information to obtain preprocessed image information; The feature extraction module calculates the signal-to-noise ratio between the preprocessed audio and the original audio to obtain audio features; Extracting facial features from preprocessed image information to obtain image features; Analysis module, which generates final control instructions based on audio features and image features; An in-vehicle multimedia module, adjusting the volume of the in-vehicle multimedia according to the final control instruction; Generate a first control instruction based on the in-car noise situation corresponding to the audio feature, and generate a second control instruction based on the state scene corresponding to the image feature; When the noise condition inside the vehicle changes from normal to excessive noise, the result of the first control instruction is an increase instruction; when the noise condition inside the vehicle changes from excessive noise to normal, the result of the first control instruction is a decrease instruction; When the driver is tired, the result of the second control command is an increase command; when the passenger is resting, the result of the second control command is a decrease command; Perform an AND operation on the results of the first control instruction and the second control instruction to obtain the final control instruction, with the decrease instruction being 1 and the increase instruction being 0; When the result of the first control instruction is a decrease instruction and the result of the second control instruction is an increase instruction, the final control instruction obtained is to decrease the volume; When the result of the first control instruction is a lowering instruction and the result of the second control instruction is a lowering instruction, the final control instruction obtained is to lower the volume; When the result of the first control instruction is an increase instruction and the result of the second control instruction is a decrease instruction, the final control instruction obtained is to decrease the volume; When the result of the first control instruction is an increase instruction and the result of the second control instruction is also an increase instruction, the final control instruction obtained is to increase the volume.

9. The system for automatically adjusting the volume of in-vehicle multimedia according to claim 8, characterized in that: The acquisition module includes a microphone, a camera and a digital-to-analog converter. The microphone is used to obtain audio information in the car, the camera is used to obtain image information of the face, and the digital-to-analog converter converts the information obtained by the microphone and camera into digital signals.

Citation Information

Patent Citations

  • Intelligent vehicle-mounted audio precise pushing method for driving congestion

    CN110134821A

  • Vehicle running state monitoring method, vehicle and computer readable storage medium

    CN112193250A