Volume adjustment method and system of vehicle-mounted sound equipment, electronic equipment and storage medium

By combining vehicle operation information with multimodal fusion algorithms and voice interaction systems, it is possible to automatically detect the passenger's sleep state and adjust the volume while the car is driving, solving the problems of passenger sleep experience and driving safety in existing technologies and improving the passenger's sleep experience and driving safety.

CN120792701APending Publication Date: 2025-10-17CHINA FAW CO LTD +1

Patent Information

Application Number
CN202510722482.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-30
Publication Date
2025-10-17

AI Technical Summary

Technical Problem

Existing technologies are unable to automatically detect the passenger's sleep state and adjust the volume while the car is driving, which affects the passenger's sleeping experience and may distract the driver, leading to safety risks.

Method used

By combining vehicle operation information with a multimodal fusion algorithm, infrared cameras, 3D structured light/ToF cameras, microphone arrays, contact and non-contact sensors are used to detect passenger status, generate sleep state determination signals, and automatically adjust the volume through the voice interaction system.

Benefits of technology

It enables accurate sleep status monitoring and volume adjustment without the need for active passenger operation, reducing the risk of driver distraction and improving passengers' sleeping experience and driving safety.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120792701A_ABST
    Figure CN120792701A_ABST
Patent Text Reader

Abstract

The invention discloses a volume adjusting method and system of a vehicle-mounted sound system, electronic equipment and a storage medium, and relates to the field of vehicle control, and the method comprises the steps: obtaining vehicle operation information, and starting a passenger detection system according to a vehicle motion state; passenger vision related information is obtained by starting a passenger detection system; through a multi-modal fusion algorithm, generating judgment information of the state of the passenger in the vehicle, and judging whether the output volume of the vehicle-mounted sound equipment is greater than a preset playing volume threshold value or not; if yes, judging whether the judgment information of the passenger state in the vehicle confirms that the passenger is in a sleep state; judging whether the in-car entertainment system has an active media focus or not according to the sleep state; if yes, a dialogue manager is started, and inquiry voice is directionally projected to the position of the driver for voice interaction; judging whether a multimedia volume reduction instruction is received or not through voice interaction; according to the scheme, the sleep comfort level of passengers is improved, and the distracting risk of drivers is reduced.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the field of vehicle control, in particular to a volume adjustment method of a vehicle audio, a volume adjustment system of a vehicle audio, an electronic device and a storage medium. BACKGROUND

[0002] With the development of technology and the improvement of living standards, people pay more and more attention to the entertainment property of the car, and audio is undoubtedly an important part. When playing songs loudly during the journey and the passenger falls asleep, the current car cannot detect the passenger state and actively reduce the volume, which will affect the passenger's sleep experience; if the driver actively pays attention to the passenger's state, it will be distracted, which will pose a risk to driving.

[0003] Related patent documents include:

[0004] Patent document 1 (CN115042795A) discloses a safe driving method and system based on driver sleep quality and vehicle, which reminds the driver after monitoring the driver's fatigue, and cannot adjust and control the volume of the passenger.

[0005] Patent document 2 (CN117119350A) discloses an intelligent volume adjustment method and system for a vehicle audio, which reduces the car audio volume after monitoring that the volume in the car is low, and does not involve monitoring the sleep state of the passenger. SUMMARY

[0006] Therefore, the purpose of the present application is to provide a volume adjustment method of a vehicle audio, a volume adjustment system of a vehicle audio, an electronic device and a storage medium. The present application will determine the running state of the vehicle, start the passenger detection system, perform multi-modal fusion algorithm according to the characteristic information of the passenger, generate the sleep state of the passenger, control the volume of the multimedia through the interaction with the driver car machine, and ensure that the passenger has a good sleep experience. Further, it reduces the risk of distraction of the driver, and provides a better intelligent cabin interaction experience for the user.

[0007] The present application provides the following solutions:

[0008] According to one aspect of the present application, a volume adjustment method of a vehicle audio is provided, comprising:

[0009] obtaining vehicle running information;

[0010] determining whether the vehicle is in motion according to the vehicle running information;

[0011] if yes, starting a passenger detection system;

[0012] By starting the passenger detection system, signals of visual receiving state, audio output state and physiological state of passengers in the vehicle are acquired;

[0013] According to the signals of visual receiving state, audio output state and physiological state of passengers in the vehicle, the determination information of the state of passengers in the vehicle is generated through a multi-modal fusion algorithm;

[0014] Based on the fact that the vehicle audio is in an output state, a preset threshold of playing volume is acquired;

[0015] It is judged whether the output volume of the vehicle audio is greater than the preset threshold of playing volume;

[0016] If yes, it is judged whether the determination information of the state of passengers in the vehicle confirms that the passengers are in a sleep state;

[0017] If yes, it is judged whether the infotainment system has an active media focus;

[0018] If yes, the dialogue manager is started;

[0019] According to the starting of the dialogue manager, the inquiry voice is directionally projected to the driver's position through the vehicle loudspeaker array for voice interaction;

[0020] Through the voice interaction, it is judged whether an instruction of reducing the multimedia volume is received;

[0021] If yes, the instruction of reducing the multimedia volume is executed to control the multimedia volume to be reduced.

[0022] Further, it comprises:

[0023] The passenger detection system detects the attention of the driver;

[0024] The passenger detection system detecting the attention of the driver comprises playing inquiry voice in the front row;

[0025] According to the playing of inquiry voice in the front row, the steering wheel is synchronously controlled to vibrate.

[0026] Further, it further comprises:

[0027] The dialogue manager further comprises constructing a domain-specific language model;

[0028] The construction of the domain-specific language model comprises the following steps:

[0029] Collecting public domain related corpus, vehicle system historical voice interaction log and simulated generated domain scene dialogue data to form an initial corpus;

[0030] The initial corpus comprises text data.

[0031] According to the text data in the initial corpus, synonym replacement and word order transformation processing are performed to generate an enhanced corpus;

[0032] A general pre-trained language model is selected, prompt engineering optimization is performed based on a domain scene instruction template, and parameter efficient fine-tuning technology is used to train the modules related to the key semantics of the domain in the model;

[0033] A mixed loss function including language modeling loss, intent classification loss and slot filling loss is constructed, and the adapted model is trained;

[0034] The trained model is applied to the dialog manager.

[0035] Further, comprising:

[0036] The passenger detection system comprises an in-vehicle infrared camera and a 3D structured light / ToF camera;

[0037] The in-vehicle infrared camera acquires visual information of passenger facial expressions, eye states and head postures;

[0038] The 3D structured light / ToF camera acquires facial depth information to determine head position and action amplitude.

[0039] Further, further comprising:

[0040] The passenger detection system comprises a microphone array;

[0041] The microphone array acquires in-vehicle sound information.

[0042] Further, further comprising:

[0043] The passenger detection system comprises a contact sensor and a non-contact sensor;

[0044] The contact sensor acquires passenger sitting posture changes, body movement frequency and heart rate;

[0045] The non-contact sensor acquires breathing frequency, body movement micro-displacement and body temperature changes.

[0046] Further, further comprising:

[0047] The multi-modal fusion algorithm comprises: according to the recognized visual, audio and physiological signals, feature vectors are extracted;

[0048] The feature vectors are spliced and jointly modeled;

[0049] The prediction results of each mode are fused by a weighted average algorithm, and an in-vehicle passenger state determination signal is output.

[0050] According to two aspects of the present application, a volume adjustment system of a vehicle audio is provided, the volume adjustment system of the vehicle audio comprising:

[0051] a vehicle operation monitoring module, a passenger detection starting module, a multi-modal data acquisition module, a state determination module, a media focus detection module, a voice interaction module and a volume adjustment execution module

[0052] the vehicle operation monitoring module is configured to acquire vehicle speed information;

[0053] the passenger detection starting module is configured to start a passenger detection system when the vehicle is in a moving state;

[0054] the multi-modal data acquisition module is configured to acquire visual, audio and physiological signal data of passengers in the vehicle through the passenger detection system;

[0055] the state determination module is configured to generate a state determination signal of the passengers in the vehicle according to the multi-modal data;

[0056] the media focus detection module is configured to acquire a preset threshold of a playing volume based on the vehicle audio being in an output state, and determine whether the output volume of the vehicle audio is greater than the preset threshold of the playing volume;

[0057] determine whether the determination information of the state of the passengers in the vehicle confirms that the passengers are in a sleep state, and if the passengers are in the sleep state, determine whether the infotainment system has an active media focus;

[0058] the voice interaction module is configured to start a dialogue manager and acquire a response of reducing the media volume through voice interaction;

[0059] the volume adjustment execution module is configured to perform a media volume adjustment operation according to the acquired response of reducing the media volume.

[0060] According to three aspects of the present application, an electronic device is provided, comprising a processor, a communication interface, a memory and a communication bus, wherein the processor, the communication interface and the memory complete mutual communication through the communication bus;

[0061] The memory stores a computer program, and when the computer program is executed by the processor, the processor executes the steps of the method for predicting operation of a vehicle facility based on habits of a vehicle owner.

[0062] According to four aspects of the present application, a computer readable storage medium is provided, which stores a computer program executable by an electronic device, and when the computer program runs on the electronic device, the electronic device executes the steps of the volume adjustment method of the vehicle audio.

[0063] Through the above scheme, the following beneficial technical effects are obtained:

[0064] The application realizes accurate judgment on complex scenes through the passenger detection system, which needs to integrate vehicle operation information, passenger physical data of OMS / sensors, multi-dimensional heterogeneous data such as media focus state and volume value of the vehicle machine, and the like. Through multi-modal data fusion, the passenger sleep state is monitored without active operation of the passenger, and the system automatically triggers the volume adjustment process.

[0065] The application uses voice interaction to replace manual operation through the voice interaction system, so that the driver can complete feedback without distraction, which meets the driving safety specification; meanwhile, the system actively asks before triggering adjustment, avoids excessive intervention, and balances the control right of the passenger and the driver.

[0066] The application triggers the inquiry process only when all conditions of “vehicle speed > 0”, “passenger sleep”, “music playing” and “volume > threshold value” are met simultaneously, which involves complex logic gating (AND operation) and state machine design, and needs to ensure the real-time performance of multi-thread data processing to avoid delay in adjustment due to processing delay. BRIEF DESCRIPTION OF DRAWINGS

[0067] In order to more clearly illustrate the technical solutions in the specific embodiments or the prior art of the present application, the drawings needed to be used in the specific embodiments or the prior art description will be briefly introduced. Obviously, the drawings in the following description are some embodiments of the present application, and other drawings can be obtained by those skilled in the art without creative labor.

[0068] Figure 1 is a flowchart of a volume adjustment method of a vehicle-mounted audio provided by one or more embodiments of the present application.

[0069] Figure 2 is a structural diagram of a volume adjustment system of a vehicle-mounted audio provided by one or more embodiments of the present application.

[0070] Figure 3 is a schematic diagram of a flowchart of a volume adjustment method of a vehicle-mounted audio of one specific embodiment of the present application.

[0071] Figure 4 is an electronic device structural block diagram of a volume adjustment method of a vehicle-mounted audio provided by one or more embodiments of the present application. DETAILED DESCRIPTION

[0072] The technical solutions of the present application will be described in detail below with reference to the drawings. Obviously, the described embodiments are some embodiments of the present application, not all embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative labor are within the protection scope of the present application.

[0073] Figure 1 is a flowchart of a volume adjustment method of a vehicle audio provided by one or more embodiments of the present application.

[0074] As shown in a volume adjustment method of a vehicle audio includes: Figure 1

[0075] Step S1, obtaining vehicle running information;

[0076] According to the vehicle running information, it is judged whether the vehicle is in motion;

[0077] If yes, start the passenger detection system;

[0078] Step S2, by starting the passenger detection system, obtaining the signals of the visual receiving state, audio output state and physiological state of the passengers in the vehicle;

[0079] According to the signals of the visual receiving state, audio output state and physiological state of the passengers in the vehicle, a multi-modal fusion algorithm is performed to generate passenger state determination information in the vehicle;

[0080] Step S3, based on the output state of the vehicle audio, obtaining a preset threshold of the playing volume;

[0081] It is judged whether the output volume of the vehicle audio is greater than the preset threshold of the playing volume;

[0082] If yes, it is judged whether the determination information of the passenger state confirms that the passenger is in a sleep state;

[0083] If yes, in a sleep state, it is judged whether the infotainment system has an active media focus;

[0084] If yes, start the dialogue manager;

[0085] Step S4, according to starting the dialogue manager, through the vehicle speaker array, the inquiry voice is directionally projected to the driver position for voice interaction;

[0086] Step S5, through voice interaction, it is judged whether a command to reduce the multimedia volume is received;

[0087] If yes, execute the command to reduce the multimedia volume to control the multimedia volume to be reduced.

[0088] Further, it includes:

[0089] The passenger detection system detects the driver's attention;

[0090] The passenger detection system detecting the driver's attention includes playing inquiry voice in the front row;

[0091] ​Wherein, the front row plays a query voice, and the steering wheel vibration is controlled synchronously.

[0092] Further, it also includes:

[0093] The dialogue manager further includes: constructing a domain-specific language model;

[0094] The construction of the domain-specific language model includes the following steps:

[0095] Collecting public domain related corpus, vehicle system historical voice interaction log and simulated generated domain scene dialogue data to form an initial corpus;

[0096] The initial corpus includes text data.

[0097] According to the text data in the initial corpus, synonym replacement and word order transformation processing are performed to generate an enhanced corpus;

[0098] Select a general pre-trained language model, perform prompt engineering optimization based on the domain scene instruction template, and use parameter efficient fine-tuning technology to train the modules related to the domain key semantics in the model;

[0099] A mixed loss function containing language modeling loss, intent classification loss and slot filling loss is constructed to train the adapted model;

[0100] The trained model is applied to the dialogue manager.

[0101] Further, it includes:

[0102] The passenger detection system includes: an in-vehicle infrared camera and a 3D structured light / ToF camera;

[0103] The in-vehicle infrared camera obtains visual information of passenger facial expressions, eye states and head poses;

[0104] The 3D structured light / ToF camera obtains facial depth information to determine head position and action amplitude.

[0105] Further, it also includes:

[0106] The passenger detection system includes: a microphone array;

[0107] The microphone array obtains in-vehicle sound information.

[0108] Further, it also includes:

[0109] The passenger detection system includes: a contact sensor and a non-contact sensor;

[0110] The contact sensor obtains passenger sitting posture changes, body movement frequency and heart rate;

[0111] Non-contact sensors acquire respiratory rate, body micro-displacement and body temperature changes.

[0112] Furthermore, it also includes:

[0113] The multimodal fusion algorithm includes: extracting and generating feature vectors based on the identified visual, audio and physiological signals;

[0114] splicing the feature vectors and performing joint modeling;

[0115] The prediction results of each mode are fused through the weighted average algorithm to output the passenger status judgment signal.

[0116] Among them, the weighted average algorithm assigns different weights to each modal data, and the larger the weight, the greater the impact on the result.

[0117] Specifically, by integrating vehicle operation information, including wheel speed signals, lateral acceleration signals, steering wheel angle signals, brake pedal signals, and accelerator pedal signals obtained through the VDC (Vehicle Dynamics Control) system, it can be determined whether the car is in motion.

[0118] Among them, a more accurate vehicle speed is calculated by fusing the four-wheel speeds to determine whether the vehicle is completely stationary. In cases of parking, for example, the passenger detection system will not be activated to protect privacy.

[0119] When the lateral acceleration changes suddenly (such as sudden braking or sharp turning), it can be determined that the vehicle is in a dynamic driving state, avoiding misjudgment of parking and premature triggering of the passenger sleep monitoring function.

[0120] Combining wheel speed and lateral acceleration signals to determine the driver's intention, such as preparing to change lanes or emergency avoidance, VDC adjusts the braking / driving force of each wheel accordingly to help maintain vehicle stability;

[0121] When the driver makes a sharp turn, the passenger detection system temporarily suppresses volume adjustment requests to avoid disrupting driving operations.

[0122] When the brake pedal is detected, such as before parking, the OMS camera for passenger sleep monitoring can be turned off in advance to protect privacy.

[0123] Frequent high-throttle operations can be identified as aggressive driving, and VDC automatically enhances the ESC intervention force. At the same time, the passenger detection system can suspend volume adjustment requests to prioritize driving safety.

[0124] Through the combination of vehicle speed + lateral acceleration + steering wheel angle, it can accurately determine whether the vehicle is in a stable driving state, avoiding misjudgment of passengers falling asleep during sudden acceleration / braking.

[0125] When VDC detects that the vehicle enters the automatic parking mode (vehicle speed ≤ 5 km / h and steering wheel angle changes greatly), a signal can be sent to the passenger detection system to temporarily turn off the OMS camera to prevent the collection of passenger privacy data during parking.

[0126] In conjunction with the Internet of Vehicles (T-BOX), vehicle dynamic data such as "today's number of sudden accelerations" is pushed to the owner's APP, assisting users in analyzing driving habits and indirectly improving passenger comfort, such as reducing aggressive driving.

[0127] Through the above scheme, the technical problem of how to obtain vehicle running signals and further combine with the passenger detection system based on vehicle running signals to avoid false triggering caused by single signals is solved. In the passenger detection scene, the vehicle speed and acceleration signals of VDC can be combined to more accurately judge the "vehicle running state", avoid false triggering caused by single signals, and improve system robustness and user experience through data fusion.

[0128] Among them, according to the aforementioned vehicle running signal to judge whether the vehicle is in a driving state, if yes, the passenger detection system is started, and after the system is started, the OMS camera, infrared emitter / receiver and other hardware components are initialized and calibrated. The camera automatically focuses and adjusts the white balance to ensure image clarity; the infrared sensor detects the temperature and light intensity inside the vehicle to switch the working mode (such as enabling the RGB camera during the day and switching to the infrared night vision mode at night).

[0129] The original image is denoised and distortion corrected, such as fisheye lens distortion compensation, to generate standardized input data.

[0130] Recognize the static elements in the vehicle through a deep learning model (such as YOLOv8):

[0131] Seat position, seat belt buckle, armrest, etc., establish a passenger space distribution coordinate system;

[0132] Mark the driver's seat and the rear passenger area to distinguish the monitoring priority (such as higher frequency for driver fatigue monitoring).

[0133] Use the MTCNN (Multi-Task Convolutional Neural Network) algorithm to detect the face area in the picture and output the bounding box coordinates.

[0134] Analyze the facial feature points (such as eye corner lines and jaw lines) through a pre-trained CNN model to output the age interval (such as 20-30 years old) and gender probability (such as 92% male);

[0135] Model trained based on FER-2013 dataset, recognize 6 basic emotions (happy, sad, angry, surprised, etc.), combined with micro-expression analysis (such as mouth corner arc, pupil zoom) to improve accuracy.

[0136] Use OpenPose algorithm to identify 18 key points such as passenger's shoulders, elbows, wrists, etc., to determine posture (such as leaning forward, leaning back, lying on the side);

[0137] Recognize specific gestures (such as waving, thumbs up, OK gesture) through template matching or time series model (such as LSTM), trigger corresponding car function (such as waving to adjust volume).

[0138] Generate visual feature vector according to the above data;

[0139] Among them, the detection of physiological signals includes: detection of respiratory rate; fatigue / sleep state judgment detection;

[0140] Respiratory rate detection, through the camera to capture the chest micro-fluctuation (cooperate with infrared light enhancement to enhance the accuracy in low light), or use millimeter wave radar to detect the displacement of the chest caused by breathing, calculate the respiratory rate (such as 12-20 times / minute is the normal range).

[0141] Fatigue / sleep state judgment detection, including:

[0142] Use PERCLOS algorithm to detect eyelid closure time ratio, when PERCLOS value exceeds 0.8 (80% of the time with closed eyes), judge as fatigue or sleep;

[0143] Head posture fusion: combined with head tilt angle (such as lowering head > 45° and lasting 10 seconds) and eye movement trajectory (such as fixed gaze point), confirm the passenger into sleep state.

[0144] Through the micro piezoresistive sensor embedded in the surface of the seat, detect the pressure distribution of the passenger's hips and back.

[0145] Through the capacitive displacement sensor installed in the seat slide rail or backrest shaft, measure the seat position, such as forward and backward sliding distance, backrest inclination angle.

[0146] Obtain the passenger's sitting posture changes through the above sensors, and combine the data to determine the passenger's sleep state;

[0147] Through the piezoelectric acceleration sensor embedded in the seat bottom or seat belt buckle, detect the vibration signal generated by the passenger's body movement.

[0148] Through the conductive fiber fabric woven in the seat fabric, according to the change of fabric resistance caused by human movement, perceive actions such as waving hands, turning over, and obtain the passenger's body movement frequency;

[0149] Sleep state determination: body movement frequency is lower than 5 times per minute and lasts for 10 minutes, determined as "falling asleep";

[0150] Emergency identification: sudden and violent body movement (such as convulsions) triggers emergency recording of in-vehicle camera and uploads to the cloud.

[0151] Through the PPG photoplethysmography sensor integrated in the seat armrest, safety belt buckle or steering wheel, the subcutaneous blood vessel volume change is detected by green light LED + photodiode.

[0152] Through the electrode type bioelectricity sensor embedded in the seat surface, the ECG electrocardiogram signal is collected by the contact of human skin and electrode.

[0153] Through the infrared thermopile sensor installed on the roof or rearview mirror and the in-vehicle infrared camera, the forehead temperature is measured.

[0154] Using the Seebeck effect, the thermocouple array senses temperature difference, suitable for rapid temperature screening.

[0155] The 3D structured light / ToF camera enhances the accuracy of head movement recognition by combining with the two-dimensional image of the infrared camera, and distinguishes between slight head shaking and large movements.

[0156] According to the above data, a physiological feature vector is generated;

[0157] Further, the collection of audio signals includes:

[0158] The microphone array collects sound signals and locates the sound source position through beamforming technology (such as distinguishing the driver's voice from the rear passenger's voice);

[0159] Combined with voice keyword detection, the vehicle interaction system is awakened, and the OMS camera turns to the sound source direction to capture the speaker's facial expression.

[0160] When the vehicle owner enables remote monitoring through the mobile phone APP, the OMS will transmit the encrypted real-time picture to the cloud (with user authorization), and the picture only shows the passenger's outline and action label (such as "sleeping" "waving hands"), without clear facial information, ensuring privacy and safety.

[0161] According to the above data, an audio feature vector is generated;

[0162] According to the above visual feature vector, physiological feature vector and audio feature vector, a comprehensive feature vector is generated;

[0163] Through the weighted strategy, the multi-modal prediction results are integrated to output the final determination signal.

[0164] Further, through the weighted strategy, the multi-modal prediction results are integrated to determine that the passenger falling asleep needs to be met;

[0165] Contact sensor: body motion frequency <5 times / minute and stable sitting posture (no change in pressure sensor);

[0166] Non-contact sensor: breathing frequency <12 times / minute (millimeter wave radar) and PERCLOS>0.8 (OMS camera);

[0167] Fusion logic: consensus of at least 3 sensors (2 contact + 1 non-contact), avoiding single sensor misjudgment.

[0168] When the system identifies that the vehicle is in a running state, and it is determined that the passenger in the vehicle is in a sleep state, it is identified whether the current car system has an active media focus, and whether there is audio playing is judged according to the active focus. If yes, the volume of the playing audio is continuously identified. According to the preset media volume threshold, for example, 15, when the media volume is greater than 15, the condition is established. A start dialog manager instruction is generated to form an interactive scene.

[0169] Among them, the dialog manager includes building a domain-specific language model;

[0170] Building a domain-specific language model includes the following steps:

[0171] Collecting public domain related corpus, vehicle system historical voice interaction log and simulated generated domain scene dialogue data to form an initial corpus;

[0172] Among them, the initial corpus includes text data;

[0173] According to the text data in the initial corpus, synonym replacement and word order transformation processing are performed to generate an enhanced corpus;

[0174] Select a general pre-trained language model, perform prompt engineering optimization based on the domain scene instruction template, and use parameter efficient fine-tuning technology to train the modules related to the domain key semantics in the model;

[0175] Build a hybrid loss function including language modeling loss, intent classification loss and slot filling loss, and train the adapted model;

[0176] Apply the trained model to the dialog manager.

[0177] Further, according to the application of the trained model, the dialog manager and the driver generate an interactive scene. The driver agrees according to the microphone feedback result, and if it is agreed, it is judged to be passed. After all the above judgments, the system will reduce the media volume to ensure that the passenger has a good sleep experience.

[0178] Figure 2A structural block diagram of a volume adjustment system of a vehicle audio provided by one or more embodiments of the present application.

[0179] As Figure 2 A volume adjustment system of a vehicle audio includes:

[0180] A vehicle operation monitoring module, a passenger detection starting module, a multi-modal data acquisition module, a state determination module, a media focus detection module, a voice interaction module, and a volume adjustment execution module.

[0181] The vehicle operation monitoring module is configured to acquire vehicle speed information.

[0182] The passenger detection starting module is configured to start a passenger detection system when the vehicle is in a moving state.

[0183] The multi-modal data acquisition module is configured to acquire signals of visual receiving states, audio output states, and physiological states of passengers in the vehicle through the passenger detection system.

[0184] The state determination module is configured to generate a state determination signal of the passengers in the vehicle according to the multi-modal data.

[0185] The media focus detection module is configured to acquire a preset threshold of a playing volume based on the vehicle audio being in an output state, and determine whether an output volume of the vehicle audio is greater than the preset threshold of the playing volume.

[0186] The determination information of the state of the passengers in the vehicle is used to determine whether the passengers are in a sleep state, and if the passengers are in the sleep state, it is determined whether the infotainment system has an active media focus.

[0187] The voice interaction module is configured to start a dialogue manager and acquire a response of reducing the media volume through voice interaction.

[0188] The volume adjustment execution module is configured to perform a media volume adjustment operation according to the acquired response of reducing the media volume.

[0189] It is worth noting that although the system only discloses the vehicle operation monitoring module, the passenger detection starting module, the multi-modal data acquisition module, the state determination module, the media focus detection module, the voice interaction module, and the volume adjustment execution module, it does not mean that the device is limited to the above basic function modules. On the contrary, the meaning expressed by the present application is that on the basis of the above basic function modules, a person skilled in the art can add one or more function modules to form infinite embodiments or technical solutions in combination with the prior art. That is, the system is open rather than closed, and it cannot be considered that the protection scope of the present application is limited to the above disclosed basic function modules because the present embodiment only discloses individual basic function modules.

[0190] Figure 3 A schematic diagram of a volume adjustment method of a car audio according to one embodiment of the present application.

[0191] In one embodiment, a volume adjustment method of a car audio is disclosed, and the method runs as follows:

[0192] The car is running: the vehicle speed and other signals are obtained through the VDC (Vehicle Dynamics Control), and the car can be determined to be running according to the signals.

[0193] The passenger is in a sleep state: the motion and vital sign information of the passenger are collected through OMS, microphone, sensor and other devices, and the passenger is determined to be in a sleep state through an algorithm.

[0194] Music is playing: whether music is playing can be determined by whether there is an active media focus in the car audio system.

[0195] The media volume exceeds the threshold value: a media volume threshold value can be customized according to user experience, for example, 15, and when the media volume is greater than 15, the condition is determined to be established.

[0196] Whether the volume needs to be reduced: when all the above conditions are determined to be established, the system will actively ask the driver whether the media volume needs to be reduced, and the driver can feed back the result to the system through the car audio microphone, and if the driver agrees, the condition is determined to be passed.

[0197] After all the above conditions are determined to be established, the system will perform a media volume reduction operation to ensure that the passenger has a good sleep experience.

[0198] Figure 4 An electronic device structure block diagram of a volume adjustment method of a car audio according to one or more embodiments of the present application.

[0199] As shown in Figure 4 The present application provides an electronic device, which comprises a processor, a communication interface, a memory and a communication bus, wherein the processor, the communication interface and the memory complete mutual communication through the communication bus.

[0200] The memory stores a computer program, and when the computer program is executed by the processor, the processor executes the steps of the volume adjustment method of the car audio.

[0201] The present application also provides a computer readable storage medium which stores a computer program executable by an electronic device, and when the computer program runs on the electronic device, the electronic device executes the steps of the volume adjustment method of the car audio.

[0202] For the method embodiments, the description is made in a series of action combinations for simplicity and clarity, but those skilled in the art should know that the present application is not limited to the action sequence described, because according to the present application, certain steps can be performed in other sequences or at the same time. In addition, those skilled in the art should know that the embodiments described in the specification are all preferred embodiments, and the actions involved are not necessarily essential to the present application.

[0203] From the above description of the embodiments, those skilled in the art can clearly understand that the present application can be implemented by means of software and the necessary general hardware platform. Based on such understanding, the technical solutions of the present application can be embodied in the form of a software product, which can be stored in a storage medium, such as a ROM / RAM, a magnetic disk, an optical disk, etc., and includes a number of instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in the various embodiments or some parts of the embodiments of the present application.

[0204] Finally, it should be noted that: the above embodiments are only used to illustrate the technical solutions of the present application, and not to limit them; although the present application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that: it can still modify the technical solutions recorded in the foregoing embodiments, or make equivalent replacement for part or all of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the scope of the technical solutions of the embodiments of the present application.

Claims

1. A method for adjusting the volume of a car audio system, characterized in that: The following steps are involved: Obtain vehicle operation information; determining whether the vehicle is in motion according to the vehicle operation information; If,yes, then start the passenger detection system; By activating the passenger detection system, signals of the visual reception status, audio output status and physiological status of passengers in the vehicle are obtained; Based on the signals of the visual reception state, audio output state and physiological state of the passengers in the car, the multimodal fusion algorithm is used to generate the judgment information of the passenger status in the car; Based on the car audio being in output state, obtain the preset threshold value of the playback volume; Determine whether the output volume of the car audio is greater than the preset playback volume threshold; If , is greater than , then it is determined whether the determination information of the passenger status in the vehicle confirms that the passenger is in a sleeping state; If,it is in sleep state, then determine whether the vehicle system has an active media focus; If,yes, then start the dialogue manager; According to the startup dialogue manager, the inquiry voice is projected to the driver's position through the vehicle speaker array for voice interaction; Determine through voice interaction whether a command to lower the multimedia volume has been received; If received, the instruction to lower the multimedia volume is executed to control the multimedia volume to be lowered.

2. The method for adjusting the volume of a car audio system according to claim 1, wherein: include: The passenger detection system detects driver attention; The passenger detection system detects driver attention, including playing an inquiry voice in the front seat; Among them, the steering wheel vibration is controlled synchronously according to the inquiry voice played in the front row.

3. The method for adjusting the volume of a car audio system according to claim 1, wherein: include: The dialog manager further includes: building a domain-specific language model; The construction of the domain-specific language model includes the following steps: Collect relevant corpora in the public domain, historical voice interaction logs of the vehicle system, and simulated domain scenario dialogue data to form an initial corpus; Wherein, the initial corpus includes text data; Perform synonym replacement and word order transformation processing on the text data in the initial corpus to generate an enhanced corpus; Select a general pre-trained language model, perform prompt engineering optimization based on domain scenario instruction templates, and use efficient parameter fine-tuning technology to conduct targeted training on modules in the model that are related to key domain semantics; Construct a hybrid loss function that includes language modeling loss, intent classification loss, and slot filling loss to train the adapted model; Apply the trained model to the dialog manager.

4. The method for adjusting the volume of a car audio system according to claim 1, wherein: include: The passenger detection system includes: an in-vehicle infrared camera and a 3D structured light / ToF camera; The infrared camera in the vehicle acquires visual information of the passenger's facial expression, eye state and head posture; The 3D structured light / ToF camera obtains facial depth information and determines the head position and movement amplitude.

5. The method for adjusting the volume of a car audio system according to claim 1, wherein: include: The passenger detection system includes: a microphone array; The microphone array acquires sound information inside the vehicle.

6. The method for adjusting the volume of a car audio system according to claim 1, wherein: include: The passenger detection system includes: a contact sensor and a non-contact sensor; The contact sensor acquires the passenger's sitting posture changes, body movement frequency and heart rate; The non-contact sensor acquires respiratory frequency, body micro-displacement and body temperature changes.

7. The method for adjusting the volume of a car audio system according to claim 1, wherein: include: The multimodal fusion algorithm includes: extracting and generating feature vectors based on the identified visual, audio and physiological signals; splicing the feature vectors and performing joint modeling; The prediction results of each modality are fused through the weighted average algorithm to output the passenger status judgment signal.

8. A volume adjustment system for a car audio system, characterized in that: include: Vehicle operation monitoring module, passenger detection and activation module, multimodal data acquisition module, state determination module, media focus detection module, voice interaction module and volume adjustment execution module; Vehicle operation monitoring module, used to obtain vehicle speed information; A passenger detection start module is used to start the passenger detection system when it is determined that the vehicle is in motion; A multimodal data acquisition module is used to collect signals of the visual reception status, audio output status and physiological status of passengers in the vehicle through the passenger detection system; A state determination module is used to generate a passenger state determination signal in the vehicle based on the multimodal data; The media focus detection module is used to obtain a preset playback volume threshold based on the vehicle audio being in an output state; and to determine whether the output volume of the vehicle audio is greater than the preset playback volume threshold; Determining whether the information regarding the passenger status in the vehicle confirms that the passenger is in a sleeping state; If it is in sleep mode, determine whether the vehicle system has an active media focus; A voice interaction module is used to start the dialogue manager and obtain a response to lower the media volume through voice interaction; The volume adjustment execution module is used to execute the corresponding media volume adjustment operation according to the response of reducing the media volume.

9. An electronic device, characterized in that: include: A processor, a communication interface, a memory, and a communication bus, wherein the processor, the communication interface, and the memory communicate with each other via the communication bus; The memory stores a computer program, and when the processor executes the computer program, the processor executes the steps of the method for adjusting the volume of a vehicle audio system according to any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that It stores a computer program that can be executed by an electronic device. When the computer program is run on the electronic device, the electronic device executes the steps of the volume adjustment method of the vehicle audio according to any one of claims 1 to 7.

Citation Information

Patent Citations

  • Safe driving method and system based on sleep quality of driver and vehicle

    CN115042795A

  • Intelligent volume adjusting method and system for vehicle-mounted audio system

    CN117119350A

Cited By

  • Vehicle volume adjusting method and device, vehicle and storage medium

    CN121469453A