An electric vehicle in-vehicle active sound generation dynamic adjustment method and system
By integrating multimodal perception and intelligent analysis, the in-vehicle active sound system of electric vehicles can identify vehicle status and driver characteristics in real time and dynamically adjust sound source parameters, solving the problem that existing systems cannot adapt to complex environments and individual differences, and realizing personalized and safe sound prompts.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- JILIN UNIVERSITY
- Filing Date
- 2026-03-26
- Publication Date
- 2026-05-29
AI Technical Summary
Existing active sound systems in electric vehicles lack the ability to dynamically adapt to complex driving environments and individual driver differences. They cannot achieve coordinated adjustments based on multi-dimensional factors, resulting in sound prompts that do not meet actual needs and affecting driving safety and driver perception.
By integrating multimodal perception, intelligent analysis, and dynamic adjustment mechanisms, the system can identify vehicle status, environmental conditions, and driver characteristics in real time, and dynamically adjust the audio source and parameters, including information based on geographical location, driver mood, vehicle speed, and background noise, to achieve a unified balance between personalized and safe audio sources.
It achieves precise control and personalized adaptation of in-vehicle sound output, improves the clarity and comfort of sound in different driving situations, has the ability to respond instantly to sudden noise and complex situations, and enhances safety prompts and emotional support functions.
Smart Images

Figure CN121893862B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of automotive intelligent control and human-computer interaction technology, and in particular to a method and system for dynamic adjustment of active in-vehicle sound generation in electric vehicles, which falls under the technical scope of intelligent driving assistance and in-vehicle acoustic control systems. Background Technology
[0002] In recent years, with the rapid development of the new energy vehicle industry, electric vehicles have been widely accepted and rapidly popularized due to their advantages such as low carbon emissions, environmental friendliness, low operating costs, and convenient maintenance. The quietness of their power systems has significantly improved driving comfort, but it has also brought a series of new technological challenges. Compared to the engine noise of traditional fuel vehicles, electric vehicles are almost silent at low speeds. While this helps reduce noise pollution, it also weakens the sound alerts inside and outside the vehicle during driving, thus affecting driving safety and driver perception.
[0003] To compensate for the "sound deficiency" caused by silence, some electric vehicles are equipped with in-vehicle active sound systems to play simulated driving sounds or warning sounds in specific scenarios. However, existing in-vehicle active sound solutions mostly use preset fixed sound sources and static control parameters, lacking the ability to dynamically adapt to complex driving environments and individual driver differences. Especially under different road conditions, weather conditions, or changes in driver mood, existing systems often fail to adjust the sound in a timely and appropriate manner, resulting in sound strategies that do not meet the actual needs of the moment. This can increase driving interference and fail to effectively provide warnings.
[0004] Furthermore, most current voice control logics are based on single sensor data and lack the ability to fuse and process multi-source information, making it impossible to achieve coordinated adjustments based on multiple dimensions such as driver status, vehicle operating parameters, and the external environment. This makes it difficult for the system to judge situations such as high-risk driving scenarios (e.g., driving through school zones, rainy or snowy weather, low visibility environments) or special driving groups (e.g., elderly drivers), thus limiting the further development of voice prompts in terms of accuracy, safety, and emotional interaction.
[0005] At the same time, the use of in-vehicle acoustic systems also faces the challenge of striking a balance between safety alerts and auditory comfort. On the one hand, the sound should have sufficient penetration to cope with background noise and environmental interference; on the other hand, overly abrupt or frequent sound output may cause driver fatigue or annoyance.
[0006] Therefore, how to dynamically control the content, intensity, and style of voice output based on driving scenarios, driver emotions, and behavioral characteristics has become a key issue in the human-machine interaction design of electric vehicles. Against this backdrop, there is an urgent need for an active in-vehicle voice system solution for electric vehicles that can integrate multi-dimensional information input and possess dynamic adjustment capabilities. Summary of the Invention
[0007] To address the aforementioned technical problems, this invention provides a method and system for dynamic adjustment of active sound generation within an electric vehicle. This method breaks through the traditional static sound source control mode and possesses real-time recognition and intelligent response capabilities based on vehicle status, environmental conditions, and driver characteristics, thereby achieving a coordinated and unified sound generation system in terms of safety, personalized experience, and acoustic comfort.
[0008] Specifically, the technical solution provided by this invention is as follows:
[0009] A method for dynamically adjusting active sound generation inside an electric vehicle includes the following steps:
[0010] S1. Select to activate the in-vehicle active sound function when starting the vehicle;
[0011] S2. Obtain the current vehicle's latitude and longitude based on the geographic location information acquisition module;
[0012] S3. Real-time acquisition of driver's facial features based on in-vehicle cameras to identify driver's emotions;
[0013] S4. Select a basic audio source from the audio source database based on the current vehicle's latitude and longitude and the driver's mood;
[0014] S5. Based on the vehicle's real-time speed information, make initial adjustments to the sound parameters of the basic sound source;
[0015] S6. Based on the real-time background noise information inside the vehicle, the sound parameters of the basic sound source are adjusted again;
[0016] S7. Determine the sound parameters of the basic sound source and make an active sound production decision. If sound production is required, play the sound source.
[0017] Furthermore, in S3: a face detection algorithm is used to locate the facial region in the driver video frames captured by the camera, and a face feature extraction algorithm is used to obtain facial features. Then, the facial features are input into the emotion recognition neural network to obtain an emotion index that represents the intensity of the driver's emotions.
[0018] Furthermore, in S4: the sound source database contains several sound sources, each of which is attached with a corresponding regional tag and emotion tag, and the regional tag and emotion tag of each sound source are encoded as a tag semantic vector; the regional tag is used to indicate the regional style to which the sound source belongs, and the emotion tag is used to indicate the emotional tone expressed by the sound source.
[0019] The selection of the basic sound source includes the following steps:
[0020] Based on the obtained current vehicle latitude and longitude and driver emotion index, natural language prompts are constructed and input into a pre-trained large language model. The large language model is used to recommend suitable regional music styles and emotional tones based on the input prompts.
[0021] The regional musical styles and emotional tones output by the large language model are encoded into semantic vectors, and these vectors are matched with the semantic vectors of the corresponding labels of each sound source in the sound source database. The sound source with the highest similarity is selected as the candidate sound source.
[0022] If the similarity of a candidate audio source is greater than the set minimum threshold, the candidate audio source will be used as the base audio source; otherwise, it indicates that a suitable audio source has not been successfully matched and the default audio source will be used as the base audio source.
[0023] Furthermore, in S5: based on the acquired vehicle speed and acceleration, the frequency, amplitude, and duration of the basic sound source are controlled by combining translation and pitch shifting algorithms and sound pressure level increase / decrease algorithms; in high-speed driving conditions, the frequency of the basic sound source is adjusted to the low-frequency band, and in low-speed driving conditions, the frequency of the basic sound source is adjusted to the mid-frequency and high-frequency range; if the vehicle is in a constant-speed high-speed cruising state, the frequency and rhythm density of the basic sound source are reduced; in the case of rapid acceleration, rapid deceleration, and sudden turning, the sound pressure level of the basic sound source is automatically increased to ensure its perceptibility in sudden situations.
[0024] Furthermore, in S6: the sound pressure level of the background noise inside the vehicle is calculated. If it is higher than the set threshold, the sound pressure level of the basic sound source is increased; if it is lower than the set threshold, the sound pressure level of the basic sound source is decreased.
[0025] The instantaneous change rate of sound pressure level is calculated. If it exceeds the set threshold, it is determined to be a sudden noise event, and the active playback of the sound source is temporarily suspended to avoid interfering with the driver's auditory concentration.
[0026] The Fourier transform algorithm is used to obtain the main frequency distribution range of the background noise inside the vehicle, and the translation and pitch shifting algorithm is used to shift the frequency of the basic sound source as a whole to mask the background noise inside the vehicle.
[0027] Furthermore, in S7: a logical judgment is made on whether to actively produce sound. The judgment conditions include: whether there is a system mute command or user-defined suppression setting, and whether the last playback time interval has met the set minimum time interval; if the sound production conditions are met, the basic sound source output value after adjusting the sound parameters is turned on to the speaker to achieve active sound production.
[0028] Preferably, to address special scenarios, when readjusting the sound parameters of the basic sound source, it is also necessary to simultaneously adjust the sound parameters based on the driver's age, driving area, and weather conditions; including:
[0029] The obtained driver facial features are input into a pre-trained age recognition model to obtain the driver's age. If the driver is identified as an elderly driver, the sound pressure amplitude of the high-frequency part of the basic sound source is actively reduced to avoid sharp tones causing auditory discomfort to older drivers.
[0030] The acquired geographic location information is linked with an embedded map engine or a network map API. If the system detects that a vehicle has entered a school, hospital, residential area, or high-speed-limit sensitive area, it will actively control the sound pressure level of the basic sound source below the set threshold and increase the energy proportion of the set frequency band, thereby enhancing the warning effectiveness without interfering with the environment.
[0031] Based on geographic location information, the system obtains the weather conditions and visibility of the current driving area. In low visibility weather, it actively reduces the sound pressure level and low frequency energy ratio of the basic sound source, and generates differential frequency stereo prompts or expands the main frequency components of the basic sound source through the left and right channels.
[0032] An in-vehicle active sound dynamic adjustment system based on the above method includes a system control and communication module, a sound perception module, a driver state recognition module, an environment and driving state perception module, a sound source management and matching module, a sound source dynamic adjustment module, and a sound output and execution module.
[0033] The system control and communication module is used to activate the system and establish data communication connections with other modules when the vehicle starts to complete data acquisition and synchronization; the sound perception module is used to collect and output in-vehicle background noise information in real time through in-vehicle sound sensors; the driver state recognition module is used to collect driver facial images in real time through in-vehicle cameras and output driver emotion and age; the environment and driving state perception module is used to obtain GPS positioning, speed, acceleration and map scene information and output vehicle operating status.
[0034] The audio source management and matching module is used to match and determine the basic audio source in the local multi-dimensional tagged audio source database based on the multi-source data output by the sound perception module, driver status recognition module and environment and driving status perception module; the audio source dynamic adjustment module is used to dynamically adjust the frequency, sound pressure level, playback duration and rhythm of the basic audio source; the sound output and execution module is used to output the audio source processed by the audio source dynamic adjustment module through the vehicle audio processing unit and the speaker.
[0035] Furthermore, the system control and communication module is deployed in the vehicle main control unit, and establishes a high-speed data channel through CAN bus, UART or MIPI communication interface to ensure that the modules cooperate with each other; the sound perception module is equipped with a capacitive microphone or an array MEMS microphone and is deployed in the driver's operating area.
[0036] This invention constructs an active in-vehicle sound system for electric vehicles that integrates multimodal perception, intelligent analysis, and dynamic adjustment mechanisms, achieving precise control and personalized adaptation of in-vehicle sound output. Compared to existing solutions that rely on fixed sound sources and static playback logic, this invention can perceive the in-vehicle sound environment, driver facial features, external driving scenarios, and vehicle operating status in real time, establishing a complete state perception model. It then dynamically selects sound sources and adaptively adjusts audio parameters, ensuring that the sound output is clear, perceptible, and does not disturb comfort in different driving situations.
[0037] The technical advantages of this invention are not only reflected in improving the environmental adaptability and individual matching of sound playback, but also in its active adjustment and closed-loop control capabilities. The system can make feedforward judgments and respond instantly to complex situations such as sudden noise, special road areas, and driver stress. By adjusting key parameters such as frequency, sound pressure level, playback rhythm, and sound field direction, it can achieve more targeted safety prompts and emotional support functions. Attached Figure Description
[0038] The accompanying drawings are provided to further illustrate the invention and form part of the specification. They are used together with the embodiments of the invention to explain the invention and do not constitute a limitation thereof.
[0039] Figure 1 This is a schematic diagram of the technical framework provided in an embodiment of the present invention;
[0040] Figure 2 This is a schematic diagram of a method flow provided in an embodiment of the present invention. Detailed Implementation
[0041] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, other embodiments obtained by those skilled in the art without creative effort are all within the scope of protection of the present invention.
[0042] Example 1
[0043] This embodiment provides a method for dynamically adjusting the active sound output inside an electric vehicle. By integrating vehicle status, environmental conditions, and driver characteristic data, and with intelligent optimization of acoustic performance as the core objective, the method dynamically adjusts the sound source and parameters of the active sound output inside the vehicle to provide the driver with a personalized driving sound.
[0044] like Figure 1 and Figure 2 As shown, the method mainly includes the following steps:
[0045] S1. Select to activate the in-vehicle active sound function when starting the vehicle;
[0046] S2. Obtain the current vehicle's latitude and longitude based on the GPS positioning module;
[0047] S3. Real-time acquisition of driver's facial features based on in-vehicle cameras to identify driver's emotions;
[0048] S4. Select a basic audio source from the audio source database based on the current vehicle's latitude and longitude and the driver's mood;
[0049] S5. Based on the vehicle's real-time speed parameters, make initial adjustments to the sound parameters of the basic sound source;
[0050] S6. Based on the real-time background noise parameters inside the vehicle, readjust the sound parameters of the basic sound source and play it.
[0051] In some embodiments, to address special driving scenarios, when readjusting the sound parameters of the basic sound source, it is also necessary to make targeted parameter adjustments based on the driver's age, driving area attributes, and weather conditions.
[0052] Specifically, for step S1, after the vehicle is powered on and started, the electric vehicle's central control system will automatically enter the initialization phase, and the in-vehicle system interface will prompt the driver whether to activate the in-vehicle active sound function. This function aims to provide a personalized sound feedback experience through multimodal perception and real-time acoustic adjustment. Therefore, its activation can be chosen by the driver or preset to be enabled by default when the vehicle is started. In actual deployment, this function can be embedded into the vehicle control unit (ECU) or the in-vehicle infotainment system, or it can be deployed independently in an embedded subsystem with edge computing capabilities, such as a dedicated audio control module equipped with a high-performance AI chip, thereby independently completing multi-source data processing and sound strategy decision-making while ensuring response speed.
[0053] Once the active sound function is activated, the relevant hardware sensing modules will be initialized sequentially through the vehicle communication network (such as CAN bus or serial communication interface, including UART, I2C or SPI). These modules include, but are not limited to: vehicle camera module for capturing driver facial features, GPS positioning module for acquiring geographical location information, vehicle speed and acceleration sensors for recording vehicle dynamic parameters, and high-sensitivity sound sensor array (such as capacitive microphone or MEMS microphone array) for monitoring in-vehicle background noise.
[0054] After initialization, the system enters a listening and waiting state, at which point the in-vehicle active sound module is capable of responding to changes in the environment and driving conditions. When specific sound emission conditions are detected, the system immediately calls upon the sound source database and executes an intelligent matching and parameter adjustment mechanism to complete the active acoustic response.
[0055] For step S2, after the vehicle starts and completes initialization, the built-in GPS positioning module continuously acquires the vehicle's current latitude and longitude coordinates. This module uses a high-precision GNSS positioning chip and supports the collaborative operation of BeiDou, GPS, and GLONASS, ensuring stable positioning performance under various complex geographical and weather conditions. The sampling frequency is no less than 1Hz to meet the requirements of continuous and real-time location data during vehicle operation.
[0056] For step S3, to achieve intelligent and personalized adjustment of the in-vehicle active sound function, the onboard camera is invoked in real time after the vehicle starts to continuously capture and analyze the driver's facial features to obtain their current emotional state. The camera used is a high-resolution CMOS image sensor with 1920×1080 pixels and a frame rate of no less than 30fps, ensuring clear and stable facial images even under high-speed movement and complex lighting conditions. To improve recognition accuracy in low-light environments, the camera integrates infrared illumination and a dynamic exposure control module, and its installation position is selected above the center console or in the steering wheel pillar area to ensure that its field of view completely covers the driver's face.
[0057] The acquired image data is transmitted to an in-vehicle embedded computing platform, such as a Jetson Nano or RK3588 module, via MIPI or USB 3.0 interface for initial data preprocessing and feature extraction at the edge. The preprocessing includes Gaussian filtering to remove image noise, histogram equalization to enhance image contrast, and ROI cropping to focus on the facial area, thereby reducing background interference. The preprocessed image then enters the face detection and localization process, using a lightweight model (such as YOLOv5-Face) to identify and locate key facial regions, ensuring accurate extraction of facial contours and feature points under different lighting, occlusion, and angle conditions.
[0058] After facial localization, key features of facial features, including the corners of the eyes, eyebrows, tip of the nose, and corners of the mouth, are further extracted and input into the emotion recognition neural network. This model is typically trained and optimized based on lightweight convolutional network architectures such as MobileNetV3 or EfficientNet-B0, and incorporates publicly available emotion recognition datasets such as FER2013 and AffectNet to improve its generalization ability in complex real-world driving environments. The final model outputs an emotion index E, ranging from 0 to 1, to quantify the driver's current emotional fluctuation. A value closer to 1 indicates that the driver is in a more unstable state, such as tension, anger, or fatigue, while a value closer to 0 indicates a calmer and more relaxed state.
[0059] For step S4, after acquiring the vehicle's geographical location and recognizing the driver's emotion, these two types of core perception data are used as input variables to drive the audio source selection module to retrieve the most suitable basic audio source for the current driving situation from the local or cloud-based audio source database. The audio source database is designed with a multi-dimensional tagging system. Each audio source file is semantically annotated according to its emotional attributes and regional style, and is further tagged with regional tags (such as Northeast, Southwest, Jiangnan, etc., matching the driver's regional cultural preferences) and emotional tags (such as calm, mild tension, anger, etc., reflecting the emotional tone of the audio source). These tags are not limited to structured keywords but are further encoded into a unified semantic vector format using natural language processing technology, enabling a higher level of semantic understanding and fuzzy search capabilities during audio source matching. Each audio source file in the database is encoded in high-quality PCM or MP3 format to ensure clear and stable playback quality under various speaker devices and acoustic environments.
[0060] Unlike traditional rule-based matching methods, this embodiment employs a semantic reasoning and vectorized matching mechanism driven by a Large Language Model (LLM). Specifically, it first combines geographic location information (latitude and longitude) and the driver's current emotional state to automatically generate natural language prompts, which are then input into a pre-trained Large Language Model, such as ChatGLM, BERT, or its multi-label generation variants. These models are fine-tuned for multi-label reasoning tasks, enabling them to output a set of target semantic labels covering dimensions such as emotion and regional style within a known driving state context. This set of labels is then uniformly encoded into a state vector and matched against a pre-built audio vector library in the database, performing line-by-line similarity matching.
[0061] Vector similarity calculations employ metrics such as cosine similarity, Manhattan distance, or Mahalanobis distance to ensure effective matching in the high-dimensional semantic space. The sound source with the highest matching degree is initially selected as the current base sound source and sent to the subsequent dynamic adjustment module for sound parameters. To ensure robustness and responsiveness, this embodiment also sets a similarity threshold mechanism: if the highest similarity is still lower than the set threshold, it indicates that the current semantic space has failed to be successfully mapped to the existing set of sound source labels, and it will revert to the default base sound source to avoid driving interference caused by improper selection.
[0062] By introducing the semantic generation capabilities of a language model and a high-dimensional vector matching mechanism, this embodiment no longer relies on fixed rules or predefined templates, but instead possesses the ability to generalize reasoning to new states and scenarios. Even if a certain combination of geographical location and sentiment index has not been encountered before, labels can be generated through the language model and a reasonable match can be achieved. This open semantic-driven sound source matching mechanism greatly enhances the intelligence level of the in-vehicle audio system, enabling it not only to respond optimally to known states, but also to provide near-optimal solutions in new situations.
[0063] For step S5, to obtain real-time speed information, a perception path for the vehicle's motion state can be established in various ways, including using the wheel speed sensors in the ABS system to read the rotational speed signal per second, obtaining the speed data output by the vehicle's ECU through the OBD on-board diagnostic interface, or combining the ground displacement change rate estimated by the GPS module as an auxiliary compensation signal. Based on the acquired speed information, the system further collects the vehicle's three-axis acceleration values, including instantaneous speed, acceleration vector, whether the vehicle is in an acceleration or deceleration phase, and whether it is currently entering a curve or undergoing continuous steering, among other discrimination information.
[0064] After acquiring the speed information, it is input into the audio source parameter adjustment module. The base audio source, selected in the previous stage, now requires initial dynamic adjustment of its core parameters based on actual driving conditions to ensure suitability and non-interference within the spatiotemporal context. Specifically, at high speeds, a low-frequency, stable, and continuous sound pattern is preferred to correspond to the sense of speed, reduce psychological interference, and avoid spectral conflicts with common wind and road noise during high-speed driving. At low speeds, a rhythmic sound pattern in the mid-to-high frequency range is more suitable to improve perception clarity and warning effect.
[0065] Meanwhile, regarding the adjustment of rhythm density, the system determines whether the vehicle is in cruise control or in a frequent start-stop or gear-shifting phase based on its acceleration characteristics. If the vehicle is in a constant-speed, high-speed cruise, the frequency and rhythm density of audio playback will be appropriately reduced to mitigate interference from redundant information. In low-speed, dense traffic environments, the frequency and rhythm of warning sounds will be appropriately increased to ensure the driver's attention remains at an appropriate level. Furthermore, in transient situations such as rapid acceleration, rapid deceleration, and sudden steering, the sound pressure level of the audio source will be automatically increased to ensure its perceptibility in sudden situations, providing sufficient penetration and warning effect.
[0066] In step S6, after setting the initial parameters of the basic sound source, the sound source will be further dynamically adjusted based on the real-time acoustic conditions of the in-vehicle environment to ensure that the sound output can effectively penetrate background noise without causing auditory interference to the driver. To this end, a high-sensitivity sound acquisition module is deployed in the electric vehicle, typically consisting of a condenser microphone or MEMS microphone array positioned near the driver's head or in the center console area, with a sampling rate of at least 48kHz, thereby acquiring complete in-vehicle sound environment data in both the time and frequency domains. These sensors will continuously perform high-frequency sampling; the acquired analog signals are first converted into digital signals via analog-to-digital converter (ADC) and immediately sent to the embedded audio processing unit for real-time analysis.
[0067] During noise analysis, the sound pressure level (SPL) of the current background noise is first calculated. The calculation method is based on the root mean square (RMS) value of the sampled signal, combined with an A-weighted filter curve, to obtain an actual noise level that closely approximates human subjective perception, measured in dBA. When the detected SPL value is higher than 60 dBA, it indicates a relatively noisy in-vehicle environment, and an automatic boosting mechanism for the warning sound's SPL level will be triggered. Conversely, when the background noise is lower than 45 dBA, the SPL level of the played audio source will be appropriately reduced to avoid abrupt sounds in quiet environments, thus ensuring auditory comfort.
[0068] Beyond overall perception, a deeper spectral analysis of in-vehicle background noise is conducted. The Fast Fourier Transform (FFT) algorithm is used to unfold the spectral structure of the sound signal, identifying the main frequency distribution range of the noise. Based on this identification, frequency shifting processing is applied to the playing sound source. A pitch shift algorithm is used to shift the active audio spectrum up or down by 50-300Hz to mask the main noise frequency band, thereby improving the actual perceived clarity of the sound source to the human ear.
[0069] To further enhance the ability to handle sudden acoustic interference, this embodiment dynamically calculates the instantaneous rate of change of SPL (ΔSPL) every 100 milliseconds. When the ΔSPL value exceeds a preset threshold (e.g., greater than 8 dB / ms), it is determined that a sudden acoustic event has occurred, such as a door slamming shut or an object falling from inside the vehicle—a discontinuous noise source. Once such an event is identified, the playback of the active sound source will be immediately paused, or its sound pressure level will be significantly reduced to avoid sound overlap between the playing sound source and the sudden event, which could affect the driver's concentration and environmental perception.
[0070] After completing the secondary parameter adjustments based on the acoustic environment, the system will determine whether to execute the final control command for audio source playback based on several conditions. These criteria include, but are not limited to: whether the current time since the last audio source playback has exceeded the minimum time interval (e.g., 10 seconds) to prevent continuous and dense sound output; whether the system is in the user-set mute state; or whether it has received a command to prohibit sound from other in-vehicle systems.
[0071] Provided that all the above conditions are met, a playback control command will be generated, and the audio source signal with the parameters already adjusted will be input into the DSP digital signal processing module for final audio encoding, decoding and enhancement processing, and then output to the vehicle speaker system, thereby forming an in-vehicle active sound response with personalization, dynamism and context adaptability.
[0072] In certain special scenarios, in order to further enhance the intelligent adaptability and personalized performance of the in-vehicle active sound system, when readjusting the sound parameters of the basic sound source, it does not only rely on real-time variables such as vehicle speed and background noise, but also introduces multi-dimensional static or semi-static factors such as driver individual attributes, geographical area characteristics and weather conditions for comprehensive correction, thereby achieving a more delicate and context-appropriate sound optimization output.
[0073] First, considering the driver's age, the system uses facial feature information extracted through facial recognition algorithms in previous steps, combined with a pre-trained age recognition model, to determine the driver's age group. For example, if the driver is identified as a middle-aged or elderly person over 50 years old, the sound pressure level in the high-frequency range of the audio playback parameters will be actively reduced, especially in the frequency band above 3000Hz, typically attenuating by about 20%. This processing mechanism is based on the conclusion in psychoacoustic research that older adults have decreased sensitivity to high-frequency sounds, while also taking into account their lower tolerance for excessively high sound pressure changes. Through this adaptation process, sharp tones can avoid causing auditory discomfort or abruptness for older drivers, improving the overall comfort and acceptance of the user.
[0074] Secondly, for dynamic identification of the geographical area where the vehicle is located, real-time GPS positioning information is used in conjunction with an embedded map engine or online map API for analysis. If the vehicle is identified as entering an area with special attributes, such as a school, hospital, residential area, or high-speed-limit sensitive area, an audio playback strategy bound to the area label will be triggered. For example, when entering a school zone, a child safety warning sound mode will be forcibly activated, the active sound pressure level will be controlled below 65dB, and the energy proportion of the 500-1000Hz frequency band will be increased, thereby enhancing the warning effectiveness without interfering with the environment.
[0075] Secondly, regarding the external dynamic factor of weather conditions, real-time weather API data is accessed through the vehicle network interface to obtain meteorological conditions and visibility information for the current driving area. When low-visibility environments such as rain, snow, or fog are detected, structural adjustments are made to the playback parameters to reduce sound propagation attenuation in humid, particulate-rich air. Specifically, the low-frequency components of 200-400Hz in the active sound source are appropriately reduced to avoid muddy reverberation in enclosed or humid environments, while the energy density of the mid-to-high frequency bands is increased to enhance audio penetration and far-field perception. Furthermore, a 10-20Hz differential frequency stereo structure is designed between the left and right channels, giving the prompts a slight dynamic drift characteristic in binaural perception, thereby improving their recognizability and psychological attention in the context of rain, snow, and noise. In addition, under extreme weather conditions, the bandwidth of the alert tone can be automatically increased, for example, by ±500Hz of the main frequency components, to adapt to unstable transmission conditions in the external environment, so that the sound remains well perceptible in noisy, multi-reflective or highly absorbent environments.
[0076] To facilitate understanding of the implementation process of this embodiment, a specific case is given below.
[0077] Step 1: When a middle-aged male driver starts a regular family electric vehicle, the in-vehicle active sound system is activated. Simultaneously, a multi-dimensional tagged sound source database containing fundamental frequencies from 100-5000Hz is loaded, and the sampling frequencies of each sensor are initialized: CAN bus 1kHz, GPS 1Hz, microphone 48kHz, and 1080p resolution camera 60Hz.
[0078] Step Two: The embedded device in the vehicle is powered on. Its core components consist of an STM32 microcontroller and a Jeston Nano, which receive corresponding data through various serial ports. The STM32 receives data from the speed sensor, accelerometer, microphone (collecting in-vehicle background noise), and GPS via serial ports. The real-time vehicle speed is 45 km / h, acceleration is +0.8g, and in-vehicle background noise (A-weighted) is 55 dB. GPS positioning shows the vehicle is located in Nanguan District, Changchun City (125.34°E, 43.85°N). Simultaneously, the Jeston Nano receives video data from the camera via serial port. The STM32 transmits the received speed, acceleration, background noise, and GPS data to the Jeston Nano, where all subsequent algorithm processing is completed.
[0079] Step 3: The driver's facial features, including emotion, gender, and age, are extracted from the video footage captured by the camera using a lightweight CNN. The extracted features are then classified using a support vector machine (SVM). The final result is: the driver is a middle-aged male and the emotion index corresponds to a calm state.
[0080] Step 4: Based on the GPS latitude and longitude obtained in Step 2 (corresponding to the "Changchun" regional label) and the calm mood index obtained in Step 3, select soothing Northeast folk music sound sources with Northeast regional and calm mood labels from the multi-dimensional labeled sound source library, and use them as the base sound sources for subsequent parameter adjustments.
[0081] Step 5: Use the CAN bus to obtain the vehicle speed (45 km / h) and acceleration (+0.8g) data from Step 2, and adjust the sound parameters by combining the translation and pitch shifting algorithm and the sound pressure level increase / decrease algorithm.
[0082] Step Six: Using the 55dB in-vehicle background noise level obtained by the microphone, combined with the in-vehicle masking effect and ANC open-loop noise reduction algorithm, the sound parameters adjusted in Step Five are optimized a second time, and the sound amplitude is dynamically compensated to 70dB to ensure the audibility of driving sounds.
[0083] Step 7: Use GPS signal to connect to the network and obtain the road type (ordinary urban road) and environmental visibility (non-low visibility) of the current vehicle location. Combined with the middle-aged driver information obtained in Step 3, determine that the vehicle is not currently in a high-risk scenario (such as a school area) or a special scenario (such as rainy or snowy weather, or an elderly driver driving scenario). Therefore, there is no need to make any special adjustments to the sound parameters and no special scenario feedforward control is triggered.
[0084] Step 8: Execute the sound production decision algorithm on the final sound parameters determined in Step 7 to determine the current sound output. Then, control the sound production module to output a soothing in-car driving sound in the style of Northeast folk music at a frequency of 800±200Hz and an amplitude of 70dB. After the sound output is completed, return to the initial state and repeat the above process of "information acquisition - state analysis - sound source selection - sound parameter adjustment - sound output" with a period of 100ms, entering the next round of active sound control in a continuous closed loop.
[0085] Example 2
[0086] Based on the above method, this embodiment provides an active in-vehicle sound dynamic adjustment system for electric vehicles. The system is deployed inside the cockpit of an electric vehicle and its overall structure mainly includes a system control and communication module, a sound perception module, a driver state recognition module, an environment and driving state perception module, a sound source management and matching module, a sound source dynamic adjustment module, and a sound output and execution module. The modules communicate at high speed with each other through an on-board control bus (such as CAN, LIN, or Ethernet) to form a closed-loop control process of data acquisition, state modeling, strategy judgment, and sound playback.
[0087] The system control and communication module serves as the command center of the entire system, responsible for the initialization, signal synchronization, and communication scheduling management of each hardware module. Deployed in the vehicle's main control unit, this module supports data interaction with various devices such as microphones, cameras, GPS, and speed and acceleration sensors. It establishes high-speed data channels through communication interfaces such as CAN bus, UART, and MIPI to ensure stable operation and coordinated cooperation among all modules.
[0088] The sound perception module is used to perceive the background noise environment inside the vehicle in real time and use this information as the basis for formulating active sound generation strategies. This module is equipped with high-sensitivity condenser microphones or array MEMS microphones, deployed in key areas of the cockpit, such as above the dashboard or in the center console area, capable of acquiring high-fidelity audio signals at a sampling rate of ≥48kHz. The system performs A-weighted sound pressure level calculation, spectrum analysis, masking threshold assessment, and transient noise identification on the acquired signals to determine whether the current acoustic environment is suitable for audio source playback and the playback frequency and sound pressure level range that need to be adjusted. The output of this module is an acoustic environment feature vector, which is an important reference parameter for dynamic adjustment of the audio source.
[0089] The driver state recognition module uses high-definition cameras mounted on the steering wheel pillar or above the center console to capture real-time images of the driver's face, extracting key information such as emotional state, age group, and gender. The system uses lightweight neural network models such as MobileNet or MTCNN to process the images, combined with support vector machines or Softmax classifiers, to recognize five basic emotions, including calmness, tension, and fatigue. It also estimates the driver's age group (e.g., youth, middle-aged, elderly) and gender. The final driver features are then used in conjunction with other data dimensions for sound source matching and voice parameter adjustment.
[0090] The environmental and driving status perception module primarily utilizes multi-source information, including GPS, IMU (Inertial Measurement Unit), vehicle speed sensor, and network map interface, to acquire the vehicle's current geographical location, speed, acceleration, heading angle, and road type. By fusing latitude and longitude coordinates with map data, the system can determine whether the vehicle is in a special area (such as a school, residential area, or tunnel) and analyze the IMU output to determine if the vehicle is experiencing rapid acceleration, deceleration, cornering, or going uphill / downhill. Simultaneously, by combining weather API interfaces and auxiliary signals such as vehicle lighting and windshield wipers, it can identify high-risk scenarios such as rain, snow, nighttime, and low visibility.
[0091] The audio source management and matching module has a built-in localized, multi-dimensional, tagged audio source database, with each audio source accompanied by regional and emotional tags. Based on the current vehicle location and driver's emotional state, the system retrieves the most suitable basic audio source from the database using vector matching or a weighted similarity model. If no audio source meets the preset matching threshold, the system will automatically use a general-purpose soothing audio source as the default playback content to ensure the continuity and adaptability of the prompts.
[0092] After matching the base audio source, the dynamic audio source adjustment module performs multiple rounds of parameter optimization to adapt to the current in-vehicle sound environment and driving conditions. This module adjusts the audio source frequency using a pitch-shifting algorithm to avoid the main background noise frequency band, while simultaneously using sound pressure level control logic to keep the output volume within a suitable range, ensuring clear and perceptible sound without disturbing the driver. The system can also dynamically adjust the playback rhythm and duration based on driving behavior characteristics, increasing the rhythm density in curves or complex road conditions and decreasing the rhythm frequency during constant-speed cruising. Furthermore, this module can control the direction of the audio source's playback sound field, such as emitting sound from the front or left speakers to awaken the driver's attention, or generating a two-channel difference frequency effect in specific scenarios to improve acoustic wake-up effectiveness.
[0093] The sound output and execution module is responsible for the final audio signal output. The system sends the adjusted audio source parameters to the onboard DSP processing unit, which performs decoding and signal enhancement, and then achieves high-fidelity audio playback through the vehicle's multi-channel speaker array. The module supports zone control, dynamic volume adjustment, and sound field positioning, which can accurately convey prompts to the driver's target auditory area.
[0094] The above system can execute the in-vehicle active sound dynamic adjustment method described in Embodiment 1, and has the corresponding functional modules and beneficial effects of the method. For technical details not described in detail in this embodiment, please refer to the in-vehicle active sound dynamic adjustment method provided in Embodiment 1 of the present invention.
[0095] Through the above description of the embodiments, those skilled in the art can clearly understand that each embodiment can be implemented using software plus a general-purpose hardware platform, or of course, using hardware. Based on this understanding, the above technical solutions, in essence or the parts that contribute to the related technology, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute the methods described in the various embodiments or some parts of the embodiments.
[0096] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; under the concept of the present invention, the technical features of the above embodiments or different embodiments can also be combined, the steps can be implemented in any order, and there are many other variations of different aspects of the present invention as described above, which are not provided in detail for the sake of brevity; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of this application.
Claims
1. A method for dynamically adjusting active sound generation inside an electric vehicle, characterized in that, Including the following steps: S1. Select to activate the in-vehicle active sound function when starting the vehicle; S2. Obtain the current vehicle's latitude and longitude based on the geographic location information acquisition module; S3. Real-time acquisition of driver's facial features based on in-vehicle cameras to identify driver's emotions; S4. Select a basic audio source from the audio source database based on the current vehicle's latitude and longitude and the driver's mood; S5. Based on the vehicle's real-time speed information, make initial adjustments to the sound parameters of the basic sound source; S6. Based on the real-time background noise information inside the vehicle, the sound parameters of the basic sound source are adjusted again; When readjusting the sound parameters of the base sound source, it is also necessary to consider the driver's age, driving area, and weather conditions, and adjust the sound parameters accordingly; including: The obtained driver facial features are input into a pre-trained age recognition model to obtain the driver's age. If the driver is identified as an elderly driver, the sound pressure amplitude of the high-frequency part of the basic sound source is actively reduced to avoid sharp tones causing auditory discomfort to older drivers. The acquired geographic location information is linked with an embedded map engine or a network map API. If the system detects that a vehicle has entered a school, hospital, residential area, or high-speed-limit sensitive area, it will actively control the sound pressure level of the basic sound source below the set threshold and increase the energy proportion of the set frequency band, thereby enhancing the warning effectiveness without interfering with the environment. Based on geographic location information, the weather conditions and visibility of the current driving area are obtained. In low visibility weather, the sound pressure level and energy proportion of the basic sound source are actively reduced, and differential frequency stereo prompts or the main frequency components of the basic sound source are generated through the left and right channels. S7. Determine the sound parameters of the basic sound source and make an active sound production decision. If sound production is required, play the sound source.
2. The method for dynamic adjustment of active sound generation inside an electric vehicle as described in claim 1, characterized in that, In S3: The face detection algorithm is used to locate the facial region in the driver's video frames captured by the camera, and the face feature extraction algorithm is used to obtain the facial features. Then, the facial features are input into the emotion recognition neural network to obtain an emotion index that represents the intensity of the driver's emotions.
3. The method for dynamic adjustment of active sound generation inside an electric vehicle as described in claim 2, characterized in that, In S4: The audio source database contains several audio sources, each with a corresponding regional tag and emotion tag. The regional tag and emotion tag of each audio source are encoded as a tag semantic vector. The regional tag is used to indicate the regional style of the audio source, while the emotion tag is used to indicate the emotional tone expressed by the audio source. The selection of the basic sound source includes the following steps: Based on the obtained current vehicle latitude and longitude and driver emotion index, natural language prompts are constructed and input into a pre-trained large language model. The large language model is used to recommend suitable regional music styles and emotional tones based on the input prompts. The regional musical styles and emotional tones output by the large language model are encoded into semantic vectors, and these vectors are matched with the semantic vectors of the corresponding labels of each sound source in the sound source database. The sound source with the highest similarity is selected as the candidate sound source. If the similarity of a candidate audio source is greater than the set minimum threshold, the candidate audio source will be used as the base audio source; otherwise, it indicates that a suitable audio source has not been successfully matched and the default audio source will be used as the base audio source.
4. The method for dynamic adjustment of active sound generation inside an electric vehicle as described in claim 1, characterized in that, In S5: Based on the acquired vehicle speed and acceleration, the frequency, amplitude, and duration of the basic sound source are controlled by combining translation and pitch shifting algorithms and sound pressure level increase / decrease algorithms. When driving at high speed, the frequency of the basic sound source is adjusted to the low frequency band, and when driving at low speed, the frequency of the basic sound source is adjusted to the mid and high frequency range. If the vehicle is in a constant speed high-speed cruising state, the frequency and rhythm density of the basic sound source are reduced. Under conditions of rapid acceleration, rapid deceleration, and sudden turning, the sound pressure level of the basic sound source is automatically increased to ensure its perceptibility in emergency situations.
5. The method for dynamic adjustment of active sound generation inside an electric vehicle as described in claim 1, characterized in that, In S6: Calculate the sound pressure level of the background noise inside the vehicle. If it is higher than the set threshold, increase the sound pressure level of the base sound source. If it is lower than the set threshold, decrease the sound pressure level of the base sound source. The instantaneous change rate of sound pressure level is calculated. If it exceeds the set threshold, it is determined to be a sudden noise event, and the active playback of the sound source is temporarily suspended to avoid interfering with the driver's auditory concentration. The Fourier transform algorithm is used to obtain the main frequency distribution range of the background noise inside the vehicle, and the translation and pitch shifting algorithm is used to shift the frequency of the basic sound source as a whole to mask the background noise inside the vehicle.
6. The method for dynamic adjustment of active sound generation inside an electric vehicle as described in claim 1, characterized in that, In S7: The system makes a logical judgment on whether to actively produce sound. The judgment conditions include: whether there is a system mute command or user-defined suppression setting, and whether the last playback time interval has met the set minimum time interval. If the sound conditions are met, the system adjusts the basic sound source output value of the sound parameters and opens the speaker to achieve active sound production.
7. A vehicle-mounted active sound dynamic adjustment system based on the method of any one of claims 1 to 6, characterized in that, It includes a system control and communication module, a sound perception module, a driver status recognition module, an environment and driving status perception module, a sound source management and matching module, a sound source dynamic adjustment module, and a sound output and execution module; The system control and communication module is used to activate the system when the vehicle starts and establish data communication connections with other modules to complete data acquisition and synchronization; The sound perception module is used to collect and output in-vehicle background noise information in real time through in-vehicle sound sensors; The driver state recognition module is used to collect driver facial images in real time through the in-vehicle camera and output the driver's emotions and age; the environment and driving state perception module is used to acquire GPS positioning, speed, acceleration and map scene information and output the vehicle's operating status. The audio source management and matching module is used to match and determine the basic audio source in the local multi-dimensional tagged audio source database based on the multi-source data output by the sound perception module, driver status recognition module and environment and driving status perception module; the audio source dynamic adjustment module is used to dynamically adjust the frequency, sound pressure level, playback duration and rhythm of the basic audio source; the sound output and execution module is used to output the audio source processed by the audio source dynamic adjustment module through the vehicle audio processing unit and the speaker.
8. The in-vehicle active sound dynamic adjustment system as described in claim 7, characterized in that, The system control and communication module is deployed in the vehicle main control unit and establishes a high-speed data channel through CAN bus, UART or MIPI communication interface to ensure that the modules work together; the sound perception module is equipped with a capacitive microphone or an array MEMS microphone and is deployed in the driver's operating area.
Citation Information
Patent Citations
System and method for external sound synthesis of a vehicle
CN107018467A
Route search device, control method, program and storage medium
JP2016053504A