A system, method, and electronic device for enhancing a sense of presence

CN122511274APending Publication Date: 2026-08-04CHINA FAW CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
CHINA FAW CO LTD
Filing Date
2026-04-03
Publication Date
2026-08-04

AI Technical Summary

Technical Problem

[0007]为弥补上述现有技术中的不足,本发明提出一种增强临场音效的系统、方法及电子设备,解决了车载音响临场感不足、体验感不真实的问题;有效降低采集失真,适配车载封闭空间;通过模拟不同音乐厅场景的声学特性,提升基础音频播放的真实感与包裹感;实现回音消除和背景噪声抑制,提取纯净语音;通过语音增强及混响参数调节,提高对话的空间层次感;通过多声道扬声器阵列输出,实现音乐厅音效与语音混响音效的融合播放,增强整体沉浸式体验

Benefits of technology

[0037] The present invention application proposes a system, method and electronic device for enhancing the immersive sound effect. By adding a sound effect adjustment module, it can accurately simulate the acoustic characteristics of a concert hall, break the limitation of existing car audio systems that can only perform basic sound effect adjustments, significantly improve the immersiveness and presence of basic audio playback, and allow users to obtain a similar professional music listening experience in the closed space of a car.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122511274A_ABST
    Figure CN122511274A_ABST
Patent Text Reader

Abstract

This invention discloses a system, method, and electronic device for enhancing immersive sound effects, belonging to the field of in-vehicle audio technology. The system includes core control, audio acquisition, sound effect adjustment, echo cancellation, speech processing, and audio output modules. The core control module incorporates an audio scene recognition and separation unit, which jointly determines the mixed audio based on sound source location, spectral characteristics, and energy temporal information, separating the base audio from the in-vehicle dialogue audio and routing them separately. The method achieves concert hall sound effect simulation and speech clarity enhancement and reverberation enhancement through mixed audio acquisition, scene recognition and separation, dual-path differential processing, and adaptive fusion output. This solution can improve the immersiveness of music playback and the spatial layering of dialogue in an in-vehicle environment, ensuring that the two audio paths do not interfere with each other, adapting to the enclosed acoustic environment of an in-vehicle setting, and demonstrating strong practicality.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of vehicle audio technology, specifically relating to a system, method, and electronic device for enhancing immersive sound effects. Background Technology

[0002] With the continuous development of the automotive industry and audio technology, car audio has become a standard feature of cars. Users' demand for the immersive sound experience of car audio is increasing, especially in the closed space of a car. Users want to obtain an immersive music listening experience similar to a professional concert hall through car audio, and also require a realistic voice reverberation experience in a concert hall environment when talking with passengers in the car, enhancing the spatial sense and clarity of speech.

[0003] In existing technologies, the sound effect adjustment functions of car audio systems are mostly limited to basic parameter adjustments such as volume, treble and bass, left and right balance, and front and rear balance. They lack parameter adjustment modules specifically designed for the acoustic environment of concert halls, making it impossible to accurately match and simulate parameters according to the acoustic characteristics of concert halls. This makes it difficult to reproduce the live sound effects of concert halls and fails to meet users' needs for a high-quality, immersive music experience.

[0004] Meanwhile, the in-vehicle space is a typical enclosed acoustic environment with a large number of hard reflective surfaces. When users are talking in the car, the voice signal is easily reflected multiple times between the reflective surfaces, forming an echo, which seriously affects the clarity of the conversation. Moreover, the existing echo cancellation technology of in-vehicle audio can only simply filter the fixed background noise in the environment. It cannot effectively separate the conversation voice from the mixed audio and eliminate the echo, nor can it perform targeted reverberation adjustment on the extracted pure voice, resulting in a poor auditory experience for in-vehicle conversations and a lack of spatial layering in the voice.

[0005] Furthermore, the in-vehicle environment contains both basic audio such as music and radio and user conversation audio. Existing technologies cannot effectively separate and differentiate these two types of audio signals. When playing basic audio and conducting in-vehicle conversations at the same time, the two types of audio signals are prone to mutual interference. Moreover, there is a lack of effective audio fusion strategies, making it impossible to achieve dynamic balance output of the two types of signals, which further reduces the overall listening experience of the in-vehicle audio system.

[0006] In summary, existing in-vehicle audio systems suffer from poor simulation of ambient sound effects, significant echo interference during in-vehicle conversations, insufficient spatial awareness of speech, and easy interference between multiple audio signals. They can no longer meet users' demands for high-quality and multifunctional in-vehicle audio systems. There is an urgent need for an in-vehicle audio system and its implementation method that can accurately simulate the acoustic characteristics of a concert hall, eliminate echoes and adjust reverberation during in-vehicle conversations, and achieve dynamic fusion and output of multiple audio signals. Summary of the Invention

[0007] To overcome the shortcomings of the existing technology, this invention proposes a system, method, and electronic device for enhancing immersive sound effects, solving the problems of insufficient immersion and unrealistic experience in in-vehicle audio systems; effectively reducing acquisition distortion and adapting to the enclosed space of an in-vehicle environment; enhancing the realism and immersiveness of basic audio playback by simulating the acoustic characteristics of different concert hall scenarios; achieving echo cancellation and background noise suppression to extract pure speech; improving the spatial hierarchy of dialogue through speech enhancement and reverberation parameter adjustment; and achieving the fusion of concert hall sound effects and speech reverberation effects through multi-channel speaker array output, enhancing the overall immersive experience.

[0008] To achieve the above objectives, the present invention adopts the following technical solution:

[0009] In a first aspect, the present invention provides a system for enhancing immersive sound effects, comprising: a core control module, an audio acquisition module, a sound effect adjustment module, an echo cancellation module, a voice processing module, and an audio output module;

[0010] The core control module is used to receive and process signals transmitted by each module, parse user operation commands, and control the coordinated operation of each module.

[0011] The audio acquisition module is used to acquire audio signals in the vehicle environment, convert the audio signals into digital audio signals, and then transmit them to the core control module.

[0012] The sound effect adjustment module is used to receive the basic audio signal routed by the core control module and adjust the sound effect of the basic audio signal.

[0013] The echo cancellation module is used to receive the in-vehicle conversation audio signal routed by the core control module, and to perform echo cancellation processing on the in-vehicle conversation audio signal to output a pure voice signal.

[0014] The voice processing module is used to receive the pure voice signal and perform voice enhancement processing and reverberation adjustment on the pure voice signal.

[0015] The audio output module is used to receive the basic audio signal processed by the sound effect adjustment module and the pure voice signal processed by the voice processing module, and output the two signals after fusion processing.

[0016] The core control module includes an audio scene recognition and separation unit. This unit is used to make a joint decision on the acquired mixed audio signal based on the sound source location information, spectral feature information, and energy time domain information, so as to separate the basic audio signal and the in-vehicle dialogue audio signal in the mixed audio signal and route them separately.

[0017] Optionally, the audio acquisition module includes a multi-channel microphone array.

[0018] Optionally, the sound effect adjustment module includes a reverb processing unit; the reverb processing unit is used to adjust the reverb parameters, reflected sound parameters, channel delay parameters, and bass attenuation parameters of the basic audio signal;

[0019] The reverberation processing unit employs at least one of the following: convolutional reverberation algorithm, algorithmic reverberation algorithm, or reverberation algorithm based on feedback delay network.

[0020] Optionally, the echo cancellation module includes a decorrelation processing unit and an adaptive filtering unit. The decorrelation processing unit is used to perform decorrelation processing on the in-vehicle dialogue audio signal, and the adaptive filtering unit is used to estimate the echo path impulse response and generate a simulated echo signal.

[0021] Optionally, the audio output module includes multiple vehicle speakers, and the audio output module is used to perform synchronization alignment, dynamic range control, automatic ducking and linear mixing processing on the basic audio signal processed by the sound effect adjustment module and the pure voice signal processed by the voice processing module.

[0022] Secondly, the present invention provides a method for enhancing immersive sound effects, comprising:

[0023] Collect mixed audio signals from the vehicle environment and convert the mixed audio signals into digital audio signals;

[0024] The digital audio signal is subjected to scene recognition and separation processing. Based on the sound source location information, spectrum feature information and energy time domain information, the mixed audio signal is jointly determined to obtain the basic audio signal and the in-vehicle dialogue audio signal.

[0025] Adjust the sound effects of the basic audio signal;

[0026] Echo cancellation processing is performed on the in-vehicle conversation audio signal to obtain a pure speech signal;

[0027] The pure speech signal is subjected to speech enhancement processing, and the enhanced pure speech signal is subjected to reverberation adjustment; the adjusted basic audio signal and the processed pure speech signal are then fused together and output.

[0028] Optionally, the scene recognition and separation processing of the digital audio signal includes:

[0029] Audio signals from different directions are collected using a multi-channel microphone array;

[0030] Determine the direction of the sound source based on beamforming results;

[0031] Feature analysis of the audio signal is performed based on Mel frequency cepstral characteristics;

[0032] The presence of a speech activity interval in the audio signal is determined by combining the speech activity detection results.

[0033] Based on the determination results of the sound source direction, the spectral characteristics, and the speech activity range, the mixed audio signal is classified and routed.

[0034] Optionally, the sound effect adjustment of the basic audio signal includes: adjusting the reverberation parameter, reflection parameter, channel delay parameter and bass attenuation parameter, and the sound effect adjustment adopts at least one of the convolutional reverberation algorithm, algorithmic reverberation algorithm or reverberation algorithm based on feedback delay network, and configuring the reverberation parameter, reflection parameter, channel delay parameter and bass attenuation parameter according to a preset mapping relationship.

[0035] Optionally, the echo cancellation processing of the in-vehicle dialogue audio signal includes: performing decorrelation processing on the in-vehicle dialogue audio signal, estimating the echo path impulse response based on an adaptive filter and generating a simulated echo signal, and performing differential operation on the in-vehicle dialogue audio signal and the simulated echo signal to obtain the pure speech signal; the fusion processing includes: performing synchronization alignment and dynamic range control on the adjusted base audio signal and the processed pure speech signal, performing automatic ducking processing on the base audio signal when the in-vehicle dialogue audio signal is detected, and performing linear mixing and anti-clipping and amplitude limiting processing on each processed audio signal before outputting.

[0036] Compared with the closest prior art, the present invention has the following beneficial effects:

[0037] The present invention application proposes a system, method and electronic device for enhancing the immersive sound effect. By adding a sound effect adjustment module, it can accurately simulate the acoustic characteristics of a concert hall, break the limitation of existing car audio systems that can only perform basic sound effect adjustments, significantly improve the immersiveness and presence of basic audio playback, and allow users to obtain a similar professional music listening experience in the closed space of a car.

[0038] This application sets up an echo cancellation module and a voice processing module. First, the echo and background noise of the in-vehicle conversation are eliminated by the algorithm, and the pure conversation voice is extracted. Then, the pure voice is enhanced and refined reverberation is adjusted. While ensuring the clarity of the conversation, the spatial layering of the voice is effectively enhanced, which solves the problems of large echo interference and dry voice without spatial sense in the closed space of the vehicle.

[0039] This application incorporates an audio scene recognition and separation unit within its core control module. Based on multi-dimensional information such as sound source location, spectral characteristics, and energy time domain, it makes joint decisions, enabling precise separation and routing of basic audio and dialogue audio in complex in-vehicle acoustic environments. This provides a basis for differentiated processing of the two types of audio and fundamentally avoids mutual interference between the two types of audio signals.

[0040] This application constructs a dual-path audio processing architecture. Basic audio and dialogue audio undergo differentiated processing through dedicated processing paths. Then, the adaptive dynamic fusion strategy of the audio output module achieves harmonious fusion output of the two types of signals. Combined with the automatic dodging mechanism that prioritizes voice, the dynamic balance of the two types of audio signals is achieved, ensuring that music playback and in-vehicle conversations do not interfere with each other.

[0041] The audio system modules in this application work together to achieve intelligent control of the entire process from audio acquisition, separation, differential processing to fusion output. The audio acquisition module adopts a multi-channel microphone array to adapt to the acoustic characteristics of the enclosed space of the vehicle. The audio output module adopts a multi-channel speaker array to achieve immersive playback. The overall system has strong adaptability, high audio processing efficiency, and convenient operation.

[0042] The method of this application is clear and logically rigorous, with each step closely linked. It ensures normal operation of the equipment through system initialization, improves signal purity through noise reduction preprocessing, enhances separation accuracy through multi-dimensional joint decision-making, simulates acoustic characteristics and adjusts voice reverberation through fine parameter adjustment, and ensures output effect through adaptive dynamic fusion. The whole method is highly feasible, has stable processing effect, and can be effectively applied to in-vehicle audio systems. Attached Figure Description

[0043] To more clearly illustrate the specific embodiments of the present invention or the technical solutions in the prior art, the accompanying drawings used in the description of the specific embodiments or the prior art will be briefly introduced below. In all the drawings, similar elements or parts are generally identified by similar reference numerals. In the drawings, the elements or parts are not necessarily drawn to scale.

[0044] Figure 1 This is a structural block diagram of a system for enhancing immersive sound effects provided in an embodiment of the present invention;

[0045] Figure 2 This is a structural block diagram of the multimodal fusion cockpit collaborative recognition and intent prediction control system provided by the present invention;

[0046] Figure 3 This is a flowchart of the audio scene recognition and separation process provided by the present invention;

[0047] Figure 4This is an internal structural diagram of the electronic device provided by the present invention. Detailed Implementation

[0048] The embodiments of the technical solution of the present invention will now be described in detail with reference to the accompanying drawings. These embodiments are only used to more clearly illustrate the technical solution of the present invention and are therefore merely examples, and should not be construed as limiting the scope of protection of the present invention.

[0049] It should be noted that, unless otherwise stated, the technical or scientific terms used in this application should have the ordinary meaning as understood by one of ordinary skill in the art to which this invention pertains.

[0050] This invention provides a system, method, and electronic device for enhancing immersive sound effects, enabling echo cancellation, pure speech extraction, and reverberation adjustment for in-vehicle voice recordings. This solves the problems of insufficient immersion and unrealistic experience in in-vehicle audio systems, improving the user's auditory immersion and practicality. Embodiments of this invention are described below with reference to the accompanying drawings.

[0051] Example 1: See Figure 1 This embodiment 1 provides a system for enhancing immersive sound effects, including: a core control module 210, an audio acquisition module 220, a sound effect adjustment module 230, an echo cancellation module 240, a voice processing module 250, and an audio output module 260;

[0052] The core control module 210 is used to receive and process signals transmitted by each module, parse user operation commands, and control the coordinated operation of each module.

[0053] The audio acquisition module 220 is used to acquire audio signals in the vehicle environment, convert the audio signals into digital audio signals, and then transmit them to the core control module.

[0054] The sound effect adjustment module 230 is used to receive the basic audio signal routed by the core control module and adjust the sound effect of the basic audio signal.

[0055] The echo cancellation module 240 is used to receive the in-vehicle dialogue audio signal routed by the core control module, and to perform echo cancellation processing on the in-vehicle dialogue audio signal to output a pure voice signal.

[0056] The voice processing module 250 is used to receive the pure voice signal and perform voice enhancement processing and reverberation adjustment on the pure voice signal.

[0057] The audio output module 260 is used to receive the basic audio signal processed by the sound effect adjustment module and the pure voice signal processed by the voice processing module, and output the two signals after fusion processing.

[0058] The core control module 210 includes an audio scene recognition and separation unit. This unit is used to make a joint decision on the acquired mixed audio signal based on the sound source location information, spectral feature information, and energy time domain information, so as to separate the basic audio signal and the in-vehicle dialogue audio signal in the mixed audio signal and route them separately.

[0059] In the above embodiments, the audio acquisition module includes a multi-channel microphone array.

[0060] In the above embodiments, the sound effect adjustment module includes a reverb processing unit; the reverb processing unit is used to adjust the reverb parameters, reflected sound parameters, channel delay parameters, and bass attenuation parameters of the basic audio signal;

[0061] The reverberation processing unit employs at least one of the following: convolutional reverberation algorithm, algorithmic reverberation algorithm, or reverberation algorithm based on feedback delay network.

[0062] In the above embodiments, the echo cancellation module includes a decorrelation processing unit and an adaptive filtering unit. The decorrelation processing unit is used to perform decorrelation processing on the in-vehicle dialogue audio signal, and the adaptive filtering unit is used to estimate the echo path impulse response and generate a simulated echo signal.

[0063] In the above embodiments, the audio output module includes multiple vehicle speakers, and the audio output module is used to perform synchronization alignment, dynamic range control, automatic ducking and linear mixing processing on the basic audio signal processed by the sound effect adjustment module and the pure voice signal processed by the voice processing module.

[0064] Accordingly, Example 2: Based on the same technical concept, this example also provides a method for enhancing immersive sound effects corresponding to Example 1 above, such as... Figure 2 As shown, the method includes:

[0065] S101 acquires mixed audio signals from the vehicle environment and converts the mixed audio signals into digital audio signals;

[0066] S102 performs scene recognition and separation processing on the digital audio signal, and performs joint determination on the mixed audio signal based on the sound source location information, spectrum feature information and energy time domain information to obtain the basic audio signal and the in-vehicle dialogue audio signal;

[0067] S103 adjusts the sound effects of the basic audio signal;

[0068] S104 performs echo cancellation processing on the in-vehicle dialogue audio signal to obtain a pure voice signal;

[0069] S105 performs speech enhancement processing on the pure speech signal and adjusts the reverberation of the enhanced pure speech signal; then, it merges the adjusted basic audio signal with the processed pure speech signal and outputs the result.

[0070] In step S101 above, the scene recognition and separation processing of the digital audio signal includes:

[0071] Audio signals from different directions are collected using a multi-channel microphone array;

[0072] Determine the direction of the sound source based on beamforming results;

[0073] Feature analysis of the audio signal is performed based on Mel frequency cepstral characteristics;

[0074] The presence of a speech activity interval in the audio signal is determined by combining the speech activity detection results.

[0075] Based on the determination results of the sound source direction, the spectral characteristics, and the speech activity range, the mixed audio signal is classified and routed.

[0076] In step S103 above, the sound effect adjustment of the basic audio signal includes adjusting the reverberation parameters, reflection parameters, channel delay parameters and bass attenuation parameters. The sound effect adjustment adopts at least one of the convolutional reverberation algorithm, algorithmic reverberation algorithm or reverberation algorithm based on feedback delay network, and configures the reverberation parameters, reflection parameters, channel delay parameters and bass attenuation parameters according to a preset mapping relationship.

[0077] In step S104 above, the echo cancellation processing of the in-vehicle dialogue audio signal includes: performing decorrelation processing on the in-vehicle dialogue audio signal, estimating the echo path impulse response based on an adaptive filter and generating a simulated echo signal, and performing differential operation on the in-vehicle dialogue audio signal and the simulated echo signal to obtain the pure speech signal.

[0078] The fusion processing in step S105 above includes: synchronizing and dynamically aligning the adjusted base audio signal and the processed pure voice signal; performing automatic ducking processing on the base audio signal when the in-vehicle dialogue audio signal is detected; and outputting the processed audio signals after linear mixing and anti-clipping and amplitude limiting processing.

[0079] Example 3: Example 3 of the present invention provides a system and method for enhancing immersive sound effects, applicable to the field of vehicle audio technology. It can accurately simulate the acoustic characteristics of a concert hall, achieve echo cancellation and reverberation adjustment for in-vehicle conversations, and complete the effective separation and dynamic fusion output of basic audio and conversation audio, thereby enhancing the immersive sound experience of vehicle audio systems.

[0080] This embodiment provides a system for enhancing immersive sound effects, such as... Figure 1 As shown, the system includes a core control module, a sound effect adjustment module, an audio acquisition module, an echo cancellation module, a voice processing module, and an audio output module. The modules are connected and transmit signals to each other through the vehicle communication bus. The system is compatible with the vehicle's 12V / 24V power supply system and meets the usage requirements of the vehicle environment.

[0081] The core control module is the main control unit of the entire audio system. It uses a high-performance automotive microprocessor as its core and integrates an audio scene recognition and separation unit, as well as a signal format conversion unit and a command parsing unit. The command parsing unit receives and parses user commands issued via the in-vehicle central control screen, physical buttons, and voice commands, including commands for parameter adjustment, mode switching, and sound effect activation / deactivation. The signal format conversion unit handles the format conversion and transmission control of digital and analog audio signals between different modules. The audio scene recognition and separation unit makes a joint judgment based on three dimensions of information: sound source location, spectral characteristics, and energy time domain. It analyzes the collected mixed audio signals, accurately distinguishes between basic audio signals (music, radio, etc.) and in-vehicle user dialogue audio signals, and routes the two types of signals to the corresponding processing modules, achieving automatic separation and routing. The core control module can also receive real-time operating status signals from each module, enabling coordinated control of each module and ensuring stable system operation.

[0082] The audio acquisition module is electrically connected to the core control module and employs a multi-channel microphone array. The microphone array is evenly distributed across the vehicle's center console, doors, and roof, enabling precise capture of audio signals from different directions within the vehicle environment. This adapts to the acoustic characteristics of the enclosed vehicle space, reducing audio signal distortion. The core function of this module is to acquire mixed audio signals from the vehicle environment, including basic audio signals and in-vehicle user conversation audio signals. It also integrates an analog-to-digital converter to convert the acquired analog audio signals into digital audio signals before transmitting them to the core control module for further processing.

[0083] The sound effect adjustment module is electrically connected to the core control module and receives the basic audio signal routed by the core control module. This module has a built-in concert hall sound effect parameter adjustment unit with a preset concert hall acoustic characteristic model, allowing for the adjustment of concert hall-specific sound effect parameters. These parameters include reverberation time, reflection time, reflection gain, sound delay of each channel, bass attenuation coefficient, and timbre fidelity. Users can adjust these parameters independently through the user interface or select a preset concert hall mode (such as a small concert hall, a large symphony hall, or a chamber concert hall). The system automatically matches the parameters, and through the coordinated adjustment of multiple parameters, it accurately simulates the acoustic characteristics of different types of concert halls, enhancing the sense of presence and immersion in the basic audio.

[0084] The echo cancellation module is electrically connected to the core control module and receives the in-vehicle user dialogue audio signal routed by the core control module. This module uses a variable-proportion least mean square algorithm as its core processing algorithm, which can effectively eliminate echo signals in the dialogue audio signal and filter in-vehicle background noise, including various interference noises in the in-vehicle environment such as music noise, engine noise, and wind noise. Through algorithmic processing, this module generates a pure dialogue voice signal without background noise or echo, and transmits this signal to the voice processing module for further processing.

[0085] The voice processing module is electrically connected to the echo cancellation module and receives the pure dialogue voice signal transmitted from the echo cancellation module. This module integrates a voice enhancement unit and a reverberation parameter adjustment unit. The voice enhancement unit enhances the pure dialogue voice signal through gain adjustment and secondary noise filtering to improve voice clarity. The reverberation parameter adjustment unit can adjust reverberation parameters such as reverberation intensity, reverberation decay time, and background silence threshold, adding appropriate reverberation effects to the pure dialogue voice signal according to user needs, simulating the voice reverberation experience in a concert hall environment, enhancing the spatial layering of the voice. The processed dialogue voice signal is then transmitted to the audio output module.

[0086] The audio output module is electrically connected to the sound effect adjustment module and the voice processing module, and includes a multi-channel speaker array composed of multiple in-vehicle speakers. It also integrates a digital-to-analog converter (DAC) and an adaptive dynamic fusion processing unit. The DAC converts the received digital audio signal into an analog audio signal to match the speaker's playback requirements. The adaptive dynamic fusion processing unit performs adaptive dynamic fusion mixing on the base audio signal processed by the sound effect adjustment module and the dialogue voice signal processed by the voice processing module. This adaptive dynamic fusion processing includes signal synchronization alignment, dynamic range control, automatic ducking with voice priority, linear mixing, and anti-clipping limiting. This module achieves immersive audio signal playback through the multi-channel speaker array, enhancing the presence of the audio signal.

[0087] Example 4: This example provides a method for enhancing immersive sound effects, applied to the system for enhancing immersive sound effects described in Example 1. The method includes the following steps:

[0088] S1, System Initialization

[0089] When a user activates the concert hall immersive sound effect mode of this invention via the vehicle's central control screen, physical buttons, or voice command, the core control module immediately initiates a system self-test program. This program comprehensively checks the hardware status, communication connection status, and software operation status of the audio acquisition module, echo cancellation module, sound effect adjustment module, voice processing module, and audio output module. After confirming that each module is functioning normally, the core control module initializes the operating parameters of each module to the system default values, preparing for subsequent audio processing. If any module is detected to be malfunctioning, the core control module will issue a fault warning via the vehicle's central control screen and simultaneously stop the activation of the sound effect mode.

[0090] S2, Audio Acquisition and Separation Routing

[0091] The audio acquisition module continuously acquires mixed audio signals from the in-vehicle environment using a multi-channel microphone array. These mixed audio signals include the basic audio signals played by the in-vehicle audio system (music, radio, etc.) and the audio signals of conversations between users inside the vehicle, as well as various background noises from the in-vehicle environment. The audio acquisition module converts the acquired analog audio signals into digital audio signals using an analog-to-digital converter before transmitting them to the core control module.

[0092] The core control module performs noise reduction preprocessing on the received digital audio signal, filtering out minor environmental interference signals in the vehicle environment through methods such as fixed threshold filtering and smoothing filtering to ensure the purity of the digital audio signal. Subsequently, the audio scene recognition and separation unit built into the core control module performs multi-dimensional joint judgment on the noise-reduced digital audio signal, based on the following criteria:

[0093] 1. Sound source location identification: Through beamforming technology, the main location of the sound source of the audio signal is identified. If the sound source comes from the vehicle speaker, it is initially determined to be the basic audio signal; if the sound source comes from the direction of the driver, front passenger or rear passenger seat, it is initially determined to be the audio signal of the user conversation in the vehicle.

[0094] 2. Spectral Feature Analysis: The spectral characteristics of the audio signal are analyzed using the MFCC feature analysis method. If the signal has the continuous spectral characteristics of music, it is determined to be a basic audio signal; if the signal has the fast time-varying spectral characteristics of speech, it is determined to be an in-vehicle user dialogue audio signal.

[0095] 3. Voice Activity Detection: The VAD algorithm is used to detect whether there is a voice activity range in the audio signal. If it exists, it is determined to be the audio signal of the in-vehicle user conversation; if it does not exist, it is determined to be the basic audio signal.

[0096] Based on the joint decision results of the above multi-dimensional information, the audio scene recognition and separation unit routes the basic audio signal to the sound effect adjustment module and the in-vehicle user dialogue audio signal to the echo cancellation module, thereby achieving accurate separation and routing of the two types of audio signals.

[0097] S3, Concert Hall Sound Effect Adjustment

[0098] The core control module receives and parses user operation commands in real time. When the user triggers a concert hall sound effect adjustment command (adjust parameters independently or select a preset concert hall mode) through the vehicle's central control screen, the core control module transmits the command to the sound effect adjustment module, which then controls the sound effect adjustment module to adjust the concert hall-specific sound effect parameters of the received basic audio signal.

[0099] The sound effect adjustment module adjusts parameters such as reverberation time, reflection time, reflection gain, sound delay of each channel, bass attenuation coefficient, and timbre fidelity in a coordinated manner according to user instructions. Based on the preset acoustic characteristic model of the concert hall, it processes the basic audio signal to accurately simulate the acoustic characteristics of the concert hall, enhance the sense of presence and immersion of the basic audio, and transmits the processed basic audio signal to the audio output module.

[0100] S4. Voice Echo Cancellation and Pure Speech Extraction

[0101] The echo cancellation module receives the in-vehicle user dialogue audio signal routed by the core control module and processes it using a variable-scale demodulation least mean square algorithm. The specific processing procedure is as follows:

[0102] 1. First, the autocorrelation of the audio signal of the in-vehicle user conversation is reduced by decorrelation filtering, thereby reducing interference from the signal itself;

[0103] 2. The echo path impulse response is dynamically estimated using an adaptive filter, and a simulated echo signal matching the actual echo signal is generated based on the estimation results;

[0104] 3. Subtract the original in-vehicle user dialogue audio signal from the analog echo signal to eliminate the echo signal;

[0105] 4. While eliminating echoes, the algorithm filters in-vehicle background noise, including music noise, engine noise, wind noise and other types of interference noise, and finally generates a pure dialogue voice signal without background noise or echoes.

[0106] The echo cancellation module transmits the generated pure conversational speech signal to the speech processing module for further processing.

[0107] S5, Voice Reverb Adjustment

[0108] The speech processing module receives the pure conversational speech signal transmitted by the echo cancellation module. First, it performs speech enhancement processing on the speech signal through the speech enhancement unit. The strength of the speech signal is increased by gain adjustment, and noise is filtered again to further improve the clarity of the speech.

[0109] Subsequently, the voice processing module adjusts the reverberation parameters of the pure dialogue voice signal according to the user's operation instructions through the reverberation parameter adjustment unit. The reverberation parameters include reverberation intensity, reverberation decay time, and background silence threshold. By finely adjusting each parameter, a suitable reverberation effect is added to the pure dialogue voice signal to simulate the voice reverberation experience in a concert hall environment, effectively enhancing the spatial layering of the voice. The processed dialogue voice signal is then transmitted to the audio output module.

[0110] S6, Adaptive Dynamic Fusion Output

[0111] The audio output module simultaneously receives the basic audio signal processed in step S3 and the dialogue voice signal processed in step S5. The adaptive dynamic fusion processing unit performs adaptive dynamic fusion processing on the two types of signals. The specific processing procedure is as follows:

[0112] 1. Signal synchronization alignment: The time axes of the two types of audio signals are synchronized to avoid time differences in signal playback and ensure a smooth listening experience;

[0113] 2. Dynamic range control: Adjusts the dynamic range of the two types of audio signals to match their volume ranges and avoid sudden volume changes;

[0114] 3. Voice-first automatic ducking: The system adopts a voice-first automatic ducking strategy. The audio detection unit detects the presence of dialogue voice signals in real time. If a dialogue voice signal is detected, the gain of the basic audio signal is automatically reduced to ensure the clarity of the dialogue. If no dialogue voice signal is detected, the gain of the basic audio signal is smoothly restored within a preset time to avoid auditory discomfort caused by sudden volume changes.

[0115] 4. Linear Mixing and Anti-Clipping Limiting: The two types of audio signals that have undergone the above processing are linearly mixed, and anti-clipping limiting is performed to avoid clipping distortion in the mixed audio signal and ensure signal integrity.

[0116] 5. Multi-channel distribution: The mixed audio signal is distributed according to the multi-channel playback requirements and transmitted to the corresponding speakers.

[0117] Finally, the audio output module converts the digital audio signal into an analog audio signal through a digital-to-analog converter, which is then played by a multi-channel speaker array to achieve an immersive, real-time audio experience.

[0118] Example 5: This example is a specific application scenario. The system and method for enhancing immersive sound effects of the present invention are applied to the in-car audio system of a family car. The user plays classical symphony in the car and talks with the passenger in the front seat, hoping to experience a concert hall-like sound effect and clear and spatial conversation at the same time. The specific implementation process is as follows:

[0119] 1. When the user activates the "Concert Hall Realistic Sound Effect Mode" through the vehicle's central control screen, the core control module initiates a system self-test. After confirming that each module is working properly, it initializes the parameters of each module to their default values.

[0120] 2. The multi-channel microphone array of the audio acquisition module continuously acquires mixed sound field signals in the vehicle, including classical symphony signals played by the car audio system, conversation signals between the user and the front passenger, and background noise such as engine and wind noise. After analog-to-digital conversion, the signals are transmitted to the core control module.

[0121] 3. For example Figure 3 After the core control module performs noise reduction preprocessing on the digital audio signal, the audio scene recognition and separation unit makes a multi-dimensional joint decision: beamforming technology identifies the sound source as mainly coming from the vehicle speakers (classical symphony) and the driver / passenger seat direction (dialogue); MFCC feature analysis determines that the signal from the speaker direction has continuous spectral characteristics of music, while the signal from the passenger direction has fast time-varying spectral characteristics of speech; the VAD algorithm confirms that the signal from the passenger direction has a speech activity range. Based on the joint decision results, the system routes the classical symphony signal to the sound effect adjustment module and the dialogue signal to the echo cancellation module.

[0122] 4. When the user selects the "Large Symphony Hall" preset mode on the vehicle's central control screen, the core control module controls the sound effect adjustment module to adjust the sound effect parameters of the classical symphony signal to simulate the acoustic characteristics of a large symphony hall.

[0123] 5. The echo cancellation module uses a variable scaling least mean square algorithm to process the separated dialogue signal, effectively filtering out the echo of classical symphony music as background noise, as well as interference noise such as engine and wind noise, and outputting a pure dialogue voice signal.

[0124] 6. After receiving a clean dialogue voice signal, the voice processing module first performs voice enhancement processing, and then adds an appropriate reverb effect according to the system's default reverb parameters, so that the dialogue voice is both clear and has a natural spatial feel similar to that of a concert hall.

[0125] 7. The audio output module receives the processed classical symphony signal and the dialogue voice signal, and performs adaptive dynamic fusion processing: when the system detects that the user has started speaking, the automatic ducking mechanism is activated, and the volume of the symphony is smoothly reduced within a preset time; when the user stops speaking, the volume of the symphony is smoothly restored within a preset time. Subsequently, the two types of signals are linearly mixed and anti-clipping limited, and then transmitted to the in-vehicle speaker array through multi-channel distribution;

[0126] 8. The in-vehicle speaker array plays the processed audio signal, allowing end users to enjoy both the grand and immersive sound effects of a large symphony hall and clear, comfortable, and spacious conversations with the front passenger. The two types of audio signals do not interfere with each other, resulting in an excellent overall listening experience.

[0127] In one embodiment, an electronic device is provided, which may be a terminal, and its internal structure diagram may be as follows: Figure 4 As shown. The electronic device includes a processor, memory, communication interface, display screen, and input device connected via a system bus. The processor provides computing and control capabilities. The memory includes a non-volatile storage medium and internal memory. The non-volatile storage medium stores an operating system and computer programs. The internal memory provides an environment for the operation of the operating system and computer programs in the non-volatile storage medium. The communication interface is used for wired or wireless communication with external terminals; wireless communication can be achieved through Wi-Fi, mobile cellular networks, NFC (Near Field Communication), or other technologies. When the computer program is executed by the processor, it implements the method for enhancing immersive sound effects as described in any one of steps S101 to S105. The display screen can be a liquid crystal display (LCD) or an e-ink display. The input device can be a touch layer covering the display screen, buttons, a trackball, or a touchpad mounted on the device's casing, or an external keyboard, touchpad, or mouse.

[0128] Those skilled in the art will understand that Figure 4 The structure shown is merely a block diagram of a portion of the structure related to the present application and does not constitute a limitation on the computer device to which the present application is applied. Specific computer devices may include more or fewer components than those shown in the figure, or combine certain components, or have different component arrangements.

[0129] Those skilled in the art will understand that embodiments of this application can be provided as methods, systems, or computer program products. Therefore, this application can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, this application can take the form of a computer program product embodied on one or more computer-usable storage media (including, but not limited to, disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0130] This application is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of this application. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart... Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.

[0131] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.

[0132] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.

[0133] The above are merely embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention are included within the scope of the claims of the present invention pending approval.

Claims

1. A system for enhancing a sense of presence, characterized by include: The core control module, audio acquisition module, sound effect adjustment module, echo cancellation module, voice processing module, and audio output module are all included. The core control module is used to receive and process signals transmitted by each module, parse user operation commands, and control the coordinated operation of each module. The audio acquisition module is used to acquire audio signals in the vehicle environment, convert the audio signals into digital audio signals, and then transmit them to the core control module. The sound effect adjustment module is used to receive the basic audio signal routed by the core control module and adjust the sound effect of the basic audio signal. The echo cancellation module is used to receive the in-vehicle conversation audio signal routed by the core control module, and to perform echo cancellation processing on the in-vehicle conversation audio signal to output a pure voice signal. The voice processing module is used to receive the pure voice signal and perform voice enhancement processing and reverberation adjustment on the pure voice signal. The audio output module is used to receive the basic audio signal processed by the sound effect adjustment module and the pure voice signal processed by the voice processing module, and output the two signals after fusion processing. The core control module includes an audio scene recognition and separation unit. This unit is used to make a joint decision on the acquired mixed audio signal based on the sound source location information, spectral feature information, and energy time domain information, so as to separate the basic audio signal and the in-vehicle dialogue audio signal in the mixed audio signal and route them separately.

2. The system of claim 1, wherein, The audio acquisition module includes a multi-channel microphone array.

3. The system of claim 1, wherein, The sound effect adjustment module includes a reverb processing unit; the reverb processing unit is used to adjust the reverb parameters, reflected sound parameters, channel delay parameters, and bass attenuation parameters of the basic audio signal; The reverberation processing unit employs at least one of the following: convolutional reverberation algorithm, algorithmic reverberation algorithm, or reverberation algorithm based on feedback delay network.

4. The system of claim 1, wherein, The echo cancellation module includes a decorrelation processing unit and an adaptive filtering unit. The decorrelation processing unit is used to perform decorrelation processing on the in-vehicle dialogue audio signal, and the adaptive filtering unit is used to estimate the echo path impulse response and generate a simulated echo signal.

5. The system of claim 1, wherein, The audio output module includes multiple vehicle speakers. The audio output module is used to perform synchronization alignment, dynamic range control, automatic ducking, and linear mixing processing on the basic audio signal processed by the sound effect adjustment module and the pure voice signal processed by the voice processing module.

6. A method of enhancing a sense of presence, characterized by, include: Collect mixed audio signals from the vehicle environment and convert the mixed audio signals into digital audio signals; The digital audio signal is subjected to scene recognition and separation processing. Based on the sound source location information, spectrum feature information and energy time domain information, the mixed audio signal is jointly determined to obtain the basic audio signal and the in-vehicle dialogue audio signal. Adjust the sound effects of the basic audio signal; Echo cancellation processing is performed on the in-vehicle conversation audio signal to obtain a pure speech signal; The pure speech signal is subjected to speech enhancement processing, and the enhanced pure speech signal is subjected to reverberation adjustment; the adjusted basic audio signal and the processed pure speech signal are then fused together and output.

7. The method of claim 6, wherein, The scene recognition and separation processing of the digital audio signal includes: Audio signals from different directions are collected using a multi-channel microphone array; Determine the direction of the sound source based on beamforming results; Feature analysis of the audio signal is performed based on Mel frequency cepstral characteristics; The presence of a speech activity interval in the audio signal is determined by combining the speech activity detection results. Based on the determination results of the sound source direction, the spectral characteristics, and the speech activity range, the mixed audio signal is classified and routed.

8. The method of claim 6, wherein, The sound effect adjustment of the basic audio signal includes: adjusting the reverberation parameters, reflection parameters, channel delay parameters and bass attenuation parameters, and the sound effect adjustment adopts at least one of the convolutional reverberation algorithm, algorithmic reverberation algorithm or reverberation algorithm based on feedback delay network, and configuring the reverberation parameters, reflection parameters, channel delay parameters and bass attenuation parameters according to a preset mapping relationship.

9. The method of claim 6, wherein, The echo cancellation processing of the in-vehicle dialogue audio signal includes: performing decorrelation processing on the in-vehicle dialogue audio signal, estimating the echo path impulse response based on an adaptive filter and generating a simulated echo signal, and performing differential operation on the in-vehicle dialogue audio signal and the simulated echo signal to obtain the pure speech signal; the fusion processing includes: performing synchronization alignment and dynamic range control on the adjusted base audio signal and the processed pure speech signal, performing automatic ducking processing on the base audio signal when the in-vehicle dialogue audio signal is detected, and performing linear mixing and anti-clipping and amplitude limiting processing on each processed audio signal before outputting.

10. An electronic device, comprising: The electronic device includes: At least one processor; and A memory communicatively connected to the at least one processor; wherein, The memory stores a computer program that can be executed by the at least one processor to enable the at least one processor to perform the method of any one of claims 6-8.