Sound field control system and method based on MCU
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- 镁佳(北京)科技有限公司
- Filing Date
- 2026-01-15
- Publication Date
- 2026-04-28
Smart Images

Figure CN121940689A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of audio processing technology, and more specifically to a sound field control system and method based on an MCU. Background Technology
[0002] In current audio applications, such as home theaters, professional conference rooms, or multimedia entertainment systems, precise sound field control is crucial for creating an immersive listening experience. Sound field control systems process and adjust the phase, intensity, and spatial distribution of audio signals to simulate or optimize specific acoustic environments, thereby significantly enhancing the user's sense of presence and sound quality satisfaction.
[0003] However, in related technologies, systems that implement the aforementioned sound field control functions are typically based on high-performance dedicated processors or complex peripheral hardware, resulting in high overall system costs and complex structures. Furthermore, existing solutions often employ preset, fixed processing modes, making it difficult to adaptively adjust to real-time changes in the acoustic environment or the user's personalized spatial location and listening preferences, thus lacking flexibility and intelligence. Therefore, how to construct a cost-effective, easily implementable sound field control system that can dynamically adapt to the environment and user needs has become a pressing technical problem to be solved. Summary of the Invention
[0004] In view of this, the present invention provides a sound field control system and method based on MCU, so as to at least solve the problem of how to construct a sound field control system that is cost-effective, easy to implement and can dynamically adapt to the environment and user needs.
[0005] This disclosure provides a sound field control system based on an MCU, comprising: an audio input module, a signal processing module, an audio output module, a human-machine interaction module, and a power supply module, wherein: An audio input module is used to acquire ambient sound signals and / or audio signals from external audio devices; The signal processing module, with the MCU as its core, is used to perform real-time analysis and processing of the audio signals acquired by the audio input module in order to implement sound field control. The signal processing module is configured to perform the following: frequency domain analysis of the audio signal, adaptive filtering based on the analysis results to suppress noise, and spatial sound field effect processing of the noise-reduced audio signal based on user location information or preset scene parameters to generate an output signal that conforms to the target sound field effect. The audio output module, connected to the signal processing module, is used to output processed audio signals; The human-computer interaction module, connected to the signal processing module, is used to receive user commands and provide feedback on the system status; The power module is used to supply power to the various modules within the system.
[0006] This disclosure also provides a MCU-based sound field control method, applied to the aforementioned MCU-based sound field control system, the method comprising: The audio input module acquires ambient sound signals and / or audio signals from external audio devices. The signal processing module, with the MCU as its core, performs real-time analysis and processing on the audio signals acquired by the audio input module to implement sound field control. The real-time analysis and processing includes: performing frequency domain analysis on the audio signals through the signal processing module, performing adaptive filtering based on the analysis results to suppress noise, and performing spatial sound field effect processing on the noise-reduced audio signals based on user location information or preset scene parameters to generate an output signal that conforms to the target sound field effect. The processed audio signal is output through the audio output module; The system receives user commands and provides feedback on system status through the human-computer interaction module. The power supply module provides power to all modules within the system.
[0007] This disclosure also provides an electronic device, including: Memory, used to store computer programs; A processor is used to execute computer programs to implement the steps of the MCU-based sound field control method described above.
[0008] In another aspect, this disclosure provides a computer-readable storage medium storing computer instructions for causing a computer to implement the aforementioned MCU-based sound field control method.
[0009] This disclosure also provides a computer program product, including computer instructions for causing a computer to execute the aforementioned MCU-based sound field control method.
[0010] The MCU-based sound field control system and method disclosed above significantly reduce system hardware complexity and overall cost by using a low-cost MCU as the system processing core and integrating frequency domain analysis, adaptive filtering, and spatial sound field effect processing. This ensures high-performance sound field control while maintaining high-performance sound field control functionality. By performing real-time frequency domain analysis on the audio signal and applying adaptive filtering based on the analysis results, interference can be dynamically suppressed according to changes in environmental noise, effectively improving the clarity and purity of the audio signal.
[0011] In addition, by processing the noise-reduced audio signal with spatial sound field effects based on real-time user location information or preset scene parameters, the sound field output can be dynamically adjusted and optimized to provide users with an immersive audio experience that adapts to their location and needs. Attached Figure Description
[0012] To more clearly illustrate the technical solutions in the specific embodiments or related technologies of this disclosure, the accompanying drawings used in the description of the specific embodiments or related technologies will be briefly introduced below. Obviously, the accompanying drawings described below are some embodiments of this disclosure. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0013] Figure 1 An exemplary schematic diagram of the architecture of a MCU-based sound field control system according to an embodiment of the present disclosure is shown; Figure 2 This is a schematic flowchart of a sound field control method based on an MCU provided in an embodiment of this disclosure; Figure 3 This is a schematic diagram of the structure of an electronic device according to an embodiment of the present disclosure. Detailed Implementation
[0014] The technical solutions of the embodiments of this disclosure will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this disclosure, and not all embodiments. Based on the embodiments of this disclosure, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the protection scope of this disclosure.
[0015] In many audio application scenarios such as home theaters, conference rooms, and recording studios, precise control of the sound field is crucial, as it directly affects the immersion, clarity, and spatial realism of the audio.
[0016] Currently, technical solutions for achieving high-quality sound field control typically rely on high-performance dedicated digital signal processors, complex multi-device arrays, or PC-based audio processing systems. While these solutions can achieve certain sound field processing effects, they generally suffer from complex system structures and high hardware costs, making them difficult to popularize in cost-sensitive consumer scenarios or where a simple deployment is desired.
[0017] More importantly, these traditional systems often employ pre-set, fixed sound field processing modes. They cannot perceive changes in the acoustic characteristics of the listening environment in real time (such as noise fluctuations and room reverberation), nor do they have the ability to track the user's actual position and orientation. Therefore, the system cannot dynamically optimize speech intelligibility based on ambient noise, nor can it adjust sound image localization in real time according to user movement, resulting in a rigid final audio output that fails to provide users with a consistently high-quality, personalized, and immersive experience across different usage scenarios.
[0018] Therefore, there is an urgent need in this field for a solution that can overcome the above-mentioned defects: that is, a new type of sound field control system that can achieve controllable system cost, simplified structure, and environmental adaptability and user perception capabilities while ensuring excellent sound field control performance.
[0019] To address the aforementioned issues, various embodiments of this disclosure provide an MCU-based sound field control system. The system includes: an audio input module, a signal processing module, an audio output module, a human-machine interface module, and a power supply module. The audio input module is used to acquire ambient sound signals and / or audio signals from external audio devices. The signal processing module, with the MCU as its core, is used to perform real-time analysis and processing of the audio signals acquired by the audio input module to implement sound field control. The signal processing module is configured to perform: frequency domain analysis of the audio signals; adaptive filtering based on the analysis results to suppress noise; and spatial sound field effect processing of the noise-reduced audio signals based on user location information or preset scene parameters to generate an output signal that conforms to the target sound field effect. The audio output module, connected to the signal processing module, is used to output the processed audio signals. The human-machine interface module, connected to the signal processing module, is used to receive user commands and provide feedback on the system status. The power supply module supplies power to all modules within the system.
[0020] To make the objectives, technical solutions, and advantages of the embodiments of this disclosure clearer, the technical solutions of the embodiments of this disclosure will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this disclosure, and not all embodiments. Based on the embodiments of this disclosure, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this disclosure.
[0021] Please refer to Figure 1 , Figure 1 An exemplary schematic diagram of the architecture of a MCU-based sound field control system according to an embodiment of this disclosure is shown. Figure 1 As shown, the MCU-based sound field control system 100 (hereinafter referred to as system 100) includes: an audio input module 101, a signal processing module 102, an audio output module 103, a human-machine interaction module 104, and a power supply module 105, wherein: The audio input module 101 is used to acquire ambient sound signals and / or audio signals from external audio devices.
[0022] In this embodiment, collecting ambient sound signals can refer to picking up natural sounds in the physical space where the system 100 is located through acoustic sensors such as microphone arrays.
[0023] Acquiring audio signals from external audio devices can refer to receiving audio electrical signals transmitted via wired or wireless means from devices such as smartphones, computers, and televisions.
[0024] The audio input module 101 can realize composite input of sound field signal sources, thereby providing raw audio data containing environmental information and program source content for subsequent processing steps.
[0025] Specifically, the audio input module 101 may include a high-fidelity microphone, a preamplifier, and corresponding interface circuitry to ensure that the acquired audio signal has sufficient signal-to-noise ratio and dynamic range to meet the requirements of high-quality sound field processing.
[0026] Preferably, the high-fidelity microphone can be a condenser microphone array with a frequency response range of 20 Hz to 20 kHz and a sensitivity of -40 dBV / Pa. The preamplifier has a gain of 1000.
[0027] The signal processing module 102, with the MCU as its core, is used to perform real-time analysis and processing of the audio signals acquired by the audio input module 101 in order to implement sound field control. The signal processing module 102 is configured to perform the following: frequency domain analysis of the audio signal, adaptive filtering based on the analysis results to suppress noise, and spatial sound field effect processing of the noise-reduced audio signal based on user location information or preset scene parameters to generate an output signal that conforms to the target sound field effect.
[0028] In this embodiment, a microcontroller unit (MCU) is selected as the main control and computing carrier of the signal processing module 102. The MCU executes dedicated program code or firmware stored within it to coordinate and complete a series of real-time digital signal processing tasks to achieve intelligent sound field control.
[0029] Specifically, real-time analysis and processing of audio signals can refer to converting the acquired time-series audio signals to the frequency domain using digital signal processing algorithms, and then analyzing the spectral composition, energy distribution, and phase relationship of each frequency component of the audio signal, thereby providing characteristic basis for subsequent processing.
[0030] Furthermore, adaptive filtering based on the analysis results to suppress noise can refer to the system 100 dynamically adjusting the coefficients of the digital filter (for example, by using an adaptive algorithm such as minimum mean square error) based on the noise spectrum characteristics obtained in real time through frequency domain analysis, thereby selectively filtering out or reducing steady-state and non-steady-state noise interference in the environment and improving the signal-to-noise ratio and purity of the signal.
[0031] Furthermore, the signal processing module 102 can obtain user location information through input from the human-computer interaction module 104 or through real-time detection by the system's built-in sensors; the preset scene parameters can correspond to preset acoustic environment model parameters such as "cinema mode" and "concert hall mode".
[0032] Spatial sound field effect processing can be a collective term for a series of algorithms. Its core purpose is to render mono or stereo audio signals in a spatial dimension based on user location information or preset scene parameters.
[0033] The rendering described above can be achieved by simulating the physical characteristics of sound waves in three-dimensional space, such as propagation, reflection, and reverberation. Specifically, it can include algorithms such as using head-related transfer functions to locate virtual sound sources, generating surround sound fields through multi-channel mixing and time-delay technology, and superimposing reverberation effects that conform to the acoustic characteristics of the scene.
[0034] Finally, the signal processing module 102 can output a set of multi-channel audio signals that have been intelligently analyzed and rendered, which can carry the target sound field characteristics that conform to the user's intention and listening environment.
[0035] The audio output module 103 is connected to the signal processing module 102 and is used to output the processed audio signal.
[0036] In this embodiment, the audio output module 103 can be used to convert the processed audio signal generated by the signal processing module 102 into sound waves and reproduce them in order to achieve the target sound field effect in the listening space.
[0037] Specifically, the audio output module 103 may include one or more speaker units and an audio power amplifier corresponding to each speaker unit.
[0038] The type, quantity, and spatial layout of the loudspeaker units can be configured according to the requirements of the target sound field. For example, a combination of different types of units such as full-range loudspeakers, high-frequency loudspeakers, and low-frequency loudspeakers can be selected. The layout can be uniformly distributed, arranged at a specific angle, or arranged in an array.
[0039] The audio power amplifier can be used to receive audio electrical signals from the signal processing module 102 and amplify them to drive the corresponding speaker unit.
[0040] Furthermore, the audio output module 103, through the aforementioned components, can ultimately restore the electrical signal carrying spatial sound field information into a sound with a specific sense of direction and immersion, thereby providing the user with an immersive auditory experience.
[0041] Preferably, the audio output module 103 can consist of four 50W full-range speakers evenly distributed at a 45° angle, with each speaker connected to a 20W audio amplifier.
[0042] The human-computer interaction module 104 is connected to the signal processing module 102 and is used to receive user commands and provide feedback on the system status.
[0043] In this embodiment, the human-computer interaction module 104 can serve as an interface for information exchange between the user and the system 100. Its core function is to receive the user's control intentions and convert them into instructions that the system 100 can recognize, while intuitively feeding back the working status of the system 100 to the user.
[0044] Specifically, the human-computer interaction module 104 may receive user commands in ways including but not limited to: receiving remote control commands from mobile terminal applications such as smartphones and tablets through integrated or external local input devices such as touch screens, physical buttons, and knobs; receiving remote control commands from smartphones, tablets, and other mobile terminal applications through wireless communication units that support protocols such as Bluetooth and Wi-Fi; or receiving and parsing user voice commands through an integrated voice recognition unit.
[0045] User commands can include, but are not limited to, comprehensive control over sound field mode selection (such as "music" or "movie"), effect parameter adjustment (such as surround intensity and reverberation level), audio source switching, and system power on / off.
[0046] Meanwhile, the human-computer interaction module 104 can be used to provide feedback on the status of the system 100 in ways including but not limited to: displaying the current sound field mode, volume level, connection status, and other information through the touch screen or independent indicator lights or digital tubes; sending status data to a mobile terminal application for graphical display through the wireless communication unit; or broadcasting status information through audio prompts or speech synthesis.
[0047] The human-computer interaction module 104 can forward various user commands received to the signal processing module 102 as one of the key input parameters (such as scene parameters) for real-time audio processing and sound field control, thereby forming a closed-loop interaction between user intervention and system response, so that the sound field effect can flexibly adapt to the user's subjective preferences and real-time needs.
[0048] Power module 105 is used to supply power to various modules within the system.
[0049] In this embodiment, the power module 105 can be used to provide a stable and reliable power supply to all functional modules within the system 100.
[0050] Specifically, the power module 105 can adopt corresponding power supply design and power management strategies according to different application scenarios and product forms. For example, in portable application scenarios, the power module 105 can have a built-in rechargeable battery (such as a lithium-ion battery) as the main power source, and be equipped with corresponding battery management circuitry to handle charging control, power monitoring, and discharge protection, in order to meet the needs of mobile use. In fixed installation application scenarios, the power module 105 can use a power adapter or a built-in AC-DC conversion circuit to convert external AC mains power into the low-voltage DC power required by the system 100 for power supply.
[0051] Furthermore, the circuit design of the power module 105 can include functional units such as filtering, voltage regulation and protection (e.g., overvoltage and overcurrent protection) to ensure that the power supply delivered to each functional module is of pure quality and stable voltage, and to avoid interference to sensitive audio signal processing circuits caused by power supply noise or fluctuations.
[0052] Through its adaptable design and stable output, the power module 105 ensures that the audio input module 101, signal processing module 102, audio output module 103, and human-computer interaction module 104 can work together stably, thereby guaranteeing the realization of the overall functions of the system 100 and the user experience.
[0053] The MCU-based sound field control system and method disclosed above significantly reduce system hardware complexity and overall cost by using a low-cost MCU as the system processing core and integrating frequency domain analysis, adaptive filtering, and spatial sound field effect processing. This ensures high-performance sound field control while maintaining high-performance sound field control functionality. Real-time frequency domain analysis of the audio signal followed by adaptive filtering based on the analysis results dynamically suppresses interference according to changes in environmental noise, effectively improving the clarity and purity of the audio signal. Spatial sound field effect processing is applied to the denoised audio signal based on real-time user location information or preset scene parameters, dynamically adjusting and optimizing the sound field output to provide users with an immersive audio experience tailored to their location and needs.
[0054] In one possible implementation of the above embodiments, the signal processing module 102 is configured as follows: Frequency domain analysis of audio signals is performed using the Fast Fourier Transform algorithm; Based on the results of frequency domain analysis, an adaptive filtering algorithm is invoked to filter the audio signal; The sound field rendering parameters are determined based on the user's location information, and virtual surround sound and reverberation processing are performed on the filtered audio signal according to the sound field rendering parameters.
[0055] In this embodiment, the signal processing module 102 can be specifically configured to execute a multi-stage algorithm flow to achieve intelligent analysis and spatial rendering of the audio signal. This algorithm flow can be decomposed into the following steps: 1. Frequency domain analysis stage.
[0056] In this stage, the signal processing module 102 can convert the time-domain audio signal to the frequency domain using a specific digital signal transformation algorithm, so as to facilitate feature extraction and analysis at the spectral level.
[0057] As one possible implementation, this embodiment can use the Fast Fourier Transform (FFT) algorithm to efficiently complete this transformation. Through FFT, the amplitude and phase information of the audio signal at various frequency points can be obtained, thus providing crucial spectral characteristics for subsequent processing.
[0058] 2. Adaptive filtering stage.
[0059] Based on the spectrum analysis results obtained in the previous stage, such as the identified specific noise frequency bands and their energy, the signal processing module 102 calls the adaptive filtering algorithm to process the original audio signal in real time.
[0060] The core of this algorithm lies in its dynamic adjustment capability. For example, it can use the Least Mean Square (LMS) algorithm or its variants to continuously update the filter coefficients based on the input signal and noise estimation, thereby selectively suppressing or eliminating noise components in the environment and improving the purity and intelligibility of the signal.
[0061] 3. Spatial sound field rendering stage.
[0062] This stage allows you to apply spatial effects to the noise-reduced audio signal to create an immersive listening experience.
[0063] The system determines sound field rendering parameters based on user location information. This user location information can be input through the human-computer interaction module 104 or obtained in real time by built-in sensors such as UWB and cameras. These parameters will guide how to spatially shape the sound.
[0064] Subsequently, the signal processing module 102 performs virtual surround sound processing and reverberation processing on the signal based on the determined sound field rendering parameters. For example, Head Related Transfer Functions (HRTF) technology can be used to simulate the subtle differences in sound reaching the human ear from different directions, rendering ordinary two-channel or multi-channel signals into virtual surround sound with a clear three-dimensional spatial positioning sense; at the same time or subsequently, digital reverberation algorithms can be applied, such as using convolutional reverberation or artificial reverberation models, to add reflections and reverberations that conform to the characteristics of the target acoustic environment (such as a small room or concert hall), enhancing the spatial realism and sense of immersion.
[0065] Ultimately, the signal processing module 102 outputs a clear and immersive multi-channel audio signal that carries precise spatial information, thereby achieving intelligent and adaptive control from the original input to the target sound field.
[0066] The MCU-based sound field control system and method disclosed above utilizes the FTT algorithm for frequency domain analysis, thereby achieving efficient and accurate extraction of the spectral characteristics of audio signals, providing a reliable basis for subsequent noise identification and filtering. By invoking an adaptive filtering algorithm and dynamically adjusting the filter based on the frequency domain analysis results, targeted and real-time suppression of steady-state and non-steady-state noise in the environment is achieved, significantly improving the clarity and purity of the audio signal. By determining the sound field rendering parameters based on real-time user location information and sequentially performing virtual surround sound processing and reverberation processing on the audio signal according to these parameters, dynamic adaptation and personalized rendering of the sound field effect to the user's actual listening position are achieved, effectively enhancing immersion and spatial realism.
[0067] In one possible implementation of the above embodiments, the signal processing module 102 is configured as follows: Based on the user location information obtained by the spatial positioning algorithm, the channel gain balance in the virtual surround sound processing and / or the reverberation time in the reverberation processing are dynamically adjusted so that the sound field center perceived by the user is relatively fixed with respect to the user's position.
[0068] In this embodiment, the signal processing module 102 can be further configured to perform adaptive optimization functions, which are designed to solve the sound field center drift problem caused by user movement and ensure a stable immersive experience for the user.
[0069] Specifically, the system 100 continuously tracks the user's precise location using a spatial positioning algorithm. This positioning algorithm can be implemented using various technologies, such as infrared sensor arrays deployed in the environment, visual recognition and tracking using cameras, or wireless positioning technologies such as ultra-wideband, to obtain the user's coordinates within the listening area in real time.
[0070] The signal processing module 102 uses the real-time user location information as the core input to dynamically adjust the key parameters of the subsequent audio rendering environment.
[0071] Specifically, the system can dynamically calculate and adjust the channel gain balance based on the real-time offset of the user's position relative to the preset sound image of each virtual speaker. For example, when the user moves to the left of the listening area, the system will fine-tune the gain of the right front virtual channel and slightly increase the gain of the left rear surround channel. Through this real-time electronic balancing operation, the main sound image remains stably positioned in front of the user (such as in the center of the screen), while the surround sound maintains the correct sense of immersion.
[0072] The system can also dynamically adjust parameters such as reverberation time in reverberation processing based on the user's relative position to the room's optimal acoustic listening position. For example, when the user is far from the center of the room (which usually has more balanced acoustic characteristics) and close to the wall (which may lead to excessive early reflections), the system can appropriately shorten the reverberation time to maintain the clarity of the sound and avoid the sound becoming muddy due to changes in position.
[0073] Through the aforementioned coordinated adjustments, the system can achieve dynamic binding between the sound field center and the user's position. This means that no matter how the user moves within the listening area, the sound field center perceived by the user remains stable, thus maintaining the continuity and realism of the immersive audio experience in physically changing spaces.
[0074] The MCU-based sound field control system and method of the above embodiments of this disclosure dynamically adjust the channel gain balance in virtual surround sound processing based on real-time user position information obtained by a spatial positioning algorithm. This allows the sound image center perceived by the user to move with the user and maintain a relatively fixed spatial position, effectively solving the sound field offset problem caused by changes in listening position. By dynamically adjusting parameters such as reverberation time in reverberation processing according to the user position, the spatial reflection characteristics of the sound are matched with the user's actual listening environment, avoiding a decrease in sound clarity or spatial distortion caused by changes in user position, and optimizing listening consistency at different positions. Through the above-mentioned coordinated dynamic adjustment of virtual surround sound and reverberation parameters, real-time adaptation of the sound field effect to the user position is achieved, significantly improving the user experience and immersion stability of the system in non-fixed listening scenarios.
[0075] In one possible implementation of the above embodiments, the signal processing module 102 is configured as follows: The degree of noise suppression by the adaptive filtering algorithm is used as a control parameter input into the virtual surround sound processing and / or reverberation processing. When the noise suppression degree is high, the rendering intensity of the corresponding processing is enhanced; when the noise suppression degree is low, the rendering intensity of the corresponding processing is weakened.
[0076] In this embodiment, the signal processing module 102 can be configured to implement an intelligent sound field rendering intensity adjustment strategy based on the noise environment.
[0077] The core of this strategy lies in using the performance evaluation results of the pre-stage noise processing stage as the basis for dynamic control of the post-stage sound field rendering stage, thereby achieving adaptive matching between the rendering effect and the clarity of the current auditory environment.
[0078] Specifically, while running the adaptive filtering algorithm, the signal processing module 102 can evaluate its processing effect in real time and generate a control parameter characterizing the current noise suppression level. This parameter can be quantized as follows: Signal-to-noise ratio (SNR) improvement: Calculated by comparing the change in SNR of the signal before and after filtering in a specific frequency band or the entire frequency band.
[0079] Noise attenuation estimation: The energy of suppressed noise is estimated by analyzing the update coefficients of the adaptive filter or the energy of the error signal.
[0080] Preset threshold comparison result: The current ambient noise level is compared with a preset threshold representing a quiet environment, and a ratio or level parameter is output.
[0081] Furthermore, the noise suppression parameters obtained from the above evaluation can be transmitted in real time to the virtual surround sound processing unit and / or reverberation processing unit as a key input for adjusting its rendering intensity.
[0082] The system can preset or dynamically generate a set of mapping relationships through algorithms, for example: When the evaluation indicates a high level of noise suppression (i.e., current ambient noise is effectively filtered out and the audio signal purity is high), the system determines that the current auditory background is clear and has the conditions to support rich spatial effects. Therefore, the spatial diffusion and envelopment parameters in virtual surround sound processing are enhanced, and / or the reverberation time and reflection intensity in reverberation processing are increased, thereby making the output sound field effect more open and maximizing the immersive experience.
[0083] When the assessment indicates low noise suppression (i.e., strong ambient noise or limited filtering effect, resulting in relatively low signal purity), the system determines that the current auditory background is noisy, and excessive spatial effects may mask the main sound, reducing clarity. Therefore, the intensity of virtual surround sound processing is reduced, and / or the amount of reverberation added is decreased (e.g., by shortening the reverberation time), prioritizing the prominence and intelligibility of core audio content such as speech and main melody.
[0084] Understandably, the above mechanism can form an intelligent feedback loop from environmental perception (noise suppression effect) to experience optimization (rendering intensity), which can ensure that the system not only purifies audio at the signal level, but also intelligently manages limited auditory attention resources at the perception level, so as to output the most suitable sound field effect that balances clarity and immersion for users in various noise environments.
[0085] The MCU-based sound field control system and method of the above embodiments of this disclosure dynamically adjust the rendering intensity of virtual surround sound and / or reverberation processing by using the noise suppression level evaluated in real time by the adaptive filtering algorithm as a control parameter. This enables the system to automatically suppress complex sound field effects in noisy environments to ensure the clarity of the core audio, and to fully release the rendering potential in quiet environments to provide a deep sense of immersion, thus achieving intelligent adaptation of sound field performance to the auditory environment. By establishing a negative correlation mapping relationship between noise suppression level and rendering intensity, a synergistic balance mechanism is established between signal purification and effect enhancement, avoiding muddy sound quality caused by over-rendering in noisy backgrounds and a thin spatial experience caused by insufficient rendering under clean signals, thereby optimizing the overall listening quality under different signal-to-noise ratio conditions.
[0086] In one possible implementation of the above embodiments, the signal processing module 102 is configured as follows: Based on the spectral characteristics of the audio signal obtained by the Fast Fourier Transform algorithm, the frequency-related parameters in the reverberation processing are adjusted to apply differentiated reverberation effects to audio signals of different frequency bands.
[0087] In this embodiment, the signal processing module 102 can also be configured to implement an intelligent reverberation parameter adaptation strategy based on the spectral characteristics of audio content.
[0088] The core of this strategy lies in using the frequency composition of the audio signal itself revealed by frequency domain analysis to guide the reverberation algorithm to perform refined and differentiated processing, so that the generated reverberation effect is more in line with the physical characteristics of sound and listening requirements.
[0089] Specifically, the signal processing module 102 can continuously analyze the currently processed audio signal in real time using the FFT algorithm to obtain its spectral characteristics. These characteristics may include: energy distribution, spectral centroid, and energy ratio of a specific frequency band.
[0090] Among these, obtaining energy distribution can refer to identifying the frequency bands where the signal energy is mainly concentrated; obtaining the spectral centroid can refer to calculating the energy-weighted average frequency of the signal spectrum to determine whether the overall sound is thick or bright; obtaining the energy ratio of a specific frequency band, for example, calculating the ratio of low-frequency (e.g., 20-200Hz) energy to the total frequency band energy, or the prominence of mid-frequency (e.g., 300Hz-3kHz) energy.
[0091] Furthermore, based on the aforementioned real-time extracted spectral features, the signal processing module 102 dynamically adjusts the frequency-related parameters within the reverberation algorithm, rather than using a uniform reverberation setting across the entire frequency band. This differentiated adjustment can be manifested as follows: For low-frequency components, when significant low-frequency energy is detected in the audio signal (such as explosions in movies or drumbeats in music), the system can automatically extend the low-frequency decay time in the reverberation time and may enhance the reflection intensity of low frequencies. This simulates the natural characteristics of large physical spaces where low-frequency sound waves are absorbed less and retained for a longer time, thereby enhancing the warmth, richness, and spatial grandeur of the sound field.
[0092] For mid-frequency components, when the signal is identified as predominantly mid-frequency (such as speech, the tonic range of most musical instruments), the system can finely control the reverberation time and early reflection density of the mid-frequency components, aiming to maintain the clarity and intelligibility of the main sound while adding appropriate spatial atmosphere. For example, for speech, a shorter reverberation time and sparser early reflections may be used.
[0093] For high-frequency components, when the signal contains rich high-frequency details (such as cymbals, bell sounds), the system can appropriately shorten the high-frequency reverberation decay time and may reduce the high-frequency reverberation level. This helps prevent excessive high-frequency reverberation from accumulating and causing the sound to sound harsh, noisy, or lacking in detail, ensuring that the overall sound field is clear and transparent.
[0094] In addition, to avoid sudden changes in reverberation parameters caused by instantaneous switching of audio content, the system can introduce smoothing filters or gradation algorithms to make parameter adjustments transition naturally and ensure the continuity of the listening experience.
[0095] The MCU-based sound field control system and method of the above embodiments of this disclosure dynamically and differentially adjust the frequency-related parameters in the reverberation processing based on the spectral characteristics of the audio signal obtained by real-time FTT analysis. This allows the generated reverberation effect to intelligently adapt to the physical characteristics of different audio content, significantly improving the realism and naturalness of the simulated acoustic environment. By specifically managing the reverberation energy and time of different frequency bands, the system enhances the sense of spatial atmosphere while effectively avoiding problems such as muddiness or loss of detail that may occur due to uniform reverberation processing across the entire frequency band, thus optimizing the clarity, layering, and listening experience of the overall sound quality.
[0096] In one possible implementation of the above embodiments, the human-computer interaction module 104 is used to receive an acoustic environment scene command selected by the user; The signal processing module 102 is configured to: synchronously adjust the parameter set of at least two of the processing processes, namely virtual surround sound processing, spatial positioning algorithm and reverberation processing, according to the acoustic environment scene selected by the user, so as to simulate the sound field characteristics of the selected acoustic environment as a whole.
[0097] In this embodiment, user commands can trigger the coordinated configuration of parameters of multiple core audio processing algorithms, thereby efficiently and accurately reconstructing the overall auditory characteristics of the target acoustic space.
[0098] Specifically, the human-computer interaction module 104 provides an intuitive interface (such as a touchscreen menu, mobile app options, or a voice command list) for the user to select a target acoustic environment scene. Common scene commands may include, but are not limited to, "concert hall," "cinema," "recording studio," "church," "jazz club," or "stadium." After receiving the user's selection, this module sends the corresponding scene identifier to the signal processing module 102.
[0099] The signal processing module 102 has a pre-stored or dynamically loaded scene-parameter mapping database. For each acoustic environment scene identifier, this database is associated with a preset set of parameters. This set of parameters simultaneously covers and optimizes the key parameters of at least two core processing steps in virtual surround sound processing, spatial positioning algorithms, and reverberation processing, ensuring that they work in a coordinated manner to jointly approximate the physical characteristics of the target sound field.
[0100] Based on the received scene identifier, the signal processing module 102 can retrieve the corresponding collaborative parameter set from the database and simultaneously configure the parameters of the relevant algorithms in one go. For example: When a user selects the "Cathedral" scene, the system may simultaneously: enhance the sound field width and ceiling reflection simulation of the virtual surround sound processing; adjust the rendering weight of the spatial positioning algorithm on the sense of vertical direction; significantly extend the decay time of the reverberation processing across the entire frequency range (especially the low frequencies) and increase the density of later reflected sound. These adjustments work together to create a spacious overall listening experience.
[0101] When a user selects the "small recording studio" scene, the system may simultaneously: tighten the sound field range of the virtual surround sound processing to make it closer to the listening experience of "near-field monitoring"; optimize the spatial positioning algorithm to highlight the performance of direct sound; significantly shorten the reverberation time and adopt drier, more direct early reflection characteristics. These adjustments work together to produce a clear, highly detailed acoustic environment.
[0102] The MCU-based sound field control system and method of the above embodiments of this disclosure receive high-level scene commands from users and simultaneously adjust the set of coordinated parameters for at least two processes in virtual surround sound processing, spatial positioning algorithms, and reverberation processing according to these commands. This simplifies the originally complex and professional multi-parameter optimization process into a single operation, greatly improving the convenience of user operation and the ease of use of the system. By abstracting sound field control into scene modes oriented towards auditory experience, the technical threshold for users is lowered, enabling non-professional users to obtain high-quality immersive sound field effects with a single click, significantly improving the universality of the product and user experience satisfaction.
[0103] Further reference Figure 2 , Figure 2 This is a flowchart illustrating a sound field control method based on an MCU provided in this disclosure embodiment, applied to the above-mentioned... Figure 1 In any of the MCU-based sound field control systems 100 shown, the method may include the following steps: Step S201: Acquire ambient sound signals and / or audio signals from external audio devices through the audio input module.
[0104] Step S202: The signal processing module, with the MCU as the core, performs real-time analysis and processing on the audio signal collected by the audio input module to implement sound field control. The real-time analysis and processing includes: performing frequency domain analysis on the audio signal through the signal processing module, performing adaptive filtering based on the analysis results to suppress noise, and performing spatial sound field effect processing on the noise-reduced audio signal based on user location information or preset scene parameters to generate an output signal that conforms to the target sound field effect.
[0105] Step S203: Output the processed audio signal through the audio output module.
[0106] Step S204: Receive user commands and provide feedback on system status through the human-computer interaction module.
[0107] Step S205: Power is supplied to each module in the system through the power module.
[0108] The MCU-based sound field control system and method disclosed above significantly reduce system hardware complexity and overall cost by using a low-cost MCU as the system processing core and integrating frequency domain analysis, adaptive filtering, and spatial sound field effect processing. This ensures high-performance sound field control while maintaining high-performance sound field control functionality. Real-time frequency domain analysis of the audio signal followed by adaptive filtering based on the analysis results dynamically suppresses interference according to changes in environmental noise, effectively improving the clarity and purity of the audio signal. Spatial sound field effect processing is applied to the denoised audio signal based on real-time user location information or preset scene parameters, dynamically adjusting and optimizing the sound field output to provide users with an immersive audio experience tailored to their location and needs.
[0109] In one possible implementation of the above embodiments, the signal processing module performs the following steps: Frequency domain analysis of audio signals is performed using the Fast Fourier Transform algorithm; Based on the results of the frequency domain analysis, an adaptive filtering algorithm is invoked to filter the audio signal; The sound field rendering parameters are determined based on the user's location information, and virtual surround sound processing and reverberation processing are performed on the filtered audio signal according to the sound field rendering parameters.
[0110] In one possible implementation of the above embodiments, the signal processing module performs the following steps: Based on the user location information obtained by the spatial positioning algorithm, the channel gain balance in the virtual surround sound processing and / or the reverberation time in the reverberation processing are dynamically adjusted so that the sound field center perceived by the user is relatively fixed with respect to the user's position.
[0111] In one possible implementation of the above embodiments, the signal processing module performs the following steps: The degree of noise suppression by the adaptive filtering algorithm is used as a control parameter input into the virtual surround sound processing and / or reverberation processing. When the noise suppression degree is high, the rendering intensity of the corresponding processing is enhanced; when the noise suppression degree is low, the rendering intensity of the corresponding processing is weakened.
[0112] In one possible implementation of the above embodiments, the signal processing module performs the following steps: Based on the spectral characteristics of the audio signal obtained by the Fast Fourier Transform algorithm, the frequency-related parameters in the reverberation processing are adjusted to apply differentiated reverberation effects to audio signals of different frequency bands.
[0113] In one possible implementation of the above embodiments, the acoustic environment scene command selected by the user is received through the human-computer interaction module; The signal processing module performs the following steps: Based on the acoustic environment scene selected by the user, it synchronously adjusts the parameter set of at least two of the processing processes in virtual surround sound processing, spatial positioning algorithm and reverberation processing to simulate the sound field characteristics of the selected acoustic environment as a whole.
[0114] It should be noted that the MCU-based sound field control system provided in the above embodiments is only illustrated by the division of the above program modules when implementing the corresponding MCU-based sound field control method. In practical applications, the above processing can be assigned to different program modules as needed, that is, the internal structure of the above system can be divided into different program modules to complete all or part of the processing described above. In addition, the system provided in the above embodiments and the corresponding Figure 2 The embodiments of the methods shown belong to the same concept, and their specific implementation process can be found in the method embodiments, which will not be repeated here.
[0115] Figure 3 This is a schematic diagram of the structure of an electronic device according to an embodiment of the present disclosure.
[0116] The following is a detailed reference. Figure 3 The diagram illustrates a structural schematic suitable for implementing an electronic device according to embodiments of the present disclosure. The electronic device may include a processor (e.g., a central processing unit, graphics processor, etc.) 301, which can perform various appropriate actions and processes based on a program stored in read-only memory (ROM) 302 or a program loaded from memory 308 into random access memory (RAM) 303. The RAM 303 also stores various programs and data required for the operation of the electronic device. The processor 301, ROM 302, and RAM 303 are interconnected via a bus 304. An input / output (I / O) interface 305 is also connected to the bus 304.
[0117] Typically, the following devices can be connected to I / O interface 305: input devices 306 including, for example, touchscreens, touchpads, keyboards, mice, cameras, microphones, accelerometers, gyroscopes, etc.; output devices 307 including, for example, liquid crystal displays (LCDs), speakers, vibrators, etc.; memory devices 308 including, for example, magnetic tapes, hard disks, etc.; and communication devices 309. Communication device 309 allows electronic devices to communicate wirelessly or wiredly with other devices to exchange data. Although Figure 3 Electronic devices with various devices are shown, but it should be understood that it is not required to implement or have all of the devices shown, and more or fewer devices may be implemented or have instead.
[0118] In particular, according to embodiments of this disclosure, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, embodiments of this disclosure include a computer program product comprising a computer program carried on a non-transitory computer-readable medium, the computer program containing program code for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via communication device 309, or installed from memory 308, or installed from ROM 302. When the computer program is executed by processor 301, it performs the functions defined in the network data stream hardware offloading method for heterogeneous descriptor unified processing of embodiments of this disclosure.
[0119] Figure 3 The electronic device shown is merely an example and should not be construed as limiting the functionality and scope of the embodiments disclosed herein.
[0120] This disclosure also provides a computer-readable storage medium in which the methods described in this disclosure can be implemented in hardware or firmware, or implemented as recordable on a storage medium, or implemented as computer code downloaded over a network and originally stored on a remote storage medium or a non-transitory machine-readable storage medium and subsequently stored on a local storage medium. Thus, the methods described herein can be processed by software stored on a storage medium using a general-purpose computer, a dedicated processor, or programmable or dedicated hardware. The storage medium can be a magnetic disk, optical disk, read-only memory, random access memory, flash memory, hard disk, or solid-state drive, etc.; further, the storage medium can also include combinations of the above types of memory. It is understood that the computer, processor, microprocessor controller, or programmable hardware includes storage components capable of storing or receiving software or computer code. When the software or computer code is accessed and executed by the computer, processor, or hardware, the network data stream hardware offloading method for unified processing of heterogeneous descriptors shown in the above embodiments is implemented.
[0121] A portion of this disclosure can be applied to computer program products, such as computer program instructions, which, when executed by a computer, can invoke or provide methods and / or technical solutions according to this disclosure through the operation of the computer. Those skilled in the art will understand that the forms in which computer program instructions exist in a computer-readable medium include, but are not limited to, source files, executable files, and installation package files. Accordingly, the ways in which computer program instructions are executed by a computer include, but are not limited to: the computer directly executing the instructions; the computer compiling the instructions and then executing the corresponding compiled program; the computer reading and executing the instructions; or the computer reading and installing the instructions and then executing the corresponding installed program. Here, the computer-readable medium can be any available computer-readable storage medium or communication medium accessible to a computer.
[0122] Although embodiments of the present disclosure have been described in conjunction with the accompanying drawings, those skilled in the art can make various modifications and variations without departing from the spirit and scope of the present disclosure, and such modifications and variations all fall within the scope defined by the appended claims.
Claims
1. A sound field control system based on an MCU, characterized in that, The system includes: an audio input module, a signal processing module, an audio output module, a human-computer interaction module, and a power supply module, wherein: The audio input module is used to collect ambient sound signals and / or audio signals from external audio devices; The signal processing module, with an MCU as its core, is used to perform real-time analysis and processing of the audio signals acquired by the audio input module to implement sound field control. The signal processing module is configured to perform the following: frequency domain analysis of the audio signal, adaptive filtering based on the analysis results to suppress noise, and spatial sound field effect processing of the noise-reduced audio signal based on user location information or preset scene parameters to generate an output signal that conforms to the target sound field effect. The audio output module is connected to the signal processing module and is used to output the processed audio signal. The human-computer interaction module is connected to the signal processing module and is used to receive user commands and provide feedback on the system status. The power module is used to supply power to the various modules within the system.
2. The system according to claim 1, characterized in that, The signal processing module is configured as follows: Frequency domain analysis of audio signals is performed using the Fast Fourier Transform algorithm; Based on the results of the frequency domain analysis, an adaptive filtering algorithm is invoked to filter the audio signal; The sound field rendering parameters are determined based on the user's location information, and virtual surround sound processing and reverberation processing are performed on the filtered audio signal according to the sound field rendering parameters.
3. The system according to claim 2, characterized in that, The signal processing module is configured as follows: Based on the user location information obtained by the spatial positioning algorithm, the channel gain balance in the virtual surround sound processing and / or the reverberation time in the reverberation processing are dynamically adjusted so that the sound field center perceived by the user is relatively fixed with the user's position.
4. The system according to claim 2, characterized in that, The signal processing module is configured as follows: The degree of noise suppression by the adaptive filtering algorithm is used as a control parameter input into the virtual surround sound processing and / or reverberation processing. When the noise suppression degree is high, the rendering intensity of the corresponding processing is enhanced. When the noise suppression level is low, the rendering intensity of the corresponding processing is reduced.
5. The system according to claim 2, characterized in that, The signal processing module is configured as follows: Based on the spectral characteristics of the audio signal obtained by the Fast Fourier Transform algorithm, the frequency-related parameters in the reverberation processing are adjusted to apply differentiated reverberation effects to audio signals of different frequency bands.
6. The system according to any one of claims 2-5, characterized in that, The human-computer interaction module is used to receive acoustic environment scene commands selected by the user. The signal processing module is configured to: synchronously adjust the parameter set of at least two of the processing processes, namely virtual surround sound processing, spatial positioning algorithm and reverberation processing, according to the acoustic environment scene selected by the user, so as to simulate the sound field characteristics of the selected acoustic environment as a whole.
7. A sound field control method based on an MCU, applied in the sound field control system based on an MCU as described in any one of claims 1-6, characterized in that, The method includes: The audio input module acquires ambient sound signals and / or audio signals from external audio devices. The signal processing module, with the MCU as its core, performs real-time analysis and processing on the audio signals acquired by the audio input module to implement sound field control. The real-time analysis and processing includes: performing frequency domain analysis on the audio signals through the signal processing module, performing adaptive filtering based on the analysis results to suppress noise, and performing spatial sound field effect processing on the noise-reduced audio signals based on user location information or preset scene parameters to generate an output signal that conforms to the target sound field effect. The processed audio signal is output through the audio output module; The system receives user commands and provides feedback on system status through the human-computer interaction module. The power supply module provides power to all modules within the system.
8. An electronic device, characterized in that, include: Memory, used to store computer programs; A processor, configured to implement the steps of the MCU-based sound field control method as described in claim 7 when executing the computer program.
9. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program, wherein when the computer program is executed by a processor, it implements the steps of the MCU-based sound field control method as described in claim 7.
10. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by the processor, it implements the steps of the MCU-based sound field control method as described in claim 7.