An audio processing system and method based on active regulation of sound field
By generating mixed acoustic signals and constructing a three-dimensional acoustic map, the problem of the inability to perceive changes in the acoustic environment in real time in existing technologies is solved, and adaptive audio control is achieved, improving the accuracy of environmental perception and the personalization effect of the audio experience.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- SHENZHEN HUAJUXIN SEMICON CO LTD
- Filing Date
- 2026-02-28
- Publication Date
- 2026-06-02
AI Technical Summary
Existing technologies cannot accurately sense changes in the acoustic environment in real time and make adaptive adjustments, resulting in an inability to provide a high-fidelity and personalized audio experience in complex listening environments.
By generating a mixed acoustic signal containing audible audio signals and ultrasonic detection signals, and using a sound transmitting unit and a sound receiving unit to perform real-time environmental detection, a three-dimensional acoustic map is constructed, a room impulse response is generated, and an output audio signal is generated based on the environmental noise field to achieve adaptive control.
It enables real-time and accurate perception and adaptive adjustment in complex acoustic environments, improving the environmental perception accuracy of the audio experience and providing high-fidelity and personalized audio output.
Smart Images

Figure CN122138117A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of audio processing technology, and more specifically, to an audio processing system and method based on active sound field modulation. Background Technology
[0002] With the rapid development of smart homes, smart cockpits, remote conferencing, and immersive entertainment (VR / AR), users' expectations for audio experiences have shifted from simply being "audible" to "high fidelity," "immersion," and "personalization." However, real-world listening environments, especially family living rooms or open-plan offices, are extremely complex acoustic spaces filled with reverberation, standing waves, and dynamic noise.
[0003] In traditional acoustics, to ensure a "uniform" listening experience for listeners moving through a space, the mainstream approach is to deploy numerous loudspeakers in the space (such as on the ceiling and walls) and create a "uniform" sound field through sophisticated acoustic design and equalization (EQ). This approach is complex to deploy, expensive, and cannot adapt to dynamic changes (such as furniture movement or the appearance of noise).
[0004] For example, Chinese patent CN120390183B discloses an audio processing method applied to bathroom audio systems. It extracts "residual noise signals" through a microphone and "predicts bathroom spatial echo signals," then compensates for the audio by obtaining "dynamic gain" based on an "auditory masking threshold." This patent document attempts to "predict" the echo, but "prediction" is difficult, and the prediction fails once the environment (such as movement) changes. It lacks a means of "real-time measurement" of the room's RIR.
[0005] Furthermore, the "dynamic gain" in the aforementioned patent document is a form of "compensation"—that is, "using a larger target sound to cover up noise and echoes." This leads to an increase in total sound pressure level, auditory fatigue, and fails to fundamentally "eliminate" the "muddyness" of echoes and the "interference" of noise. Summary of the Invention
[0006] This invention provides an audio processing system and method based on active sound field control, which at least solves the problem in related technologies that cannot accurately sense changes in the acoustic environment in real time and adaptively control them according to these changes.
[0007] According to an embodiment of the present invention, an audio processing method based on active sound field modulation is provided, comprising: A mixed acoustic signal is generated and transmitted through a sound transmission unit. The mixed acoustic signal includes an audible audio signal in a first frequency band and an ultrasonic detection signal in a second frequency band, wherein the frequency of the second frequency band is higher than that of the first frequency band. The composite frequency band hybrid acoustic signal is formed by the propagation of the hybrid acoustic signal in the acoustic space through the synchronous acquisition of the hybrid acoustic signal by the sound receiving unit. The composite frequency band mixed acoustic signal is separated into ultrasonic echo stream and audible ambient stream; A three-dimensional acoustic map is constructed based on the ultrasonic echo flow, and a room impulse response from the sound delivery unit to a preset listener position is generated based on the three-dimensional acoustic map, wherein the three-dimensional acoustic map is used to indicate the acoustic characteristics of the acoustic space; Based on the audible environmental flow, determine the environmental noise field; Based on the room impulse response and the ambient noise field, an output audio signal is generated and played through the sound transmission unit.
[0008] In one exemplary embodiment, generating the mixed acoustic signal includes: Based on the audible audio signal, predict the vibration displacement of the transducer diaphragm in the sound transmission unit; Based on the vibration displacement, the ultrasonic detection signal is pre-distorted to generate a pre-compensated ultrasonic signal to counteract the nonlinear distortion introduced by the vibration displacement on the propagation of the ultrasonic detection signal. The audible audio signal and the pre-compensated ultrasonic signal are combined to form the mixed acoustic signal.
[0009] In one exemplary embodiment, constructing a three-dimensional acoustic map based on the ultrasonic echo flow includes: The ultrasonic echo stream is subjected to time-domain analysis to obtain the three-dimensional geometric information of the object in the acoustic space, and a geometric point cloud map is constructed based on the three-dimensional geometric information and the time-domain analysis results. Energy attenuation and dispersion characteristics of the ultrasonic echo stream are analyzed to determine the surface acoustic material of the object in the acoustic space. The acoustic material information is fused with the geometric point cloud map to obtain the three-dimensional acoustic map.
[0010] In one exemplary embodiment, generating a room impulse response from the sound delivery unit to a preset listener position based on the three-dimensional acoustic map includes: Based on the geometric information in the three-dimensional acoustic map and the sound absorption coefficient corresponding to the acoustic material, the early reflected sound response in the mid-to-high frequency band is calculated. Based on the main dimensions of the acoustic space extracted from the three-dimensional acoustic map, the standing wave mode response in the low-frequency band is calculated. The early reflected acoustic response is spliced with the standing wave mode response to generate the full-frequency room impulse response.
[0011] In one exemplary embodiment, determining the ambient noise field based on the audible ambient flow includes: Acoustic echo cancellation processing is performed on the audible ambient flow to obtain a residual ambient noise signal; Sound source localization is performed based on the residual environmental noise signal to determine the location of one or more external noise sources.
[0012] In one exemplary embodiment, generating the output audio signal based on the room impulse response and the ambient noise field includes: Calculate the inverse filter of the room impulse response, and generate an audio signal based on the inverse filter and the audible audio signal; Based on the location of the external noise source and the preset audience location, an anti-phase noise signal is generated, which is used to cancel the external noise at the preset audience location. The audio signal is superimposed with the inverted noise signal to determine the output audio signal.
[0013] In one exemplary embodiment, the step of generating an anti-phase noise signal for canceling external noise at the preset listener position includes: Based on the three-dimensional acoustic map, a first noise transmission path from the location of the microphone unit to the preset listener location and a second noise transmission path from the location of the external noise source to the preset listener location are determined. Based on the residual environmental noise signal collected by the sound receiving unit, and using the difference or mapping relationship of the transfer function between the first noise transmission path and the second noise transmission path, the virtual error noise signal at the preset listener position is determined. An adaptive filtering algorithm is used to generate the inverse noise signal with the goal of minimizing the virtual error noise signal.
[0014] In one exemplary embodiment, the method further includes: Based on the Doppler frequency shift information or inter-frame difference information in the ultrasonic echo stream, the dynamic changes of the preset audience position are tracked in real time. The room impulse response is updated in real time based on the dynamic changes.
[0015] According to another embodiment of the present invention, an audio processing system based on active sound field modulation is provided, comprising: Audio delivery unit; Radio unit; A processing unit is configured to generate a mixed acoustic signal comprising an audible audio signal in a first frequency band and an ultrasonic detection signal in a second frequency band, wherein the frequency of the second frequency band is higher than that of the first frequency band, and to control the transmitting unit to transmit the mixed acoustic signal; to synchronously acquire, via a receiving unit, a composite frequency band mixed acoustic signal formed after the mixed acoustic signal propagates within the acoustic space; to separate the composite frequency band mixed acoustic signal into an ultrasonic echo stream and an audible ambient stream; to construct a three-dimensional acoustic map based on the ultrasonic echo stream, and to generate a room impulse response from the transmitting unit to a preset listener position based on the three-dimensional acoustic map, wherein the three-dimensional acoustic map is used to indicate the acoustic characteristics of the acoustic space; to determine an ambient noise field based on the audible ambient stream; and to generate an output audio signal based on the room impulse response and the ambient noise field, and to play it through the transmitting unit.
[0016] According to yet another embodiment of the present invention, a computer-readable storage medium is also provided, wherein a computer program is stored therein, wherein the computer program is configured to perform the steps in any of the above method embodiments when executed.
[0017] This invention integrates the sound transmitting unit and the sound receiving unit into the same device, and uses composite frequency band signals to simultaneously complete music playback and environmental detection. Therefore, without the need for an external microphone or camera, it can perceive the room's geometry, material properties, and the listener's position in real time, and achieve adaptive adjustment according to environmental changes. Thus, it can solve the problem in related technologies that cannot accurately perceive changes in the acoustic environment in real time and make adaptive adjustments according to environmental changes, thereby improving the accuracy of environmental perception. Attached Figure Description
[0018] Figure 1 This is a flowchart of an audio processing method based on active sound field modulation according to an embodiment of the present invention; Figure 2 This is a structural block diagram of an audio processing system based on active sound field modulation according to an embodiment of the present invention; Figure 3 This is a specific structural block diagram of an audio processing system based on active sound field modulation according to an embodiment of the present invention. Detailed Implementation
[0019] The technical solutions of the embodiments of this application will be described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments.
[0020] In the following description, the terms "first," "second," etc., are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of technical features indicated. Therefore, a feature defined with "first," "second," etc., may explicitly or implicitly include one or more of that feature. In the description of this application, unless otherwise stated, "a plurality of" means two or more.
[0021] Furthermore, in this application, directional terms such as "upper," "lower," "left," and "right" may be defined relative to the orientation of the components shown in the accompanying drawings. It should be understood that these directional terms can be relative concepts, used for relative description and clarification, and may change accordingly depending on the orientation of the components in the accompanying drawings.
[0022] In this application, unless otherwise expressly specified and limited, the term "connection" should be interpreted broadly. For example, "connection" can be a fixed connection, a detachable connection, or an integral part; it can be a direct connection or an indirect connection through an intermediate medium. Furthermore, the term "coupled" can refer to an electrical connection that enables signal transmission.
[0023] As used herein, “about,” “approximately,” or “approximately” includes the stated value and the average value within an acceptable range of deviation from the given value, wherein the acceptable range of deviation is determined by a person skilled in the art taking into account the measurement under discussion and the error associated with the measurement of the given quantity (i.e., the limitations of the measurement system).
[0024] Example 1 like Figure 1 As shown, this embodiment provides an audio processing method based on active sound field modulation, which can be applied to... Figure 2 The audio processing system shown is implemented as an integrated smart speaker device. The system includes a sound transmission unit 111 consisting of M speakers and a sound receiving unit 112 consisting of K microphones, as well as a processing unit with a built-in heterogeneous computing architecture. The processing unit achieves adaptive optimization of audio playback and active suppression of environmental noise by actively detecting and modeling the acoustic space.
[0025] Specifically, it includes the following steps: S100: Generates a mixed acoustic signal and transmits the mixed acoustic signal through the sound transmission unit, wherein the mixed acoustic signal includes an audible audio signal in a first frequency band and an ultrasonic detection signal in a second frequency band, the frequency of the second frequency band being higher than that of the first frequency band.
[0026] In this embodiment, a mixed acoustic signal is generated by an audio digital signal processor (DSP) 122, with the first frequency band being the range of human hearing, which can be set here. to The second frequency band is the ultrasonic frequency band outside the range of human hearing. Considering air attenuation and detection resolution, it can be set here to... to Audible audio signals are what the user expects to hear, such as music or movie dialogue; ultrasonic detection signals are used to detect the acoustic characteristics of space, and this signal can be a linear frequency modulated signal with a frequency between arrive The relationship changes linearly with time.
[0027] In practical applications, the speaker diaphragm will generate significant low-frequency vibrations when the speaker unit 111 (speaker) simultaneously plays strong low-frequency audible audio (e.g., drum beats in music) and high-frequency ultrasonic detection signals. This physical displacement will produce a nonlinear modulation effect on the ultrasonic signal to be emitted, the most significant of which is the Doppler effect, causing a shift in the frequency of the emitted ultrasonic wave, as well as intermodulation distortion (IMD), generating additional stray frequency components. These distortions will severely interfere with the subsequent analysis of ultrasonic echoes and reduce the accuracy of spatial perception. To solve this problem, a nonlinear pre-distortion processing module can be integrated for processing.
[0028] The pre-distortion process can be further divided into two sub-steps, S110 and S120.
[0029] S110: Based on audible audio signals, predict the vibration displacement of the transducer diaphragm in the sound delivery unit.
[0030] The audio digital signal processor 122 internally stores the physical model of the loudspeaker. To accurately predict diaphragm behavior under large dynamic ranges, this embodiment employs a nonlinear system model based on Volterra series or Klippel nonlinear parameters (such as the variation curves of the force factor Bl(x) and suspension compliance Cms(x)); the model incorporates the input low-frequency audio signal voltage x. low (n) serves as the excitation, outputting a real-time prediction of the loudspeaker diaphragm displacement d(n). Unlike traditional linear models, this model not only characterizes the basic motor transitions but also encompasses the influence of nonlinear parameters that vary with displacement on the vibration.
[0031] For example, based on truncated discrete Volterra series, the input-output relationship of a loudspeaker system can be expressed as a superposition of linear terms and higher-order nonlinear terms: Where, x lowThe input audio signal is d[n], the predicted displacement sequence is d[n], h1 is a linear first-order convolution kernel (corresponding to the conventional impulse response), h2 is a second-order nonlinear convolution kernel (used to characterize intermodulation distortion and second harmonic distortion), and so on. The audio digital signal processor 122 can accurately reconstruct the true motion trajectory of the diaphragm by parallel computing the above multidimensional convolutions.
[0032] For example, when a large-amplitude sine wave with a frequency of 50Hz is input... low At time (t), the linear model can only predict a displacement of 50 Hz; however, the displacement sequence d[n] calculated by the nonlinear model in this embodiment not only includes the fundamental component of 50 Hz, but also the harmonic components of 100 Hz, 150 Hz, etc. contributed by h2 and h3 terms, as well as the corresponding phase shift. This d[n]0 containing precise nonlinear characteristics will be sent to the next step to calculate the precise time delay that needs to be compensated, and so on.
[0033] S120: Based on the vibration displacement, the ultrasonic detection signal is pre-distorted to generate a pre-compensated ultrasonic signal.
[0034] The purpose of this step is to analyze the raw ultrasonic detection signal. Inverse modulation is performed so that, after being modulated by the physical movement of the speaker diaphragm, it can be precisely restored to the pure signal originally intended to be emitted. Its primary compensation objective is Doppler frequency shift. When the diaphragm moves forward (at a velocity of...),... It compresses the sound waves in front, causing the frequency to increase; conversely, it causes the frequency to decrease. This change in frequency... Approximately equal to ,in It is an ultrasonic frequency. It is the speed of sound (approximately) Therefore, predistortion processing can be achieved by modulating the phase of the original ultrasonic signal.
[0035] For example, based on the foregoing example, the diaphragm velocity is Assuming the original ultrasonic signal is a single-frequency signal At this point, the predistortion module calculates a compensation time delay. That is, when the diaphragm moves forward to shorten the sound path, the electrical signal should be delayed accordingly; then, a pre-compensated signal is generated: Will Substituting the expression, we get When this predistortion signal Drive a being displaced When the diaphragm vibrates, the actual phase of the sound wave radiated into the air experiences an additional term of +k·d(t) (phase lead due to path shortening). This additional term precisely cancels out the -k·d(t) term introduced by predistortion (phase lag due to the -d(t) / c time delay). Wave number This additional feature will precisely offset the effects introduced by predistortion. This process ultimately restores the ultrasound waves generated in the far field to their pure form. .
[0036] Finally, the DSP 122 will output the audible audio signal. With pre-compensated ultrasonic signals The signals are superimposed in the digital domain to form a hybrid digital signal, which drives the sound transmission unit 111 to transmit the signal through a digital-to-analog converter (DAC) and a power amplifier (integrated in the driver circuit 113).
[0037] S200: Synchronously acquires composite frequency band mixed acoustic signals formed after the mixed acoustic signals propagate in the acoustic space through the radio unit.
[0038] In this embodiment, the mixed acoustic signal emitted by the sound transmitting unit 111 propagates within the room, undergoing reflection, scattering, and absorption by walls, furniture, and human bodies, and is ultimately captured by the sound receiving unit 112 (microphone array); in order to collect high-resolution acoustic signals without distortion... According to the Nyquist theorem, the sampling rate of the ultrasonic signal must be at least twice that of the signal. Therefore, this system employs a high-sampling-rate analog-to-digital converter (ADC), whose sampling rate is strictly locked at [value missing]. The sound receiving unit 112 consists of K (exemplarily, K>8) MEMS microphones arranged in a non-uniform three-dimensional array to achieve three-dimensional sound source localization capability. All K channels of the ADC are driven by a synchronous clock to ensure strict temporal alignment of the acquired multi-channel data. The acquired raw data stream is a composite frequency band mixed acoustic signal, which is a K-channel signal with a sampling rate of [missing information]. , bit depth is Digital signal stream.
[0039] S300: Separates composite frequency band mixed acoustic signals into ultrasonic echo stream and audible ambient stream.
[0040] In this embodiment, this step is performed at the hardware level to achieve efficient data distribution. Specifically, the data input from the ADC... The mixed signal stream first enters the hardware splitter module, which contains two sets of parallel digital filters.
[0041] The first group is a bandpass filter whose passband range matches the frequency band of the emitted ultrasonic detection signal, for example, to After the mixed signal stream passes through this filter, all audible frequency signals and high-frequency noise are filtered out, retaining only the echo information of the ultrasound, thus obtaining the ultrasound echo stream, which still maintains... The sampling rate is directly routed to the sonar coprocessor 123 (based on NPU / FPGA) dedicated to high-density computing.
[0042] The second group is a low-pass filter whose cutoff frequency is the upper limit of audible audio; for example, it can be set to... After the mixed signal stream passes through this filter, the ultrasonic component is completely filtered out, leaving only the signal in the audible frequency range. Since the highest frequency of this signal does not exceed [a certain value], [the remaining signal is omitted as it is not directly related to the filter's description]. To reduce the computational burden of subsequent processing, the signal can be downsampled; for example, the signal can be downsampled to [value missing]. This output stream, the audible ambient stream, is transmitted to the audio DSP 122 for subsequent sound field analysis. Through this hardware-level splitting and preprocessing, the system allocates computational tasks of different natures to the most suitable processing units before the data enters the processor core, maximizing the efficiency of the heterogeneous computing architecture.
[0043] S400: Based on the ultrasonic echo flow, a three-dimensional acoustic map characterizing the acoustic properties of the acoustic space is constructed, and a room impulse response from the sound delivery unit to the preset listener position is generated based on the three-dimensional acoustic map.
[0044] In this embodiment, this step is completed collaboratively by the sonar coprocessor 123 and the central processing unit (CPU) 121 to transform the unknown physical space into a computable acoustic model, specifically including the following sub-steps: S410: Construct a three-dimensional acoustic map.
[0045] This sub-step is executed on the sonar coprocessor 123 and includes the following process: S411: Construct a geometric point cloud map.
[0046] The coprocessor 123 first performs time-of-flight (ToF) analysis on the ultrasonic echo to determine the transmission time of the ultrasonic detection signal (e.g., the chirp signal). For each received echo peak, the system determines its arrival time by performing matched filtering with the transmitted signal. The total path length from the transmitting unit to the reflector and then to the receiving unit is At this point, by analyzing the time difference of arrival (TDOA) of the echoes received by the K microphones, and combining this with the known geometric configuration of the microphone array, the system can calculate the three-dimensional spatial coordinates of the reflection point corresponding to the echo. The system then continuously transmits detection signals and processes the echoes, obtaining a large number of three-dimensional spatial points. These points collectively form an original point cloud describing the surfaces of objects inside the room. To stitch together the local point clouds of consecutive frames into a complete global map, the system uses the Iterative Closest Point (ICP) algorithm to align the point clouds acquired at different times, thereby constructing a complete geometric point cloud map of the room. .
[0047] S412: Determine the acoustic material.
[0048] Besides their geometric location, different object surfaces exhibit vastly different sound wave reflection characteristics, primarily reflected in their sound absorption coefficients. For instance, hard surfaces (such as glass and tiles) reflect strongly and absorb little, while soft surfaces (such as curtains and carpets) show the opposite. To address this, the coprocessor 123 infers the material by analyzing the energy attenuation rate and dispersion characteristics of the echo signal. Energy attenuation rate: The energy of the echo signal for a single reflection. With the energy of the incident sound wave The relationship between them is ,in This is the sound absorption coefficient of the material; the coprocessor 123 can estimate the sound absorption coefficient by comparing the amplitude of the echo with the amplitude of the direct sound (or a reference). Therefore, it can be inferred that Size.
[0049] Dispersion characteristics: Different materials have different absorption capabilities for different frequencies. Soft materials usually absorb high-frequency sound waves more strongly. In response, the coprocessor 123 analyzes the spectral structure of the received echo signal by performing a short-time Fourier transform (STFT). If an echo signal has significantly attenuated high-frequency components compared to the transmitted signal, and the center of the spectrum shifts to lower frequencies, the system tends to label it as a soft material, and so on.
[0050] To achieve efficient and accurate material classification, the coprocessor 123 runs a lightweight convolutional neural network (CNN) model. The model's input is a feature vector extracted from each echo band, containing multiple dimensions such as energy attenuation, spectral centroid, and spectral spread. The model outputs the probability that the reflection point belongs to a predefined material category (such as "glass / concrete," "wood," or "fabric"). These material labels are then appended to the geometric point cloud map. At each point, the system ultimately generates a complete three-dimensional acoustic map with acoustic semantic information. .
[0051] For example, when an ultrasonic wave detects a wall, the coprocessor 123 first determines the three-dimensional coordinates of a series of points on the wall through Time-of-Flight (ToF) analysis, forming a point cloud; then, for one of the points, its echo is analyzed, and it is found that the echo energy is only slightly lower than the theoretical attenuation of a single reflection (due to distance), and the calculated sound absorption coefficient is obtained. Approximately Simultaneously, analysis of its spectrum revealed almost no difference from the spectrum of the emitted chirp signal; subsequently, the CNN model received these features and output the "glass / concrete" category with a high probability; while when the ultrasonic wave detected a curtain, its echo energy attenuated significantly, and the calculated sound absorption coefficient... Gundam Meanwhile, spectrum analysis shows The frequency components mentioned above almost completely disappear, so the CNN model classifies them as "fabric", and so on.
[0052] S420: Generate room impulse response (RIR).
[0053] In this embodiment, this sub-step is executed by CPU 121, which receives data from coprocessor 123. Using this as input, the room impulse response from each speaker of the delivery unit 111 to the listener's ear position is calculated through physical acoustic modeling. Because the propagation characteristics of sound waves differ greatly across different frequency bands, this embodiment employs a hybrid modeling strategy, specifically: Mid-to-high frequency band (>500Hz): In this frequency band, sound wave propagation can be approximated as a ray. For this, CPU 121 uses either the image source method or the ray tracing method. Taking the image source method as an example, for each reflecting surface in the room, CPU 121 calculates a "mirror source" of the sound source about that surface; the reflected sound received by the listener can be equivalent to sound propagating in a straight line from this mirror source; subsequently, CPU 121 recursively calculates higher-order image sources (secondary, tertiary reflections, etc.). For each path from the (mirror) source to the listener, its delay is determined by the path length, and its amplitude is attenuated by the inverse square of the path length and the wall absorption coefficient at each reflection (from...). The early reflection portion of the RIR is determined by the sum of the contributions of all paths (delayed pulses multiplied by the corresponding amplitude).
[0054] Low frequency band (<500Hz): In this frequency band, the wavelength of the sound wave is comparable to the room size, exhibiting significant fluctuations and forming standing waves. At this point, CPU 121, based on... Extracted room length, width, and height dimensions Modal analysis was used for analysis, where the resonant frequencies (modal frequencies) of the room were determined by the following formula: in It is a non-negative integer; finally, the CPU calculates all significant low-frequency modal frequencies in the room and their corresponding spatial distributions, thereby constructing the low-frequency part of the RIR.
[0055] Next, CPU 121 smoothly stitches the calculated high-frequency and low-frequency responses together in the frequency domain to generate a room impulse response model covering the entire frequency band. Where j represents the j-th speaker and i represents the listener. And so on.
[0056] S500: Determines the ambient noise field based on audible ambient flow.
[0057] In this embodiment, this step is performed by the audio DSP 122 to identify all external noise other than the system's own playback sound, specifically including the following sub-steps: S510: Self-echo cancellation.
[0058] The audible ambient noise stream includes both external noise (such as air conditioner noise) and echoes from the audible audio signal played by the system itself, reflected off the room. To separate out the pure noise, the DSP 122 first performs Acoustic Echo Cancellation (AEC). The AEC module uses the audible audio signal that the system is about to play as a reference signal and employs an adaptive filter (such as the NLMS algorithm) to simulate the acoustic path from the speaker to the microphone. Then, by subtracting this simulated echo signal from the signal received from the microphone, the residual ambient noise signal can be obtained. .
[0059] S520: Noise source location.
[0060] After obtaining a clean ambient noise signal, the DSP 122 uses the multi-microphone array of the microphone unit 112 to perform sound source localization (SSL), that is, by analyzing the time difference or phase difference of the noise signal arriving at different microphones to calculate the location of the noise source; commonly used algorithms include generalized cross-correlation (GCC-PHAT) or multiple signal classification (MUSIC) algorithms. Ultimately, the system locates the external noise source. three-dimensional position This will not be elaborated upon here.
[0061] S600: Based on the room impulse response and ambient noise field, it generates an output audio signal that is actively controlled by the sound field and plays it through the sound delivery unit.
[0062] In this embodiment, this step is performed on the audio DSP 122. Specifically, the DSP 122 performs active room correction (ARC) and virtual sensing active noise reduction (ANC) in parallel: S610: Active Room Correction (ARC).
[0063] The goal is to counteract the "pollution" of the original audio signal by room reverberation, allowing listeners to hear a "dry" sound as if they were in an anechoic chamber. This is achieved through the calculation of the room's impulse response. ( Inverse filter of Z-transform Theoretically However, in reality, the RIR may have deep valleys in the frequency domain, and directly inverting it would lead to excessive filter gain and ringing effect. Therefore, this embodiment uses regularized least squares to solve for a stable inverse filter: in yes conjugate, It is a small positive constant used to limit the maximum gain of the filter and prevent over-equalization. The DSP 122 compares the original, clean audible audio signal with this calculated inverse filter. Real-time convolution is performed on this pre-distorted signal as it passes through the room. After its spread, its effect is This means that the room's influence is canceled out, and the listener receives pure audio without reverberation.
[0064] S620: Virtual Sensing Active Noise Cancellation (ANC).
[0065] Traditional ANC headphones measure and cancel noise directly by placing microphones inside and outside the earcups. This system, however, uses virtual sensing technology to create a "silent zone" around the listener's head without placing any devices at the listener's location. Specifically, it estimates the virtual error signal near the listener's ear. To achieve this.
[0066] DSP122 utilizes the obtained residual noise signal mic (n) (measured at the device microphone) is used to calculate the virtual error signal. Due to noise source N... j For external locations (such as exhaust fans), directly applying the microphone-to-user transfer function ignores the propagation process from the noise source to the microphone, leading to phase errors. Therefore, the noise source location Pos(N) based on the aforementioned localization... j The acoustic responses of two critical paths are extracted from the 3D acoustic map: the impulse response (RIR) from the noise source to the receiver unit. Source→MicAnd the impulse response RIR from the noise source to the preset audience position. Source→User .
[0067] Based on these two sets of responses, the system constructs a relative transfer function filter or performs a two-step operation: first, it eliminates the RIR. Source→Mic For e mic The effect of (n) (e.g., through inverse filtering or Wiener filtering) is used to estimate the original transmitted signal at the noise source, and then this original signal is compared with the RIR. Source→User Convolution is performed to accurately infer the virtual error signal e at the listener's ear. virtua l(n). In the frequency domain, this relationship can be approximated as: .
[0068] After obtaining the objective function The system then uses the Filtered-x LMS (FxLMS) adaptive filtering algorithm to generate an inverse noise signal. During this process, the FxLMS algorithm continuously adjusts the coefficients of the noise reduction filter W, ensuring that the signal is transmitted through the speaker and then travels along the acoustic path from the speaker to the listener. (Also provided by the acoustic map) The back-phase noise after propagation can be canceled out to the greatest extent. Ultimately, the entire process is closed-loop and adaptive, creating a quiet area around the listener's head.
[0069] S630: Signal synthesis and beamforming projection.
[0070] Finally, the DSP 122 superimposes the generated ARC-corrected audio signal with the ANC phase-inverting noise signal to form the final drive signal. This signal is then precisely projected onto the listener. To ensure optimal placement and avoid disturbing other areas of the room (e.g., another family member resting), the system employs beamforming technology. Specifically, the acoustic contrast control algorithm (ACC) is used to solve for the complex weight vectors of the M speakers in a set of driving units 111. This maximizes sound energy in the target audience area while minimizing sound energy in non-target areas. in and These are the spatial correlation matrices of the target area and the silent area, respectively. These matrices can also be calculated based on the acoustic map. By solving this problem, the optimal weight vector is obtained. The DSP 122 multiplies the synthesized signal by this weight vector and drives M speakers respectively. In this way, the system creates a high-fidelity and interference-free listening "bubble" for the listener.
[0071] Example 2: like Figure 3 As shown, the hardware architecture of an audio processing system based on active sound field modulation provided in this embodiment includes: A highly integrated processing unit, in the form of a System-on-a-Chip (SoC), employs a heterogeneous computing architecture and contains three cooperating cores: Sonar coprocessor 123: Implemented based on NPU (Neural Processing Unit) or FPGA (Field Programmable Gate Array); it is specifically responsible for processing signals from the ADC. Ultrasonic echo flow in high sampling rate signals ( Its internal hardware-based parallel computing unit is well-suited for performing point cloud registration and Time-of-Flight (ToF) calculations in SLAM algorithms, as well as running CNN models for material inference; furthermore, it tracks the position of dynamic targets (the audience) in real time by analyzing the Doppler shift and inter-frame difference of the echoes. This unit will calculate a three-dimensional acoustic map. The real-time coordinates of the audience are sent to the CPU via a high-speed internal bus.
[0072] Audio Digital Signal Processor (DSP) 122: Used to receive downsampled audio signals. The audible ambient stream; the processor includes: IMD pre-compensation module: Before signal transmission, it performs pre-distortion processing on the ultrasonic signal based on the speaker model provided by the CPU.
[0073] ANI identification module: Executes the AEC algorithm to eliminate autoechoes and runs the SSL algorithm to locate noise sources.
[0074] DAC control module: Runs the FxLMS adaptive filtering algorithm of ARC inverse filter convolution and ANC in parallel to generate the final corrected audio and inverse noise.
[0075] Beamforming module: Executes the ACC algorithm and calculates the drive weights of the speaker array. The DSP's efficient instruction set and low latency ensure that all these complex audio processing tasks can be completed within each sampling cycle.
[0076] Central Processing Unit (CPU) 121: Responsible for high-level logic control and non-real-time but computationally complex physics modeling tasks; this processor includes: Hybrid RIR modeling engine: Receives acoustic maps generated by sonar coprocessor 123, runs physical models such as mirror source method and modal analysis, calculates high-precision RIR, and sends RIR model parameters (such as inverse filter coefficients) to DSP.
[0077] Model and Task Management Module: Manages and stores the acoustic map database and RIR database; when a significant change in the environment is detected (e.g., the coprocessor reports a large-scale change in the point cloud), the CPU triggers a complete environment rescan and model update.
[0078] State machine control module: manages various working states of the system, such as the initial "environmental holographic scanning" mode, the running "dynamic tracking and control" mode, and the "hardware preheating" mode in low temperature environments.
[0079] The system also includes peripheral units: The composite acoustic transceiver front end includes a transmitting unit 111 (an array of M wideband loudspeakers) and a receiving unit 112 (an array of K high signal-to-noise ratio MEMS microphones). The loudspeakers and microphones are integrated into the same device, enabling "shared aperture multiplexing" and closed-loop sensing.
[0080] The drive and acquisition circuit 113 includes a high-precision, strictly synchronized 96kHz ADC and DAC, as well as a multi-channel Class D power amplifier. This circuit also integrates a temperature sensor to monitor the speaker voice coil temperature, preventing overheating due to prolonged ultrasonic wave emission, and feeds the temperature data back to the CPU for thermal management.
[0081] The following explanation uses real-world application scenarios.
[0082] Scenario 1: Single-user, high reverberation, localized noise scenario (e.g., bathroom) Environment: User L1 is showering in a tiled bathroom with a reverberation time RT60 > 1.5s; the exhaust fan N1 in the corner is emitting a continuous low-frequency noise.
[0083] System actions: SMR+: After the system is powered on, the emitted ultrasonic waves detect extremely strong echoes from the surrounding walls, which attenuate slowly. The coprocessor 123 classifies the material as a "hard surface" with a high sound absorption coefficient. .
[0084] RIR Correction: The CPU calculates the RIR with an extremely long reverberation tail based on this material parameter and the measured room dimensions.
[0085] ARC processing: The DSP generates an "aggressive" inverse filter based on the RIR. When music is played, the DSP outputs a signal that has undergone deep "anti-reverb" preprocessing. This signal sounds very "dry," but when it propagates through the strong reverberation space of the bathroom and reaches the user's ears, the reverberation effect is precisely canceled out, restoring a normal listening experience.
[0086] Virtual ANC: Simultaneously, the DSP's ANI module locks the exhaust fan's position via SSL. The ANC module estimates the waveform of the exhaust fan noise at the user's L1 head position based on the acoustic map, and generates a precise anti-phase sound wave, which is then projected through beamforming.
[0087] Results: Users heard clear, echo-free music in the bathroom, while the noise from the exhaust fan was significantly suppressed.
[0088] Scenario 2: Multi-user, multi-content, sound field zoned scenario (e.g., living room) Environment: User L1 is watching a football game on the sofa (requires exciting sound effects), while User L2 is listening to light music at the dining table (requires quiet).
[0089] System actions: Multi-target tracking: The sonar coprocessor 123 simultaneously locks onto and continuously tracks the three-dimensional coordinates of L1 and L2. and .
[0090] Acoustic contrast control: The DSP receives two audio streams (sports game and light music). The ACC module solves for a MIMO (Multiple-Input Multiple-Output) beamforming weighting matrix. Its optimization objective is twofold: Maximize the audio stream from the game to The energy, while minimizing its energy to Energy (in) (A zero-depression is formed at the location).
[0091] Maximize the streaming of light music audio to The energy, while minimizing its energy to Energy.
[0092] Dynamic adjustment: The coprocessor updates in real time when the user (L1) gets up and moves around. The DSP recalculates the weight matrix in milliseconds. This ensures that the sound beam directed at L1 always follows his movement.
[0093] Effect: Achieved "different audio in the same room". Two people in the same open space can enjoy their own audio content without wearing headphones, without disturbing each other.
[0094] Through the above description of the embodiments, those skilled in the art can clearly understand that the methods according to the above embodiments can be implemented by means of software plus necessary general-purpose hardware platforms. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) and includes several instructions to cause a terminal device (which may be a mobile phone, computer, server, or network device, etc.) to execute the methods described in the various embodiments of the present invention.
[0095] This embodiment also provides an apparatus for implementing the above embodiments and preferred embodiments; details already described will not be repeated. As used below, the term "module" can refer to a combination of software and / or hardware that performs a predetermined function. Although the apparatus described in the following embodiments is preferably implemented in software, hardware implementation, or a combination of software and hardware, is also possible and contemplated.
[0096] Figure 2 This is a structural block diagram of an apparatus according to an embodiment of the present invention, such as... Figure 2 As shown, the device includes: It should be noted that the above modules can be implemented by software or hardware. For the latter, they can be implemented in the following ways, but are not limited to: all the above modules are located in the same processor; or, the above modules are located in different processors in any combination.
[0097] Embodiments of the present invention also provide a computer-readable storage medium storing a computer program, wherein the computer program is configured to perform the steps in any of the above method embodiments when executed.
[0098] In one exemplary embodiment, the aforementioned computer-readable storage medium may include, but is not limited to, various media capable of storing computer programs, such as a USB flash drive, read-only memory (ROM), random access memory (RAM), portable hard disk, magnetic disk, or optical disk.
[0099] Embodiments of the present invention also provide an electronic device including a memory and a processor, the memory storing a computer program and the processor being configured to run the computer program to perform the steps in any of the above method embodiments.
[0100] In one exemplary embodiment, the electronic device may further include a transmission device and an input / output device, wherein the transmission device is connected to the processor and the input / output device is connected to the processor.
[0101] Through the above description of the embodiments, those skilled in the art can clearly understand that, for the sake of convenience and brevity, only the division of the above functional modules is used as an example. In actual applications, the above functions can be assigned to different functional modules as needed, that is, the internal structure of the device can be divided into different functional modules to complete all or part of the functions described above.
[0102] In the several embodiments provided in this application, it should be understood that the disclosed apparatus and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of modules or units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another apparatus, or some features may be ignored or not executed. Furthermore, the mutual coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between apparatuses or units may be electrical, mechanical, or other forms.
[0103] The units described as separate components may or may not be physically separate. A component shown as a unit can be one or more physical units; that is, it can be located in one place or distributed in multiple different locations. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.
[0104] Furthermore, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.
[0105] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a readable storage medium. Based on this understanding, the technical solutions of the embodiments of this application, in essence, or the parts that contribute to the prior art, or all or part of the technical solutions, can be embodied in the form of a software product. This software product is stored in a storage medium and includes several instructions to cause a device (which may be a microcontroller, chip, etc.) or processor to execute all or part of the steps of the methods of the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0106] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any changes or substitutions within the technical scope disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.
Claims
1. An audio processing method based on active sound field modulation, characterized in that, include: A mixed acoustic signal is generated and transmitted through a sound transmission unit. The mixed acoustic signal includes an audible audio signal in a first frequency band and an ultrasonic detection signal in a second frequency band, wherein the frequency of the second frequency band is higher than that of the first frequency band. The composite frequency band hybrid acoustic signal is formed by the propagation of the hybrid acoustic signal in the acoustic space through the synchronous acquisition of the hybrid acoustic signal by the sound receiving unit. The composite frequency band mixed acoustic signal is separated into ultrasonic echo stream and audible ambient stream; A three-dimensional acoustic map is constructed based on the ultrasonic echo flow, and a room impulse response from the sound delivery unit to a preset listener position is generated based on the three-dimensional acoustic map, wherein the three-dimensional acoustic map is used to indicate the acoustic characteristics of the acoustic space; Based on the audible environmental flow, determine the environmental noise field; Based on the room impulse response and the ambient noise field, an output audio signal is generated and played through the sound transmission unit.
2. The method according to claim 1, characterized in that, The generation of the mixed acoustic signal includes: Based on the audible audio signal, predict the vibration displacement of the transducer diaphragm in the sound transmission unit; Based on the vibration displacement, the ultrasonic detection signal is pre-distorted to generate a pre-compensated ultrasonic signal to counteract the nonlinear distortion introduced by the vibration displacement on the propagation of the ultrasonic detection signal. The audible audio signal and the pre-compensated ultrasonic signal are combined to form the mixed acoustic signal.
3. The method according to claim 1, characterized in that, The construction of the three-dimensional acoustic map based on the ultrasonic echo flow includes: The ultrasonic echo stream is subjected to time-domain analysis to obtain the three-dimensional geometric information of the object in the acoustic space, and a geometric point cloud map is constructed based on the three-dimensional geometric information and the time-domain analysis results. Energy attenuation and dispersion characteristics of the ultrasonic echo stream are analyzed to determine the surface acoustic material of the object in the acoustic space. The acoustic material information is fused with the geometric point cloud map to obtain the three-dimensional acoustic map.
4. The method according to claim 3, characterized in that, The generation of the room impulse response from the sound delivery unit to the preset listener position based on the three-dimensional acoustic map includes: Based on the geometric information in the three-dimensional acoustic map and the sound absorption coefficient corresponding to the acoustic material, the early reflected sound response in the mid-to-high frequency band is calculated. Based on the main dimensions of the acoustic space extracted from the three-dimensional acoustic map, the standing wave mode response in the low-frequency band is calculated. The early reflected acoustic response is spliced with the standing wave mode response to generate the full-frequency room impulse response.
5. The method according to claim 1, characterized in that, Determining the environmental noise field based on the audible environmental flow includes: Acoustic echo cancellation processing is performed on the audible ambient flow to obtain a residual ambient noise signal; Sound source localization is performed based on the residual environmental noise signal to determine the location of one or more external noise sources.
6. The method according to claim 1 or 5, characterized in that, The process of generating the output audio signal based on the room impulse response and the ambient noise field includes: Calculate the inverse filter of the room impulse response, and generate an audio signal based on the inverse filter and the audible audio signal; Based on the location of the external noise source and the preset audience location, an anti-phase noise signal is generated, which is used to cancel the external noise at the preset audience location. The audio signal is superimposed with the inverted noise signal to determine the output audio signal.
7. The method according to claim 6, characterized in that, The step of generating an anti-phase noise signal for canceling external noise at the preset audience position includes: Based on the three-dimensional acoustic map, a first noise transmission path from the location of the microphone unit to the preset listener location and a second noise transmission path from the location of the external noise source to the preset listener location are determined. Based on the residual environmental noise signal collected by the sound receiving unit, and the difference or mapping relationship of the transfer function between the first noise transmission path and the second noise transmission path, the virtual error noise signal at the preset listener position is determined. An adaptive filtering algorithm is used to generate the inverse noise signal with the goal of minimizing the virtual error noise signal.
8. The method according to claim 1, characterized in that, The method further includes: Based on the Doppler frequency shift information or inter-frame difference information in the ultrasonic echo stream, the dynamic changes of the preset audience position are tracked in real time. The room impulse response is updated in real time based on the dynamic changes.
9. An audio processing system based on active sound field modulation, characterized in that, include: Audio delivery unit; Radio unit; A processing unit is configured to generate a mixed acoustic signal comprising an audible audio signal in a first frequency band and an ultrasonic detection signal in a second frequency band, wherein the frequency of the second frequency band is higher than that of the first frequency band, and to control the transmitting unit to transmit the mixed acoustic signal; to synchronously acquire, via a receiving unit, a composite frequency band mixed acoustic signal formed after the mixed acoustic signal propagates within the acoustic space; to separate the composite frequency band mixed acoustic signal into an ultrasonic echo stream and an audible ambient stream; to construct a three-dimensional acoustic map based on the ultrasonic echo stream, and to generate a room impulse response from the transmitting unit to a preset listener position based on the three-dimensional acoustic map, wherein the three-dimensional acoustic map is used to indicate the acoustic characteristics of the acoustic space; to determine an ambient noise field based on the audible ambient stream; and to generate an output audio signal based on the room impulse response and the ambient noise field, and to play it through the transmitting unit.
10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program, wherein the computer program is configured to perform the method described in any one of claims 1 to 8 when executed.