Room acoustic calibration method based on mid-frequency reverberation time and related device

By adjusting the dynamic window length based on the mid-frequency reverberation time and performing frequency correlation processing, the problem of low acoustic calibration accuracy caused by the fixed window length in the existing technology is solved, and intelligent adaptation and sound quality improvement in different rooms are achieved.

CN122027974BActive Publication Date: 2026-07-03LINKPLAY TECHNOLOGY INC NANJING
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
LINKPLAY TECHNOLOGY INC NANJING
Filing Date
2026-04-16
Publication Date
2026-07-03

AI Technical Summary

Technical Problem

Existing room acoustic calibration systems use a fixed analysis window length, which makes it difficult to adapt to the acoustic characteristics of different rooms, resulting in low accuracy of calibration results. In particular, in rooms with long or short reverberation times, this can easily lead to problems such as over-equalization or a lack of naturalness in the sound.

Method used

By acquiring the mid-frequency reverberation time of the target room, a mapping relationship is established to determine the frequency-dependent window length. Combined with the user setting mode, the actual window length is dynamically adjusted to perform time-domain windowing and frequency-dependent transfer function estimation, and filter calibration coefficients are generated for audio calibration.

Benefits of technology

It enables intelligent optimization of window length configuration in rooms with different acoustic characteristics, improving the accuracy of acoustic calibration results and listening comfort, suppressing late reverberation interference, and preserving useful information in early reflections.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122027974B_ABST
    Figure CN122027974B_ABST
Patent Text Reader

Abstract

The present application relates to the technical field of audio sound calibration, and discloses a room sound calibration method based on a mid-frequency reverberation time and related equipment, which is used for improving the accuracy of sound calibration results. The room sound calibration method based on the mid-frequency reverberation time comprises the following steps: obtaining a room impulse response of a target room and calculating a mid-frequency reverberation time; determining a recommended window length of a frequency-dependent window by querying a predefined mapping relationship according to the mid-frequency reverberation time, the mapping relationship being established based on the acoustic correlation between the room reverberation time and the analysis window length; determining an actual window length actually applied based on the recommended window length and in combination with a user setting mode; performing time-domain windowing on the room impulse response according to the actual window length and frequency-dependent transfer function estimation processing to obtain a window function weighted transfer function estimation; and generating filter calibration coefficients based on the transfer function estimation and applying the filter calibration coefficients to perform audio calibration.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of audio acoustic calibration technology, and in particular to a room acoustic calibration method and related equipment based on mid-frequency reverberation time. Background Technology

[0002] Current mainstream room acoustic calibration systems typically use a fixed time window to analyze the room's impulse response in order to estimate the room's frequency response and generate equalization correction parameters. However, the acoustic characteristics of a room (especially reverberation time) vary significantly due to differences in room size, interior materials, and layout, making a fixed analysis window length difficult to apply universally: in "active" rooms with long reverberation times, an excessively long analysis window will incorporate a large amount of late reverberation energy, leading to distortion in the transfer function estimation and resulting in over-equalization, a dry sound, or loss of detail; in "quiet" rooms with short reverberation times, an excessively short analysis window cannot fully capture the beneficial spatial information in early reflections, resulting in a thin and unnatural sound after calibration. Furthermore, most existing methods ignore the differences in sound wave attenuation characteristics across different frequency bands (e.g., long reverberation at low frequencies and rapid attenuation at high frequencies), using a uniform window length to process signals across the entire frequency range, further limiting calibration accuracy and auditory adaptability. Summary of the Invention

[0003] This invention provides a room acoustic calibration method and related equipment based on mid-frequency reverberation time, to solve the problem of low accuracy of calibration results caused by using a fixed analysis window length in the prior art for room acoustic calibration.

[0004] The first aspect of this invention provides a room acoustic calibration method based on mid-frequency reverberation time, comprising: acquiring the room impulse response of a target room and calculating the mid-frequency reverberation time; determining a recommended window length for a frequency-dependent window based on the mid-frequency reverberation time by querying a predefined mapping relationship, wherein the mapping relationship is established based on the acoustic correlation between the room reverberation time and the analysis window length; determining the actual window length for practical application based on the recommended window length and in conjunction with user setting modes; performing time-domain windowing and frequency-dependent transfer function estimation on the room impulse response based on the actual window length to obtain a window function-weighted transfer function estimate; generating filter calibration coefficients based on the transfer function estimate and applying the filter calibration coefficients for audio calibration.

[0005] In one feasible implementation, acquiring the room impulse response of the target room and calculating the mid-frequency reverberation time includes: playing a test signal in the target room through an audio playback device and acquiring the response signal of the test signal after propagation through the room through a microphone; calculating the room impulse response based on the test signal and the response signal; bandpass filtering the room impulse response to extract the impulse response component in the mid-frequency band; calculating the energy decay curve based on the impulse response component in the mid-frequency band using the Schrödinger inverse integral method, and linearly fitting the energy decay curve to estimate the mid-frequency reverberation time.

[0006] In one feasible implementation, determining the recommended window length of the frequency-related window based on the mid-frequency reverberation time by querying a predefined mapping relationship includes:

[0007] When the intermediate frequency reverberation time is not less than a preset first time threshold, the preset minimum window length is determined as the recommended window length.

[0008] When the mid-frequency reverberation time is not greater than the preset second time threshold, the preset maximum window length is determined as the recommended window length, and the first time threshold is greater than the second time threshold;

[0009] When the mid-frequency reverberation time is greater than the second time threshold and less than the first time threshold, the recommended window length is determined based on interpolation calculation.

[0010] In one feasible implementation, determining the recommended window length based on interpolation includes: calculating the offset of the mid-frequency reverberation time relative to the second time threshold, and using the difference between the first time threshold and the second time threshold as the interpolation interval length; calculating a normalization ratio based on the ratio of the offset to the interpolation interval length; calculating a floating-point value of the recommended window length using a reverse mapping based on the maximum window length, the minimum window length, and the normalization ratio, wherein a larger normalization ratio results in a smaller calculated floating-point value; and rounding the floating-point value to obtain the recommended window length.

[0011] In one feasible implementation, determining the actual window length for the actual application based on the recommended window length and the user setting mode includes: if the current mode is user automatic, then the recommended window length is used as the actual window length; if the current mode is user manual, then the user-specified window length is obtained as the actual window length; in the user manual mode, if the mid-frequency reverberation time is detected to exceed a preset risk threshold and the specified window length exceeds a preset safety limit, then a defensive interaction mechanism is triggered.

[0012] In one feasible implementation, the triggering of the defensive interaction mechanism includes: determining a risk level based on the mid-frequency reverberation time and the specified window length; and performing a corresponding interaction operation based on the risk level, wherein the interaction operation includes at least one of issuing a risk warning, limiting the window length adjustment range, or forcibly applying a recommended window length.

[0013] In one feasible implementation, after determining the actual window length for the actual application based on the recommended window length and the user setting mode, the method further includes: recording the recommended window length, the actual window length, the user mode selection, and the status information of whether the defense mechanism is triggered; and writing the status information into the calibration result data.

[0014] In one feasible implementation, determining the effective time window length corresponding to different frequency bands based on the actual window length includes: dividing the target frequency band into multiple sub-frequency bands according to a preset frequency segmentation rule; for each sub-frequency band, converting the actual window length into an effective time window length in the time domain based on its center frequency or frequency band characteristics; and fine-tuning the effective time window length according to the frequency correlation based on the difference in sensitivity of different frequency bands to direct sound and reflected sound to obtain the target effective time window length.

[0015] In one feasible implementation, the step of performing time-domain windowing on the room impulse response based on the target effective time window length to obtain the windowed impulse response includes: generating a time-domain window function sequence for each frequency band according to its corresponding target effective time window length; multiplying the time-domain window function sequence with the room impulse response in the time domain; and truncating the signal after multiplication, retaining the data within the target effective time window length to form the windowed impulse response.

[0016] In one feasible implementation, the step of performing frequency domain transformation on the windowed impulse response to obtain a window function-weighted transfer function estimate includes: performing a fast Fourier transform on the windowed impulse response to transform it to the frequency domain; smoothing or averaging the frequency domain-transformed signal; and using the processed frequency domain signal as a window function-weighted transfer function estimate.

[0017] In one feasible implementation, generating filter calibration coefficients based on the transfer function estimate and applying the filter calibration coefficients to perform audio calibration includes: comparing the transfer function estimate with a preset target frequency response to obtain the frequency response difference to be compensated; calculating filter calibration coefficients based on the frequency response difference; and applying the filter calibration coefficients to the filters in the audio processing link to calibrate the played audio signal.

[0018] A second aspect of the present invention provides a room acoustic calibration device based on mid-frequency reverberation time, comprising: an acquisition module for acquiring the room impulse response of a target room and calculating the mid-frequency reverberation time; a first determination module for determining a recommended window length of a frequency-related window based on the mid-frequency reverberation time by querying a predefined mapping relationship, wherein the mapping relationship is established based on the acoustic correlation between the room reverberation time and the analysis window length; a second determination module for determining an actual window length for practical application based on the recommended window length and in conjunction with a user setting mode; a processing module for performing time-domain windowing and frequency-related transfer function estimation processing on the room impulse response based on the actual window length to obtain a window function-weighted transfer function estimate; and a calibration module for generating filter calibration coefficients based on the transfer function estimate and applying the filter calibration coefficients to perform audio calibration.

[0019] In one feasible implementation, the acquisition module is specifically used to: play a test signal in the target room through an audio playback device, and acquire the response signal of the test signal after it propagates through the room through a microphone; calculate the room impulse response based on the test signal and the response signal; bandpass filter the room impulse response to extract the impulse response component in the mid-frequency band; calculate the energy decay curve based on the impulse response component in the mid-frequency band using the Schroeder inverse integral method, and perform linear fitting on the energy decay curve to estimate the mid-frequency reverberation time.

[0020] In one feasible implementation, the first determining module includes: a first determining unit, configured to determine a preset minimum window length as a recommended window length when the mid-frequency reverberation time is not less than a preset first time threshold; a second determining unit, configured to determine a preset maximum window length as a recommended window length when the mid-frequency reverberation time is not greater than a preset second time threshold, wherein the first time threshold is greater than the second time threshold; and a third determining unit, configured to determine the recommended window length based on interpolation calculation when the mid-frequency reverberation time is greater than the second time threshold and less than the first time threshold.

[0021] In one feasible implementation, the third determining unit is specifically used to: calculate the offset of the mid-frequency reverberation time relative to the second time threshold, and use the difference between the first time threshold and the second time threshold as the interpolation interval length; calculate the normalization ratio based on the ratio of the offset to the interpolation interval length; calculate the recommended window length using a reverse mapping based on the maximum window length, the minimum window length, and the normalization ratio, wherein the larger the normalization ratio, the smaller the calculated floating-point value; and round the floating-point value to obtain the recommended window length.

[0022] In one feasible implementation, the second determining module includes: a processing unit, configured to use the recommended window length as the actual window length if the current mode is user automatic mode; a specifying unit, configured to obtain the specified window length specified by the user as the actual window length if the current mode is user manual mode; and a triggering unit, configured to trigger a defensive interaction mechanism if, in the user manual mode, the mid-frequency reverberation time exceeds a preset risk threshold and the specified window length exceeds a preset safety limit.

[0023] In one feasible implementation, the triggering unit is specifically used to: determine a risk level based on the mid-frequency reverberation time and the specified window length; and perform a corresponding interactive operation based on the risk level, the interactive operation including at least one of issuing a risk warning, limiting the window length adjustment range, or forcibly applying a recommended window length.

[0024] In one feasible implementation, the room acoustic calibration device based on mid-frequency reverberation time further includes: a recording module for recording the recommended window length, the actual window length, the user mode selection, and the status information of whether the defense mechanism is triggered; and writing the status information into the calibration result data.

[0025] In one feasible implementation, the processing module includes: a fourth determining unit, configured to determine the target effective time window length corresponding to different frequency bands based on the actual window length; a windowing unit, configured to perform time-domain windowing processing on the room impulse response based on the target effective time window length to obtain the windowed impulse response; and a conversion unit, configured to perform frequency-domain conversion on the windowed impulse response to obtain a window function-weighted transfer function estimate.

[0026] In one feasible implementation, the fourth determining unit is specifically used to: divide the target frequency band into multiple sub-frequency bands according to a preset frequency segmentation rule; for each sub-frequency band, convert the actual window length into an effective time window length in the time domain according to its center frequency or frequency band characteristics; and fine-tune the effective time window length according to the frequency correlation based on the difference in sensitivity of different frequency bands to direct sound and reflected sound, so as to obtain the target effective time window length.

[0027] In one feasible implementation, the windowing unit is specifically used to: generate a time-domain window function sequence for each frequency band according to its corresponding target effective time window length; multiply the time-domain window function sequence with the room impulse response in the time domain; truncate the signal after multiplication, retain the data within the target effective time window length, and form a windowed impulse response.

[0028] In one feasible implementation, the conversion unit is specifically used to: perform a fast Fourier transform on the windowed impulse response to convert it to the frequency domain; smooth or average the frequency domain signal; and use the processed frequency domain signal as a transfer function estimate weighted by the window function.

[0029] In one feasible implementation, the calibration module is specifically used to: compare the estimated transfer function with a preset target frequency response to obtain the frequency response difference to be compensated; calculate the filter calibration coefficient based on the frequency response difference; and apply the filter calibration coefficient to the filter in the audio processing link to calibrate the played audio signal.

[0030] A third aspect of the present invention provides an electronic device, comprising: a memory and at least one processor, wherein the memory stores instructions; the at least one processor invokes the instructions in the memory to cause the electronic device to perform the above-described room acoustic calibration method based on mid-frequency reverberation time.

[0031] A fourth aspect of the present invention provides a computer-readable storage medium storing instructions that, when executed on a computer, cause the computer to perform the above-described room acoustic calibration method based on mid-frequency reverberation time.

[0032] In the technical solution provided by this invention, the room impulse response of the target room is obtained and the mid-frequency reverberation time is calculated; based on the mid-frequency reverberation time, a recommended window length for a frequency-related window is determined by querying a predefined mapping relationship, wherein the mapping relationship is established based on the acoustic correlation between the room reverberation time and the analysis window length; based on the recommended window length and combined with the user setting mode, the actual window length for practical application is determined; based on the actual window length, the room impulse response is subjected to time-domain windowing and frequency-related transfer function estimation processing to obtain a window function-weighted transfer function estimate; based on the transfer function estimate, filter calibration coefficients are generated and applied to perform audio calibration. In this embodiment of the invention, by introducing an adaptive matching mechanism between mid-frequency reverberation time and frequency-dependent window length, intelligent optimization configuration of the analysis window length is achieved in rooms with different acoustic characteristics: in rooms with longer reverberation times, a shorter window length is automatically selected to suppress the interference of late reverberation on transfer function estimation; in rooms with shorter reverberation times, a longer window length is automatically selected to fully extract the beneficial acoustic information in early reflections. At the same time, combined with frequency-dependent processing, the impulse response of different frequency bands is windowed differently to further improve the accuracy of full-band acoustic feature extraction. Ultimately, the generated filter calibration coefficients can more accurately compensate for the frequency response defects of the room itself, thereby effectively improving the clarity and balance of sound quality while maintaining the natural spatial sense and listening comfort of the sound. Attached Figure Description

[0033] Figure 1 This is a schematic diagram of an embodiment of the room acoustic calibration method based on mid-frequency reverberation time in this invention.

[0034] Figure 2 This is a schematic diagram of another embodiment of the room acoustic calibration method based on mid-frequency reverberation time in this invention;

[0035] Figure 3 This is a schematic diagram of an embodiment of a room acoustic calibration device based on mid-frequency reverberation time according to the present invention;

[0036] Figure 4 This is a schematic diagram of another embodiment of the room acoustic calibration device based on mid-frequency reverberation time in this invention;

[0037] Figure 5 This is a schematic diagram of one embodiment of the electronic device in this invention. Detailed Implementation

[0038] This invention provides a room acoustic calibration method and related equipment based on mid-frequency reverberation time. By establishing an adaptive mapping mechanism between mid-frequency reverberation time and analysis window length and combining it with frequency-related window length fine-tuning, intelligent adaptation to rooms with different acoustic characteristics is achieved, thereby improving the accuracy of acoustic calibration results.

[0039] The terms "first," "second," "third," "fourth," etc. (if present) in the specification, claims, and accompanying drawings of this invention are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" or "having" and any variations thereof are intended to cover a non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.

[0040] It is understood that the executing entity of this invention can be a room acoustic calibration device based on mid-frequency reverberation time, or it can be a terminal or a server; no specific limitation is made here. This embodiment of the invention will be described using a server as an example.

[0041] For ease of understanding, the specific process of the embodiments of the present invention is described below. Please refer to [link / reference]. Figure 1 One embodiment of the room acoustic calibration method based on mid-frequency reverberation time in this invention includes:

[0042] 101. Obtain the room impulse response of the target room and calculate the mid-frequency reverberation time;

[0043] The mid-frequency range contains the components that human hearing is most sensitive to in terms of timbre perception, especially the frequency bands where speech intelligibility and the core harmonics of music are located. Its reverberation characteristics best reflect the overall absorption and diffusion capacity of a room for sound energy. Compared with the low frequency, which is dominated by room patterns and fluctuates wildly, and the high frequency, which is significantly affected by air absorption and material scattering, the mid-frequency reverberation time has better stability and representativeness.

[0044] Professional measurement microphones are deployed at typical listening positions in the target room and connected to a playback system via an audio interface. An optimized logarithmic sweep sine wave or maximum length sequence is played as the excitation signal. The acquisition system synchronously records the room response. In the digital signal processing unit, the recorded response signal and the original excitation signal are precisely timed-aligned, and frequency domain division (or time domain deconvolution) is performed to extract the complete room impulse response, including direct sound, early reflections, and late reverberation. A linear-phase digital bandpass filter with a passband of 500 Hz to 2000 Hz is applied to the acquired impulse response to extract the mid-frequency impulse response component. The Schroeder inverse integration method is used to calculate the energy decay curve of this component. A linear interval from -5 dB to -35 dB is selected in the decay curve for least-squares linear fitting. By calculating the time required for the fitted line to decrease by 60 dB, the mid-frequency reverberation time (RT60) of the room is finally calculated.

[0045] 102. Based on the mid-frequency reverberation time, the recommended window length of the frequency-related window is determined by querying the predefined mapping relationship. The mapping relationship is established based on the acoustic correlation between the room reverberation time and the analysis window length.

[0046] A lookup table or mapping function based on an acoustic model and experimental data is pre-defined, defining a non-linear relationship from "mid-frequency reverberation time" to "recommended window length for frequency-related windows." When the measured mid-frequency reverberation time value is input, it is compared with two preset key thresholds: if the reverberation time is greater than the active room threshold (e.g., greater than 0.6 seconds), a shorter minimum recommended window length (e.g., 50 milliseconds) is directly output to quickly truncate strong reverberation interference; if the reverberation time is less than the silent room threshold (e.g., less than 0.3 seconds), a longer maximum recommended window length (e.g., 200 milliseconds) is output to fully utilize acoustic energy. For reverberation times between these two thresholds, the system initiates a linear or logarithmic interpolation algorithm: calculating the normalized position of the current reverberation time within the threshold range, and using this ratio to perform mapping calculations between the minimum and maximum window lengths to obtain an intermediate value.

[0047] 103. Based on the recommended window length and combined with the user setting mode, determine the actual window length for practical applications;

[0048] A graphical user interface is provided, allowing users to switch between "fully automatic mode" and "manual mode." In automatic mode, the calculated recommended window length is used as the final actual window length. In manual mode, users can specify a desired window length via a slider or numeric input box. Simultaneously, two key parameters are continuously monitored: the measured intermediate frequency reverberation time and the user-specified window length. A risk is identified when the reverberation time exceeds a preset risk threshold (e.g., 0.8 seconds) and the user-specified window length simultaneously exceeds a safety limit (e.g., 150% of the recommended window length). The risk level is calculated based on the degree of exceedance, triggering a tiered response mechanism: a prompt box appears for low-risk situations; for medium-risk situations, the maximum value of the slider is dynamically limited to prevent users from setting more dangerous window lengths; for high-risk situations, a forced confirmation dialog box appears. If the user insists, the selection is recorded, but a clear warning is given, or in extreme cases, the system automatically reverts to the recommended window length to ensure the basic functionality of the calibration system is not compromised.

[0049] 104. Based on the actual window length, perform time-domain windowing and frequency-dependent transfer function estimation on the room impulse response to obtain a window function-weighted transfer function estimate;

[0050] The full audio frequency band (e.g., 20Hz-20kHz) is divided into several overlapping or non-overlapping sub-bands (e.g., divided by 1 / 3 octave). The analysis window length is then finely and adaptively adjusted based on the distinct acoustic characteristics of different frequency bands. According to the laws of acoustic physics, the low-frequency band, due to its long wavelength, strong diffraction ability, and significant influence from room modes (standing waves), typically has a longer energy decay (reverberation time). If a uniform short window is used for all frequency bands, the pulse response tail containing important low-frequency standing wave and modal information will be prematurely truncated, leading to severe distortion in low-frequency response estimation. Conversely, high-frequency sound waves are highly directional and easily absorbed by air and materials, resulting in a shorter reverberation time and rapid sound energy decay. Therefore, based on the reference window length determined by the mid-frequency reverberation time, frequency-related intelligent fine-tuning is performed: for example, for low-frequency sub-bands, the target window length is appropriately extended from the reference value (e.g., increased by 20%-50%) to fully capture the necessary acoustic information that determines the low-frequency sound quality and fullness; for high-frequency sub-bands, the window length can be appropriately shortened (e.g., reduced by 10%-30%) to better focus on early reflections and avoid including unnecessary noise, taking into account their rapid decay characteristics.

[0051] For each sub-band, its center frequency and acoustic characteristics (e.g., the typical reverberation decay rate of that band) are used as inputs. A fine-tuning function is called to scale the actual window length, resulting in a unique target effective time window length for that sub-band. Subsequently, a time-domain window function (e.g., a Hanning window) of corresponding length is generated for each sub-band. The original room impulse response is multiplied point-by-point by the window function of each sub-band to achieve time-domain windowing, and impulse response segments within the non-zero intervals of the window function are extracted. Finally, a Fast Fourier Transform is performed on each windowed time-domain segment to obtain its complex spectrum. These spectra are the window function-weighted transfer function estimates, which together constitute a three-dimensional response surface of frequency-amplitude-phase. The estimate of each frequency point is mainly derived from the acoustic energy information within its most favorable time window, thereby maximally suppressing late reverberation and noise pollution.

[0052] 105. Generate filter calibration coefficients based on transfer function estimation and apply the filter calibration coefficients for audio calibration.

[0053] The transfer function estimate is compared to a pre-defined target frequency response curve. In the frequency domain, the required compensation gain and phase adjustment for each sub-band are calculated using complex division or logarithmic subtraction, thus generating a "desired corrected filter" frequency response. Next, a filter design algorithm (e.g., frequency sampling, least squares, or adaptive iterative algorithm) is used to convert this frequency response target into a set of digital filter coefficients that can be processed in real time. For complex corrections, this might involve cascading multiple second-order section filters or a high-order finite-length unit impulse response filter. Finally, these calibration coefficients are loaded into digital filter modules (e.g., parametric equalizers, convolution engines, or graphic equalizers) in the audio signal path via software APIs or hardware registers. Subsequently, all playback signals passing through the audio system are processed by these filters in a real-time audio thread, dynamically correcting their frequency components to compensate for acoustic defects in the target room and reproduce a sound closer to the original recording intent or more in line with listening expectations.

[0054] In this embodiment of the invention, by establishing an intelligent mapping relationship between mid-frequency reverberation time and analysis window length, adaptive matching of the window length parameter to the room's acoustic characteristics is achieved: the window length is automatically shortened in active rooms with long reverberation to avoid late-stage reverberation interference, and the window length is automatically extended in quiet rooms with short reverberation to retain more beneficial acoustic information. Furthermore, a frequency-dependent fine-tuning mechanism is used to differentiate the window length settings for attenuation characteristics in different frequency bands. This design enables the calibration system to more accurately extract the transfer function estimate reflecting the true acoustic defects of the room from the impulse response, thereby generating filter calibration coefficients that better match the actual room conditions. Ultimately, while suppressing the effects of harmful reverberation, it preserves a natural spatial listening experience, significantly improving the universality, accuracy, and auditory experience of acoustic calibration in different types of rooms.

[0055] Please see Figure 2 Another embodiment of the room acoustic calibration method based on mid-frequency reverberation time in this invention includes:

[0056] 201. Obtain the room impulse response of the target room and calculate the mid-frequency reverberation time;

[0057] The test signal is played in the target room using an audio playback device, and the response signal after the test signal propagates through the room is collected using a microphone. The room impulse response is calculated based on the test signal and the response signal. The room impulse response is bandpass filtered to extract the impulse response component in the mid-frequency band. Based on the impulse response component in the mid-frequency band, the energy decay curve is calculated using the Schroder inverse integral method, and the energy decay curve is linearly fitted to estimate the mid-frequency reverberation time.

[0058] A pre-designed broadband test audio signal is played in the target room, and the complete response signal after reflection is acquired via a microphone. Using signal processing techniques, the room impulse response is obtained by deconvolving the acquired response signal with the source test signal. This impulse response is then filtered using a digital bandpass filter covering the mid-frequency band to extract the mid-frequency impulse response component. The extracted mid-frequency impulse response component is then applied to the Schroeder inverse integration method, i.e., the energy is cumulatively integrated backward from the tail of the response to generate a smooth energy decay curve. A decay interval with the best linearity is selected from this energy decay curve for linear fitting. By calculating the time interval required for the fitted line to decrease by a certain decibel value, the mid-frequency reverberation time of the room is finally determined.

[0059] 202. Based on the mid-frequency reverberation time, the recommended window length of the frequency-related window is determined by querying the predefined mapping relationship. The mapping relationship is established based on the acoustic correlation between the room reverberation time and the analysis window length.

[0060] When the mid-frequency reverberation time is not less than a preset first time threshold, the preset minimum window length is determined as the recommended window length; when the mid-frequency reverberation time is not greater than a preset second time threshold, the preset maximum window length is determined as the recommended window length, and the first time threshold is greater than the second time threshold; when the mid-frequency reverberation time is greater than the second time threshold and less than the first time threshold, the recommended window length is determined based on interpolation calculation.

[0061] A longer reverberation time means that the energy of later reflected sound decays slowly. If an analysis window that is too long is used, it will include a large amount of late reverberation energy, leading to distortion of the transfer function estimation. Therefore, when the reverberation time exceeds the first time threshold representing an "active room", the minimum window length should be used to quickly cut off reverberation interference. Conversely, a shorter reverberation time indicates that the room has strong sound absorption and rapid sound energy decay, so it is safe to use a longer analysis window to obtain more complete early reflected sound information. Therefore, when the reverberation time is lower than the second time threshold representing a "quiet room", the maximum window length should be used to make full use of acoustic information. For normal rooms that fall between the two, interpolation calculations are used to smoothly transition the window length between the minimum and maximum, achieving an adaptive match between the window length and the actual acoustic activity level of the room.

[0062] The specific steps for determining the recommended window length based on interpolation calculation are as follows: calculate the offset of the intermediate frequency reverberation time relative to the second time threshold, and use the difference between the first time threshold and the second time threshold as the length of the interpolation interval; calculate the normalization ratio based on the ratio of the offset to the length of the interpolation interval; calculate the floating-point value of the recommended window length using reverse mapping based on the maximum window length, the minimum window length, and the normalization ratio, wherein the larger the normalization ratio, the smaller the calculated floating-point value; and round the floating-point value to obtain the recommended window length.

[0063] The time difference between the reverberation time and the second time threshold is used as a positive offset. Simultaneously, the absolute time span between the first and second time thresholds is calculated to define the complete reverberation interval range requiring dynamic window length adjustment, i.e., the interpolation interval length. The calculated offset is divided by the interpolation interval length to obtain a normalized scaling factor, which quantifies the relative position of the current room reverberation level between the two acoustic extremes of silence and activity. The system then calls the preset maximum and minimum analysis window length parameters and uses a linear inverse mapping algorithm to calculate the initial floating-point value of the recommended window length based on the calculated normalized scaling factor. The core design of this algorithm ensures that the larger the normalized scaling factor (i.e., the longer the reverberation time), the smaller the calculated floating-point window length value, thus effectively avoiding late-stage reverberation energy. This floating-point value is then optimized and rounded, typically to the nearest integer or power of two that meets the efficiency requirements of digital signal processing. This final integer value is then used as the system's automatically recommended analysis window length output, thus obtaining the recommended window length.

[0064] Calculate the normalization ratio:

[0065]

[0066] Recommended floating-point window length for S-shaped nonlinear inverse mapping calculation:

[0067]

[0068] in, T1 is the measured mid-frequency reverberation time (seconds); T2 is the first time threshold (upper limit for active rooms, e.g., 0.6s); T2 is the second time threshold (lower limit for quiet rooms, e.g., 0.3s). The normalized ratio represents the current reverberation position relative to the silent-active axis; The maximum recommended window length (e.g., 200ms); The minimum recommended window length (e.g., 50ms); k is the shape parameter of the sigmoid function, controlling the steepness of the transition interval (generally taken as 4-8); and the above... Rounding down yields the final recommended window length. .

[0069] Taking the living room as an example, the first time threshold T1 is set to 0.6s, the second time threshold T2 is set to 0.3s, and the maximum recommended window length is set. 200ms minimum recommendation window length The reverberation time is 50ms, and the shape parameter k is 6. (This is based on the measured mid-frequency reverberation time.) When the value is 0.45s, first calculate the normalized proportion. = (0.45 - 0.3) / (0.6 - 0.3) = 0.5, substitute into the S-shaped nonlinear interpolation formula to obtain the recommended floating-point window length value. = 50 + (200 - 50) × (1 / (1 + e^(6×(0.5-0.5)))) =50 + 150 × 0.5 = 125ms. After rounding, the recommended window length is 125ms, which is exactly at the midpoint between the minimum and maximum values, meeting acoustic expectations. If the measured reverberation time increases to 0.5s, then... = 0.667, calculated as follows ≈91ms, which is about 27% shorter than the former, effectively avoiding late reverberation interference in active rooms.

[0070] Compared to linear interpolation, the S-curve tends to saturate at both ends (extremely quiet / extremely active rooms), effectively preventing the calculated window length from exceeding the reasonable range under extreme room conditions; while retaining sufficient transition sensitivity in the middle section, making the window length change more smoothly and naturally with reverberation. The shape parameter k can be flexibly adjusted according to the tuning style of different products, making the adaptive mapping of the window length more controllable in engineering.

[0071] 203. Based on the recommended window length and combined with the user settings, determine the actual window length for practical applications;

[0072] If the current mode is user automatic, the recommended window length will be used as the actual window length; if the current mode is user manual, the specified window length will be obtained as the actual window length; in user manual mode, if the mid-frequency reverberation time exceeds the preset risk threshold and the specified window length exceeds the preset safety limit, a defensive interaction mechanism will be triggered.

[0073] The specific steps for triggering the defensive interaction mechanism can be as follows: determine the risk level based on the mid-frequency reverberation time and the specified window length; and perform the corresponding interaction operation based on the risk level. The interaction operation includes issuing a risk warning, limiting the window length adjustment range, or forcibly applying the recommended window length, at least one of these.

[0074] Furthermore, record the recommended window length, actual window length, user mode selection, and status information such as whether the defense mechanism is triggered; write the status information into the calibration result data.

[0075] The system reads the currently configured user mode flag. If the user mode flag is set to automatic mode, the calculated recommended window length is directly set as the actual window length for final application. If the user mode flag is set to manual mode, the system retrieves the user-defined and confirmed window length from the slider or numerical input box in the graphical user interface and uses it as the actual window length. In manual mode, a continuously running security monitoring background process is started. This process compares the measured intermediate frequency reverberation time with a preset risk reverberation time threshold in real time, and simultaneously compares the user-specified window length with another preset safe window length upper limit. When the intermediate frequency reverberation time exceeds the risk threshold and the specified window length also exceeds the safe upper limit, a parameter conflict risk is immediately identified, and a multi-layered defensive interaction mechanism is triggered. This mechanism first calculates a quantified risk level based on a weighted combination of the degree of reverberation time exceeding the risk threshold and the degree of window length exceeding the risk threshold. The level can be divided into three levels: low, medium, and high. Subsequently, corresponding interactive operations are executed based on the risk level: low-risk situations only display a mild text prompt on the interface, informing the user that the current settings may affect the calibration effect; medium-risk situations display a warning while dynamically limiting the maximum draggable value of the window length adjustment slider within a safe limit, preventing the user from setting more dangerous parameters; high-risk situations display a mandatory confirmation dialog box explaining the danger, and automatically restore the actual window length to the system's recommended value if the user ignores the suggestion, ensuring that basic calibration functions are not compromised. Furthermore, all key data throughout the decision-making and application process, including the system's recommended window length, the final actual window length, the user's mode selection, risk monitoring results, and whether the defense mechanism is triggered and its specific execution actions, are encapsulated in real time into a structured status log. This log is ultimately written into the result data file of this room acoustic calibration, providing complete data support for subsequent calibration effect analysis, problem backtracking, and algorithm optimization.

[0076] 204. Determine the effective time window length for different frequency bands based on the actual window length;

[0077] According to the preset frequency segmentation rules, the target frequency band is divided into multiple sub-bands; for each sub-band, the actual window length is converted into the effective time window length in the time domain based on its center frequency or frequency band characteristics; based on the difference in sensitivity of different frequency bands to direct sound and reflected sound, the effective time window length is finely adjusted according to frequency correlation to obtain the target effective time window length.

[0078] According to pre-defined frequency segmentation rules, such as using equal ratios or equal widths, the complete audible frequency band is divided into several consecutive sub-bands. Common divisions include low-frequency band, mid-low-frequency band, mid-frequency band, mid-high-frequency band, and high-frequency band. For each sub-band, its center frequency is read or its acoustic characteristics of frequency band coverage are analyzed. Then, based on a pre-defined window length frequency conversion model, the actual window length is initially converted into the effective time window length corresponding to that sub-band in the time domain. This conversion process usually considers the relationship between the sound wave period and time resolution to ensure that each frequency band has sufficient period to be analyzed. Subsequently, a frequency-dependent fine-tuning coefficient matrix is ​​introduced. This matrix is ​​pre-set based on the physical characteristics of sound waves propagating in a room at different frequency bands. For example, low-frequency sound waves have strong diffraction capabilities and typically longer reverberation times, making them more sensitive to later reflections. Therefore, their fine-tuning coefficient may be less than one, thus appropriately shortening the initially calculated effective time window length to cut off potentially interfering low-frequency reverberation tails earlier. Conversely, high-frequency sound waves are highly directional and decay rapidly, making them more sensitive to the details of early reflections. Their fine-tuning coefficient may be greater than one, thus appropriately extending the window length to capture more useful early spatial information. Through this coefficient weighting based on sub-band characteristics, the initial window length for each frequency band is individually fine-tuned, ultimately outputting a set of target effective time window lengths for different frequencies.

[0079] The fine-tuning formula is:

[0080]

[0081] Upper and lower limit constraints:

[0082]

[0083] in, The effective time window length (ms) for the kth sub-band; The actual window length (ms) is determined by the user mode. is the center frequency (Hz) of the kth sub-band; The intermediate frequency reference frequency (taken as 1000 Hz, compared with RT) 60 (Measurement band alignment) To analyze the upper limit of the frequency band (e.g., 16000 Hz); To analyze the lower limit of the frequency band (e.g., 62.5 Hz); This is the frequency sensitivity coefficient, which controls the fine-tuning intensity (generally taken as 0.3~0.6). The target window length is set as a lower limit protection value (to prevent insufficient frequency resolution due to excessively short window length). The upper limit protection value for the target window length (to prevent excessive reverberation energy from being included due to excessive length); clip(·) limits the value to [T floor , Tceil The truncation function for the interval.

[0084] For example, setting the actual window length 120ms, intermediate frequency reference frequency 1000Hz, upper limit of analysis frequency band 16000Hz, lower limit The frequency sensitivity coefficient is 62.5 Hz. The value is 0.4, covering a total of 8 octaves. Taking three typical sub-bands as examples: In the 125Hz low-frequency band, since it is below the reference frequency, substituting into the formula T... eff = 120 × (1-0.4×log2(125 / 1000)) = 120×(1-0.4×(-3)) = 120×2.2 = 264ms. After upper and lower limit constraints, we get about 138ms, and the window length is extended. In the 1000Hz mid-frequency band, since it is consistent with the reference frequency, log2(1000 / 1000)=0, and the calculation result remains unchanged at 120ms. In the 8000Hz high-frequency band, since it is higher than the reference frequency, T eff = 120×(1-0.4×log2(8000 / 1000))= 120×(1-0.4×3) = 120×(-0.2) = -24ms. After upper and lower limit constraints, we get about 102ms, which significantly shortens the window length. The results show that the target window length decreases monotonically from low frequency to high frequency, which is consistent with the physical law of long reverberation time at low frequency and fast decay at high frequency in room acoustics.

[0085] Using the octave logarithm as a metric aligns with the logarithmic characteristics of frequency perception by the human ear and the physical decay of room reverberation with frequency, thus avoiding the problem of excessive high-frequency compression caused by linear frequency axis fine-tuning. The measurement frequency band directly corresponds to the mid-frequency range, ensuring the mid-frequency remains stable as a reference anchor point. The fine-tuning amounts for low and high frequencies naturally increase with the octave distance from the reference frequency, providing clear physical meaning and simple calculation. Coefficients and upper / lower limit constraints jointly ensure the algorithm's robustness, adapting to the calibration needs of rooms of different sizes and product types.

[0086] 205. Based on the target effective time window length, perform time-domain windowing on the room impulse response to obtain the windowed impulse response;

[0087] For each frequency band, a time-domain window function sequence is generated based on its corresponding target effective time window length; the time-domain window function sequence is multiplied with the room impulse response in the time domain; the signal after multiplication is truncated, retaining the data within the target effective time window length to form the windowed impulse response.

[0088] After calculating the target effective time window length for all sub-bands, a dedicated time-domain window function sequence is generated for each independent sub-band. The total number of sampling points for the window is determined based on the target effective time window length corresponding to that band. The corresponding mathematical expression is selected according to the window function type, such as using the discrete calculation formulas for the Hanning window or Blackman window. A symmetrical window sequence with a length equal to the target window length and a value that gradually rises from zero to a peak and then smoothly decreases back to zero is generated in the time domain. The initially measured complete room impulse response time-domain signal is read, and the window function sequence generated for the specific sub-band is multiplied point-by-point with the impulse response signal at each time-domain sampling point. This operation allows the portion of the impulse response within the effective support region of the window function to retain its original amplitude, while the portions outside the window are attenuated to near zero, thus achieving selective windowing of the impulse response in that band on the time axis. A precise truncation operation is performed on the mixed signal after multiplication, retaining only the core data segment that starts from the beginning of the impulse response and has a length exactly equal to the effective time window of the target frequency band, while discarding all sampling point data outside this time window range. This results in a windowed impulse response segment with a regular length that contains only acoustic information within the target time range of the frequency band.

[0089] 206. Perform frequency domain transformation on the windowed impulse response to obtain the window function weighted transfer function estimate;

[0090] Perform a Fast Fourier Transform on the windowed impulse response to convert it to the frequency domain; smooth or average the frequency domain signal; and use the processed frequency domain signal as the transfer function estimate with window function weighting.

[0091] The Fast Fourier Transform (FFT) process employs a number of transform points adapted to the window length to completely transform the windowed impulse response in the time domain to the frequency domain, obtaining its complex spectral representation. This spectral representation is preprocessed to improve the stability and practicality of the estimation. For example, a frequency smoothing filter is applied to smooth the amplitude curve of the spectrum, or continuous spectral lines are grouped according to a preset frequency band division rule, and the average energy within each band is calculated, thus obtaining a set of data points representing the characteristics of each frequency band. After these processing steps, the physical meaning of the resulting frequency domain signal has changed. It is no longer the complete frequency response of the original room impulse response, but rather a transfer function estimated by weighting and truncation over a specific time length using a specific window function. This transfer function estimate effectively suppresses late reverberation and noise components outside the selected time window and is recorded as a window-weighted transfer function estimate.

[0092] 207. Generate filter calibration coefficients based on transfer function estimation and apply the filter calibration coefficients for audio calibration.

[0093] The transfer function estimate is compared with the preset target frequency response to obtain the frequency response difference that needs to be compensated; the filter calibration coefficient is calculated based on the frequency response difference; the filter calibration coefficient is applied to the filter in the audio processing link to calibrate the played audio signal.

[0094] A preset target frequency response curve is read, and combined with transfer function estimation, a point-by-point or band-by-band comparison is performed on the two sets of data in the frequency domain. Specifically, the amplitude value corresponding to the target response curve is divided by the amplitude value estimated by the transfer function, or the two are subtracted in the logarithmic domain. This allows for the precise calculation of the gain difference that needs to be filled or reduced at each frequency point, thus obtaining the frequency response difference that needs to be compensated across the entire frequency band. The target frequency response curve typically represents the flat response expected to be achieved under ideal acoustic conditions or a certain optimized listening curve. Based on this frequency response difference, a filter design algorithm is invoked, such as using the least squares method or frequency sampling method to design a finite-length unit impulse response filter, or a set of parametric equalizers. The frequency response difference is used as the objective function, and the digital filter coefficients that can achieve the compensation effect are calculated through iterative optimization; these are the filter calibration coefficients. Finally, these calculated filter calibration coefficients are loaded into the real-time digital filter hardware or software module in the audio processing link through a specific communication interface or by writing to a designated memory register. When any audio signal flows through the processing link, the filter will reshape the frequency components of the signal in real time according to these coefficients, thereby offsetting the distortion introduced by the room itself and achieving acoustic calibration of the played audio signal.

[0095] In this embodiment of the invention, an adaptive mapping mechanism between mid-frequency reverberation time and analysis window length is established to achieve intelligent parameter configuration for different room acoustic characteristics: the window length is automatically shortened in active rooms with long reverberation to avoid late reverberation interference, and the window length is automatically extended in quiet rooms with short reverberation to fully extract early acoustic information. Combined with a frequency-related window length fine-tuning mechanism, different frequency bands (low-frequency extension, high-frequency shortening) are processed differently. This design enables the system to more accurately separate the transfer function estimate reflecting the real defects of the room from the impulse response, thereby generating filter calibration coefficients that are more consistent with the actual acoustic environment. Ultimately, while effectively compensating for the room frequency response, it suppresses the "over-equalization" problem caused by harmful reverberation and preserves the natural spatial listening experience, significantly improving the accuracy, adaptability, and auditory experience of acoustic calibration in different types of rooms.

[0096] The above describes the room acoustic calibration method based on mid-frequency reverberation time in the embodiments of the present invention. The following describes the room acoustic calibration device based on mid-frequency reverberation time in the embodiments of the present invention. Please refer to [link / reference]. Figure 3 One embodiment of the room acoustic calibration device based on mid-frequency reverberation time in this invention includes:

[0097] The acquisition module 301 is used to acquire the room impulse response of the target room and calculate the mid-frequency reverberation time;

[0098] The first determining module 302 is used to determine the recommended window length of the frequency-related window based on the mid-frequency reverberation time by querying a predefined mapping relationship. The mapping relationship is established based on the acoustic correlation between the room reverberation time and the analysis window length.

[0099] The second determining module 303 is used to determine the actual window length of the actual application based on the recommended window length and the user setting mode.

[0100] Processing module 304 is used to perform time-domain windowing and frequency-related transfer function estimation on the room impulse response based on the actual window length, so as to obtain a window function weighted transfer function estimate;

[0101] The calibration module 305 is used to generate filter calibration coefficients based on the transfer function estimate and to apply the filter calibration coefficients for audio calibration.

[0102] In this embodiment of the invention, by establishing an intelligent mapping relationship between mid-frequency reverberation time and analysis window length, adaptive matching of the window length parameter to the room's acoustic characteristics is achieved: the window length is automatically shortened in active rooms with long reverberation to avoid late-stage reverberation interference, and the window length is automatically extended in quiet rooms with short reverberation to retain more beneficial acoustic information. Furthermore, a frequency-dependent fine-tuning mechanism is used to differentiate the window length settings for attenuation characteristics in different frequency bands. This design enables the calibration system to more accurately extract the transfer function estimate reflecting the true acoustic defects of the room from the impulse response, thereby generating filter calibration coefficients that better match the actual room conditions. Ultimately, while suppressing the effects of harmful reverberation, it preserves a natural spatial listening experience, significantly improving the universality, accuracy, and auditory experience of acoustic calibration in different types of rooms.

[0103] Please see Figure 4 Another embodiment of the room acoustic calibration device based on mid-frequency reverberation time in this invention includes:

[0104] The acquisition module 301 is used to acquire the room impulse response of the target room and calculate the mid-frequency reverberation time;

[0105] The first determining module 302 is used to determine the recommended window length of the frequency-related window based on the mid-frequency reverberation time by querying a predefined mapping relationship. The mapping relationship is established based on the acoustic correlation between the room reverberation time and the analysis window length.

[0106] The second determining module 303 is used to determine the actual window length of the actual application based on the recommended window length and the user setting mode.

[0107] Processing module 304 is used to perform time-domain windowing and frequency-related transfer function estimation on the room impulse response based on the actual window length, so as to obtain a window function weighted transfer function estimate;

[0108] The calibration module 305 is used to generate filter calibration coefficients based on the transfer function estimate and to apply the filter calibration coefficients for audio calibration.

[0109] Optionally, the acquisition module 301 can be specifically used for:

[0110] The test signal is played in the target room using an audio playback device, and the response signal after the test signal propagates through the room is collected using a microphone. The room impulse response is calculated based on the test signal and the response signal. The room impulse response is bandpass filtered to extract the impulse response component in the mid-frequency band. Based on the impulse response component in the mid-frequency band, the energy decay curve is calculated using the Schroder inverse integral method, and the energy decay curve is linearly fitted to estimate the mid-frequency reverberation time.

[0111] Optionally, the first determining module 302 includes:

[0112] The first determining unit 3021 is used to determine the preset minimum window length as the recommended window length when the mid-frequency reverberation time is not less than the preset first time threshold.

[0113] The second determining unit 3022 is used to determine the preset maximum window length as the recommended window length when the mid-frequency reverberation time is not greater than the preset second time threshold, and the first time threshold is greater than the second time threshold.

[0114] The third determining unit 3023 is used to determine the recommended window length based on interpolation calculation when the mid-frequency reverberation time is greater than the second time threshold and less than the first time threshold.

[0115] Optionally, the third determining unit 3023 can be specifically used for:

[0116] The offset of the mid-frequency reverberation time relative to the second time threshold is calculated, and the difference between the first and second time thresholds is used as the interpolation interval length. The normalization ratio is calculated based on the ratio of the offset to the interpolation interval length. The recommended window length is calculated using a reverse mapping method based on the maximum window length, the minimum window length, and the normalization ratio. The larger the normalization ratio, the smaller the calculated floating-point value. The floating-point value is then rounded to obtain the recommended window length.

[0117] Optionally, the second determining module 303 includes:

[0118] The processing unit 3031 is used to use the recommended window length as the actual window length if the current mode is user automatic.

[0119] The specified unit 3032 is used to obtain the specified window length specified by the user as the actual window length if the current mode is manual.

[0120] The triggering unit 3033 is used to trigger a defensive interaction mechanism in user manual mode if the intermediate frequency reverberation time exceeds a preset risk threshold and the specified window length exceeds a preset safety limit.

[0121] Optionally, the trigger unit 3033 can be specifically used for:

[0122] The risk level is determined based on the mid-frequency reverberation time and the specified window length; based on the risk level, the corresponding interactive operation is performed, including at least one of issuing a risk warning, limiting the window length adjustment range, or forcibly applying the recommended window length.

[0123] Optionally, the room acoustic calibration device based on mid-frequency reverberation time also includes:

[0124] The recording module 306 is used to record the recommended window length, the actual window length, the user mode selection, and the status information of whether the defense mechanism is triggered; and writes the status information into the calibration result data.

[0125] Optionally, the processing module 304 includes:

[0126] The fourth determining unit 3041 is used to determine the effective time window length of the target for different frequency bands based on the actual window length.

[0127] The windowing unit 3042 is used to perform time-domain windowing processing on the room impulse response based on the target effective time window length to obtain the windowed impulse response;

[0128] The conversion unit 3043 is used to perform frequency domain conversion on the windowed impulse response to obtain a window function weighted transfer function estimate.

[0129] Optionally, the fourth determining unit 3041 can be specifically used for:

[0130] According to the preset frequency segmentation rules, the target frequency band is divided into multiple sub-bands; for each sub-band, the actual window length is converted into the effective time window length in the time domain based on its center frequency or frequency band characteristics; based on the difference in sensitivity of different frequency bands to direct sound and reflected sound, the effective time window length is finely adjusted according to frequency correlation to obtain the target effective time window length.

[0131] Optionally, the windowing unit 3042 can be specifically used for:

[0132] For each frequency band, a time-domain window function sequence is generated based on its corresponding target effective time window length; the time-domain window function sequence is multiplied with the room impulse response in the time domain; the signal after multiplication is truncated, retaining the data within the target effective time window length to form the windowed impulse response.

[0133] Optionally, the conversion unit 3043 can be specifically used for:

[0134] Perform a Fast Fourier Transform on the windowed impulse response to convert it to the frequency domain; smooth or average the frequency domain signal; and use the processed frequency domain signal as the transfer function estimate with window function weighting.

[0135] Optionally, the calibration module 305 can be specifically used for:

[0136] The transfer function estimate is compared with the preset target frequency response to obtain the frequency response difference that needs to be compensated; the filter calibration coefficient is calculated based on the frequency response difference; the filter calibration coefficient is applied to the filter in the audio processing link to calibrate the played audio signal.

[0137] In this embodiment of the invention, an adaptive mapping mechanism between mid-frequency reverberation time and analysis window length is established to achieve intelligent parameter configuration for different room acoustic characteristics: the window length is automatically shortened in active rooms with long reverberation to avoid late reverberation interference, and the window length is automatically extended in quiet rooms with short reverberation to fully extract early acoustic information. Combined with a frequency-related window length fine-tuning mechanism, different frequency bands (low-frequency extension, high-frequency shortening) are processed differently. This design enables the system to more accurately separate the transfer function estimate reflecting the real defects of the room from the impulse response, thereby generating filter calibration coefficients that are more consistent with the actual acoustic environment. Ultimately, while effectively compensating for the room frequency response, it suppresses the "over-equalization" problem caused by harmful reverberation and preserves the natural spatial listening experience, significantly improving the accuracy, adaptability, and auditory experience of acoustic calibration in different types of rooms.

[0138] above Figure 3 and Figure 4 The room acoustic calibration device based on mid-frequency reverberation time in the embodiments of the present invention will be described in detail from the perspective of modular functional entities. The electronic devices in the embodiments of the present invention will be described in detail from the perspective of hardware processing.

[0139] See Figure 5 As shown, the electronic device includes a processor 500 and a memory 501. The memory 501 stores machine-executable instructions that can be executed by the processor 500. The processor 500 executes the machine-executable instructions to implement the above-described room acoustic calibration method based on mid-frequency reverberation time.

[0140] Furthermore, Figure 5 The electronic device shown also includes a bus 502 and a communication interface 503. The processor 500, the communication interface 503 and the memory 501 are connected via the bus 502.

[0141] The memory 501 may include high-speed random access memory (RAM) and may also include non-volatile memory, such as at least one disk storage device. Communication between this system network element and at least one other network element is achieved through at least one communication interface 503 (which can be wired or wireless), such as the Internet, wide area network, local area network, metropolitan area network, etc. The bus 502 may be an ISA bus, PCI bus, or EISA bus, etc. The bus can be divided into address bus, data bus, control bus, etc. For ease of representation, Figure 5 The symbol is represented by a single double-headed arrow, but this does not mean that there is only one bus or one type of bus.

[0142] The processor 500 may be an integrated circuit chip with signal processing capabilities. In implementation, each step of the above method can be completed by the integrated logic circuitry in the hardware of the processor 500 or by instructions in software form. The processor 500 may be a general-purpose processor, including a central processing unit (CPU), a network processor (NP), etc.; it may also be a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components. It can implement or execute the methods, steps, and logic block diagrams disclosed in the embodiments of this disclosure. The general-purpose processor may be a microprocessor or any conventional processor. The steps of the methods disclosed in the embodiments of this disclosure can be directly embodied in the execution of a hardware decoding processor, or executed by a combination of hardware and software modules in the decoding processor. The software modules may reside in random access memory, flash memory, read-only memory, programmable read-only memory, electrically erasable programmable memory, registers, or other mature storage media in the art. The storage medium is located in memory 501. The processor 500 reads the information in memory 501 and, in conjunction with its hardware, completes the method steps of the aforementioned embodiment.

[0143] The present invention also provides an electronic device, the computer device including a memory and a processor, the memory storing computer-readable instructions, which, when executed by the processor, cause the processor to perform the steps of the room acoustic calibration method based on mid-frequency reverberation time in the above embodiments.

[0144] The present invention also provides a computer-readable storage medium, which may be a non-volatile computer-readable storage medium or a volatile computer-readable storage medium, wherein the computer-readable storage medium stores instructions that, when executed on a computer, cause the computer to perform the steps of the room acoustic calibration method based on mid-frequency reverberation time.

[0145] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working processes of the systems, devices, and units described above can be referred to the corresponding processes in the foregoing method embodiments, and will not be repeated here.

[0146] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0147] The above-described embodiments are only used to illustrate the technical solutions of the present invention, and are not intended to limit it. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A method for room acoustic calibration based on mid-frequency reverberation time, characterized in that, The room acoustic calibration method based on mid-frequency reverberation time includes: Obtain the room impulse response of the target room and calculate the mid-frequency reverberation time; Based on the mid-frequency reverberation time, the recommended window length of the frequency-related window is determined by querying a predefined mapping relationship, which is established based on the acoustic correlation between the room reverberation time and the analysis window length. Based on the recommended window length and the user setting mode, the actual window length for the actual application is determined. Based on the actual window length, the room impulse response is subjected to time-domain windowing and frequency-related transfer function estimation to obtain a window function-weighted transfer function estimate. Based on the transfer function, the filter calibration coefficients are estimated and generated, and the filter calibration coefficients are applied to perform audio calibration. The step of determining the recommended window length of the frequency-related window based on the mid-frequency reverberation time by querying a predefined mapping relationship includes: when the mid-frequency reverberation time is not less than a preset first time threshold, determining a preset minimum window length as the recommended window length; when the mid-frequency reverberation time is not greater than a preset second time threshold, determining a preset maximum window length as the recommended window length, wherein the first time threshold is greater than the second time threshold; and when the mid-frequency reverberation time is greater than the second time threshold and less than the first time threshold, determining the recommended window length based on interpolation calculation.

2. The room acoustic calibration method based on mid-frequency reverberation time according to claim 1, characterized in that, The process of acquiring the room impulse response of the target room and calculating the mid-frequency reverberation time includes: A test signal is played in the target room using an audio playback device, and the response signal after the test signal propagates through the room is collected using a microphone. The room impulse response is calculated based on the test signal and the response signal; The room impulse response is bandpass filtered to extract the impulse response component in the intermediate frequency band; Based on the impulse response components of the mid-frequency band, the energy decay curve is calculated using the Schroeder inverse integral method, and the energy decay curve is linearly fitted to estimate the mid-frequency reverberation time.

3. The room acoustic calibration method based on mid-frequency reverberation time according to claim 1, characterized in that, The determination of the recommendation window length based on interpolation calculation includes: Calculate the offset of the mid-frequency reverberation time relative to the second time threshold, and use the difference between the first time threshold and the second time threshold as the length of the interpolation interval; The normalization ratio is calculated based on the ratio of the offset to the length of the interpolation interval; The recommended window length is calculated using a reverse mapping method based on the maximum window length, the minimum window length, and the normalization ratio. The larger the normalization ratio, the smaller the calculated floating-point value. The floating-point value is rounded down to obtain the recommended window length.

4. The room acoustic calibration method based on mid-frequency reverberation time according to claim 1, characterized in that, The process of determining the actual window length for the actual application based on the recommended window length and the user settings includes: If the current mode is user automatic, the recommended window length will be used as the actual window length. If the current mode is manual, the specified window length is obtained as the actual window length. In the user manual mode, if the intermediate frequency reverberation time exceeds a preset risk threshold and the specified window length exceeds a preset safety limit, a defensive interaction mechanism is triggered.

5. The room acoustic calibration method based on mid-frequency reverberation time according to claim 4, characterized in that, The triggering of the defensive interaction mechanism includes: The risk level is determined based on the mid-frequency reverberation time and the specified window length; Based on the risk level, perform the corresponding interactive operation, which includes at least one of issuing a risk warning, limiting the window length adjustment range, or forcibly applying the recommended window length.

6. The room acoustic calibration method based on mid-frequency reverberation time according to claim 5, characterized in that, After determining the actual window length for the actual application based on the recommended window length and the user setting mode, the process also includes: Record the recommended window length, the actual window length, the user mode selection, and the status information of whether the defense mechanism is triggered; The status information is written into the calibration result data.

7. The room acoustic calibration method based on mid-frequency reverberation time according to claim 1, characterized in that, The step of performing time-domain windowing and frequency-related transfer function estimation on the room impulse response based on the actual window length to obtain a window-weighted transfer function estimate includes: Based on the actual window length, determine the target effective time window length corresponding to different frequency bands; Based on the target effective time window length, the room impulse response is subjected to time-domain windowing to obtain the windowed impulse response; The windowed impulse response is frequency-domain transformed to obtain a window-weighted transfer function estimate.

8. The room acoustic calibration method based on mid-frequency reverberation time according to claim 7, characterized in that, The step of determining the effective time window length corresponding to different frequency bands based on the actual window length includes: According to the preset frequency segmentation rules, the target frequency band is divided into multiple sub-frequency bands; For each sub-band, the actual window length is converted into an effective time window length in the time domain based on its center frequency or frequency band characteristics. Based on the differences in sensitivity to direct and reflected sound in different frequency bands, the effective time window length is finely adjusted according to frequency correlation to obtain the target effective time window length.

9. The room acoustic calibration method based on mid-frequency reverberation time according to claim 7, characterized in that, The step of performing time-domain windowing on the room impulse response based on the target effective time window length to obtain the windowed impulse response includes: For each frequency band, a sequence of time-domain window functions is generated based on the length of its corresponding target effective time window; The time-domain window function sequence is multiplied with the room impulse response in the time domain; The signal after multiplication is truncated, and the data within the effective time window of the target is retained to form a windowed impulse response.

10. The room acoustic calibration method based on mid-frequency reverberation time according to claim 7, characterized in that, The step of performing a frequency domain transformation on the windowed impulse response to obtain a window function-weighted transfer function estimate includes: Perform a Fast Fourier Transform on the windowed impulse response to convert it to the frequency domain; Smooth or average the frequency band of the signal after frequency domain transformation; The processed frequency domain signal is used as a window function weighted transfer function estimate.

11. The room acoustic calibration method based on mid-frequency reverberation time according to claim 1, characterized in that, The step of estimating and generating filter calibration coefficients based on the transfer function and applying the filter calibration coefficients to perform audio calibration includes: The estimated transfer function is compared with the preset target frequency response to obtain the frequency response difference that needs to be compensated. Calculate the filter calibration coefficients based on the frequency response difference; The filter calibration coefficients are applied to the filters in the audio processing link to calibrate the played audio signal.

12. A room acoustic calibration device based on mid-frequency reverberation time, characterized in that, The room acoustic calibration device based on mid-frequency reverberation time includes: The acquisition module is used to acquire the room impulse response of the target room and calculate the mid-frequency reverberation time; The first determining module is used to determine the recommended window length of the frequency-related window based on the mid-frequency reverberation time by querying a predefined mapping relationship, wherein the mapping relationship is established based on the acoustic correlation between the room reverberation time and the analysis window length; The second determining module is used to determine the actual window length of the actual application based on the recommended window length and the user setting mode. The processing module is used to perform time-domain windowing and frequency-related transfer function estimation on the room impulse response based on the actual window length, so as to obtain a window function-weighted transfer function estimate. The calibration module is used to estimate and generate filter calibration coefficients based on the transfer function and apply the filter calibration coefficients to perform audio calibration. The first determining module includes: a first determining unit, configured to determine a preset minimum window length as a recommended window length when the mid-frequency reverberation time is not less than a preset first time threshold; a second determining unit, configured to determine a preset maximum window length as a recommended window length when the mid-frequency reverberation time is not greater than a preset second time threshold, wherein the first time threshold is greater than the second time threshold; and a third determining unit, configured to determine a recommended window length based on interpolation calculation when the mid-frequency reverberation time is greater than the second time threshold and less than the first time threshold.

13. An electronic device, characterized in that, The electronic device includes: a memory and at least one processor, wherein the memory stores instructions; The at least one processor invokes the instructions in the memory to cause the electronic device to perform the room acoustic calibration method based on mid-frequency reverberation time as described in any one of claims 1-11.

14. A computer-readable storage medium storing instructions thereon, characterized in that, When the instructions are executed by the processor, they implement the room acoustic calibration method based on mid-frequency reverberation time as described in any one of claims 1-11.

Citation Information

Patent Citations

  • Microphone array sound source localization method and system based on cross-correlation-beam forming closed-loop optimization

    CN121385803A

  • Processing audio signals

    WO2018234618A1