An environment-adaptive sound field calibration method, device, equipment and storage medium

By acquiring signals in real time and calculating acoustic characteristic deviations when playing program audio on multi-channel audio devices, and dynamically adjusting calibration parameters, the problem of low matching degree between calibration parameters and real scene in existing technologies is solved, and more efficient sound field calibration is achieved.

CN122227175APending Publication Date: 2026-06-16MALANSHAN AUDIO & VIDEO LABORATORY
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
MALANSHAN AUDIO & VIDEO LABORATORY
Filing Date
2026-05-08
Publication Date
2026-06-16

AI Technical Summary

Technical Problem

In existing sound field calibration methods, the calibration parameters do not match the actual playback scene well, resulting in poor sound field calibration results.

Method used

When playing program audio on a multi-channel audio device, the sound signal is acquired in real time. By calculating the acoustic characteristic deviation between the acquired signal and the historical reference signal, the calibration parameters, including channel delay compensation, EQ parameters and gain adjustment, are dynamically adjusted to achieve sound field calibration.

Benefits of technology

It improves the matching degree between calibration parameters and actual playback scenarios, can adapt to environmental changes, save computing resources, and avoids invalid recalibration.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122227175A_ABST
    Figure CN122227175A_ABST
Patent Text Reader

Abstract

This application relates to the field of sound field calibration technology, and discloses an environment-adaptive sound field calibration method, apparatus, device, and storage medium. The method includes: extracting transient strong signals and independent signals for each channel from the original digital audio signal; acquiring a multi-channel reference signal; acquiring the current-moment sound signal of the multi-channel audio device while the original digital audio signal is played on the multi-channel audio device, and recording the current-moment sound signal as the acquired signal; acquiring the acoustic characteristics of the acquired signal based on the acquired signal and the multi-channel reference signal; acquiring a first deviation; if the first deviation is greater than a deviation threshold, calculating the optimal calibration parameters for each channel based on the acquired signal, and calibrating each channel of the multi-channel audio device based on the optimal calibration parameters. The method of this invention can effectively improve the matching degree between the calibration results and the actual playback scenario.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of sound field calibration technology, and in particular to an environmental adaptive sound field calibration method, apparatus, device and storage medium. Background Technology

[0002] With the increasing popularity of multi-channel audio systems in home theaters, car audio systems, and other scenarios, sound field calibration technology has become an important means to overcome room acoustic defects (such as reflection interference, inconsistent time delay, frequency response distortion, etc.) and restore a realistic and immersive sound field.

[0003] Currently, active sound field calibration solutions are widely used in the industry. The implementation logic of this type of solution is as follows: before calibration, the system controls the audio equipment to stop playing program audio, forcing the audio equipment to actively emit pre-set dedicated test tones (such as logarithmic sine sweep signals, pink noise, or impulse noise) through each channel speaker. Then, microphones are used to receive reflected signals from the room, thereby resolving the acoustic parameters and fixing them in the system. However, the dedicated test tones often differ significantly from the actual audio of movies or music played by the user in terms of frequency distribution, dynamic range, and other acoustic characteristics. This measurement method, which deviates from the actual sound source characteristics, results in a low degree of matching between the calculated calibration parameters and the actual playback scenario. Summary of the Invention

[0004] To address the technical problem of poor matching between calibration parameters and actual playback scenarios in existing sound field calibration methods, this invention provides technical solutions in the following aspects.

[0005] In a first aspect, the present invention provides an environment-adaptive sound field calibration method, comprising: Acquire the raw digital audio signal played by a multi-channel audio device, and extract transient strong signals and independent signals of each channel from the raw digital audio signal; The original digital audio signal, the transient strong signal, and each of the channel-independent signals are time-aligned and data-encapsulated to obtain a multi-channel reference signal. When the original digital audio signal is played by the multi-channel audio device, the current sound signal of the multi-channel audio device is collected, and the current sound signal is recorded as the collected signal. The acoustic characteristics of the acquired signal are obtained based on the acquired signal and the multi-channel reference signal; The first deviation between the acoustic characteristics of the acquired signal and the acoustic characteristics of a reference sound signal acquired within a historical time period is obtained; the historical time period is the time period after the moment when the multi-channel audio device starts playing the original digital audio signal and before the current moment. If the first deviation is greater than the deviation threshold, the optimal calibration parameters for each channel are calculated based on the acquired signal, and each channel of the multi-channel audio device is calibrated based on the optimal calibration parameters.

[0006] Preferably, the acoustic characteristics include room frequency response and direct sound delay of each channel, and the calculation method of the first deviation includes: obtaining the absolute value of the first difference between the direct sound delay of the i-th channel corresponding to the acquired signal and the direct sound delay of the i-th channel of the reference sound signal, and determining the ratio of the absolute value of the first difference to the direct sound delay of the i-th channel of the reference sound signal as the relative deviation rate of the i-th channel in the time domain; The absolute value of the second difference between the room frequency response response corresponding to the acquired signal and the room frequency response response corresponding to the reference sound signal is obtained, and the ratio of the absolute value of the second difference to the room frequency response response corresponding to the reference sound signal is determined as the frequency domain relative deviation amplitude. The frequency domain relative deviation amplitude is integrally calculated within a preset frequency range to obtain the frequency domain relative deviation integral value. The relative deviation rate in the time domain of the i-th channel is added to the integral value of the relative deviation in the frequency domain to obtain the cumulative deviation value of the i-th channel. The first deviation is obtained by averaging the cumulative deviation values ​​of each channel.

[0007] Preferably, the optimal calibration parameters include channel delay compensation values ​​for each channel, and the calculation of the optimal calibration parameters for each channel based on the acquired signal includes: Calculate the second deviation between the direct sound delay and the reference delay for each channel corresponding to the acquired signal, and select the largest deviation from all the second deviations. The reference delay is the center channel delay or the main channel delay. Calculate the target difference between the maximum deviation and the second deviation of the i-th channel, and use the target difference as the channel delay compensation value of the i-th channel.

[0008] Preferably, the optimal calibration parameters include the EQ parameters for each target frequency band, and the calculation of the optimal calibration parameters for each channel based on the acquired signal further includes: For the k-th target frequency band, extract the frequency range that falls within the room frequency response corresponding to the acquired signal. The amplitude of each discrete frequency point within the target frequency band is calculated, and the average amplitude of all discrete frequency points is calculated to obtain the average frequency response gain of the k-th target frequency band. The average frequency response gain is logarithmically transformed to obtain its decibel value, and the decibel value is inverted to obtain the initial EQ correction value for the k-th target frequency band. Determine whether the initial EQ correction value belongs to the preset safe adjustment frequency range. If it does not, truncate and limit the initial EQ correction value to obtain the EQ parameters of the kth target frequency band.

[0009] Preferably, the optimal calibration parameters include the gain adjustment values ​​for each channel, and the calculation of the optimal calibration parameters for each channel based on the acquired signal includes: The acquisition signal for each channel is obtained based on the acquisition signal and the independent signal for each channel; Extract the steady-state amplitude of the acquired signal from each channel, and calculate the average steady-state amplitude of the acquired signal from each channel, using it as the target equalization amplitude. The ratio of the target equalization amplitude to the steady-state amplitude of the acquired signal of the i-th channel is used as the gain adjustment value of the i-th channel.

[0010] Preferably, the optimal calibration parameters include the phase correction values ​​for each channel, and the calculation of the optimal calibration parameters for each channel based on the acquired signal includes: Obtain the phase shift of the acquired signal of the i-th channel relative to the independent signal of the i-th channel; Calculate the difference between the reference channel phase and the phase offset, and use it as the phase correction amount; The phase correction amount is mapped to the interval [-π, π] to obtain the phase correction value.

[0011] Preferably, calibrating each channel of the multi-channel audio device according to the optimal calibration parameters includes: The calibration parameters for the next time step are calculated based on the calibration parameters for the first time step; the calibration parameters for the next time step are greater than the calibration parameters for the first time step; the initial first time step is the current time step. After calculating the calibration parameters for the next time moment, the next time moment is taken as the new first time moment; The above steps are executed iteratively until the calculated calibration parameter at the next moment is greater than or equal to the optimal calibration parameter. The calibration parameters at each future moment within the calibration time period are obtained, and each channel of the multi-channel audio device is calibrated based on the calibration parameters at each future moment within the calibration time period.

[0012] In a second aspect, the present invention provides an environment-adaptive sound field calibration device, comprising: The reference signal extraction module is used to acquire the original digital audio signal played by the multi-channel audio device, and extract the transient strong signal and the independent signal of each channel from the original digital audio signal; and perform timing alignment and data encapsulation on the original digital audio signal, the transient strong signal and the independent signal of each channel to obtain the multi-channel reference signal. The signal acquisition module is used to acquire the current sound signal of the multi-channel audio device when the original digital audio signal is played by the multi-channel audio device, and record the current sound signal as the acquisition signal. An environmental acoustic feature extraction module is used to obtain the acoustic features of the acquired signal based on the acquired signal and the multi-channel reference signal; An acoustic feature comparison module is used to obtain a first deviation between the acoustic features of the acquired signal and the acoustic features of a reference sound signal acquired within a historical time period; the historical time period is the time period after the moment when the multi-channel audio device starts playing the original digital audio signal and before the current moment. The sound field calibration module is used to calculate the optimal calibration parameters for each channel based on the acquired signal when the first deviation is greater than the deviation threshold, and to calibrate each channel of the multi-channel audio device based on the optimal calibration parameters.

[0013] In a third aspect, the present invention provides an electronic device including a memory and a processor, wherein the memory stores computer program instructions that, when executed by the processor, implement the environmental adaptive sound field calibration method of the present invention.

[0014] In a fourth aspect, the present invention provides a computer-readable storage medium storing computer program instructions that, when executed by a processor, implement the environmental adaptive sound field calibration method of the present invention.

[0015] The embodiments of this application have the following beneficial effects: The method of the present invention collects the sound signals of the multi-channel audio equipment in the room in real time during the playback of the audio equipment, and calculates the calibration parameters based on the collected signals and reference signals to perform sound field calibration on the audio equipment, instead of calculating the calibration parameters by playing professional test sounds, thereby effectively improving the matching degree between the calibration results and the actual playback scenario.

[0016] Furthermore, by calculating the first deviation between the acoustic features of the acquired signal and the historical reference acoustic features, and triggering calibration parameter calculation and sound field calibration only when the deviation is greater than the deviation threshold, this invention constructs a real-time closed-loop dynamic tracking mechanism. This mechanism can accurately adapt to changes in the physical environment (such as opening and closing of doors and windows, and movement of people), while avoiding invalid recalibration caused by minor disturbances, thus saving computing resources of the underlying processor. Attached Figure Description

[0017] To more clearly illustrate the technical solutions of this application, the accompanying drawings used in the embodiments will be briefly described below. It should be understood that the following drawings only show some embodiments of this application and therefore should not be considered as a limitation on the scope of protection of this application. For those skilled in the art, other related drawings can be obtained based on these drawings without creative effort.

[0018] Figure 1 An environmental adaptive sound field calibration method according to an embodiment of the present invention is shown; Figure 2 A schematic diagram of the smooth gradient calibration method according to an embodiment of the present invention is shown; Figure 3 A schematic diagram of the structure of an environmental adaptive sound field calibration device according to an embodiment of the present invention is shown; Figure 4 A schematic diagram of an electronic device structure according to an embodiment of the present invention is shown. Detailed Implementation

[0019] The technical solutions in the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments.

[0020] The components of the embodiments of this application described and illustrated in the accompanying drawings can be arranged and designed in a variety of different configurations. Therefore, the following detailed description of the embodiments of this application provided in the drawings is not intended to limit the scope of the claimed application, but merely to illustrate selected embodiments of the application. All other embodiments obtained by those skilled in the art based on the embodiments of this application without inventive effort are within the scope of protection of this application.

[0021] Furthermore, the terms "first," "second," and "third" are used only to distinguish descriptions and should not be interpreted as indicating or implying relative importance.

[0022] Unless otherwise specified, all terms used herein (including technical and scientific terms) shall have the same meaning as commonly understood by one of ordinary skill in the art to which the various embodiments of this application pertain. Terms (such as those defined in commonly used dictionaries) shall be interpreted as having the same meaning as in their contextual meaning in the relevant technical field and shall not be construed as having an idealized or overly formal meaning, unless clearly defined in the various embodiments of this application.

[0023] The following detailed description of some embodiments of this application is provided in conjunction with the accompanying drawings. Unless otherwise specified, the following embodiments and features can be combined with each other.

[0024] Example of an environmental adaptive sound field calibration method: When calibrating multi-channel audio devices, the existing active sound field calibration method suffers from a low degree of matching between the calculated calibration parameters and the actual playback scene. The environmental adaptive sound field calibration method of this embodiment acquires the collected signal while the multi-channel audio device is playing program audio, and determines whether sound field calibration is needed based on the deviation between the acoustic characteristics of the collected signal and the acoustic characteristics of the reference sound signal collected in the historical time period. When sound field calibration is needed, the calibration parameters are directly calculated based on the collected signal. The entire calibration process does not play any dedicated test tone, but uses the sound emitted by the audio device playing program audio as a virtual probe signal, thereby effectively improving the matching degree between the calculated calibration parameters and the actual playback scene.

[0025] like Figure 1 As shown, this embodiment provides an environmental adaptive sound field calibration method, including: S101. Extracting transient strong signals and independent signals for each channel, specifically: acquiring the original digital audio signal played by the multi-channel audio device, and extracting transient strong signals and independent signals for each channel from the original digital audio signal.

[0026] In this embodiment, the original digital audio signal can be movie audio, music audio, etc. The independent signals for each channel include multiple signals such as the left channel independent signal, right channel independent signal, center channel independent signal, left surround channel independent signal, right surround channel independent signal, and panoramic channel independent signal. Transient strong signals refer to audio signals with high signal-to-noise ratio and obvious temporal characteristics. Extracting transient strong signals can improve the accuracy of environmental feature extraction. The methods for extracting transient strong signals include: (1) Perform sliding frame division processing on the original digital audio signal and calculate the short-time energy of each audio frame; the formula for calculating the short-time energy is: ; In the formula, This represents the short-time energy of the original digital audio signal in the nth frame, where N represents the frame length and m represents the local time offset within the nth audio frame. This represents the signal amplitude of the original digital audio signal at the (n+m)th position on the overall time axis.

[0027] (2) Obtain the average energy of a preset number of historical audio frames before the current audio frame, and initially mark audio frames whose short-term energy is greater than the product of the average energy and a preset multiple as high-energy candidate frames; In this embodiment, the preset quantity is 10, but in other embodiments, it can be other suitable quantities. The preset multiple ranges from 3 to 5 times.

[0028] (3) Determine whether each high-energy candidate frame belongs to a valid transient signal segment; for the nth high-energy candidate frame, the methods for determining whether it belongs to a valid transient signal segment include: First, a first-order difference operation is performed on the discrete points within the nth high-energy candidate frame to obtain the temporal difference magnitude between adjacent sampling points; the expression for calculating the temporal difference magnitude is as follows: ; In the formula, and These represent the temporal differential amplitude values ​​at the k-th sampling point of the n-th high-energy candidate frame, respectively. This represents the signal amplitude at the k-th sampling point and the (k-1)-th sampling point.

[0029] Then, in response to the rising trend of the temporal differential amplitude of the sampling points of a consecutive preset number within the nth high-energy candidate frame, it is determined that the nth high-energy candidate frame has transient impact characteristics; The time-domain difference magnitude of the i-th sampling point refers to the time-domain difference magnitude between the i-th sampling point and the (i-1)-th sampling point.

[0030] Next, the short-time zero-crossing rate (ZCR) of the nth high-energy candidate frame is calculated; the expression for calculating the ZCR is: ; In the formula, This represents the amplitude of the k-th sampling point within the n-th high-energy candidate frame. This represents the amplitude of the (k-1)th sampling point within the nth high-energy candidate frame. For a sign function, when If the value is greater than or equal to 0, the value is 1; otherwise, the value is 0.

[0031] Finally, in response to the short-time zero-crossing rate of the k-th high-energy candidate frame being lower than the preset zero-crossing rate threshold, the k-th high-energy candidate frame is determined to be a valid transient signal segment.

[0032] (4) Extract high-energy candidate frames belonging to the effective transient signal segment and configure corresponding timestamps for them to obtain the transient strong signal.

[0033] When extracting transient strong signals, a sliding frame-by-frame processing method is used, with the product of historical average energy and a preset multiple serving as the judgment criterion. This achieves efficient and adaptive locking of complex audio segments. This mechanism abandons the traditional fixed energy threshold, allowing the system to dynamically adapt to program sources with different macroscopic volumes, ranging from gentle classical music to action movies. This effectively avoids continuous false triggering due to excessive overall volume or missed detection due to insufficient volume. Simultaneously, this initial screening mechanism strictly limits subsequent complex microscopic feature calculations to "high-energy candidate frames," greatly saving the system's underlying computational power. By performing first-order difference operations on adjacent sampling points within candidate frames and determining whether they exhibit a continuous upward trend, accurate capture of the sound's physical "attack" envelope is achieved. Furthermore, in complex real-world environments, the equipment is highly susceptible to interference from high-energy, high-frequency noise such as high-frequency electronic background noise and environmental friction noise. Since the waveforms of this type of high-frequency noise cross the zero-level line at extremely high frequencies (i.e., high zero-crossing rate), this invention cleverly utilizes the physical acoustic characteristic that "strong transient signals usually have low zero-crossing rates and concentrated energy" to precisely filter out these pseudo-transient high-frequency noises by setting a low zero-crossing rate threshold, thereby greatly improving the signal-to-noise ratio and purity of the extracted signal.

[0034] In this embodiment, the extraction process of independent signals for each audio channel includes: First, determine the total number of channels of the multi-channel audio device and the arrangement order of each channel based on the audio frame header or channel layout configuration. Then, the pulse code modulation data sequence stored in interleaved format is extracted from the original digital audio signal; Next, based on the total number of channels and the arrangement order of each channel, the pulse code modulation data sequence is demultiplexed at the sampling point level; wherein, for any specific channel, the total number of channels is used as the sampling interval step size, and discrete sampling points matching the arrangement order of the specific channel are periodically extracted from the pulse code modulation data sequence to construct the corresponding mono sampling point sequence. For a multi-channel audio device with seven channels arranged in the order of left front channel, right front channel, center channel, left surround channel, right surround channel, low frequency channel, top left channel, and top right channel, when extracting discrete sampling points of the left front channel, the sampling interval step size is 7. Starting from the first data point in the pulse code modulation data sequence, one sampling point is extracted every 7 sampling points to obtain the discrete sampling point sequence of the left front channel.

[0035] Finally, the mono sampling point sequences corresponding to each channel are independently buffered and continuously reassembled to generate independent signals for each channel.

[0036] S102. Obtain a multi-channel reference signal, specifically by performing timing alignment and data encapsulation on the original digital audio signal, the transient strong signal, and each of the independent channel signals to obtain a multi-channel reference signal.

[0037] S103. Acquire the acquisition signal, specifically: when the multi-channel audio device is playing the original digital audio signal, acquire the current sound signal of the multi-channel audio device, and record the current sound signal as the acquisition signal.

[0038] In this embodiment, the actual sound signal in the room can be collected using a microphone array built into or external to the audio device, serving as the current sound signal of the multi-channel audio device. The process of the microphone array collecting the actual sound signal in the room includes: (1) Preprocessing: High-pass filtering (cutoff frequency 20Hz) and low-pass filtering (cutoff frequency 20000Hz) are applied to the acquired signal to filter out interference signals that are beyond the range of human hearing.

[0039] (2) Noise reduction: The NLMS (Normalized Least Mean Square) adaptive filtering algorithm is used to suppress steady-state noise (such as air conditioner and fan noise) and non-steady-state noise (such as sudden noise) in the environment, retain the effective reflected sound signal, and avoid noise interference with acoustic feature extraction.

[0040] (3) Signal synchronization: Based on the timing of the reference signal output by the reference signal extraction module, the time deviation of the acquired signal is corrected to ensure that the acquired signal is aligned with the time domain of the reference signal, thus providing a basis for subsequent comparative analysis.

[0041] (4) Amplitude adjustment: The processed acquisition signal is normalized to make its amplitude consistent with the reference signal, which facilitates subsequent signal comparison and analysis.

[0042] S104. Obtain the acoustic characteristics of the acquired signal based on the acquired signal and the multi-channel reference signal; Typically, the acoustic characteristics of sound signals collected during audio equipment playback include steady-state amplitude, phase response, room frequency response, reverberation time, direct sound delay of each channel, and reflection path delay distribution of each channel. The direct sound delay of each channel can be calculated using a cross-correlation function. Specifically, the calculation method for the direct sound delay and reflection path delay distribution of the i-th channel is as follows: First, construct the cross-correlation function between the channel-independent signal of the i-th channel and the acquired signal, the corresponding expression of which is: ; In the formula, This represents the cross-correlation function value between the channel-independent signal of the i-th channel and the acquired signal. This represents the amplitude of the j-th sampling point of the channel-independent signal of the i-th channel. This indicates that the acquired signal is shifted along the time axis. The signal amplitude after each sampling point.

[0043] Then, the relative time offset corresponding to the maximum value of the cross-correlation function is taken as the direct sound delay of the i-th channel.

[0044] Finally, peak detection is performed on the cross-correlation function. The time delays corresponding to all peaks except the one corresponding to the direct sound are the reflection path delay distribution. This reflection path delay distribution is used to analyze the room's reflected sound characteristics and provides a basis for calculating calibration parameters.

[0045] The room frequency response is calculated as follows: with reference signal and acquisition signal Perform Fourier transforms on each signal to obtain the frequency domain signal. and The ratio of the Fourier-transformed acquired signal to the Fourier-transformed reference signal is used as the room's frequency response. The room's frequency response characterizes the attenuation or gain characteristics of a room for sounds at different frequencies. This is the core basis for calibrating frequency response deviations, similar to the calibration logic for sound pressure sensitivity in professional acoustic laboratories, ensuring the accuracy of parameter extraction.

[0046] The reverberation time extraction method is as follows: based on the room frequency response, the energy attenuation method is adopted. By detecting the energy attenuation curve of the acquired signal, the reverberation time T60 (the time required for the sound energy to attenuate by 60dB) is obtained by fitting, which reflects the room reverberation characteristics and is used to optimize the spatial sense of the sound field.

[0047] S105. Obtain the first deviation between the acoustic characteristics of the acquired signal and the reference sound signal, specifically: obtain the first deviation between the acoustic characteristics of the acquired signal and the acoustic characteristics of the reference sound signal acquired within a historical time period; the historical time period is the time period after the moment when the multi-channel audio device starts playing the original digital audio signal and before the current moment. It is understood that, since the historical time period is after the moment when the multi-channel audio device starts playing the original digital audio signal and before the current moment, the reference sound signal collected within the historical time period is the sound signal of the multi-channel audio device.

[0048] In this embodiment, the length of the historical time period can range from 30s to 60s. The reference sound signal collected within the historical time period includes multiple sound signals, and the acoustic characteristics of each sound signal are averaged to obtain the acoustic characteristics of the reference sound signal.

[0049] When the acoustic environment in a room changes, the acoustic characteristics of the sound signals from the multi-channel audio devices will change. Therefore, by calculating the deviation between the acoustic characteristics of the sound signals from the multi-channel audio devices acquired in real time (i.e., the acoustic characteristics of the acquired signals) and the acoustic characteristics of the reference sound signals acquired in historical time periods, it is helpful to measure the amount of change in the acoustic environment in the room and determine whether subsequent sound field calibration actions need to be triggered based on the amount of change in the acoustic environment in the room.

[0050] As can be seen from step S104, the acoustic characteristics of the sound signal collected during the playback of a program by the audio device include types such as steady-state amplitude, phase frequency response, room frequency response, reverberation time, direct sound delay of each channel, and reflection path delay distribution of each channel. One or more types of acoustic characteristics can be selected from these types to calculate the first deviation between the acoustic characteristics of the collected signal and the acoustic characteristics of the reference sound signal collected in the historical time period.

[0051] In one embodiment, the room frequency response and the direct sound delay of each channel are selected from the above-mentioned various types of acoustic characteristics to calculate the first deviation between the acoustic characteristics of the acquired signal and the acoustic characteristics of the reference sound signal. The specific method includes: (1) Obtain the absolute value of the first difference between the direct sound delay of the i-th channel corresponding to the acquired signal and the direct sound delay of the i-th channel of the reference sound signal, and determine the ratio of the absolute value of the first difference to the direct sound delay of the i-th channel of the reference sound signal as the relative deviation rate of the i-th channel in the time domain; (2) Obtain the absolute value of the second difference between the room frequency response response corresponding to the acquired signal and the room frequency response response corresponding to the reference sound signal, and determine the ratio of the absolute value of the second difference to the room frequency response response corresponding to the reference sound signal as the frequency domain relative deviation amplitude. Perform a definite integral operation on the frequency domain relative deviation amplitude within a preset frequency range to obtain the frequency domain relative deviation integral value. (3) Add the time-domain relative deviation rate of the i-th channel to the frequency-domain relative deviation integral value to obtain the cumulative deviation value of the i-th channel; (4) Calculate the average of the cumulative deviation values ​​of each channel to obtain the first deviation.

[0052] In this embodiment, the first deviation The calculation expression is: ; In the formula, N is the number of channels of the multi-channel audio device. This represents the direct sound delay of the i-th channel corresponding to the acquired signal. This represents the direct sound delay of the i-th channel corresponding to the reference sound signal. This indicates the room frequency response corresponding to the acquired signal. This represents the room frequency response corresponding to the reference sound signal. Indicates the frequency of sound waves, and These represent the lower and upper limits of the preset frequency range, respectively.

[0053] For sound signals below 20Hz and above 20000Hz, the human ear cannot perceive them. To accurately cover the limits of human hearing and avoid wasting computational power, system interference noise is eliminated, and the lower and upper limits of the preset frequency range are set to 20Hz and 20000Hz, respectively. Setting the lower and upper limits of the preset frequency range to 20Hz and 20000Hz is equivalent to applying an ideal bandpass filter before calculating the first deviation. This ensures that only environmental changes that truly degrade the user's subjective auditory experience are included in the deviation assessment, effectively preventing the system from being misled by occasional noise in non-auditory frequency bands, thus avoiding frequent and meaningless calibrations.

[0054] By obtaining the absolute value of the first difference in direct sound time delay and the absolute value of the second difference in room frequency response, and dividing them by the corresponding reference values ​​to convert them into time-domain relative deviation rate and frequency-domain relative deviation amplitude, parameters with different physical properties and dimensions are scientifically normalized, eliminating weight imbalance during feature fusion. In addition, by performing definite integral calculation on the frequency-domain relative deviation amplitude within a preset frequency range and calculating the average of the accumulated multi-channel deviation values, the sudden noise interference at single-point frequencies and the occasional effects of physical blockage in one direction are effectively filtered out, enhancing the global stability of determining whether the environment has significantly deteriorated.

[0055] It should be noted that when calculating the first deviation between the acoustic features of the acquired signal and the acoustic features of the reference sound signal, one type of acoustic feature can be selected from the above-mentioned multiple types of acoustic features, and the deviation between the acoustic feature of the acquired signal corresponding to that type and the acoustic feature of the reference signal corresponding to that type can be calculated and used as the first deviation between the acoustic features of the acquired signal and the acoustic features of the reference sound signal. Alternatively, multiple acoustic features can be selected from the above-mentioned types of acoustic features, and the deviations of each type of acoustic feature between the acquired signal and the reference sound signal can be calculated separately. Then, the deviations of each type of acoustic feature are normalized and weighted and fused to obtain the first deviation between the acoustic features of the acquired signal and the acoustic features of the reference sound signal.

[0056] S106. Calculate the optimal calibration parameters and calibrate each channel of the multi-channel audio device. Specifically, if the first deviation is greater than the deviation threshold, calculate the optimal calibration parameters for each channel based on the acquired signal, and calibrate each channel of the multi-channel audio device based on the optimal calibration parameters.

[0057] In this embodiment, the deviation threshold can range from 0.1 to 0.2 and can be adaptively adjusted. The optimal calibration parameters for each channel can be one or more of the following: channel delay compensation value for each channel, EQ parameters for each target frequency band, gain adjustment value for each channel, and phase correction value for each channel.

[0058] In one embodiment, the optimal calibration parameters include channel delay compensation values ​​for each channel, and the calculation of the optimal calibration parameters for each channel based on the acquired signal includes: First, calculate the second deviation between the direct sound delay and the reference delay for each channel corresponding to the acquired signal, and select the largest deviation from all the second deviations. The reference delay is the center channel delay or the main channel delay. Then, the target difference between the maximum deviation and the second deviation of the i-th channel is calculated, and this target difference is used as the channel delay compensation value for the i-th channel. The expression for calculating the channel delay compensation value of the i-th channel is: ; In the formula, This represents the channel delay compensation value for the i-th channel. This represents the second deviation between the direct sound delay of the nth channel and the reference delay. This represents the second deviation between the direct sound delay of the i-th channel and the reference delay. This represents the maximum value function.

[0059] By calculating the second deviation between the direct arrival time delay of the acquired signal and the reference time delay, and selecting the maximum deviation, and then calculating the target difference between the maximum deviation and the second deviation of each channel as the channel delay compensation value, this invention strictly follows the physical limitation that time can only be "delayed" in audio signal processing. It forces all sound-emitting nodes of all channels in the system to be aligned with the physical channel that deviates most severely from the reference delay in space, thereby completely eliminating the sound image drift phenomenon and transient response degradation caused by inconsistent arrival times, and achieving accurate spatial sound image imaging.

[0060] In another embodiment, the optimal calibration parameters include not only the channel delay compensation values ​​for each channel, but also the EQ parameters for each target frequency band. The calculation of the optimal calibration parameters for each channel based on the acquired signals further includes: (1) For the kth target frequency band, extract the frequency range that falls within the room frequency response corresponding to the acquired signal. The amplitude of each discrete frequency point within the target band is calculated, and the average amplitude of all discrete frequency points is taken to obtain the average frequency response gain of the k-th target frequency band. The corresponding calculation expression is: ; In the formula, M represents the average frequency response gain of the k-th target frequency band, and M represents the frequency range of the room frequency response corresponding to the acquired signal that falls within the specified frequency range. The number of discrete frequency points within the range Indicates the frequency of sound waves, This represents the lower frequency limit of the k-th target frequency band. This represents the upper frequency limit of the k-th target frequency band.

[0061] (2) Perform a logarithmic domain transformation on the average frequency response gain to obtain its decibel value, and invert the decibel value to obtain the initial EQ correction value of the kth target frequency band.

[0062] The expression for calculating the initial EQ correction value of the k-th target frequency band is as follows: ; In the formula, This represents the initial EQ correction value for the k-th target frequency band. This represents the average frequency response gain of the k-th target frequency band.

[0063] (3) Determine whether the initial EQ correction value belongs to the preset safe adjustment frequency range. If it does not belong, then the initial EQ correction value is truncated and limited to obtain the EQ parameters of the kth target frequency band.

[0064] In this embodiment, the entire frequency band can be divided into eight target frequency bands: 20Hz to 100Hz, 100Hz to 200Hz, 200Hz to 500Hz, 500Hz to 1kHz, 1kHz to 2kHz, 2kHz to 5kHz, 5kHz to 10kHz, and 10kHz to 20kHz. For each target frequency band, the corresponding EQ parameters are calculated using the method described above. By dividing the entire frequency band into these eight target frequency bands, the algorithm can be ensured to have sufficient resolution in the low-frequency band to suppress room standing waves and strong macroscopic trend fitting ability in the high-frequency band.

[0065] By calculating the initial EQ correction value by averaging the amplitudes of discrete frequency points within the target frequency band, this invention effectively smooths and cancels out local spectral distortion caused by room resonance. More importantly, by determining whether the initial EQ correction value falls within a preset safe adjustment frequency range and performing truncation and amplitude limiting when it does not, this invention prevents the system from forcibly allocating excessively high digital gain when encountering extremely deep acoustic standing wave troughs (destructive interference points), thereby fundamentally eliminating the hardware risks of amplifying low-level noise and burning out speaker units due to peak clipping distortion.

[0066] In another embodiment, the optimal calibration parameters include, in addition to the channel delay compensation values ​​for each channel and the EQ parameters for each target frequency band, the gain adjustment values ​​for each channel. The calculation of the optimal calibration parameters for each channel based on the acquired signals further includes: First, the acquisition signal of each channel is obtained based on the acquisition signal and the independent signal of each channel; Then, the steady-state amplitude of the acquired signal from each channel is extracted, and the average steady-state amplitude of the acquired signal from each channel is calculated, which is used as the target equalization amplitude; the calculation expression for the target equalization amplitude is: ; In the formula, Indicates the target equilibrium amplitude. represents the steady-state amplitude of the acquired signal of the i-th channel, and N represents the total number of channels.

[0067] Finally, the ratio of the target equalization amplitude to the steady-state amplitude of the acquired signal of the i-th channel is used as the gain adjustment value of the i-th channel. The corresponding calculation expression is: ; In the formula, This represents the gain adjustment value for the i-th channel. Indicates the target equilibrium amplitude. This represents the steady-state amplitude of the acquired signal for the i-th channel.

[0068] By extracting the steady-state amplitude of the acquired signals from each channel and calculating the average as a unified target equalization amplitude, and then comparing it with the steady-state amplitude of the corresponding mono channel, this invention generates a high-precision product-type gain adjustment coefficient. This mechanism, targeting the average level of global acoustic energy, precisely compensates for unilateral acoustic energy collapse caused by inconsistent speaker placement distances or asymmetrical sound-absorbing materials in the room, thus reshaping a balanced, symmetrical, and highly immersive three-dimensional sound field energy distribution in a multi-channel system.

[0069] In another embodiment, the optimal calibration parameters include, in addition to the gain adjustment value of each channel, the channel delay compensation value of each channel, and the EQ parameters of each target frequency band, the phase correction value of each channel. The calculation of the optimal calibration parameters for each channel based on the acquired signal further includes: First, obtain the phase shift of the acquired signal of the i-th channel relative to the independent signal of the i-th channel; Then, the difference between the reference channel phase and the phase offset is calculated and used as the phase correction amount; the corresponding calculation expression is: ; In the formula, This represents the phase correction amount for the i-th channel. Indicates the reference channel phase. This represents the phase shift of the acquired signal of the i-th channel relative to the independent signal of the i-th channel.

[0070] Finally, the phase correction amount is mapped to the interval [-π, π] to obtain the phase correction value.

[0071] By calculating the phase correction value between the reference channel phase and the relative offset difference of the acquired signals of each channel, and mathematically mapping it to the [-π,π] interval, this invention forcibly eliminates the comb filter interference caused by phase disorder in the cross-coverage area of ​​multiple physical sound sources, significantly improving the solidity and power of the low-frequency band. At the same time, this interval mapping mechanism perfectly matches the operating characteristics of the underlying digital all-pass filter, ensuring the accurate transmission and execution of phase control commands.

[0072] After calculating the optimal calibration parameters for each channel, the calibration of each channel of the multi-channel audio device based on the optimal calibration parameters can be performed in two ways: one-step calibration and smooth gradual calibration. One-step calibration refers to directly calibrating each channel of the multi-channel audio device based on the optimal calibration parameters. Smooth gradual calibration refers to setting multiple future calibration times within a future time period, performing calibration on each channel at each future calibration time, and having the calibration parameters corresponding to each future calibration time have an increasing or decreasing trend, thereby gradually transitioning the sound field of each channel of the multi-channel audio device from the current state to the sound field corresponding to the optimal calibration parameters.

[0073] When calibrating each channel of a multi-channel audio device according to the aforementioned optimal calibration parameters, the calibration parameters for each channel can be transmitted asynchronously to the digital signal processor (DSP) of the audio device. At calibration time, the DSP's operating parameters are updated based on the calibration parameters, thereby achieving real-time sound field calibration. Asynchronous transmission avoids consuming the audio processing resources of the DSP, ensuring uninterrupted and smooth audio playback. The method of updating the DSP's operating parameters based on the calibration parameters is existing technology and will not be elaborated upon here.

[0074] After calibrating each channel of the multi-channel audio device, a new acquisition signal can be obtained every first preset time interval, and a reference sound signal acquired within a new historical time period can be obtained every second preset time interval. The acoustic characteristics of the new acquisition signal are then compared with the acoustic characteristics of the reference sound signal acquired within the new historical time period, resulting in a first deviation. If the first deviation exceeds a deviation threshold, new optimal calibration parameters are recalculated to calibrate the multi-channel audio device. This allows for seamless real-time dynamic calibration of the multi-channel audio device.

[0075] Existing active sound field calibration methods involve playing professional test tones and using these tones to calculate calibration parameters for calibrating the sound field of audio equipment. In contrast, the method in this embodiment involves real-time acquisition of sound signals from multi-channel audio equipment in the room during program playback. The method then uses the acquired signals in conjunction with reference signals to calculate calibration parameters for sound field calibration of the audio equipment, thereby effectively improving the matching degree between the calibration results and the actual playback scenario.

[0076] Furthermore, by calculating the first deviation between the acoustic features of the acquired signal and the historical reference acoustic features, and triggering calibration parameter calculation and sound field calibration only when the deviation is greater than the deviation threshold, this invention constructs a real-time closed-loop dynamic tracking mechanism. This mechanism can accurately adapt to changes in the physical environment (such as opening and closing of doors and windows, and movement of people), while avoiding invalid recalibration caused by minor disturbances, thus saving computing resources of the underlying processor.

[0077] Furthermore, existing active sound field calibration techniques require interrupting normal audio playback before calibration. Whether watching movies or listening to music, this is disrupted by actively emitted test tones (sweep signals, impulse noise, etc.), damaging the continuity of use. Moreover, the test tones themselves contain significant noise, making calibration unsuitable for noise-sensitive scenarios such as nighttime or quiet movie watching. Since the method in this embodiment does not require playing professional test tones during calibration, it eliminates the need to interrupt normal audio playback, thus ensuring a better user experience when playing programs using audio devices.

[0078] To ensure that each channel of a multi-channel audio device is aligned in terms of time delay, balanced in terms of sound level (sound level difference ≤ 1dB), and consistent in terms of phase (phase difference ≤ 10°) during calibration, thereby improving the spatial consistency and immersion of the sound field, in one embodiment, the environmental adaptive sound field calibration method further includes: applying cooperative constraints to each channel during the calibration process. These cooperative constraints include: time delay alignment constraints, sound level balance constraints, phase consistency constraints, and sound image position cooperative constraints. The time delay alignment constraint ensures that all channel sounds arrive at the listening point simultaneously. The sound level balance constraint ensures that the sound level difference between any two channels is less than or equal to a preset sound intensity threshold. The phase consistency constraint ensures that the phase difference between any two channels is less than or equal to a phase difference threshold. The sound image position cooperative constraint ensures that the front channel maintains a positive positioning, the surround channel maintains a lateral sense of enclosure, and the sky channel maintains an upward sense of space.

[0079] In this embodiment, the preset sound intensity threshold can be 1dB or other suitable sound intensity, and the phase difference threshold can be 10 degrees or other suitable angle value.

[0080] The expression corresponding to the time delay alignment constraint is: ; In the formula, and Let $\mathbf{i}$ and $\mathbf{j}$ represent the actual transmission delays for the sound from the $i$-th channel to reach the target listening point, respectively. This indicates that the absolute value is true. This represents a sampling period.

[0081] The expression corresponding to the sound level balance constraint is: ; In the formula, and Let represent the gain adjustment value of the i-th channel and the gain adjustment value of the j-th channel, respectively, and log represents the logarithmic function.

[0082] The expression corresponding to the phase consistency constraint is:

[0083] In the formula, and These represent the phases of the i-th and j-th audio channels, respectively.

[0084] As can be seen from the above embodiments, the calibration of each channel of a multi-channel audio device according to the optimal calibration parameters can be performed in one step or in a smooth, gradual manner. To avoid sudden changes in sound and popping noises, in one embodiment, a smooth, gradual calibration method is used to calibrate the sound field of each channel of the multi-channel audio device. Figure 2 As shown, the calibration process includes: S201. Calculate the calibration parameters for the next time step after the first time step, specifically: calculate the calibration parameters for the next time step after the first time step based on the calibration parameters for the first time step; the calibration parameters for the next time step are greater than the calibration parameters for the first time step; the initial first time step is the current time step. In one embodiment, the calculation expression for calculating the calibration parameters of the next time step based on the calibration parameters of the first time step is as follows: ; In the formula, and These represent the calibration parameters at the first time point and the calibration parameters at the time point immediately following the first time point, respectively. Represents the smoothing coefficient. This represents the optimal calibration parameters. Greater than 0 and less than 1, preferably, The value range is from 0.05 to 0.1.

[0085] S202. After calculating the calibration parameters for the next time moment, the next time moment is taken as the new first time moment. S203. The calibration parameters for each future moment within the calibration time period are obtained, and each channel of the multi-channel audio device is calibrated. Specifically, the above steps are executed iteratively until the calculated calibration parameter for the next moment is greater than or equal to the optimal calibration parameter. The calibration parameters for each future moment within the calibration time period are obtained, and each channel of the multi-channel audio device is calibrated based on the calibration parameters for each future moment within the calibration time period.

[0086] By using the calibration parameters at the current moment as a starting point, iteratively calculating multiple calibration parameters at subsequent moments that gradually approach the optimal calibration parameters, and using these parameter sequences for step-by-step transition calibration, this invention abandons the traditional "one-click" hard parameter switching logic. This "gradual in and gradual out" flexible parameter distribution mechanism smoothly dilutes the adjustment process of gain, delay, and EQ on the time axis, effectively avoiding digital pops, sudden volume changes, or imaging shifts caused by parameter mutations, and ensuring that the adaptive reconstruction process is absolutely seamless and natural for the user's subjective listening experience.

[0087] Because different music or movie soundtracks may introduce slight digital domain baseline shifts during mixing, mastering, or lossy compression decoding (such as Dolby / DTS decoding) due to asymmetric clipping or improper digital filter design, these extracted signals will be encapsulated as "reference signals" in subsequent steps for cross-correlation calculations with the "acquired signals." If the digital source itself has a slight shift, it will cause an overall rise in the cross-correlation sequence, increasing the risk of errors in time delay extraction. In one embodiment, to ensure the stability of the reference signal and the accuracy of the acoustic feature calculation results of the subsequent acquired signals, the environmental adaptive sound field calibration method further includes: after extracting transient strong signals and independent signals for each channel from the original digital audio signal, performing DC component removal processing on the extracted signals, wherein DC component removal is achieved by frame-by-frame mean subtraction, and a certain extracted signal is recorded as the target audio signal. The specific steps for DC component removal of the target audio signal are as follows: S301. Obtain the current audio frame of the target audio signal, and calculate the arithmetic mean of the original signal amplitudes of all discrete sampling points contained in the current audio frame, using it as the DC offset of the current audio frame; the corresponding calculation expression is: ; In the formula, This represents the DC offset of the current audio frame, where N represents the total number of discrete points contained within the current audio frame. This represents the amplitude of the nth discrete sampling point.

[0088] S302. Subtract the DC offset from the original signal amplitude of each discrete sampling point within the current audio frame to obtain the amplitude correction value of each discrete sampling point; the corresponding calculation expression is: ; In the formula, This represents the amplitude correction value for the nth discrete sampling point. This represents the amplitude of the nth discrete sampling point. This represents the DC offset of the current audio frame.

[0089] S303. The amplitude correction values ​​of each discrete sampling point are sequentially input into a preset first-order high-pass filter for low-frequency jitter suppression processing, and the final zero-mean amplitude of each discrete sampling point is output, thereby realizing the removal of DC components from the benchmark audio signal; wherein, the calculation expression of the first-order high-pass filter is: ; In the formula, This represents the final zero-mean magnitude of the nth discrete sampling point of the first-order high-pass filter output. This represents the final zero-mean magnitude of the (n-1)th discrete sample point of the high-pass filter output. and Let represent the magnitudes of the nth discrete sampling point and the (n-1)th discrete sampling point, respectively, and let represent the dot product.

[0090] This embodiment calculates the arithmetic mean of the original signal amplitude at each sampling point within the current audio frame and performs point-by-point subtraction. This accurately and quickly eliminates the inherent static baseline shift in the original digital audio data, initially constructing a strict zero-mean signal benchmark. Based on this, the pre-debiased amplitude correction value is further input into a preset first-order high-pass filter for cascaded processing. The filter's feedback mechanism effectively suppresses any residual extremely low-frequency baseline drift and dynamic jitter in the signal sequence. Overall, this dual DC removal mechanism, combining frame-by-frame mean subtraction with first-order high-pass filtering, thoroughly eliminates non-ideal DC components and extremely low-frequency interference in the audio signal, ensuring absolute centering and stability of the signal waveform on the global time axis. This provides extremely pure underlying input data for subsequent high-precision short-time energy calculation and cross-correlation delay feature extraction, fundamentally eliminating energy misjudgment and feature matching errors caused by baseline drift, and significantly improving the measurement accuracy and algorithm robustness of the entire sound field calibration system.

[0091] To eliminate the energy magnitude difference between the reference source signal and the acquired signal, and to ensure that subsequent cross-correlation analysis and frequency response calculation depend only on waveform characteristics and spatial acoustic properties, thereby improving the robustness of the algorithm under different playback volumes, in one embodiment, the environmental adaptive sound field calibration method further includes: after extracting transient strong signals and independent signals for each channel from the original digital audio signal, normalizing the extracted signals, and designating one of the extracted signals as the target audio signal. The specific process of normalizing the target audio signal includes: S401. Obtain the current audio frame of the target audio signal, and iterate through and calculate the absolute amplitude of each discrete sampling point within the current audio frame, extracting the maximum absolute amplitude; the corresponding calculation expression is: ; In the formula, Indicates the maximum absolute amplitude. This represents the amplitude correction value for the nth discrete sampling point. Represents the maximum value function. Represents the absolute value symbol.

[0092] S402. Calculate the ratio of the preset target normalized amplitude to the maximum absolute amplitude, and use this ratio as the current calculated gain of the current audio frame.

[0093] By extracting the maximum absolute amplitude of the target audio signal frame by frame and combining it with the target normalized amplitude to calculate the current computational gain, the safety constraints of the signal dynamic range and the unification of the measurement scale are realized, effectively avoiding low-level computation overflow or weak signals being drowned out by background noise.

[0094] S403. Calculate the normalized amplitude of each sampling point of the current audio frame based on the current calculated gain; In one implementation, the calculation is performed only based on the current computational gain, and the corresponding calculation expression is: ; In the formula, This represents the normalized amplitude of the nth sample point of the current audio frame. This represents the amplitude of the nth sample point of the current audio frame before normalization. This represents the maximum absolute amplitude. This represents the preset target normalized amplitude. This indicates dot product.

[0095] In another embodiment, instead of directly calculating based on the current calculated gain, the current calculated gain and the gain of the previous audio frame are weighted and summed to obtain the current frame gain coefficient, and then the normalized amplitude of each sampling point of the current audio frame is calculated using the current frame gain coefficient.

[0096] By innovatively introducing a historical smoothing gain mechanism, the historical gain of the previous audio frame and the current calculated gain are weighted and summed using preset weights to obtain the smoothing gain coefficient of the current frame, thus completely eliminating the inter-frame amplitude discontinuity and "distorted sound" abrupt changes caused by directly applying independent frame gains.

[0097] The target audio signal normalization process provided in this embodiment extracts the maximum absolute amplitude frame by frame and calculates the current calculated gain by combining it with the preset target normalized amplitude. This achieves a safe constraint on the dynamic range of the signal and a global uniformity of the measurement scale, effectively avoiding the problems of low-level computational overflow or weak signals being submerged by background noise. At the same time, this process constructs a flexible implementation architecture when applying gain. It supports both direct use of the current calculated gain for basic single-frame linear scaling and obtaining the current frame gain coefficient by weighted summation of the current calculated gain and the gain of the previous audio frame. This inter-frame weighted smoothing mechanism can effectively eliminate the gain gap and amplitude abrupt change caused by drastic energy fluctuations between adjacent frames. Thus, while ensuring accurate standardization and alignment of the underlying data, it perfectly takes into account the smoothness and coherence of the audio envelope temporal domain transition, significantly improving the stability and anti-interference capability of subsequent acoustic feature extraction.

[0098] Example of an environmental adaptive sound field calibration device: like Figure 3 As shown, this application also provides an environmental adaptive sound field calibration device, comprising: The reference signal extraction module 110 is used to acquire the original digital audio signal played by the multi-channel audio device, and extract the transient strong signal and the independent signal of each channel from the original digital audio signal; and perform timing alignment and data encapsulation on the original digital audio signal, the transient strong signal and the independent signal of each channel to obtain the multi-channel reference signal. The signal acquisition module 120 is used to acquire the current sound signal of the multi-channel audio device when the original digital audio signal is played by the multi-channel audio device, and record the current sound signal as the acquisition signal. An environmental acoustic feature extraction module 130 is used to obtain the acoustic features of the acquired signal based on the acquired signal and the multi-channel reference signal. The acoustic feature comparison module 140 is used to obtain a first deviation between the acoustic features of the acquired signal and the acoustic features of a reference sound signal acquired within a historical time period; the historical time period is the time period after the moment when the multi-channel audio device starts playing the original digital audio signal and before the current moment. The sound field calibration module 150 is used to calculate the optimal calibration parameters for each channel based on the acquired signal when the first deviation is greater than the deviation threshold, and to calibrate each channel of the multi-channel audio device based on the optimal calibration parameters.

[0099] It is understood that the environmental adaptive sound field calibration device in this embodiment corresponds to the environmental adaptive sound field calibration method in the above embodiment. The options in the above embodiment are also applicable to this embodiment, so they will not be described again here.

[0100] Electronic device example: like Figure 4 As shown, this application also provides an electronic device, exemplary of which includes a processor and a memory, wherein the memory stores a computer program, and the processor, by running the computer program, causes the computer device to perform the functions of the various modules in the above-described environmental adaptive sound field calibration method or the above-described environmental adaptive sound field calibration device.

[0101] The processor can be an integrated circuit chip with signal processing capabilities. The processor can be a general-purpose processor, including at least one of a Central Processing Unit (CPU), Graphics Processing Unit (GPU), Network Processor (NP), Digital Signal Processor (DSP), Application-Specific Integrated Circuit (ASIC), Field-Programmable Gate Array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components. The general-purpose processor can be a microprocessor or any conventional processor, capable of implementing or executing the methods, steps, and logic block diagrams disclosed in the embodiments of this application.

[0102] The memory can be, but is not limited to, Random Access Memory (RAM), Read Only Memory (ROM), Programmable Read-Only Memory (PROM), Erasable Programmable Read-Only Memory (EPROM), Electrically Erasable Programmable Read-Only Memory (EEPROM), etc. The memory is used to store computer programs, and the processor can execute the computer programs accordingly after receiving execution instructions.

[0103] Examples of computer storage media: This application also provides a computer storage medium for storing the computer program instructions used in the aforementioned electronic device. The computer storage medium can be a readable storage medium, a non-volatile storage medium, or a volatile storage medium. For example, the computer storage medium may include, but is not limited to, various media capable of storing program code, such as a USB flash drive, a portable hard drive, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk.

[0104] In the several embodiments provided in this application, it should be understood that the disclosed apparatus and methods can also be implemented in other ways. The apparatus embodiments described above are merely illustrative. For example, the flowcharts and block diagrams in the accompanying drawings show the architecture, functionality, and operation of possible implementations of apparatus, methods, and computer program products according to various embodiments of this application. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that, in alternative implementations, the functions marked in the blocks may occur in a different order than those marked in the drawings. For example, two consecutive blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagram and / or flowchart, and combinations of blocks in the block diagram and / or flowchart, can be implemented using a dedicated hardware-based system that performs the specified function or action, or using a combination of dedicated hardware and computer instructions.

[0105] In addition, the functional modules or units in the various embodiments of this application can be integrated together to form an independent part, or each module can exist independently, or two or more modules can be integrated to form an independent part.

[0106] If the aforementioned functions are implemented as software functional modules and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a smartphone, personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this application.

[0107] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any changes or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application.

Claims

1. An environmental adaptive sound field calibration method, characterized in that, include: Acquire the raw digital audio signal played by a multi-channel audio device, and extract transient strong signals and independent signals of each channel from the raw digital audio signal; The original digital audio signal, the transient strong signal, and each of the channel-independent signals are time-aligned and data-encapsulated to obtain a multi-channel reference signal. When the original digital audio signal is played by the multi-channel audio device, the current sound signal of the multi-channel audio device is collected, and the current sound signal is recorded as the collected signal. The acoustic characteristics of the acquired signal are obtained based on the acquired signal and the multi-channel reference signal; The first deviation between the acoustic characteristics of the acquired signal and the acoustic characteristics of a reference sound signal acquired within a historical time period is obtained; the historical time period is the time period after the moment when the multi-channel audio device starts playing the original digital audio signal and before the current moment. If the first deviation is greater than the deviation threshold, the optimal calibration parameters for each channel are calculated based on the acquired signal, and each channel of the multi-channel audio device is calibrated based on the optimal calibration parameters.

2. The environmental adaptive sound field calibration method according to claim 1, characterized in that, The acoustic characteristics include room frequency response and direct sound delay of each channel. The calculation method of the first deviation includes: obtaining the absolute value of the first difference between the direct sound delay of the i-th channel corresponding to the acquired signal and the direct sound delay of the i-th channel of the reference sound signal, and determining the ratio of the absolute value of the first difference to the direct sound delay of the i-th channel of the reference sound signal as the relative deviation rate of the i-th channel in the time domain. The absolute value of the second difference between the room frequency response response corresponding to the acquired signal and the room frequency response response corresponding to the reference sound signal is obtained, and the ratio of the absolute value of the second difference to the room frequency response response corresponding to the reference sound signal is determined as the frequency domain relative deviation amplitude. The frequency domain relative deviation amplitude is integrally calculated within a preset frequency range to obtain the frequency domain relative deviation integral value. The relative deviation rate in the time domain of the i-th channel is added to the integral value of the relative deviation in the frequency domain to obtain the cumulative deviation value of the i-th channel. The first deviation is obtained by averaging the cumulative deviation values ​​of each channel.

3. The environmental adaptive sound field calibration method according to claim 2, characterized in that, The optimal calibration parameters include channel delay compensation values ​​for each channel, and the calculation of the optimal calibration parameters for each channel based on the acquired signals includes: Calculate the second deviation between the direct sound delay and the reference delay for each channel corresponding to the acquired signal, and select the largest deviation from all the second deviations. The reference delay is the center channel delay or the main channel delay. Calculate the target difference between the maximum deviation and the second deviation of the i-th channel, and use the target difference as the channel delay compensation value of the i-th channel.

4. The environmental adaptive sound field calibration method according to claim 3, characterized in that, The optimal calibration parameters include the EQ parameters for each target frequency band, and the calculation of the optimal calibration parameters for each channel based on the acquired signals further includes: For the k-th target frequency band, extract the frequency range that falls within the room frequency response corresponding to the acquired signal. The amplitude of each discrete frequency point within the target frequency band is calculated, and the average amplitude of all discrete frequency points is calculated to obtain the average frequency response gain of the k-th target frequency band. The average frequency response gain is logarithmically transformed to obtain its decibel value, and the decibel value is inverted to obtain the initial EQ correction value for the k-th target frequency band. Determine whether the initial EQ correction value belongs to the preset safe adjustment frequency range. If it does not, truncate and limit the initial EQ correction value to obtain the EQ parameters of the kth target frequency band.

5. The environmental adaptive sound field calibration method according to claim 1, characterized in that, The optimal calibration parameters include the gain adjustment values ​​for each channel, and the calculation of the optimal calibration parameters for each channel based on the acquired signal includes: The acquisition signal for each channel is obtained based on the acquisition signal and the independent signal for each channel; Extract the steady-state amplitude of the acquired signal from each channel, and calculate the average steady-state amplitude of the acquired signal from each channel, using it as the target equalization amplitude. The ratio of the target equalization amplitude to the steady-state amplitude of the acquired signal of the i-th channel is used as the gain adjustment value of the i-th channel.

6. The environmental adaptive sound field calibration method according to claim 1, characterized in that, The optimal calibration parameters include the phase correction values ​​for each channel, and the calculation of the optimal calibration parameters for each channel based on the acquired signals includes: Obtain the phase shift of the acquired signal of the i-th channel relative to the independent signal of the i-th channel; Calculate the difference between the reference channel phase and the phase offset, and use it as the phase correction amount; The phase correction amount is mapped to the interval [-π, π] to obtain the phase correction value.

7. The environmental adaptive sound field calibration method according to claim 1, characterized in that, The calibration of each channel of the multi-channel audio device based on the optimal calibration parameters includes: The calibration parameters for the next time step are calculated based on the calibration parameters for the first time step; the calibration parameters for the next time step are greater than the calibration parameters for the first time step; the initial first time step is the current time step. After calculating the calibration parameters for the next time moment, the next time moment is taken as the new first time moment; The above steps are executed iteratively until the calculated calibration parameter at the next moment is greater than or equal to the optimal calibration parameter. The calibration parameters at each future moment within the calibration time period are obtained, and each channel of the multi-channel audio device is calibrated based on the calibration parameters at each future moment within the calibration time period.

8. An environmentally adaptive sound field calibration device, characterized in that, include: The reference signal extraction module is used to acquire the original digital audio signal played by the multi-channel audio device, and extract the transient strong signal and the independent signal of each channel from the original digital audio signal; and perform timing alignment and data encapsulation on the original digital audio signal, the transient strong signal and the independent signal of each channel to obtain the multi-channel reference signal. The signal acquisition module is used to acquire the current sound signal of the multi-channel audio device when the original digital audio signal is played by the multi-channel audio device, and record the current sound signal as the acquisition signal. An environmental acoustic feature extraction module is used to obtain the acoustic features of the acquired signal based on the acquired signal and the multi-channel reference signal; An acoustic feature comparison module is used to obtain a first deviation between the acoustic features of the acquired signal and the acoustic features of a reference sound signal acquired within a historical time period; the historical time period is the time period after the moment when the multi-channel audio device starts playing the original digital audio signal and before the current moment. The sound field calibration module is used to calculate the optimal calibration parameters for each channel based on the acquired signal when the first deviation is greater than the deviation threshold, and to calibrate each channel of the multi-channel audio device based on the optimal calibration parameters.

9. An electronic device comprising a memory and a processor, wherein the memory stores computer program instructions, characterized in that, When the computer program instructions are executed by the processor, the environmental adaptive sound field calibration method according to any one of claims 1 to 7 is implemented.

10. A computer storage medium storing computer program instructions, characterized in that, When the computer program instructions are executed by the processor, the environmental adaptive sound field calibration method according to any one of claims 1 to 7 is implemented.