A speaker self-calibration method and system based on cooperation of a webpage end and a speaker end

CN122317524BActive Publication Date: 2026-08-11CHENGDU SHUIYUEYU TECHNOLOGY CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2026-06-01
Publication Date
2026-08-11

AI Technical Summary

Technical Problem

传统的音箱自校准方案通常依赖于昂贵的专业声学测量设备,如声卡、标准麦克风及消声室数据,并由专业人员进行现场调试

Benefits of technology

1、本发明提供的一种基于网页端与音箱端协同的音箱自校准方法,通过将解卷积、空间平均、麦克风补偿及PEQ参数搜索任务卸载至网页端执行,仅保留扫频播放与滤波执行功能在音箱端侧,使得低成本、低算力的消费级音箱无需搭载昂贵的嵌入式DSP芯片即可实现专业级的空间声学校准,这种任务分工不仅大幅降低了硬件成本,还通过网页端的强交互能力实现了校准过程的实时可视化与用户可控性,实现了普通用户的自助式高精度调音。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122317524B_ABST
    Figure CN122317524B_ABST
Patent Text Reader

Abstract

This invention discloses a speaker self-calibration method and system based on web-based and speaker-based collaboration, belonging to the field of speaker calibration technology. The key technical points are: sending control commands to the speaker; acquiring recording signals from a calibration microphone; performing robust deconvolution processing with regularization on the reference sweep signal and the recording signal to obtain the original frequency response curve; performing energy domain averaging on multiple original frequency response curves to generate a spatial average frequency response; performing differential compensation on the spatial average frequency response to obtain the compensated frequency response; generating PEQ filter parameters based on the compensated frequency response and the target curve using an automatic parametric equalizer search algorithm; calculating the maximum gain value after cascading the PEQ filters, and calculating the safe attenuation value of the preceding stage before sending it to the speaker. This invention achieves professional-grade spatial acoustic calibration without the need for expensive embedded DSP chips, enabling ordinary users to perform self-service high-precision sound tuning.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of speaker calibration technology, and more specifically, to a speaker self-calibration method and system based on collaboration between a web page and the speaker itself. Background Technology

[0002] With the increasing popularity of smart speakers and multimedia audio devices, users have placed higher demands on the consistency of sound quality across different listening environments. Traditional speaker self-calibration solutions typically rely on expensive professional acoustic measurement equipment, such as sound cards, standard microphones, and anechoic chamber data, and require on-site adjustments by professionals. While integrating a digital signal processor into the speaker to execute calibration algorithms is an option, this end-to-end solution places extremely high demands on the speaker's hardware computing power, resulting in high costs and an inability to adapt to complex acoustic variations in non-professional environments.

[0003] To address this, some existing technologies attempt to use the built-in microphone of the television for calibration. However, because the microphone and speaker are fixed in position and too close, they cannot reflect the true frequency response of the user's actual listening position, resulting in distortion in the calibrated sound during practical applications. Furthermore, existing calibration logic often lacks the ability to fine-tune the acoustic characteristics of the room. On the one hand, for the problem of poor consistency among mid-to-low-end speakers, the traditional method of manually adjusting equalizer (EQ) parameters one by one is inefficient and highly subjective. On the other hand, existing technologies have shortcomings in frequency response data processing. For example, directly averaging decibel values ​​can lead to excessive weighting of valley frequencies, failing to scientifically characterize the overall acoustic features of the listening area. Simultaneously, conventional automatic EQ search algorithms are prone to getting trapped in local optima and lack protection mechanisms against digital signal overflow, resulting in limited improvement in sound quality after calibration or even clipping distortion.

[0004] Therefore, researching and designing a speaker self-calibration method and system based on the collaboration between the web page and the speaker can overcome the above-mentioned defects, which is an urgent problem to be solved. Summary of the Invention

[0005] To address the shortcomings of existing technologies, the present invention aims to provide a speaker self-calibration method and system based on collaboration between a web page and the speaker. By offloading deconvolution, spatial averaging, microphone compensation, and PEQ parameter search tasks to the web page for execution, and retaining only frequency sweep playback and filtering functions on the speaker side, low-cost, low-computing-power consumer speakers can achieve professional-grade spatial acoustic calibration without the need for expensive embedded DSP chips. This task division not only significantly reduces hardware costs but also enables real-time visualization and user control of the calibration process through the strong interactivity of the web page, achieving self-service high-precision sound tuning for ordinary users.

[0006] The above-mentioned technical objective of the present invention is achieved through the following technical solution: Firstly, a speaker self-calibration method based on collaboration between a web page and the speaker is provided, executed by the web page, including the following steps: Send control commands to the speaker to control the speaker to play the reference sweep signal at a set sampling rate; Acquire the recording signal collected by the calibration microphone, wherein the recording signal is the signal collected by the calibration microphone after the reference sweep frequency signal is reflected by the listening space; A robust deconvolution process with a regularization term is performed on the reference sweep signal and the recording signal to obtain the original frequency response curve at the current measurement position; After completing the measurements at multiple preset measurement locations, the original frequency response curves are averaged in the energy domain to generate a spatial average frequency response. The pre-acquired microphone frequency response calibration curve corresponding to the calibration microphone is invoked to perform differential compensation on the spatial average frequency response, thereby obtaining the compensated frequency response; Based on the compensated frequency response and target curve, PEQ filter parameters are generated using a parametric equalizer automatic search algorithm. Calculate the maximum gain value of the cascaded PEQ filters, and calculate the safe attenuation value of the preceding stage based on the maximum gain value; The PEQ filter parameters and the preamp safety attenuation value are sent to the speaker so that the speaker applies the preamp safety attenuation value in the preamp gain control module and loads the PEQ filter parameters in the audio playback link.

[0007] Furthermore, the robust deconvolution process with regularization term applied to the reference swept frequency signal and the recorded signal specifically includes: The reference sweep signal and the recording signal are respectively padded with zeros to a preset transformation length; Perform Discrete Fourier Transform on the two zero-padded signals to obtain the reference frequency domain sequence and the recorded audio domain sequence; The complex frequency domain transfer function is calculated by using the squared modulus of the reference frequency domain sequence superimposed with a regularization term as the denominator and the product of the complex conjugate of the reference frequency domain sequence and the recorded audio domain sequence as the numerator; wherein, the regularization term is the maximum squared modulus of the reference frequency domain sequence multiplied by a preset regularization coefficient. The amplitude-frequency response at the current measurement position is extracted based on the complex frequency domain transfer function and used as the original frequency response curve.

[0008] Furthermore, the energy domain averaging process for the multiple original frequency response curves specifically includes: Convert the decibel value of each of the original frequency response curves into an energy domain value; The average energy value is obtained by averaging the energy domain values ​​corresponding to all measurement locations. The average energy value is converted back to the decibel domain to generate the spatial average frequency response.

[0009] Furthermore, after generating the spatially averaged frequency response, the method further includes: Determine whether the calibration microphone is a third-party microphone; If so, the system receives the third-party microphone calibration file uploaded by the user and parses it to generate the microphone frequency response calibration curve. If not, the microphone frequency response calibration curve corresponding to the official microphone will be automatically retrieved from local or cloud resources.

[0010] Furthermore, before invoking the pre-acquired microphone frequency response calibration curve corresponding to the calibration microphone to perform differential compensation on the spatial average frequency response, the method further includes: The microphone frequency response calibration curve is resampled according to the frequency point of the spatial average frequency response to align their frequency coordinate axes. The differential compensation is: in the decibel domain, subtract the aligned microphone frequency response calibration curve from the spatial average frequency response.

[0011] Furthermore, the step of generating PEQ filter parameters based on the compensated frequency response and target curve using an automatic parametric equalizer search algorithm specifically includes: The compensated frequency response and the target curve are resampled and reference normalized to make them on a unified frequency grid and have the same reference reference. Calculate the deviation curve between the normalized compensated frequency response and the target curve, and detect peak and valley regions based on the deviation curve; Each detected peak or valley region is mapped to an initial candidate filter; The center frequency, quality factor, and gain of the initial candidate filters are optimized through multiple rounds of iterative optimization using local grid search, resulting in an optimized filter set. Merge similar filters in the optimized filter set whose center frequencies differ by less than 1 / 6 octave on the logarithmic frequency axis and whose gain directions are consistent, and remove filters whose absolute gain value is lower than a threshold from the merged filter set; Determine if the number of effective filters has reached the expected value. If so, output the final PEQ filter parameters.

[0012] Furthermore, after calculating the deviation curve between the normalized compensated frequency response and the target curve, the method further includes: The desired compensation curve is smoothed using octave bands. In peak clipping mode only, all positive values ​​in the desired compensation curve are set to zero to limit the filters generated by subsequent searches to attenuate only the peak frequency band.

[0013] Furthermore, the calculation of the maximum gain value after the PEQ filters are cascaded, and the calculation of the safety attenuation value of the preceding stage based on the maximum gain value, specifically includes: The frequency response of the filter cascade is obtained by cascading all PEQ filters on the same frequency grid. Find the maximum gain value in the frequency response of the cascaded filters; If the maximum gain value is less than or equal to 0, then the safe attenuation value of the preceding stage is determined to be 0; If the maximum gain value is greater than 0, the maximum gain value is rounded up to a precision of 0.1dB, and then 0.1dB is added to the rounded value. The negative value of the resulting value is used as the safety attenuation value of the preamplifier.

[0014] Furthermore, before sending the control command to the speaker, the following steps are also included: Send a test signal to the speaker and receive the feedback signal collected by the calibration microphone; Based on the feedback signal, determine whether the current acquisition level is within the preset ideal range; If the volume is below the ideal range, the user is prompted to increase the playback volume or microphone gain; if the volume is above the ideal range or there is a risk of clipping, the user is prompted to decrease the playback volume or microphone gain.

[0015] Secondly, a speaker self-calibration system based on collaboration between a web page and the speaker is provided, including: The instruction sending module is used to send control instructions to the speaker to control the speaker to play the reference sweep frequency signal at a set sampling rate; The signal acquisition module is used to acquire the recording signal collected by the calibration microphone, wherein the recording signal is the signal collected by the calibration microphone after the reference sweep frequency signal is reflected by the listening space; The deconvolution processing module is used to perform robust deconvolution processing with regularization on the reference sweep signal and the recording signal to obtain the original frequency response curve at the current measurement position; The averaging module is used to perform energy domain averaging on the multiple original frequency response curves after completing the measurement at multiple preset measurement locations, thereby generating a spatial average frequency response. The differential compensation module is used to call the microphone frequency response calibration curve corresponding to the pre-acquired calibration microphone, perform differential compensation on the spatial average frequency response, and obtain the compensated frequency response; The parameter generation module is used to generate PEQ filter parameters based on the compensated frequency response and target curve using a parametric equalizer automatic search algorithm. The attenuation calculation module is used to calculate the maximum gain value after the PEQ filters are cascaded, and to calculate the safe attenuation value of the preceding stage based on the maximum gain value; The information sending module is used to send the PEQ filter parameters and the preamplifier safety attenuation value to the speaker, so that the speaker applies the preamplifier safety attenuation value in the preamplifier gain control module and loads the PEQ filter parameters in the audio playback link.

[0016] Compared with the prior art, the present invention has the following beneficial effects: 1. This invention provides a speaker self-calibration method based on collaboration between a web page and the speaker. By offloading the tasks of deconvolution, spatial averaging, microphone compensation, and PEQ parameter search to the web page, and retaining only the frequency sweep playback and filtering functions on the speaker side, low-cost, low-computing-power consumer speakers can achieve professional-grade spatial acoustic calibration without the need for expensive embedded DSP chips. This task division not only significantly reduces hardware costs, but also achieves real-time visualization and user control of the calibration process through the strong interactivity of the web page, enabling ordinary users to perform self-service high-precision sound tuning.

[0017] 2. This invention introduces a pink noise test before the formal frequency sweep and provides real-time feedback based on RMS level to guide users in adjusting the volume or gain, ensuring the purity and linearity of the input signal for subsequent deconvolution operations. This guarantees the accuracy of frequency response extraction from the source and avoids calibration failures or error accumulation caused by initial signal quality issues.

[0018] 3. This invention introduces a regularization term based on the maximum energy value into the frequency domain division, which suppresses the interference of weak room excitation modes and ambient noise on the transfer function estimation without losing effective frequency response details. This enables the extraction of smooth and realistic speaker frequency response curves even in non-anechoic chambers.

[0019] 4. This invention converts the frequency response of each point to the energy domain, averages it, and then converts it back to the decibel domain, ensuring that the physical acoustic energy contribution of each measurement point is treated fairly. The resulting spatial average frequency response can more scientifically represent the comprehensive acoustic characteristics of the entire listening area, avoiding the misleading effect of a single deep valley on subsequent EQ compensation.

[0020] 5. This invention effectively eliminates the microphone's own frequency coloration from the measurement data by subtracting the microphone calibration curve from the spatial average frequency response in the decibel domain. Users can use different types of microphones for measurement, and the system can always reproduce the true frequency response of the speaker and the room itself, thus improving the compatibility and universality of the solution.

[0021] 6. This invention, through a closed-loop optimization process of peak and valley detection, initial mapping, local mesh search, similarity merging, and residual point compensation, can generate customized filter parameters for room acoustic problems such as standing waves and reflections. Compared with the traditional simple adjustment of high and low shelves, this algorithm can more accurately combat distortion in specific frequency bands while avoiding the introduction of unnecessary phase distortion or filter redundancy.

[0022] 7. This invention smooths the desired compensation curve and limits the gain direction, forcing the filter to focus on correcting obvious peak distortion without overfilling the valleys caused by the physical characteristics of the room. This effectively prevents system instability caused by ineffective large gain boosts and ensures the listenability.

[0023] 8. This invention calculates the maximum gain after all PEQs are cascaded, and adds an extra 0.1dB margin by rounding up to a precision of 0.1dB. This provides a precise safety dynamic margin before the signal enters the EQ processing, which avoids both the decrease in signal-to-noise ratio caused by excessive attenuation and the distortion of digital signals caused by the superposition of positive gains. Attached Figure Description

[0024] The accompanying drawings, which are included to provide a further understanding of embodiments of the invention and form part of this application, do not constitute a limitation thereof. In the drawings: Figure 1 This is the overall flowchart of Embodiment 1 of the present invention; Figure 2 This is the overall execution logic diagram in Embodiment 1 of the present invention; Figure 3 This is an implementation architecture diagram of Embodiment 2 of the present invention; Figure 4 This is a system block diagram in Embodiment 2 of the present invention. Detailed Implementation

[0025] To make the objectives, technical solutions, and advantages of the present invention clearer, the present invention will be further described in detail below with reference to the embodiments and accompanying drawings. The illustrative embodiments and descriptions of the present invention are only used to explain the present invention and are not intended to limit the present invention.

[0026] Example 1: A speaker self-calibration method based on web-based and speaker-based collaboration, executed by the web-based terminal, such as... Figure 1 As shown, it includes the following steps: S1: Send a control command to the speaker to control the speaker to play the reference sweep signal at the set sampling rate; S2: Acquire the recording signal collected by the calibration microphone. The recording signal is the signal collected by the calibration microphone after the reference sweep signal is reflected by the listening space. S3: Perform robust deconvolution processing with regularization on the reference sweep signal and the recording signal to obtain the original frequency response curve at the current measurement position; S4: After completing the measurement at multiple preset measurement locations, perform energy domain averaging on multiple original frequency response curves to generate a spatial average frequency response. S5: Call the microphone frequency response calibration curve corresponding to the pre-acquired calibration microphone, perform differential compensation on the spatial average frequency response, and obtain the compensated frequency response; S6: Based on the compensated frequency response and target curve, generate PEQ filter parameters using the parametric equalizer automatic search algorithm; S7: Calculate the maximum gain after cascading the PEQ filters, and calculate the safe attenuation value of the preceding stage based on the maximum gain value; S8: Send the PEQ filter parameters and preamp safety attenuation value to the speaker end so that the speaker end applies the preamp safety attenuation value in the preamp gain control module and loads the PEQ filter parameters in the audio playback link.

[0027] PEQ stands for Parametric Equalizer, which is an audio filter that can precisely adjust the amplitude of a specific frequency band by adjusting the center frequency, quality factor, and gain.

[0028] In step S1, as Figure 2 As shown, when a user accesses the speaker calibration interface through a web browser on a personal computer, the webpage first provides two entry points: creating a new calibration and loading existing measurement results. If the user chooses to load existing measurement results, the webpage loads the historical measurement frequency response file, parses out the average frequency response and related metadata such as sampling rate, microphone information, and spatial mode, and directly proceeds to the subsequent calibration parameter setting steps, skipping the remeasurement process. If the user chooses to create a new calibration, the webpage executes the level test process to ensure that the signal-to-noise ratio of the subsequent frequency sweep acquisition is within the ideal range.

[0029] In practice, the web interface generates or calls preset pink noise audio data via the Web Audio API (Web Audio Application Programming Interface), and outputs it to the target audio playback device via the browser's audio playback interface, causing the target audio playback device to play the pink noise as a test signal. At this time, the web interface prompts the user to place the calibration microphone in the main listening position. The calibration microphone acquires the test signal and transmits the recording data to the web interface via a USB (Universal Serial Bus) interface.

[0030] After receiving the recording data, the web interface calculates the RMS (root mean square) level of the currently acquired signal in real time. The web interface has a preset ideal level range. If the calculated RMS level is lower than the lower limit of this ideal range, a prompt box appears on the web interface, indicating that the volume is too low and suggesting increasing the speaker volume or microphone gain. Conversely, if signal clipping is detected or the level is higher than the upper limit of the ideal range, the web interface indicates that the volume is too high and there is a risk of clipping, suggesting adjusting the subwoofer volume. After the user adjusts the hardware knob according to the prompts, the web interface resends the test signal and judges it again until the level is within the ideal range.

[0031] After level calibration is completed, the web interface and the speaker interface negotiate and determine the sampling rate parameters for subsequent processing. Then, the web interface sends control commands to the speaker interface, instructing it to play a reference sweep signal at the set sampling rate. This reference sweep signal is typically a logarithmic sinusoidal sweep signal from 20Hz to 20kHz, used to excite the acoustic modes of the listening space.

[0032] In step S2, after the web interface confirms that the level is within the ideal range and completes sampling rate negotiation through step S1, the web interface officially sends a control command to the speaker interface, controlling the speaker interface to play the reference sweep signal according to the set sampling rate. This reference sweep signal is a logarithmic sinusoidal sweep signal from 20Hz to 20kHz, used to excite the acoustic modes of the listening space. At the same time, the web interface controls the calibration microphone to start the recording function, acquiring the acoustic signal after room reflections, reverberation, and cabinet coloration.

[0033] In practice, the calibration microphone collects analog acoustic signals, which are then converted into digital audio data via analog-to-digital conversion. This digital audio data is transmitted to the terminal device via a USB interface. The digital audio stream from the corresponding audio input device is acquired through the media acquisition interface and then connected to the Web Audio API for real-time processing and caching to form a recording signal.

[0034] After acquiring the recording signal on the web interface, the duration of the recording signal and the reference sweep signal need to be aligned and verified. The web interface calculates the time delay difference between the recording signal and the reference sweep signal by calculating the cross-correlation function, and then truncates or pads the recording signal with zeros to ensure that the two are strictly synchronized on the time axis. The synchronized recording signal and the reference sweep signal are used together as input for the subsequent S3 step to extract the original frequency response curve at the current measurement position.

[0035] In step S3, after acquiring the recording signal in step S2, the web page performs frequency domain transformation on the recording signal and its corresponding original excitation signal to obtain the frequency domain representation of the recording signal and the frequency domain representation of the original excitation signal. Based on the ratio between the frequency domain representation of the recording signal and the frequency domain representation of the original excitation signal, the frequency domain transfer function of the system under test is determined.

[0036] First, the web page determines the transform length N required for frequency domain calculation. Let the length of the reference swept signal be... The length of the recording signal is To prevent time-domain aliasing and maintain computational efficiency, the web page calculates the transform length using the following formula: ; in, This indicates the rounding up operation; For logarithmic operations to base 2, This is the maximum value function. The formula ensures that N is an integer power of 2, not less than twice the longest signal length, to meet the base requirement of the Fast Fourier Transform (FFT). After determining N, the web interface adds zero-value samples to the end of both the reference sweep signal and the recording signal, padding both signals to length N.

[0037] Subsequently, the web page performs N-point Discrete Fourier Transform on the two zero-padded signals respectively, converting the time-domain signals to the frequency domain to obtain the reference frequency domain sequence. and recording audio domain sequence ,in Represents a discrete frequency index.

[0038] If you pass directly Calculate the transfer function when When the amplitude is extremely small, noise can be amplified infinitely. Therefore, the web page uses a robust deconvolution algorithm with regularization: ; in, for The complex conjugate of is used to correct phase information; For the reference signal at the frequency point The energy at the location, This is a regularization term. The calculation method is as follows: ; in, The regularization coefficient is taken as [value] in the embodiments of the present invention. , This represents the maximum energy of the reference signal across the entire frequency band. This design effectively suppresses numerical divergence at low-energy frequencies.

[0039] Web page based on complex frequency domain transfer function Extracting the amplitude frequency response To prevent logarithmic underflow, the formula for calculating the amplitude-frequency response is: ; in, To prevent numerical underflow, the lower limit value is set in this embodiment. Calculated This is the original frequency response curve at the current measurement location.

[0040] In addition, the web interface displays the original frequency response curve of the current measurement point to the user in real time. The user can judge the measurement quality based on whether the curve has abnormal abrupt changes, whether there is obvious clipping, and whether it is affected by environmental noise. If the result of the current measurement point is qualified, the frequency response of that point is saved and the measurement proceeds to the next measurement point; if it is unqualified, the web interface deletes the recording data and frequency response curve corresponding to that point, guiding the user to only remeasure that point without having to re-perform measurements of other completed and qualified points.

[0041] In step S4, after completing step S3 processing of all preset measurement points within the listening area, the web page has obtained the original frequency response curves corresponding to each of the multiple points. Where p represents the measurement point number and f represents the frequency. To avoid directly averaging the decibel values, which would result in excessive weighting of the deep valley frequency bands (with extremely small energy domain values) and thus mask the overall listening experience, the web version uses an energy domain averaging algorithm.

[0042] The web interface also performs stereo discrimination. It determines whether the input frequency response data is mono or stereo. If it's stereo, the web interface performs paired averaging of the left and right channels. For the same measurement location, the web interface obtains the frequency response of the left and right channels at that location. The web interface first converts the decibel values ​​of the left and right channels to the energy domain, performs an arithmetic average, and then converts them back to the decibel domain to obtain the position-level frequency response for that measurement location. If it's mono data, it skips paired averaging and directly enters multi-position energy domain averaging.

[0043] First, the webpage converts the decibel values ​​of each raw frequency response curve into energy domain values. The conversion formula is: ;in, It is the amplitude-frequency response (in dB) of the p-th point at frequency f. It converts the decibel value back to a linear energy value (i.e., power ratio). The standard conversion process.

[0044] Next, the web page performs an arithmetic average of the energy domain values ​​corresponding to all measurement locations. Assuming there are P valid measurement points, calculate the average energy value at the f-th frequency point. The formula is: .

[0045] Finally, the web page converts the calculated average energy value back to the decibel domain to generate the spatial average frequency response. : The spatial average frequency response can represent the comprehensive acoustic characteristics of the entire listening area, effectively reducing the error caused by anomalies at a single measurement point.

[0046] The web interface saves the generated spatial average frequency response and related metadata as a reusable measurement frequency response file. This file can be reloaded in subsequent calibration processes via the "Load Existing Measurement Results" entry in S1, avoiding repeated multi-point frequency sweep measurements.

[0047] After generating the spatial average frequency response, the web-based client executes the microphone compensation process. First, the web-based client determines whether the calibration microphone currently being used is a third-party microphone based on the user's selection on the front-end interface.

[0048] If the user selects an official microphone, the webpage will automatically retrieve the microphone frequency response calibration curve that uniquely corresponds to that official microphone model from local storage or the cloud server. No user intervention is required.

[0049] If the user selects a third-party microphone, the webpage provides a file upload entry to receive the user's uploaded third-party microphone calibration file, usually in TXT or CSV format, and parse the file to generate the corresponding microphone frequency response calibration curve. This curve reflects the sensitivity deviation of the specific microphone at different frequencies, and serves as the basis for subsequent differential compensation.

[0050] In step S5, the spatial average frequency response is generated in step S4. And obtain the corresponding microphone frequency response calibration curve. Then, the web page executes a microphone compensation process to eliminate the interference of the microphone's own frequency response characteristics on the measurement results.

[0051] Because microphone frequency response calibration curves typically have independent frequency sampling points, while spatial average frequency response... It also has its specific frequency distribution, and the two are often misaligned on the frequency axis. Therefore, the web page first resamples the microphone frequency response calibration curve. The web page uses the spatially averaged frequency response. Using the specified frequency as a reference, a logarithmic frequency domain linear interpolation method is employed to calibrate the microphone frequency response curve. Interpolation to On the exact same frequency grid, the aligned microphone frequency response calibration curves are obtained.

[0052] After frequency axis alignment is completed, the web interface performs differential compensation on the spatial average frequency response in the decibel domain. The compensation formula is as follows: ; in, This is the compensated frequency response after removing the influence of the microphone's own frequency response. Since subtraction in the decibel domain corresponds to division in the linear domain, this step essentially removes the microphone's own frequency response characteristics from the measured acoustic transfer function (i.e., transfer function division), thus obtaining the true frequency response that only reflects the acoustic characteristics of the speaker and the room.

[0053] In step S6, before starting the PEQ automatic search algorithm, the web interface provides a calibration parameter setting interface for users to configure the control parameters for this calibration. These control parameters include the calibration frequency range, the number of PEQ filters, search accuracy, whether valley filling is allowed, filter gain upper and lower limits, and target curve selection. The target curve can be a default flat curve, a template curve provided on the web interface, or a user-uploaded custom curve. After the user completes the parameter settings, the web interface starts the PEQ automatic search based on the compensated frequency response and the user-defined target curve.

[0054] First, the web page displays the compensated frequency response. Frequency resampling and reference normalization are performed on the target curve selected by the user. The web interface generates a set of logarithmic frequency grids based on the set calibration frequency range. Subsequently, the compensated frequency response and the target curve are interpolated onto this unified frequency grid. Next, within the preset reference frequency band, the web interface calculates the energy reference values ​​of the two curves and translates the two curves as a whole to align them with the reference within this frequency band, thus obtaining the normalized measurement frequency response. and normalization target frequency Specifically: ;in, The curve before normalization, As a reference frequency band energy standard, Used as the target benchmark.

[0055] The web-based calculation of the deviation curve between the normalized measured frequency response and the target curve : The webpage scans based on a set significance threshold, such as ±1.5dB. A region is identified as a peak when the curve is continuously above a positive threshold, and as a valley when valley filling is allowed and the curve is continuously below a negative threshold. Each region is recorded for its starting frequency, ending frequency, geometric center frequency, and mean deviation within the interval.

[0056] The web interface maps each detected peak or valley region to an initial peaking filter. The filter's center frequency is set to the geometric center frequency of the region, the quality factor Q is determined based on the ratio of the region's bandwidth to the center frequency, and is limited to the range [0.1, 10]. The gain is set to the negative of the mean deviation of the region, and is limited to the range [-18dB, 12dB]. If the number of initial filters is less than the user-defined value, the web interface supplements zero-gain placeholder filters along the logarithmic frequency direction.

[0057] The web-based application performs multiple rounds of iterative optimization of filter parameters using a local grid search. Each time, other filters are kept constant, and only the center frequency, Q value, and gain of the current filter are adjusted. The current cascaded response and the desired compensation curve are then calculated. The distance, i.e., the mean square error. If the distance decreases after adjustment, the new parameters are adopted. During the optimization process, the web interface performs octave smoothing on the desired compensation curve, such as dividing each octave into 12 equal parts and moving averages. In peak clipping mode only, all positive values ​​in the desired compensation curve are set to zero, limiting the filter to only attenuation.

[0058] The web interface merges similar filters with close center frequencies, consistent gain directions, and small gain differences. New parameters are generated by weighted averaging based on the absolute gain values. Close center frequencies are defined as those differing by less than 1 / 6 octave on the logarithmic axis. Weak filters with absolute gain values ​​below a threshold (e.g., 0.1 dB) are then removed from the merged filter set. Finally, the web interface checks if the number of effective filters meets the user's expectations. If yes, it outputs the final PEQ filter parameter set, including filter type, center frequency, Q value, and gain. If not, it calculates the current residual, adds new filters near frequencies with large residuals, and returns to the optimization step.

[0059] The web interface generates the calibrated predicted frequency response based on the compensated frequency response, the user-defined target curve, the PEQ filter parameters, and the subsequently calculated pre-stage safety attenuation value. It then displays a visual comparison of the original, target, and calibrated frequency responses on the same interface. Users can determine whether manual fine-tuning is needed based on the comparison results: if no adjustment is required, the system proceeds to S7 to calculate the pre-stage safety attenuation value; if optimization is needed, the web interface provides a filter parameter editing interface, allowing users to manually adjust the center frequency, Q value, gain, number, and type of the PEQ filters. After each adjustment, the web interface recalculates the filter cascade response and updates the visual comparison results until the final parameters are confirmed.

[0060] It should be noted that PEQ filters can also be classified as lowshelf, highshelf, lowpass, and highpass filters. Lowshelf filters provide overall gain or attenuation for frequencies below a set cutoff frequency without affecting the high-frequency range; highshelf filters provide overall gain or attenuation for frequencies above a set cutoff frequency without affecting the low-frequency range; lowpass filters allow only signals below a set cutoff frequency to pass through, filtering out high-frequency noise; and highpass filters allow only signals above a set cutoff frequency to pass through, filtering out low-frequency noise. The web interface can adaptively select the above filter type based on the deviation characteristics between the compensated frequency response and the target curve to optimize the calibration effect.

[0061] In step S7, after generating the PEQ parameter set containing M filters in step S6, the web interface performs a pre-stage safety attenuation calculation to prevent digital clipping or signal overflow at the speaker end during EQ execution. The web interface first cascades all M PEQ filters on the same frequency grid. Let the frequency response of the i-th PEQ filter at frequency f be... The total frequency response after M filters are cascaded is... The frequency response of the filter cascade is obtained by linearly superimposing (or adding in the decibel domain) the frequency responses of each filter.

[0062] The web page iterates through all frequency points to find the frequency response of the cascaded filters. The maximum gain value in is denoted as The unit is dB. This value represents the maximum positive boost that all PEQ filters can produce to the signal when working together.

[0063] Web version Calculate the pre-amplifier safety attenuation value (Pregain): ; in, This indicates the maximum gain value. Round up to the nearest 0.1 dB, for example ,but +0.1 indicates an additional safety margin.

[0064] like This indicates that the filter combination did not positively boost the signal, and there is no risk of clipping; therefore, Pregain is set to 0. The webpage will calculate the negative value Pregain (e.g.) The 2.4dB signal, along with the PEQ filter parameters, is sent to the speaker. At the speaker end, in the audio playback chain, a negative gain attenuation is first applied to the input signal, followed by PEQ filtering, thus ensuring dynamic sound quality while completely avoiding the risk of digital overflow.

[0065] In step S8, after calculating the final PEQ filter parameter set and the pre-amplifier safety attenuation value Pregain in step S7, the web interface sends the calibration parameters to the speaker via a preset communication channel. In a specific implementation, the communication method can be based on the WebHID API (Web Human Interface Device) and uses HID (Human Interface Device) communication. The web interface encapsulates the control commands as binary data in HID report format and sends them to the speaker's HID interface via the browser's WebHID interface. The speaker's microcontroller receives and parses the control commands.

[0066] First, after receiving the calibration parameters, the speaker performs parameter parsing and verification. Once the data integrity is confirmed, the preamplifier safety attenuation value (Pregain) is written into the preamplifier gain control module register in the audio playback chain. This module, located before the digital filter, performs negative gain attenuation on the signal amplitude before it enters the PEQ processing unit, thus reserving dynamic margin for any subsequent positive EQ boost and preventing digital signal overflow.

[0067] Subsequently, the speaker loads the resolved M PEQ filter parameters into the filter coefficient register of the digital signal processing (DSP) chip. During audio playback, the input audio signal first undergoes attenuation processing by the pre-amplifier gain control module, then sequentially passes through a cascaded structure composed of the M PEQ filters for frequency response correction, finally outputting a calibrated audio signal that is transmitted to the power amplifier. At this point, the collaborative calibration process between the web interface and the speaker is complete, and the speaker's frequency response in the actual listening space will closely approximate the user-defined target curve.

[0068] Example 2: A speaker self-calibration system based on web-based and speaker-based collaboration. This system is used to implement the speaker self-calibration method described in Example 1, such as... Figure 3As shown, speaker self-calibration relies on a web interface, a calibration microphone, and the speaker device. The web interface, as the primary computing and interaction interface, handles calibration process control, measurement data reception, real-time frequency response display, sweep frequency deconvolution, multi-location averaging, microphone compensation, frequency resampling, target curve normalization, PEQ parameter search, preamplifier attenuation calculation, and visualization of calibration results. The calibration microphone collects the sound signals generated at various measurement points after the speaker plays the sweep frequency signal and provides the recording data to the web interface. The calibration microphone can be an official microphone or a third-party microphone selected by the user. When using a third-party microphone, the web interface receives the corresponding microphone calibration file. The speaker device, as the end-side execution terminal, plays the reference sweep frequency signal provided by the web interface and receives the PEQ filter parameters and preamplifier attenuation values ​​from the web interface after calibration. During normal playback, the speaker device filters and controls the gain of the input audio signal based on these parameters.

[0069] Specifically, such as Figure 4 As shown, the web interface includes a command sending module, a signal acquisition module, a deconvolution processing module, an averaging processing module, a difference compensation module, a parameter generation module, an attenuation calculation module, and an information delivery module.

[0070] The system includes: a command sending module for sending control commands to the speaker to control the speaker to play the reference sweep signal at a set sampling rate; a signal acquisition module for acquiring the recording signal collected by the calibration microphone, which is the signal of the reference sweep signal after reflection through the listening space and then recorded by the calibration microphone; a deconvolution processing module for performing robust deconvolution processing with a regularization term on the reference sweep signal and the recording signal to obtain the original frequency response curve at the current measurement position; an averaging module for performing energy domain averaging processing on multiple original frequency response curves after completing measurements at multiple preset measurement positions to generate a spatial average frequency response; and a differential compensation module for... The system utilizes a pre-acquired microphone frequency response calibration curve to differentially compensate the spatial average frequency response, resulting in a compensated frequency response. A parameter generation module generates PEQ filter parameters based on the compensated frequency response and the target curve using an automatic parametric equalizer search algorithm. An attenuation calculation module calculates the maximum gain after cascading the PEQ filters and calculates the preamplifier safety attenuation value based on the maximum gain. An information delivery module sends the PEQ filter parameters and the preamplifier safety attenuation value to the speaker, enabling the speaker to apply the preamplifier safety attenuation value in the preamplifier gain control module and load the PEQ filter parameters into the audio playback link.

[0071] Working principle: This invention offloads the tasks of deconvolution, spatial averaging, microphone compensation, and PEQ parameter search to the web page, retaining only the frequency sweep playback and filtering functions on the speaker side. This allows low-cost, low-computing-power consumer speakers to achieve professional-grade spatial acoustic calibration without the need for expensive embedded DSP chips. This division of tasks not only significantly reduces hardware costs but also enables real-time visualization and user control of the calibration process through the strong interactivity of the web page, achieving self-service high-precision sound tuning for ordinary users.

[0072] Those skilled in the art will understand that embodiments of this application can be provided as methods, systems, or computer program products. Therefore, this application can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, this application can take the form of a computer program product embodied on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0073] This application is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of this application. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart... Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.

[0074] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.

[0075] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.

[0076] The specific embodiments described above further illustrate the purpose, technical solution, and beneficial effects of the present invention. It should be understood that the above description is only a specific embodiment of the present invention and is not intended to limit the scope of protection of the present invention. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the scope of protection of the present invention.

Claims

1. A speaker self-calibration method based on collaboration between a web page and the speaker, characterized in that, When executed via a web browser, the process includes the following steps: Send control commands to the speaker to control the speaker to play the reference sweep signal at a set sampling rate; Acquire the recording signal collected by the calibration microphone, wherein the recording signal is the signal collected by the calibration microphone after the reference sweep frequency signal is reflected by the listening space; A robust deconvolution process with a regularization term is performed on the reference sweep signal and the recording signal to obtain the original frequency response curve at the current measurement position; After completing the measurements at multiple preset measurement locations, the original frequency response curves are averaged in the energy domain to generate a spatial average frequency response. The pre-acquired microphone frequency response calibration curve corresponding to the calibration microphone is invoked to perform differential compensation on the spatial average frequency response, thereby obtaining the compensated frequency response; Based on the compensated frequency response and target curve, PEQ filter parameters are generated using a parametric equalizer automatic search algorithm. Calculate the maximum gain value of the cascaded PEQ filters, and calculate the safe attenuation value of the preceding stage based on the maximum gain value; The PEQ filter parameters and the preamp safety attenuation value are sent to the speaker so that the speaker applies the preamp safety attenuation value in the preamp gain control module and loads the PEQ filter parameters in the audio playback link. The step of generating PEQ filter parameters based on the compensated frequency response and target curve using a parametric equalizer automatic search algorithm specifically includes: The compensated frequency response and the target curve are resampled and reference normalized to make them on a unified frequency grid and have the same reference reference. Calculate the deviation curve between the normalized compensated frequency response and the target curve, and detect peak and valley regions based on the deviation curve; Each detected peak or valley region is mapped to an initial candidate filter; The center frequency, quality factor, and gain of the initial candidate filters are optimized through multiple rounds of iterative optimization using local grid search, resulting in an optimized filter set. Merge similar filters in the optimized filter set whose center frequencies differ by less than 1 / 6 octave on the logarithmic frequency axis and whose gain directions are consistent, and remove filters whose absolute gain value is lower than a threshold from the merged filter set; Determine if the number of effective filters has reached the expected value. If so, output the final PEQ filter parameters.

2. The speaker self-calibration method based on web-based and speaker-based collaboration according to claim 1, characterized in that, The robust deconvolution process with regularization applied to the reference swept frequency signal and the recorded signal specifically includes: The reference sweep signal and the recording signal are respectively padded with zeros to a preset transformation length; Perform Discrete Fourier Transform on the two zero-padded signals to obtain the reference frequency domain sequence and the recorded audio domain sequence; The complex frequency domain transfer function is calculated by using the squared modulus of the reference frequency domain sequence superimposed with a regularization term as the denominator and the product of the complex conjugate of the reference frequency domain sequence and the recorded audio domain sequence as the numerator; wherein, the regularization term is the maximum squared modulus of the reference frequency domain sequence multiplied by a preset regularization coefficient. The amplitude-frequency response at the current measurement position is extracted based on the complex frequency domain transfer function and used as the original frequency response curve.

3. The speaker self-calibration method based on web-based and speaker-based collaboration according to claim 1, characterized in that, The energy domain averaging process for the multiple original frequency response curves specifically includes: Convert the decibel value of each of the original frequency response curves into an energy domain value; The average energy value is obtained by averaging the energy domain values ​​corresponding to all measurement locations. The average energy value is converted back to the decibel domain to generate the spatial average frequency response.

4. The speaker self-calibration method based on web-based and speaker-based collaboration according to claim 1, characterized in that, After generating the spatially averaged frequency response, the method further includes: Determine whether the calibration microphone is a third-party microphone; If so, the system receives the third-party microphone calibration file uploaded by the user and parses it to generate the microphone frequency response calibration curve. If not, the microphone frequency response calibration curve corresponding to the official microphone will be automatically retrieved from local or cloud resources.

5. The speaker self-calibration method based on web-based and speaker-based collaboration according to claim 1, characterized in that, Before invoking the pre-acquired microphone frequency response calibration curve corresponding to the calibration microphone and performing differential compensation on the spatial average frequency response, the method further includes: The microphone frequency response calibration curve is resampled according to the frequency point of the spatial average frequency response to align their frequency coordinate axes. The differential compensation is: in the decibel domain, subtract the aligned microphone frequency response calibration curve from the spatial average frequency response.

6. The speaker self-calibration method based on web-based and speaker-based collaboration according to claim 1, characterized in that, After calculating the deviation curve between the normalized compensated frequency response and the target curve, the method further includes: The desired compensation curve is smoothed using octave bands. In peak clipping mode only, all positive values ​​in the desired compensation curve are set to zero to limit the filters generated by subsequent searches to attenuate only the peak frequency band.

7. The speaker self-calibration method based on web-based and speaker-based collaboration according to claim 1, characterized in that, The calculation of the maximum gain value after the PEQ filters are cascaded, and the calculation of the safe attenuation value of the preceding stage based on the maximum gain value, specifically includes: The frequency response of the filter cascade is obtained by cascading all PEQ filters on the same frequency grid. Find the maximum gain value in the frequency response of the cascaded filters; If the maximum gain value is less than or equal to 0, then the safe attenuation value of the preceding stage is determined to be 0; If the maximum gain value is greater than 0, the maximum gain value is rounded up to a precision of 0.1dB, and then 0.1dB is added to the rounded value. The negative value of the resulting value is used as the safety attenuation value of the preamplifier.

8. The speaker self-calibration method based on web-based and speaker-based collaboration according to claim 1, characterized in that, Before sending the control command to the speaker, the following is also included: Send a test signal to the speaker and receive the feedback signal collected by the calibration microphone; Based on the feedback signal, determine whether the current acquisition level is within the preset ideal range; If the volume is below the ideal range, the user is prompted to increase the playback volume or microphone gain; if the volume is above the ideal range or there is a risk of clipping, the user is prompted to decrease the playback volume or microphone gain.

9. A speaker self-calibration system based on collaboration between a web page and the speaker, characterized in that, include: The instruction sending module is used to send control instructions to the speaker to control the speaker to play the reference sweep frequency signal at a set sampling rate; The signal acquisition module is used to acquire the recording signal collected by the calibration microphone, wherein the recording signal is the signal collected by the calibration microphone after the reference sweep frequency signal is reflected by the listening space; The deconvolution processing module is used to perform robust deconvolution processing with regularization on the reference sweep signal and the recording signal to obtain the original frequency response curve at the current measurement position; The averaging module is used to perform energy domain averaging on the multiple original frequency response curves after completing the measurement at multiple preset measurement locations, thereby generating a spatial average frequency response. The differential compensation module is used to call the microphone frequency response calibration curve corresponding to the pre-acquired calibration microphone, perform differential compensation on the spatial average frequency response, and obtain the compensated frequency response; The parameter generation module is used to generate PEQ filter parameters based on the compensated frequency response and target curve using a parametric equalizer automatic search algorithm. The attenuation calculation module is used to calculate the maximum gain value after the PEQ filters are cascaded, and to calculate the safe attenuation value of the preceding stage based on the maximum gain value; The information sending module is used to send the PEQ filter parameters and the preamplifier safety attenuation value to the speaker end, so that the speaker end applies the preamplifier safety attenuation value in the preamplifier gain control module and loads the PEQ filter parameters in the audio playback link; The step of generating PEQ filter parameters based on the compensated frequency response and target curve using a parametric equalizer automatic search algorithm specifically includes: The compensated frequency response and the target curve are resampled and reference normalized to make them on a unified frequency grid and have the same reference reference. Calculate the deviation curve between the normalized compensated frequency response and the target curve, and detect peak and valley regions based on the deviation curve; Each detected peak or valley region is mapped to an initial candidate filter; The center frequency, quality factor, and gain of the initial candidate filters are optimized through multiple rounds of iterative optimization using local grid search, resulting in an optimized filter set. Merge similar filters in the optimized filter set whose center frequencies differ by less than 1 / 6 octave on the logarithmic frequency axis and whose gain directions are consistent, and remove filters whose absolute gain value is lower than a threshold from the merged filter set; Determine if the number of effective filters has reached the expected value. If so, output the final PEQ filter parameters.

Citation Information

Patent Citations

  • Sound field correction method and device and audio playing equipment

    CN120238803A

  • Self-calibrating loudspeaker system

    US20100272270A1