Sound sound field low-frequency increasing and decreasing method and system based on volume amplitude recognition

By using a volume amplitude recognition method, a sliding time window algorithm, and a lightweight feedforward neural network to dynamically adjust the speaker gain, the problem of inconsistent sound perception and equipment damage in traditional methods of increasing or decreasing low frequencies in the sound field of audio equipment is solved, achieving optimal sound field reproduction and equipment safety across the entire volume range.

CN121967969APending Publication Date: 2026-05-01深圳市立平科技有限公司
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
深圳市立平科技有限公司
Filing Date
2026-03-31
Publication Date
2026-05-01

AI Technical Summary

Technical Problem

Traditional methods of increasing or decreasing low frequencies in audio systems cannot provide adaptive processing based on the loudness characteristics of the human ear, resulting in a lack of low-frequency sound at low volumes and a risk of speaker damage at high volumes, affecting the consistency of sound and the safety of the equipment.

Method used

A volume amplitude recognition-based method is adopted, which extracts instantaneous amplitude envelope features through a sliding time window algorithm and short-time discrete Fourier transform. Combined with a lightweight feedforward neural network and speaker physical parameters, a dynamic safety cutoff threshold is constructed, and the gain of the digital shelving filter is dynamically adjusted to achieve adaptive amplitude limiting.

Benefits of technology

It achieves optimal sound field reproduction across the entire volume range, solves the problem of insufficient low-frequency sound at low volumes, and prevents speaker overload distortion at high volumes, ensuring equipment safety.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121967969A_ABST
    Figure CN121967969A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of sound field control, in particular to a sound equipment sound field low-frequency increasing and decreasing method and system based on volume amplitude recognition, and the method comprises the following steps: collecting an original audio stream, and extracting instantaneous amplitude features; calculating a low-frequency compensation coefficient according to the nonlinear auditory model; constructing a dynamic safety threshold and correcting a coefficient in combination with the physical limit of the equipment; and adjusting the low-frequency component based on the corrected parameter, and generating an optimized sound field signal. According to the invention, by establishing a dynamic mapping and constraint mechanism between the instantaneous volume amplitude and the low-frequency gain, the technical defects that the low-frequency hearing feeling of a traditional sound system is deficient under the low volume and overload distortion is extremely easy to occur under the high volume are thoroughly solved, and under the premise of strictly guaranteeing the physical safe operation of loudspeaker hardware, the sound quality is greatly improved. And full-dynamic-range sound field adaptive equalization and restoration conforming to human ear psychological acoustic characteristics are realized.
Need to check novelty before this filing date? Find Prior Art

Description

A method and system for low-frequency addition and subtraction of sound field based on volume amplitude recognition Technical Field

[0001] This invention relates to the field of sound field control technology, and in particular to a method and system for increasing or decreasing the low frequency of an audio field based on volume amplitude recognition. Background Technology

[0002] The field of sound field control technology involves adjusting the frequency response, dynamic range, and spatial distribution of audio signals to reproduce the optimal listening experience in a specific listening environment. This technology is widely used in home theaters, professional sound reinforcement systems, and smart speaker devices. It aims to solve the matching problem between the physical characteristics of speakers and human auditory perception, ensuring that sound signals maintain high fidelity and a good sense of spatial immersion during transmission and playback. It is a key underlying technology for improving the subjective listening experience of modern audio systems.

[0003] Traditional methods for adjusting low-frequency frequencies in audio systems primarily rely on manual or preset low-frequency gain adjustments using a fixed graphic equalizer (GEQ) or a simple shelving filter. This approach typically sets a fixed center frequency and gain value, or linearly alters the low-frequency response simply by adjusting a volume knob, lacking the ability to adapt to real-time dynamic changes in the input signal.

[0004] Traditional methods of adjusting low frequencies in audio systems employ fixed gain or linear linkage processing modes, which result in insufficient low-frequency compensation based on the loudness characteristics of the human ear at low volumes. This leads to a dry and unenhancing sound. Furthermore, at high volumes with large dynamic ranges, these methods fail to effectively limit large-amplitude low-frequency vibrations, which can easily cause speaker voice coils to bottom out, distortion, or even permanent physical damage. This severely impacts the consistency of sound across the entire volume range and the safety of the equipment. Summary of the Invention

[0005] The purpose of this invention is to overcome the shortcomings of the existing technology and to propose a method and system for increasing or decreasing the low frequency of the sound field based on volume amplitude recognition.

[0006] To achieve the above objectives, the present invention adopts the following technical solution: a method for increasing or decreasing the low frequency of a sound field based on volume amplitude recognition, comprising the following steps:

[0007] S1: Acquire the input audio data stream, perform a short-time discrete Fourier transform on the input audio data stream using a sliding time window algorithm, reduce spectral leakage using a Hanning window function, perform calculations based on the root mean square energy distribution logic in the time domain, and extract instantaneous amplitude envelope features;

[0008] S2: Input the instantaneous amplitude envelope feature into a lightweight feedforward neural network constructed based on the human ear equal loudness curve standard, establish a nonlinear mapping relationship between the input amplitude and the perceived loudness through a polynomial fitting algorithm, and calculate the basic low-frequency gain coefficient.

[0009] S3: Construct a safety response boundary based on the speaker's physical stroke parameters, voice coil thermal time constant, and mechanical damping characteristics. Combine the instantaneous amplitude envelope characteristics with the historical peak statistics of the time series within a set period for weighted evaluation to generate a dynamic safety cutoff threshold.

[0010] S4: The basic low-frequency gain coefficient is dynamically compressed and adaptively limited using the dynamic safety cutoff threshold. The gain parameter of the digital shelving filter is controlled according to the processed coefficient to adjust the low-frequency components in the input audio data stream and generate an optimized sound field driving signal.

[0011] As a further aspect of the present invention, step S1 specifically comprises:

[0012] S11: Construct a time sliding window of a preset length, perform frame processing on the input audio data stream according to a predetermined overlap rate, perform weighted smoothing on each frame of audio data through the Hanning window function, and perform short-time discrete Fourier transform to obtain the spectral amplitude vector in the frequency domain.

[0013] S12: Traverse the effective frequency point data in the spectrum amplitude vector, extract the short-time energy value of the current frame through the sum of squares and mean calculation logic, and perform recursive smoothing operation in combination with the energy data of adjacent historical frames to generate the root mean square energy distribution sequence.

[0014] S13: Perform peak preservation and attenuation processing on the root mean square energy distribution sequence, extract the envelope curve that can characterize the dynamic change trend of the signal, and generate the instantaneous amplitude envelope feature.

[0015] As a further aspect of the present invention, step S2 specifically comprises:

[0016] S21: Retrieve the preset lightweight feedforward neural network, obtain the sound pressure level value corresponding to the instantaneous amplitude envelope feature, and retrieve the perceived loudness difference of the low frequency band relative to the reference frequency under the current sound pressure level according to the human ear equal loudness curve standard.

[0017] S22: Construct a mapping function between the perceived loudness difference and the compensation gain using a polynomial fitting algorithm, and perform a nonlinear transformation on the current input sound pressure level value through the mapping function to calculate the original compensation amount;

[0018] S23: Normalize and smooth the original compensation amount to remove abrupt gain noise and generate the basic low-frequency gain coefficient.

[0019] As a further aspect of the present invention, step S3 specifically comprises:

[0020] S31: Obtain the maximum linear displacement parameters, rated power handling capacity, and mechanical damping coefficient of the loudspeaker. Based on physical limit logic, construct a safe response space using displacement boundary and thermal load boundary.

[0021] S32: Establish a time series historical peak statistical buffer, store the instantaneous amplitude envelope feature data within a set period in real time, calculate the peak factor and average power density of the data in the buffer, and perform weighted evaluation to quantify the current cumulative thermal stress state.

[0022] S33: Dynamically match the safety response space with the current cumulative thermal stress state, and adjust the maximum allowable gain limit in real time according to the matching result to generate the dynamic safety cutoff threshold.

[0023] As a further aspect of the present invention, step S4 specifically comprises:

[0024] S41: Compare the basic low-frequency gain coefficient with the dynamic safety cutoff threshold. When the coefficient exceeds the threshold, start the hard inflection point compression logic. When the coefficient does not exceed the threshold, maintain linear propagation, thereby calculating the actual execution gain.

[0025] S42: Calculate the coefficient matrix of the digital shelving filter based on the actual execution gain, and dynamically update the center frequency, quality factor, and gain-bandwidth product parameters of the filter to ensure that the filter response curve meets the target sound field requirements.

[0026] S43: The input audio data stream is imported into the digital shelving filter with updated parameters, and the signal components below the specified low-frequency cutoff frequency are amplitude modulated while the mid-to-high frequency components are kept through, thereby generating the optimized sound field drive signal.

[0027] As a further aspect of the present invention, the process for obtaining the basic low-frequency gain coefficient includes a nonlinear operation performed according to the following formula:

[0028] ;

[0029] in, This represents the basic low-frequency gain coefficient. This represents the reference sound pressure level threshold in the standard equal-loudness curve. This represents the currently monitored real-time sound pressure level value. The normalized value representing the instantaneous amplitude envelope feature, This represents the hearing compensation sensitivity factor. The exponent represents the order of a nonlinear polynomial. The basic boost bias constant representing low-frequency energy.

[0030] As a further aspect of the present invention, the process of setting the dynamic security truncation threshold specifically includes a weighted calculation performed according to the following logic:

[0031] ;

[0032] in, This represents the dynamic security cutoff threshold. This represents the maximum allowable linear displacement of the speaker's voice coil. Represents the gain conversion coefficient based on the damping characteristics of the mechanical system. Represents the thermal stress weighting factor of the voice coil. This represents the total number of sampling points within a set period. Representing the Historical power values ​​at each sampling point This represents the value of the exponentially weighted function that decays over time.

[0033] As a further aspect of the present invention, the application process of the Hanning window function specifically includes:

[0034] Obtain the length value of the current frame, generate a Hanning window coefficient sequence of the corresponding length, and perform point-by-point multiplication operation on the coefficient sequence with the time domain sampling data of the current frame to reduce the amplitude at both ends of the signal and smooth the discontinuity at the splicing point.

[0035] The windowed data frame is padded with zeros to a length equal to an integer power of 2. A fast Fourier transform is then performed to calculate the magnitude of the transform result and discard the phase information, thus obtaining spectral data containing only amplitude information.

[0036] As a further aspect of the present invention, the process of updating the parameters specifically includes:

[0037] The current sampling rate and target cutoff frequency of the digital shelving filter are obtained, and intermediate variables are calculated in combination with the actual execution gain. The feedforward coefficients and feedback coefficients of the second-order IIR filter are derived using the bilinear transform method.

[0038] The stability of the detection coefficient update process is checked. If the calculated coefficient causes the poles to exceed the unit circle, the filter is forcibly reset to the safe coefficient state of the previous frame. Otherwise, the new coefficients are written into the filter register to complete the parameter update.

[0039] A low-frequency boosting / reducing system for an audio field based on volume amplitude recognition, the system being used to implement the aforementioned low-frequency boosting / reducing method for an audio field based on volume amplitude recognition, the system comprising:

[0040] The feature extraction module is used to acquire the input audio data stream, perform short-time discrete Fourier transform on the input audio data stream using a sliding time window algorithm, reduce spectral leakage using a Hanning window function, and perform calculations based on the root mean square energy distribution logic in the time domain to extract instantaneous amplitude envelope features.

[0041] The compensation calculation module is used to input the instantaneous amplitude envelope features into a lightweight feedforward neural network constructed based on the human ear equal loudness curve standard, establish a nonlinear mapping relationship between the input amplitude and the perceived loudness through a polynomial fitting algorithm, and calculate the basic low-frequency gain coefficient.

[0042] The safety assessment module is used to construct a safety response boundary based on the speaker's physical stroke parameters, voice coil thermal time constant, and mechanical damping characteristics. It then performs a weighted assessment by combining the instantaneous amplitude envelope characteristics with historical peak statistics of the time series within a set period to generate a dynamic safety cutoff threshold.

[0043] The control execution module is used to dynamically compress and adaptively limit the basic low-frequency gain coefficient using the dynamic safety truncation threshold, control the gain parameter of the digital shelving filter based on the processed coefficient, adjust the low-frequency band component in the input audio data stream, and generate an optimized sound field drive signal.

[0044] Compared with the prior art, the advantages and positive effects of the present invention are as follows:

[0045] In this invention, by establishing a dynamic mapping and constraint mechanism between instantaneous volume amplitude and low-frequency gain, this innovative solution utilizes real-time extracted amplitude features to drive a lightweight feedforward neural network, accurately calculating the low-frequency requirements that conform to the characteristics of human hearing, thus solving the problem of insufficient low-frequency sound at low volumes in traditional technologies. Simultaneously, by introducing a dynamic safety cutoff threshold based on the physical limits of the device, this solution can automatically smooth and limit the gain coefficient and correct it when large dynamic range signals are input, effectively solving the risk of overload distortion and hardware damage at high volumes, and achieving optimal sound field reproduction across the entire volume range. Attached Figure Description

[0046] Figure 1 is a flowchart of the method for increasing or decreasing the low frequency of the sound field in this invention.

[0047] Figure 2 is a flowchart of the instantaneous amplitude envelope feature extraction process of the present invention;

[0048] Figure 3 is a flowchart of the basic low-frequency gain coefficient calculation of the present invention;

[0049] Figure 4 is a flowchart of the dynamic safety truncation threshold generation process of the present invention;

[0050] Figure 5 is a flowchart of the optimized sound field driving signal generation process of the present invention. Detailed Implementation

[0051] To make the objectives, technical solutions, and advantages of this invention clearer, the software-based technical solution is described in detail below with reference to system architecture diagrams and embodiments. It should be understood that the specific embodiments described herein are only for explaining the technical solutions of this invention and do not constitute a limitation on the scope of protection.

[0052] In the description of this invention, the system architecture relationships or data processing flows indicated by terms such as "layer," "module," "interface," "data flow," "client," and "server" are all defined based on the architecture diagram or flowchart corresponding to the embodiments. This way of describing is only used to clearly illustrate the logical relationships between the elements in the technical solution, and not to limit the physical deployment form. The term "multiple" includes two or more technical units, including but not limited to multiple data nodes, processing threads, service instances, or functional components and other scalable elements. The specific number is determined according to the actual business scenario and needs to be specifically specified.

[0053] Please refer to Figures 1 and 2. This invention provides a technical solution: a method for increasing or decreasing the low frequency of an audio field based on volume amplitude recognition, comprising the following steps:

[0054] S1: Acquire the input audio data stream, perform a short-time discrete Fourier transform on the input audio data stream using a sliding time window algorithm, reduce spectral leakage using the Hanning window function, perform calculations based on the root mean square energy distribution logic in the time domain, and extract instantaneous amplitude envelope features.

[0055] Step S1 is as follows:

[0056] S11: Construct a time sliding window of a preset length, perform frame processing on the input audio data stream according to a predetermined overlap rate, perform weighted smoothing on each frame of audio data through the Hanning window function, and perform short-time discrete Fourier transform to obtain the spectral amplitude vector in the frequency domain.

[0057] S12: Traverse the effective frequency data in the spectrum amplitude vector, extract the short-time energy value of the current frame through the sum of squares and mean calculation logic, and perform recursive smoothing operation in combination with the energy data of adjacent historical frames to generate the root mean square energy distribution sequence.

[0058] S13: Perform peak preservation and attenuation processing on the root mean square energy distribution sequence, extract the envelope curve that can characterize the dynamic change trend of the signal, and generate instantaneous amplitude envelope features.

[0059] The application process of the Hanning window function specifically includes:

[0060] Obtain the length value of the current frame, generate a Hanning window coefficient sequence of the corresponding length, and perform point-by-point multiplication of the coefficient sequence with the time-domain sampled data of the current frame to reduce the amplitude at both ends of the signal and smooth the discontinuity at the splicing point.

[0061] The windowed data frame is padded with zeros to a length equal to an integer power of 2. A fast Fourier transform is then performed to calculate the magnitude of the transform result and discard the phase information, thus obtaining spectral data containing only amplitude information.

[0062] The system acquires the input audio data stream, configures the audio sampling interface to receive digital signals at a sampling rate of 48kHz and a bit depth of 24bit, and establishes a circular buffer capable of accommodating 4096 sampling points. When the amount of data in the buffer reaches the preset frame length of 2048 sampling points, the frame synchronization mechanism is activated. The frame movement step size is set to 1024 sampling points, i.e., the overlap rate is set to 50%, to ensure the continuity of the signal in the time domain and capture transient changes. Based on the current frame length... This generates the corresponding Hanning window coefficient sequence. Hanning window coefficients The generation follows the logic of cosine functions, for each point in the sequence (From 0 to 2047), calculate The generated coefficient sequence is multiplied point-by-point with the 2048 original audio samples currently read from the buffer. Specifically, the multiplication operation is performed on the first... The amplitude value of the sampling point and the first sampling point The coefficients of the window functions are multiplied to obtain the weighted time-domain data. This process significantly reduces the amplitude at both ends of the data frame to near zero, eliminating the spectral discontinuities caused by truncation, thereby suppressing sidelobe effects in subsequent frequency domain transformations.

[0063] The Hanning window function mentioned above refers to a window function used for signal processing. It is usually a type of raised cosine window, with a main lobe width that is twice that of a rectangular window, but with significant side lobe attenuation. It can effectively reduce spectral leakage and is often used for spectral analysis of random signals.

[0064] For the weighted 2048-point data frame, 2048 zero-value sampling points are padded at the end to bring its total length to 4096 points (i.e., This satisfies the optimal operational conditions for the radix-2 Fast Fourier Transform. The Fast Fourier Transform core is invoked to convert the 4096-point real number sequence into a frequency domain complex number sequence. For the transform result, the real part of each frequency point is extracted. and the virtual part Using the formula Calculate the magnitude at each frequency point. Since the input is a real signal, its spectrum has conjugate symmetry. Only the amplitude data from point 0 to point 2048 (corresponding to frequencies from 0Hz to 24000Hz) is retained, while the phase information is discarded. Finally, a spectrum amplitude vector containing 2049 elements is generated.

[0065] Traverse the spectral amplitude vector to pinpoint the low-frequency region of interest, specifically the effective frequency points between 20Hz and 250Hz. Assume that this frequency range contains... For each frequency point, its amplitude value is read. Using a sum-of-squares and mean-of-squares calculation logic, the amplitude values ​​of all valid frequency points are squared, summed, and then divided by the number of frequency points. Finally, the square root is taken to extract the short-time root-mean-square energy value of the current frame. To smooth energy fluctuations between frames, a first-order recursive smoothing algorithm is introduced. A smoothing factor is set. The value is 0.85; the smoothed energy value calculated in the previous frame is read. The instantaneous root mean square value of the current frame. Substitute into the formula: For example, if the smoothed energy of the previous frame was 0.04, and the instantaneous root mean square value of the current frame is 0.06, then the smoothed energy result of the current frame is... The process iterates frame by frame, generating a continuous root mean square energy distribution sequence.

[0066] Dynamic envelope extraction based on time constants is performed on the root mean square energy distribution sequence. The attack time constant is set to 10ms, and the release time constant to 150ms, corresponding to the rapidly rising transient response and the slowly falling reverberation decay, respectively. When processing each new energy sample, the current sample value is compared with the envelope value held at the previous moment. If the current sample value is greater than the envelope value at the previous moment, the signal is determined to be on the rising edge, and the envelope is updated using an attack coefficient to make the envelope line quickly track the signal peak. If the current sample value is less than the envelope value at the previous moment, the signal is determined to be in the decay phase, and the envelope value is slowly decayed exponentially using a release coefficient. This logic simulates the diode detection and capacitor charging / discharging process in an analog circuit. After this peak-holding and decay processing, the original sawtooth energy sequence is transformed into a smooth envelope curve that closely follows the top of the signal. This curve data is normalized to the 0-1 interval to generate the final instantaneous amplitude envelope feature, which accurately characterizes the perceived dynamic fluctuations of the audio stream in the low-frequency range.

[0067] Please refer to Figures 1 and 3. S2: Input the instantaneous amplitude envelope features into a lightweight feedforward neural network constructed based on the human ear equal loudness curve standard. Establish the nonlinear mapping relationship between the input amplitude and the perceived loudness through a polynomial fitting algorithm, and calculate the basic low-frequency gain coefficient.

[0068] Step S2 is as follows:

[0069] S21: Retrieve the preset lightweight feedforward neural network, obtain the sound pressure level value corresponding to the instantaneous amplitude envelope feature, and retrieve the perceived loudness difference of the low frequency band relative to the reference frequency under the current sound pressure level according to the human ear equal loudness curve standard.

[0070] The "Lightweight Feedforward Neural Network" is a pre-built model in memory, constructed primarily from the ISO 226:2003 standard equal-loudness curve dataset. It is mainly used to establish a non-linear mapping between the amplitude of an input sound and the perceived loudness. In its actual workflow, the network first receives normalized instantaneous amplitude envelope features as input and converts them into actual physical sound pressure level (PSL) values ​​(e.g., given that a full-scale digital scale corresponds to a 100dB PSL, a feature value of 0.5 is converted to a logarithmic level of 94dB). Subsequently, the network locks onto the system-defined low-frequency enhancement center frequency (e.g., 60Hz) and reference mid-frequency (1000Hz), and through bilinear interpolation retrieval in the model database, accurately calculates the perceived loudness difference of the low-frequency band relative to the reference frequency at the current specific PSL. Since the human ear's sensitivity to low frequencies decays much faster than to mid- and high frequencies when the sound pressure level decreases, this neural network can dynamically reflect this physiological characteristic and ultimately output the theoretical compensation difference required to maintain auditory balance consistent with the reference loudness at the current sound pressure level. This provides a crucial data foundation for subsequent calculation of the basic low-frequency gain coefficient using polynomial fitting.

[0071] S22: A mapping function between the perceived loudness difference and the compensation gain is constructed using a polynomial fitting algorithm. The current input sound pressure level value is nonlinearly transformed through the mapping function to calculate the original compensation amount.

[0072] S23: Normalize and smooth the original compensation amount, remove abrupt gain noise, and generate the basic low-frequency gain coefficient.

[0073] The process for determining the basic low-frequency gain coefficient involves nonlinear calculations performed according to the following formula:

[0074] ;

[0075] in, Represents the basic low-frequency gain coefficient. This represents the reference sound pressure level threshold in the standard equal-loudness curve. This represents the currently monitored real-time sound pressure level value. Normalized values ​​representing the instantaneous amplitude envelope characteristics. This represents the hearing compensation sensitivity factor. The exponent represents the order of a nonlinear polynomial. The basic boost bias constant representing low-frequency energy.

[0076] A lightweight feedforward neural network is retrieved from the memory. The core of this model is built upon the ISO226:2003 standard equal-loudness curve dataset. First, the normalized instantaneous amplitude envelope feature output from step S1 is converted into a sound pressure level (SPL) value. The system's digital full-scale (0dBFS) corresponds to a physical SPL of 100dB. If the normalized value of the current instantaneous amplitude envelope feature is 0.5, the corresponding logarithmic level is -6dB, meaning the actual physical SPL is 94dB. The system's low-frequency enhancement center frequency of 60Hz and reference mid-frequency of 1000Hz are locked. Bilinear interpolation is performed in the model database. For example, at a SPL of 94dB, the perceived loudness of a 60Hz sound differs from that of a 1000Hz sound. If the currently monitored real-time SPL drops to 60dB, according to the equal-loudness curve, the human ear's sensitivity to low frequencies decays much faster than to mid- and high-frequency frequencies. Therefore, the perceived loudness difference at 60Hz will significantly increase. The model outputs the theoretical compensation difference required to maintain auditory balance consistent with the reference loudness at the current sound pressure level.

[0077] The aforementioned ISO 226:2003 standard refers to the acoustic standard published by the International Organization for Standardization regarding the equal loudness curves of normal human ears. It describes the relationship between sound pressure levels required for the human ear to perceive the same loudness at different frequencies, reflecting the nonlinear frequency response characteristics of human hearing.

[0078] The original compensation amount is calculated using a mapping function constructed using a polynomial fitting algorithm. The currently monitored real-time sound pressure level value is then obtained. and the normalized values ​​of the instantaneous amplitude envelope features Set the reference sound pressure level threshold in the standard equal-loudness curve. The threshold is 85dB, which is the optimal listening reference point determined through subjective evaluation experiments. The listening compensation sensitivity factor is then set. The exponent of the order of the nonlinear polynomial is 0.45. The basic boost bias constant for low-frequency energy is 1.2. The value is 2.5. Substituting the above parameters into the nonlinear calculation formula: .

[0079] in, This represents the basic low-frequency gain coefficient, used to quantify the low-frequency compensation gain value required for the current frame. The reference sound pressure level threshold in the standard equal loudness curve serves as the benchmark operating point for auditory compensation. This represents the currently monitored real-time sound pressure level, reflecting the current actual playback volume level. The normalized value representing the instantaneous amplitude envelope characteristic characterizes the real-time dynamic strength of the signal. This represents the hearing compensation sensitivity factor, used to adjust the response rate of the compensation amount as the sound pressure level difference changes; The order exponent of the nonlinear polynomial determines the curvature of the compensation curve to fit the characteristics of human hearing. The basic boost bias constant, representing low-frequency energy, is used to introduce fine-tuning gain that is dynamically related to the signal.

[0080] Taking actual monitoring data as an example, let's assume the current real-time sound pressure level value... The normalized value of the instantaneous amplitude envelope feature is 65 dB. The value is 0.017. Substituting into the formula, the first step is to calculate the sound pressure level difference. Perform exponential operations ; Calculate the result of the first term Next, calculate the logarithmic part. ; Calculate the result of the second term Finally, the basic low-frequency gain coefficient is obtained by summing. The result is approximately 16.40 dB. This calculation indicates that at the current low sound pressure level of 65 dB, the system needs to introduce approximately 16.40 dB of low-frequency gain to compensate for the human ear's insensitivity to low frequencies. The exponential term in the formula ensures that the gain increases rapidly and non-linearly as the volume decreases, consistent with the characteristics of human hearing; the logarithmic term introduces fine-tuning related to signal dynamics, preventing excessive noise amplification at quiet locations.

[0081] The calculated raw compensation value is normalized to limit it within the system's allowed hardware gain range (0dB to 18dB). If the calculation result exceeds the upper limit, it is truncated to 18dB. This gain value is then input into a moving average filter of length 5 to eliminate gain abrupt noise caused by inter-frame calculation errors or minor signal jitter. For example, if the gain sequence of five consecutive frames is 16.2, 16.4, 18.1, 16.3, 16.4, where 18.1 is an abnormal abrupt change, after moving average processing, a smoothed base low-frequency gain coefficient is output to ensure the stability of the final control signal and avoid unnatural fluctuations in sound.

[0082] Please refer to Figures 1 and 4. S3: Construct a safety response boundary based on the speaker's physical stroke parameters, voice coil thermal time constant, and mechanical damping characteristics. Combine the instantaneous amplitude envelope characteristics with the historical peak statistics of the time series within a set period for weighted evaluation to generate a dynamic safety cutoff threshold.

[0083] The S3 steps are as follows:

[0084] S31: Obtain the maximum linear displacement parameters, rated power handling capacity, and mechanical damping coefficient of the loudspeaker. Based on physical limit logic, construct a safe response space using displacement boundary and thermal load boundary.

[0085] S32: Establish a historical peak statistical buffer for time series, store instantaneous amplitude envelope feature data within a set period in real time, calculate the peak factor and average power density of the data in the buffer, and perform weighted evaluation to quantify the current cumulative thermal stress state.

[0086] S33: Dynamically match the safety response space with the current cumulative thermal stress state, and adjust the maximum allowable gain limit in real time based on the matching result to generate a dynamic safety cutoff threshold.

[0087] The process of setting a dynamic safety truncation threshold specifically includes a weighted calculation based on the following logic:

[0088] ;

[0089] in, Represents the dynamic safety cutoff threshold. This represents the maximum allowable linear displacement of the speaker's voice coil. Represents the gain conversion coefficient based on the damping characteristics of the mechanical system. Represents the thermal stress weighting factor of the voice coil. This represents the total number of sampling points within a set period. Representing the Historical power values ​​at each sampling point This represents the value of the exponentially weighted function that decays over time.

[0090] Access the physical parameter database of the loudspeaker unit to obtain key electromechanical specifications. Set the maximum allowable linear displacement value of the loudspeaker voice coil. The displacement is 5.5mm (unidirectional), with a rated power handling capacity of 50W and a mechanical damping coefficient of 2.4kg / s. Based on physical limit logic, a two-dimensional safety response space is constructed: the horizontal axis represents the voice coil displacement, and the vertical axis represents the input power. The displacement boundary is defined as ±5.5mm, and the thermal load boundary is defined as a continuous power of 50W. Through calibration experiments of the laser displacement sensor, actual displacement data at different frequencies and voltages are measured, and a voltage-displacement conversion model is established to map the physical displacement limit to the corresponding voltage amplitude limit. For example, at 60Hz, the peak voltage required to achieve a 5.5mm displacement is 18V.

[0091] A circular buffer is allocated in memory as a buffer for statistical analysis of historical peak values ​​in the time series data. The buffer length is set to [value missing]. Each sampling point (corresponding to a 1-second duration, sampling rate 48kHz) stores the instantaneous amplitude envelope feature data within the past second in real time. Whenever new data is stored, the oldest data is removed. The crest factor of the data in the buffer, i.e., the ratio of the peak value to the root mean square value, is calculated to determine whether the signal contains large dynamic impulses. Simultaneously, the average power density is calculated. An exponentially weighted function that decays over time is introduced. ,in For sample index, The sampling rate is used. This weighting logic assigns higher weight to the most recent data to accurately reflect the current thermal accumulation state of the voice coil. The current cumulative thermal stress state value is calculated by weighted summation, quantifying the temperature rise trend of the voice coil relative to the ambient temperature.

[0092] The crest factor mentioned above refers to the ratio of the peak value of a waveform to its effective value (root mean square value). It is used to describe the extreme degree of a signal waveform and is often used in the audio field to measure the dynamic range of a signal.

[0093] The calculated cumulative thermal stress state is dynamically matched with the safety response space, and a dynamic safety cutoff threshold is generated using weighted calculation logic. Speaker parameters are then obtained. Set the gain conversion coefficient based on the damping characteristics of the mechanical system. (Unit conversion and normalization factor), voice coil thermal stress weighting coefficient Assuming that after the statistics in step S32, the historical power weighted sum within the period is set. The calculated result is 15.6 (normalized power units). Substituting the parameters into the formula: Calculate the numerator ; Calculate the weighted terms in the denominator ; Calculate the total value in the denominator ; Calculate the final dynamic safety cutoff threshold The value is approximately 4.40. This result indicates that, due to the current high thermal stress accumulation state of the loudspeaker (increased denominator), the maximum permissible linear gain is dynamically compressed to 4.40. Compared to the threshold in the cold state (assuming a denominator of 1, the threshold is 9.9), the system actively lowers the safety limit to prevent overheating damage. Table 1 shows the experimental data of the system's dynamically adjusted safety cutoff threshold under different accumulated thermal stress states.

[0094] Table 1 Comparison of Speaker Thermal Stress State and Dynamic Threshold

[0095] The experimental group's cumulative thermal stress weighted and thermal state description calculated safe cutoff threshold, measured voice coil temperature rise (°C), and whether overload distortion occurred: 12.5 Cold state or light load 8.255 No 28.0 Warm state or medium load 6.0435 No 315.6 Hot state or heavy load 4.4082 No 425.0 Limiting state 3.30110 No surface

[0096] Experimental data show that as the accumulated thermal stress increases, the calculated safe cutoff threshold decreases inversely, effectively controlling the voice coil temperature rise. Under extreme conditions (Group 4), overload distortion was successfully avoided, verifying the formula's... The rationality of parameter settings.

[0097] Please refer to Figures 1 and 5. S4: Dynamic compression and adaptive limiting are performed on the basic low-frequency gain coefficient using a dynamic safety truncation threshold. The gain parameters of the digital shelving filter are controlled based on the processed coefficients to adjust the low-frequency components in the input audio data stream and generate an optimized sound field drive signal.

[0098] The S4 steps are as follows:

[0099] S41: Compare the basic low-frequency gain coefficient with the dynamic safety cutoff threshold. When the coefficient exceeds the threshold, the hard inflection point compression logic is activated. When the coefficient does not exceed the threshold, linear propagation is maintained, thereby calculating the actual execution gain.

[0100] S42: Calculate the coefficient matrix of the digital shelving filter based on the actual execution gain, and dynamically update the filter's center frequency, quality factor, and gain-bandwidth product parameters to ensure that the filter response curve meets the target sound field requirements.

[0101] S43: Imports the input audio data stream into the updated digital shelving filter, performs amplitude modulation on the signal components below the specified low-frequency cutoff frequency, while keeping the mid-to-high frequency components pass-through, and generates an optimized sound field drive signal.

[0102] The process of updating parameters specifically includes:

[0103] Obtain the current sampling rate and target cutoff frequency of the digital shelving filter, calculate intermediate variables based on the actual execution gain, and derive the feedforward coefficients and feedback coefficients of the second-order IIR filter using the bilinear transform method.

[0104] The stability of the detection coefficient update process is checked. If the calculated coefficient causes the poles to exceed the unit circle, the filter is forcibly reset to the safe coefficient state of the previous frame. Otherwise, the new coefficients are written into the filter register to complete the parameter update.

[0105] The "basic low-frequency gain coefficient" generated in step S2 (approximately 16.40 dB, which translates to a linear multiple of approximately 6.61) is compared with the "dynamic safety cutoff threshold" (4.40) generated in step S3. The decision logic is as follows: if the basic gain coefficient (6.61) is greater than the dynamic safety cutoff threshold (4.40), then the hard-knee compression logic is triggered. At this time, the actual executed gain is forcibly clamped to the threshold 4.40, corresponding to approximately 12.87 dB. If the basic gain coefficient is less than the threshold, linear propagation is maintained, and the actual executed gain equals the basic gain coefficient. In this example, since 6.61 is greater than 4.40, the system determines that there is an overload risk, so the output actual executed gain is 4.40. This process achieves millisecond-level adaptive limiting, prioritizing the physical safety of the speaker while maximizing low-frequency loudness within a safe range.

[0106] Calculate the coefficient matrix of the second-order digital shelving filter based on the actual execution gain (linear value 4.40). Set the center frequency of the filter. Quality Factor Sampling rate Calculate intermediate variables using the bilinear transform method: calculate the angular frequency parameters. Set the gain parameter Calculate the feedforward coefficients according to the formula for the coefficients of a second-order IIR filter. , , and feedback coefficient , After each parameter update calculation, a stability check is performed immediately: the characteristic equation is solved. The root (pole). If the magnitude of any pole is greater than or equal to 1, it indicates that the filter is unstable. In this case, the currently calculated coefficients are forcibly discarded and the filter is reset to the safe coefficient state of the previous frame. Otherwise, the new coefficients are written to the DSP register.

[0107] The aforementioned bilinear transformation method refers to a mathematical transformation method used to convert the transfer function of a continuous-time system (analog filter) into the transfer function of a discrete-time system (digital filter). Its characteristic is that it can nonlinearly map the analog frequency axis to the digital frequency axis, thus avoiding frequency aliasing.

[0108] The input audio data stream is fed into a digital shelving filter with updated parameters. The filter is implemented using a direct type II transpose structure to minimize quantization noise accumulation. For each input sample point... Perform difference equation operations: ; ; This calculation modulates the signal components below 60Hz (amplified by 4.40 times in this example) while maintaining the mid-to-high frequency components above 200Hz with 0dB gain. Table 2 shows a comparison of total harmonic distortion and subjective listening scores before and after applying this technology at different input sound pressure levels.

[0109] Table 2 Comparison of Sound Quality Optimization Effects

[0110] Test Conditions (Input Sound Pressure Level) Traditional Solution Total Harmonic Distortion (%) This Solution Total Harmonic Distortion (%) Traditional Solution Low-Frequency Listening Score (1 to 10) This Solution Low-Frequency Listening Score (1 to 10) 60dB (Low Volume) 0.05 0.063 885dB (Standard Volume) 0.50 0.52 7898dB (High Volume) 12.40 1.8047 surface

[0111] Experimental data show that at low volume (60dB), this solution significantly improves the low-frequency listening experience through nonlinear compensation (score increases from 3 to 8) without introducing obvious distortion; under the extreme condition of high volume (98dB), thanks to the effect of the dynamic safety cutoff threshold, this solution greatly reduces the total harmonic distortion from 12.40% to 1.80%, generating an optimized sound field driving signal that meets the listening requirements while remaining within the physical safety boundary.

[0112] A low-frequency enhancement / reduction system for an audio field based on volume amplitude recognition is used to execute the aforementioned low-frequency enhancement / reduction method for an audio field based on volume amplitude recognition. The system includes:

[0113] The feature extraction module is used to acquire the input audio data stream, perform short-time discrete Fourier transform on the input audio data stream using a sliding time window algorithm, reduce spectral leakage using the Hanning window function, and perform calculations based on the root mean square energy distribution logic in the time domain to extract instantaneous amplitude envelope features.

[0114] The compensation calculation module is used to input the instantaneous amplitude envelope features into a lightweight feedforward neural network constructed based on the human ear equal loudness curve standard, and to establish a nonlinear mapping relationship between the input amplitude and the perceived loudness through a polynomial fitting algorithm, and to calculate the basic low-frequency gain coefficient.

[0115] The safety assessment module is used to construct a safety response boundary based on the speaker's physical stroke parameters, voice coil thermal time constant, and mechanical damping characteristics. It combines the instantaneous amplitude envelope characteristics with historical peak statistics of the time series within a set period for weighted evaluation to generate a dynamic safety cutoff threshold.

[0116] The control execution module is used to dynamically compress and adaptively limit the basic low-frequency gain coefficient using a dynamic safety truncation threshold, control the gain parameters of the digital shelving filter based on the processed coefficient, adjust the low-frequency components in the input audio data stream, and generate an optimized sound field drive signal.

[0117] The above embodiments illustrate preferred embodiments of the present invention. Any equivalent adjustments to the technical solution based on software engineering methods are within the scope of protection, including but not limited to: implementing algorithm logic using different programming languages, refactoring functional modules into services, adjusting data interaction protocols, and optimizing resource scheduling strategies. Any implementation scheme derived from reasonable modifications to the data processing flow, service call chain, or system architecture layer without departing from the core technology of the present invention should be considered within the protection scope defined by the technical solution of the present invention.

Claims

1. A method for low-frequency addition and subtraction of sound field based on volume amplitude recognition, characterized in that, Includes the following steps: S1: Acquire the input audio data stream, perform a short-time discrete Fourier transform on the input audio data stream using a sliding time window algorithm, reduce spectral leakage using a Hanning window function, and perform calculations based on the root mean square energy distribution logic in the time domain to extract instantaneous amplitude envelope features; S2: Input the instantaneous amplitude envelope features into a lightweight feedforward neural network constructed based on the human ear equal loudness curve standard, establish a nonlinear mapping relationship between input amplitude and perceived loudness using a polynomial fitting algorithm, and calculate the basic low-frequency gain coefficient; S3: Construct a safe response boundary based on the speaker's physical stroke parameters, voice coil thermal time constant, and mechanical damping characteristics, and perform a weighted evaluation based on the historical peak statistics of the instantaneous amplitude envelope features within a set period to generate a dynamic safe cutoff threshold; S4: Use the dynamic safe cutoff threshold to perform dynamic compression and adaptive limiting processing on the basic low-frequency gain coefficient, control the gain parameters of the digital shelving filter based on the processed coefficient, adjust the low-frequency components in the input audio data stream, and generate an optimized sound field driving signal.

2. The method for low-frequency addition and subtraction of sound field based on volume amplitude recognition according to claim 1, characterized in that, The steps of S1 are as follows: S11: Construct a time sliding window of a preset length, perform frame processing on the input audio data stream according to a predetermined overlap rate, perform weighted smoothing on each frame of audio data through the Hanning window function, and perform short-time discrete Fourier transform to obtain the spectral amplitude vector in the frequency domain; S12: Traverse the effective frequency point data in the spectral amplitude vector, extract the short-time energy value of the current frame through the sum of squares and mean calculation logic, and perform recursive smoothing operation in combination with the energy data of adjacent historical frames to generate the root mean square energy distribution sequence; S13: Perform peak preservation and attenuation processing on the root mean square energy distribution sequence, extract the envelope curve that can characterize the dynamic change trend of the signal, and generate the instantaneous amplitude envelope feature.

3. The method for low-frequency increase / decrease of sound field based on volume amplitude recognition according to claim 1, characterized in that, The steps in S2 are as follows: S21: Retrieve a preset lightweight feedforward neural network to obtain the sound pressure level value corresponding to the instantaneous amplitude envelope feature, and retrieve the perceived loudness difference of the low-frequency band relative to the reference frequency at the current sound pressure level according to the human ear equal loudness curve standard; S22: Construct a mapping function between the perceived loudness difference and the compensation gain using a polynomial fitting algorithm, and perform a nonlinear transformation on the current input sound pressure level value through the mapping function to calculate the original compensation amount; S23: Normalize and smooth the original compensation amount, remove abrupt gain noise, and generate the basic low-frequency gain coefficient.

4. The method for low-frequency increase / decrease of sound field based on volume amplitude recognition according to claim 1, characterized in that, The steps in S3 are as follows: S31: Obtain the maximum linear displacement parameters, rated power handling capacity, and mechanical damping coefficient of the loudspeaker. Based on physical limit logic, construct a safe response space using displacement boundaries and thermal load boundaries; S32: Establish a time series historical peak statistical buffer, store the instantaneous amplitude envelope feature data within a set period in real time, calculate the peak factor and average power density of the data in the buffer, and perform weighted evaluation to quantify the current cumulative thermal stress state; S33: Dynamically match the safe response space with the current cumulative thermal stress state, adjust the maximum allowable gain limit in real time based on the matching result, and generate the dynamic safety cutoff threshold.

5. The method for low-frequency addition and subtraction of sound field based on volume amplitude recognition according to claim 1, characterized in that, The steps in S4 are as follows: S41: Compare the basic low-frequency gain coefficient with the dynamic safety cutoff threshold. When the coefficient exceeds the threshold, activate the hard inflection point compression logic. When the coefficient does not exceed the threshold, maintain linear propagation to calculate the actual execution gain. S42: Calculate the coefficient matrix of the digital shelving filter based on the actual execution gain. Dynamically update the center frequency, quality factor, and gain-bandwidth product parameters of the filter to ensure that the filter response curve meets the target sound field requirements. S43: Import the input audio data stream into the digital shelving filter with updated parameters. Amplify the signal components below the specified low-frequency cutoff frequency while maintaining the mid-to-high frequency components to generate the optimized sound field driving signal.

6. The method for low-frequency addition and subtraction of sound field based on volume amplitude recognition according to claim 3, characterized in that, The process for obtaining the basic low-frequency gain coefficient involves a nonlinear calculation performed according to the following formula: ;in, This represents the basic low-frequency gain coefficient. This represents the reference sound pressure level threshold in the standard equal-loudness curve. This represents the currently monitored real-time sound pressure level value. The normalized value representing the instantaneous amplitude envelope feature, This represents the hearing compensation sensitivity factor. The exponent represents the order of a nonlinear polynomial. The basic boost bias constant representing low-frequency energy.

7. The method for low-frequency addition and subtraction of sound field based on volume amplitude recognition according to claim 4, characterized in that, The process of setting the dynamic safety truncation threshold specifically includes a weighted calculation performed according to the following logic: ;in, This represents the dynamic security truncation threshold. This represents the maximum allowable linear displacement of the speaker's voice coil. Represents the gain conversion coefficient based on the damping characteristics of the mechanical system. Represents the thermal stress weighting factor of the voice coil. This represents the total number of sampling points within a set period. Representing the Historical power values ​​at each sampling point This represents the value of the exponentially weighted function that decays over time.

8. The method for low-frequency addition and subtraction of sound field based on volume amplitude recognition according to claim 2, characterized in that, The application process of the Hanning window function specifically includes: obtaining the length value of the current frame, generating a Hanning window coefficient sequence of corresponding length, multiplying the coefficient sequence with the time-domain sampled data of the current frame point by point to reduce the amplitude at both ends of the signal and smooth the discontinuity at the splicing point; padding the windowed data frame with zeros to an integer power of 2 length, performing a fast Fourier transform operation, calculating the magnitude of the transform result and discarding the phase information, thereby obtaining spectral data containing only amplitude information.

9. The method for low-frequency addition and subtraction of sound field based on volume amplitude recognition according to claim 5, characterized in that, The parameter update process specifically includes: obtaining the current sampling rate and target cutoff frequency of the digital shelving filter, calculating intermediate variables in conjunction with the actual execution gain, deriving the feedforward coefficients and feedback coefficients of the second-order IIR filter using the bilinear transform method; detecting the stability during the coefficient update process, if the calculated coefficients cause the poles to exceed the unit circle, then forcibly resetting to the safe coefficient state of the previous frame, otherwise writing the new coefficients into the filter register to complete the parameter update.

10. A low-frequency boosting / decrease system for a sound field based on volume amplitude recognition, characterized in that, The system is used to implement the method for low-frequency addition and subtraction of sound field based on volume amplitude recognition as described in any one of claims 1-9. The system includes: a feature extraction module, used to acquire input audio data streams, perform short-time discrete Fourier transform on the input audio data streams using a sliding time window algorithm, reduce spectral leakage using a Hanning window function, and perform calculations based on the root mean square energy distribution logic in the time domain to extract instantaneous amplitude envelope features; and a compensation calculation module, used to input the instantaneous amplitude envelope features into a lightweight feedforward neural network constructed based on the human ear equal loudness curve standard, and establish a nonlinear relationship between the input amplitude and perceived loudness using a polynomial fitting algorithm. The system employs a mapping relationship to calculate the basic low-frequency gain coefficient; a safety assessment module to construct a safety response boundary based on the speaker's physical stroke parameters, voice coil thermal time constant, and mechanical damping characteristics, and to perform a weighted evaluation based on the instantaneous amplitude envelope characteristics within a set period of historical peak statistics to generate a dynamic safety cutoff threshold; and a control execution module to dynamically compress and adaptively limit the basic low-frequency gain coefficient using the dynamic safety cutoff threshold, control the gain parameters of the digital shelving filter based on the processed coefficients, adjust the low-frequency components in the input audio data stream, and generate an optimized sound field drive signal.