Sound gain control method and device, electronic equipment and storage medium

By using dual-channel signal acquisition and signal-to-noise ratio processing via external and in-ear microphones, and dynamically adjusting the sound gain, the problem of incorrect volume adjustment caused by ambient sound fluctuations in existing technologies is solved, achieving a clear and comfortable audio experience in smart headphones and hearing aids.

CN120856085APending Publication Date: 2025-10-28SHENZHEN PHICOUSTIC SYST DEV CO LTD +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510862458.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-25
Publication Date
2025-10-28

AI Technical Summary

Technical Problem

Existing adaptive volume solutions are prone to accidentally increasing volume or having a delayed response when faced with ambient sound fluctuations, making it difficult to achieve appropriate sound gain adjustment under different noise environments.

Method used

By acquiring signals from both an external microphone and an in-ear microphone, and combining scene reference thresholds and real-time signal-to-noise ratio, the device's sound gain is dynamically adjusted. The external microphone captures ambient sound field characteristics, while the in-ear microphone focuses on the user's voice. The dual-microphone array accurately distinguishes between ambient noise and user speech, and gain adjustment is performed based on the real-time signal-to-noise ratio.

Benefits of technology

It achieves adaptive adjustment in different environments, improves the accuracy of signal-to-noise ratio calculation, and provides a clearer, more comfortable, and power-optimized listening experience, making it particularly suitable for smart headphones and hearing aids.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120856085A_ABST
    Figure CN120856085A_ABST
Patent Text Reader

Abstract

The embodiment of the invention discloses a sound gain control method and device, electronic equipment and a computer readable storage medium, and relates to the technical field of audio processing, and the method comprises the steps: obtaining an environment sound signal collected through an external microphone, and a target sound signal collected through an in-ear microphone; based on the environment sound signal, obtaining a reference threshold value corresponding to a scene where the equipment is located; determining a real-time signal-to-noise ratio based on the environment sound signal and the target sound signal; and adjusting the sound gain of the device based on the reference threshold and the real-time signal-to-noise ratio. Therefore, adaptive adjustment in different environments is realized, the accuracy of signal-to-noise ratio calculation is improved through dual-microphone signal processing, clearer and more comfortable auditory experience with optimized power consumption is finally brought to a user, and the method is particularly suitable for scenes needing real-time audio enhancement, such as intelligent earphones and hearing-aid equipment.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates to the field of audio processing technology, specifically to a sound gain control method, apparatus, electronic device, and computer-readable storage medium. Background Technology

[0002] Most existing adaptive volume solutions only detect the sound pressure level of the external environment. When there are brief fluctuations in ambient sound (wind noise, clothing friction, covering oneself with a blanket before sleeping, etc.), the algorithm is prone to mistakenly triggering an increase in volume, while it is slow to react when faced with continuous noise.

[0003] Therefore, how to dynamically determine the appropriate sound gain for the device is an urgent problem to be solved. Summary of the Invention

[0004] This disclosure provides a sound gain control method, apparatus, electronic device, and computer-readable storage medium, which aims to at least partially solve one of the technical problems in the related art.

[0005] In a first aspect, embodiments of this disclosure provide a sound gain control method, the method comprising:

[0006] Acquire ambient sound signals collected via an external microphone, and target sound signals collected via an in-ear microphone;

[0007] Based on the ambient sound signal, obtain a reference threshold corresponding to the scene where the device is located;

[0008] The real-time signal-to-noise ratio is determined based on the ambient sound signal and the target sound signal;

[0009] The sound gain of the device is adjusted based on the reference threshold and the real-time signal-to-noise ratio.

[0010] Secondly, embodiments of this disclosure also provide a sound gain control device, the device comprising:

[0011] The first acquisition module is used to acquire ambient sound signals collected by an external microphone and target sound signals collected by an in-ear microphone.

[0012] The second acquisition module is used to acquire a reference threshold corresponding to the scene where the device is located based on the ambient sound signal;

[0013] The first determining module is used to determine the real-time signal-to-noise ratio based on the ambient sound signal and the target sound signal;

[0014] An adjustment module is used to adjust the sound gain of the device based on the reference threshold and the real-time signal-to-noise ratio.

[0015] Thirdly, this disclosure also provides an electronic device, which includes a memory, a processor, and a computer program stored in the memory and executable on the processor. When the computer program is executed by the processor, it implements the steps in the above-described sound gain control method.

[0016] Fourthly, embodiments of this disclosure also provide a computer-readable storage medium storing a computer program, which, when executed by a processor, implements the steps in the aforementioned sound gain control method.

[0017] Fifthly, embodiments of this disclosure also provide a computer program product or computer program including computer instructions stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium and executes the computer instructions, causing the computer device to perform the methods provided in various optional implementations of embodiments of this disclosure.

[0018] In this embodiment, the system first acquires ambient sound signals collected by an external microphone and target sound signals collected by an in-ear microphone. Then, based on the ambient sound signals, a reference threshold corresponding to the scene in which the device is located is obtained. Next, based on the ambient sound signals and the target sound signals, the real-time signal-to-noise ratio (SNR) is determined. Finally, based on the reference threshold and the real-time SNR, the device's sound gain is adjusted. By acquiring signals from both the external and in-ear microphones, combined with the collaborative processing of scene reference thresholds and real-time SNR, multi-dimensional auditory experience optimization can be achieved. The dual-microphone array can accurately distinguish between ambient noise and user speech. The external microphone captures ambient sound field characteristics to match the reference threshold for the corresponding scene, while the in-ear microphone focuses on near-field sounds from the user, improving the accuracy of target signal pickup. Dynamically adjusting the gain based on the real-time SNR can ensure speech clarity while avoiding noise amplification. For example, when ambient noise suddenly increases, the system automatically enhances noise reduction and appropriately increases speech gain based on the scene threshold. Furthermore, the near-field pickup characteristics of the in-ear microphone can reduce echo and wind noise interference. This solution utilizes scene classification to achieve adaptive adjustment in different environments, and improves the accuracy of signal-to-noise ratio calculation through dual-microphone signal processing. Ultimately, it brings users a clearer, more comfortable, and power-optimized auditory experience, making it particularly suitable for scenarios such as smart headphones and hearing aids that require real-time audio enhancement.

[0019] It is to be understood that the foregoing general description and the following detailed description are exemplary and explanatory only and are not restrictive of the disclosure. Attached Figure Description

[0020] To more clearly illustrate the technical solutions in this disclosure, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0021] Figure 1 This is a schematic flowchart of the sound gain control method provided in the first embodiment of this disclosure;

[0022] Figure 2 This is a schematic flowchart of the sound gain control method provided in the second embodiment of this disclosure;

[0023] Figure 3 This is a schematic diagram of the structure of the sound gain control device provided in the embodiments of this disclosure;

[0024] Figure 4 This is a schematic diagram of the structure of the electronic device provided in the embodiments of this disclosure. Detailed Implementation

[0025] Some embodiments of this disclosure will now be described in detail, examples of which are illustrated in the accompanying drawings. When the following description refers to the drawings, unless otherwise indicated, the same numbers in different drawings denote the same or similar elements. Various changes, modifications, and equivalents of the methods, apparatus, and / or systems described herein will become apparent upon understanding this disclosure. For example, the order of operations described herein is merely illustrative and is not limited to those orders set forth herein, but can be changed as will become apparent upon understanding this disclosure, except for operations that must be performed in a particular order. Furthermore, for clarity and brevity, descriptions of features known in the art may be omitted.

[0026] The embodiments described in the following examples of this disclosure are not representative of all embodiments consistent with this disclosure. Rather, they are merely examples of apparatuses and methods consistent with some aspects of this disclosure as detailed in the appended claims.

[0027] It should be noted that the execution subject of the sound gain control method in this embodiment can be a sound gain control device, which can be configured in any type of electronic device, such as headphones, without limitation.

[0028] In this embodiment, the "sound gain control device" will be used as the execution subject to describe the "sound gain control method", and no limitation will be made here.

[0029] It should be noted that the order of description of the following embodiments is not intended to limit the priority of the embodiments.

[0030] Figure 1 This is a schematic flowchart of the sound gain control method provided according to the first embodiment of this disclosure.

[0031] like Figure 1 As shown, the method includes:

[0032] Step 101: Acquire the ambient sound signal collected by the external microphone and the target sound signal collected by the in-ear microphone.

[0033] Among them, the external microphone (ambient microphone) is usually located on the outside of the headphones (such as the outside of the headphone stem or earcups), facing the external environment. It can collect background sounds in the environment (such as human voices, traffic noise, wind noise, etc.) and cover a wide frequency range in order to capture various environmental noises.

[0034] The in-ear microphone is located close to the inside of the ear canal and near the eardrum. It collects the user's own voice signal (target sound) or residual noise in the ear canal for call noise reduction, voice assistant wake-up, or feedback adjustment of noise reduction effect. It is closer to the sound source (human voice), which can reduce the interference of environmental noise and focus on the target audio.

[0035] The target sound signal is the sound signal captured by the in-ear microphone, which may include a composite signal of audio played by the headphones and the sound signal of current ambient noise entering the ear canal.

[0036] Among them, the ambient sound signal is the sound signal collected by the external microphone.

[0037] Step 102: Based on the ambient sound signal, obtain the reference threshold corresponding to the scene where the device is located.

[0038] The equipment can be located in various settings such as subways, plazas, supermarkets, classrooms, cafes, and airports.

[0039] Optionally, features can be extracted from the ambient sound signal to obtain multidimensional Mel frequency cepstral coefficient features. Then, a pre-trained scene classification model can be used to classify the Mel frequency cepstral coefficient features to obtain the scene where the device is located. After that, based on the mapping relationship between the preset scene and scene threshold combination, the scene threshold combination corresponding to the scene where the device is located is used as the reference threshold.

[0040] The scene threshold combination may include at least the following thresholds: Tlow (low gain threshold), Thigh (low gain threshold), Δimpulse (impulse noise threshold), and Gmax (upper gain limit).

[0041] The reference threshold can be a combination of scene thresholds corresponding to the scene in which the current device is located.

[0042] As an example, a scene classification model can use MobileNetV1-1D (one-dimensional convolutional neural network, 1DCNN) to process temporal audio signals. Depthwise separable convolutions reduce the number of parameters, making it suitable for low-computing-power devices (such as headphone chips). By converting floating-point parameters to 8-bit integers, the model size is compressed to 45kB (original parameter count ≈ 38k), reducing computational complexity and power consumption, making it suitable for embedded deployments.

[0043] Specifically, the ambient sound signal (acquired by an external microphone) can be processed by framing (e.g., 25ms) and windowing (Hamming window). A Mel filter bank is used to convert the linear spectrum to a Mel spectrum, and then a Discrete Cosine Transform (DCT) is performed to obtain the 20-dimensional Mel frequency cepstral coefficient features (MFCC features), simulating the human ear's perception of sound. MFCC is sensitive to the spectral envelope characteristics of ambient noise and can effectively distinguish the acoustic characteristics of different scenarios (such as reverberation in a shopping mall or low-frequency rumble in a subway).

[0044] For example, it can be pre-trained on public audio datasets (such as ESC-50 and UrbanSound8K) to learn feature patterns for six scene categories (quiet indoor, shopping mall, street, car, subway, and pulsating wind). An audio data segment (corresponding to approximately 10 frames of MFCC features) is input every 320ms. Deep features are extracted using MobileNetV1-1D, and the classification probability is output through a fully connected layer to determine the current scene.

[0045] The table below shows the typical acoustic features and MFCC features for each scenario.

[0046]

[0047]

[0048] The table below shows the mapping relationship between reference thresholds and scenes:

[0049]

[0050] It should be noted that the above examples are merely illustrative and are not intended to limit the embodiments of this application.

[0051] Step 103: Determine the real-time signal-to-noise ratio based on the ambient sound signal and the target sound signal.

[0052] The signal-to-noise ratio (SNR) is expressed in this application as the real-time SNR, which is the difference between the playback power of the target sound signal and the ambient noise power of the ambient sound signal.

[0053] Optionally, the ambient sound signal can be segmented into frames based on a preset frame length and frame shift. The segmented ambient sound signal can be subjected to a first weighted filtering process to obtain the first sound power corresponding to each frame of audio signal. The first weighting is a frequency weighting that simulates the hearing characteristics of the human ear. The power of the target sound signal can be calculated to obtain the second sound power. The difference between the second sound power and the first sound power is used as the real-time signal-to-noise ratio.

[0054] The frame length is the duration of the audio signal processed each time, commonly measured in milliseconds (ms), with typical values ​​of 20–50 ms (corresponding to 320–800 sampling points at a 16 kHz sampling rate). The frame shift is the interval between the start positions of two adjacent frames, typically 50%–75% of the frame length (e.g., a 25 ms frame length with a 10 ms frame shift), ensuring frame overlap for smooth processing of timing signals. A preset frame length of 32 ms and a frame shift of 16 ms are available.

[0055] The first weighting method can be A-weighting, a frequency-weighted filtering method based on the characteristics of human hearing. By applying different gains to sound signals of different frequencies, it makes the measurement results closer to the subjective perception of loudness, originating from IEC 61672 (International Electrotechnical Commission acoustic standard). The principle behind A-weighting's filtering curve is that the human ear is most sensitive to mid-range frequencies (2-5kHz) and less sensitive to low frequencies (<500Hz) and high frequencies (>8kHz). A-weighting attenuates low-frequency and high-frequency signals, making the measured values ​​more consistent with subjective loudness perception.

[0056] As one embodiment, dual-track energy tracking processing can be performed on the first sound power corresponding to each frame of audio signal based on preset fast trajectory time constant and slow trajectory time constant to obtain the sound power difference corresponding to each frame of audio signal, determine the impulse noise threshold included in the reference threshold, and then determine that the audio signal is impulse noise if the sound power difference corresponding to the audio signal is greater than or equal to the impulse noise threshold; otherwise, the audio signal is not impulse noise, and the audio signal is discarded if it is impulse noise.

[0057] Among them, the impulse noise threshold is the fast trajectory difference (P fast The difference between the slow trajectory and the slow trajectory (P) slow The decision threshold is used to identify impulse noise and avoid false triggering.

[0058] Among them, the fast trajectory P fast Using a relatively large smoothing coefficient αf = 0.2 and a time constant of 64–80 ms, it is sensitive to real-time power changes and quickly follows power spikes (such as impulse noise). The formula is: P fast[n] =P fast[n-1] +αf(P frame[n] -P fast[n-1]).

[0059] Among them, P frame[n] This represents the sound power of the nth frame after the first weighting process.

[0060] Among them, the slow trajectory P slow A relatively small smoothing coefficient αs = 0.02 and a time constant of 0.8s are used to reflect the long-term trend of power, maintain the background noise baseline, and suppress short-term fluctuations. The formula is: P slow[n] =P slow[n-1] +αs(P frame[n] -P slow[n-1] The smoothing process can be equivalent to a first-order RC low-pass filter, with a time constant τ≈(1 / α-1)*Tshift, where Tshift=16ms is the frame shift time. It should be noted that a fast trajectory can follow a spike within tens of milliseconds, while a slow trajectory catches up to the same height in about 0.8s, sufficient to maintain memory of the background baseline.

[0061] The formula for calculating the sound power difference is as follows:

[0062] ΔP=∣P fast -P slow |

[0063] If ΔP ≥ Δimpulse, then the current frame is impulse noise and needs to be removed.

[0064] The power calculation steps for the first sound power can be as follows: perform a fast Fourier transform (FFT) on the framed audio signal of the ambient sound signal to convert it to the frequency domain, apply a weighting filter (such as the frequency response function of an A-weighting filter) to weight each frequency component, then calculate the power spectrum of the weighted frequency domain signal, and accumulate to obtain the first sound power (Pout) of each frame.

[0065] The calculation steps for the second sound power can be as follows: directly perform time-domain power calculation (such as root mean square power) on the target sound signal, or convert it to the frequency domain and calculate the power spectrum to obtain the second sound power (Pplay), which has the same unit as the first sound power to ensure comparability.

[0066] The formula for calculating the real-time signal-to-noise ratio is: SNRinst = Pplay - Pout.

[0067] Step 104: Adjust the sound gain of the device based on the reference threshold and the real-time signal-to-noise ratio.

[0068] Optionally, the reference thresholds may also include an upper gain threshold, a low gain threshold, and a high gain threshold.

[0069] Optionally, if the real-time signal-to-noise ratio remains below the low gain threshold for a first preset time period, the sound gain of the device can be increased in preset steps, provided that the sound gain of the device does not exceed the upper limit of the gain value.

[0070] As an example, the preset step size can be 1.5dB.

[0071] The first preset time (hold_time) can be 0.5 seconds to avoid overreacting to short-term noise fluctuations.

[0072] The low gain threshold can be -10dB. Below this threshold, it means "cannot hear clearly" and triggers an increase in gain. A negative value means that the target signal needs to be stronger than the noise in order to hear clearly.

[0073] In this embodiment, a dual-threshold hysteresis comparison mechanism can be adopted. When the real-time signal-to-noise ratio is lower than Tlow (low gain threshold) and continues to exceed hold_time, the gain is gradually increased (in 1.5dB steps) until Gmax (upper gain limit) is reached.

[0074] Optionally, if the real-time signal-to-noise ratio remains above the high gain threshold for a second preset time period, the sound gain of the device can be reduced by a preset step size, wherein the first preset time period is shorter than the second preset time period.

[0075] The second preset time, release_time, can be 2 seconds to ensure that the environment is truly quiet before reducing the gain and avoiding frequent volume fluctuations.

[0076] The high-gain threshold can be 5dB; values ​​above this threshold indicate a "quiet environment," resulting in a reduced trigger gain. A positive value indicates that the target signal is significantly stronger than the noise.

[0077] The gain limit can be set to 30dB, which represents the maximum gain limit to prevent excessive volume in extremely quiet environments.

[0078] Specifically, if a high signal-to-noise ratio (SNR) is maintained (in a quiet environment), when the real-time SNR exceeds Thigh (high gain threshold) and continues to exceed release_time, the gain is gradually reduced (in 1.5 dB steps) until it returns to its initial state. The purpose of this design is to avoid overreacting to transient noise or quiet environments, providing a more natural listening experience.

[0079] For example, assuming the input real-time signal-to-noise ratio (SNR) sequence is: [-15, -12, -18, -20, -10, 0, 8, 10, 12, 5, 0, -5], running the above algorithm, the gain adjustment process is as follows: For the first 4 frames, the real-time SNR is less than the low gain threshold (-10dB), and the low SNR counter increments to 4 (not triggered because it is less than the first preset time (approximately 15 frames)). In the 5th frame, the real-time SNR is -10dB, the counter is reset, and the system remains normal. From the 6th to the 10th frames, the real-time SNR is greater than the high gain threshold (5dB), and the high SNR counter increments to 5 (not triggered because it is less than the second preset time (approximately 60 frames)). In subsequent frames, the real-time SNR drops, and the system continues to maintain normal operation with no change in gain. If the triggering condition is continuously met, the gain will be smoothly adjusted in 1.5dB steps until the target value is reached.

[0080] In this embodiment, the system first acquires ambient sound signals collected by an external microphone and target sound signals collected by an in-ear microphone. Then, based on the ambient sound signals, a reference threshold corresponding to the scene in which the device is located is obtained. Next, based on the ambient sound signals and the target sound signals, the real-time signal-to-noise ratio (SNR) is determined. Finally, based on the reference threshold and the real-time SNR, the device's sound gain is adjusted. By acquiring signals from both the external and in-ear microphones, combined with the collaborative processing of scene reference thresholds and real-time SNR, multi-dimensional auditory experience optimization can be achieved. The dual-microphone array can accurately distinguish between ambient noise and user speech. The external microphone captures ambient sound field characteristics to match the reference threshold for the corresponding scene, while the in-ear microphone focuses on near-field sounds from the user, improving the accuracy of target signal pickup. Dynamically adjusting the gain based on the real-time SNR can ensure speech clarity while avoiding noise amplification. For example, when ambient noise suddenly increases, the system automatically enhances noise reduction and appropriately increases speech gain based on the scene threshold. Furthermore, the near-field pickup characteristics of the in-ear microphone can reduce echo and wind noise interference. This solution utilizes scene classification to achieve adaptive adjustment in different environments, and improves the accuracy of signal-to-noise ratio calculation through dual-microphone signal processing. Ultimately, it brings users a clearer, more comfortable, and power-optimized auditory experience, making it particularly suitable for scenarios such as smart headphones and hearing aids that require real-time audio enhancement.

[0081] Figure 2 This is a schematic flowchart of the sound gain control method provided according to the second embodiment of this disclosure.

[0082] like Figure 2 As shown, the method includes:

[0083] Step 201: Acquire the ambient sound signal collected by the external microphone and the target sound signal collected by the in-ear microphone.

[0084] Step 202: Based on the ambient sound signal, obtain the reference threshold corresponding to the scene where the device is located.

[0085] It should be noted that the specific implementation methods of steps 201 and 202 can refer to the above embodiments, and will not be repeated here.

[0086] Step 203: Detect the user's motion status.

[0087] Optionally, triaxial data can be acquired from accelerometers and gyroscopes to calculate motion characteristics (such as acceleration amplitude, angular velocity, and vibration frequency). For example, motion can be categorized into six types (such as lying down, sitting, walking slowly, walking briskly, running, and jumping). Based on the current motion state, three key thresholds—Tlow, Thigh, and Δimpulse—are adjusted in real time. The adjusted thresholds are then used for more precise audio enhancement and noise suppression.

[0088] Step 204: When the user's motion state is in a moving state, update the reference threshold to the first threshold.

[0089] The first threshold can be the updated upper gain limit Thigh, the low gain threshold Tlow, and the high gain threshold Δimpulse, and the first threshold is higher than the reference threshold.

[0090] It's important to note that if the user is moving, such as walking or running, transient interference such as footsteps and wind noise increases. If the reference threshold Tlow is too low, the system may misinterpret this noise as "unintelligible" and increase the volume unnecessarily. Therefore, during significant movement, Tlow should be increased (e.g., from -10dB to -8dB or even higher) to update to the first threshold and reduce false triggers. Environmental noise (such as wind and footsteps) fluctuates frequently. If the reference threshold Thigh is too low, the system will immediately lower the volume due to a brief decrease in noise, resulting in frequent volume jumps. Therefore, during movement, Thigh should be increased (e.g., from 5dB to 7dB) to update to the first threshold, ensuring the environment is truly quiet before lowering the volume, thus improving auditory comfort. Wind noise or vibration generated by head movements during running or rapid head nodding may be misinterpreted as "pulse signals." If the reference threshold Δimpulse is too low, the system may mistakenly interpret it as "too many pulses" and block normal speech frames. Therefore, during large movements, Δimpulse should be increased (e.g., from 3.0 to 3.5) to reduce false shielding caused by wind noise interference.

[0091] Step 205: When the user's movement state is a resting state, update the reference threshold to the second threshold.

[0092] The second threshold can be the updated upper gain limit Thigh, the low gain threshold Tlow, and the high gain threshold Δimpulse, and the second threshold is lower than the reference threshold.

[0093] It's important to note that if the user is in a lying-down position, indicating minimal physical movement and stable ambient noise, the Tlow threshold should be lowered (e.g., from -10dB to -12dB). This way, even if the SNR is slightly low, the system will recognize it as potentially inaudible and promptly increase the volume to ensure speech clarity in quiet environments. In a lying-down position, ambient noise fluctuations are small, so the Thigh threshold should be lowered (e.g., from 5dB to 3dB) to allow the system to reduce the volume more quickly in quiet environments, preventing excessive volume. In a lying-down position, there is little head or body movement, resulting in less wind noise and vibration interference. Therefore, the Δimpulse threshold should be lowered (e.g., from 3.0 to 2.5) to make the system more sensitive to speech impulses and avoid missing normal speech frames.

[0094] It's important to note that motion state recognition avoids misinterpreting transient noises such as footsteps and wind as "inaudible," reducing volume fluctuations. It more sensitively captures faint speech when the user is lying still and adjusts the volume more conservatively during movement, aligning with the auditory needs of the human ear in different states. When transitioning from movement to rest, the threshold doesn't immediately revert but adjusts gradually, avoiding audio fluctuations caused by sudden threshold changes.

[0095] For example, when a user is running, the IMU detects high-frequency acceleration fluctuations. The system automatically increases Tlow from -10dB to -6dB, Thigh from 5dB to 8dB, and Δimpulse from 3.0 to 3.5. Even if wind noise causes a brief drop in SNR, the system will not immediately increase the volume. When the user stops to rest, the IMU detects reduced movement, and the thresholds gradually adjust back to restore sensitivity to faint speech. When a user is lying down listening to music, Tlow is reduced to -12dB, Thigh to 3dB, and Δimpulse to 2.5. If ambient noise (such as air conditioning) causes the SNR to drop to -11dB, the system will consider it "possibly inaudible" and moderately increase the volume. However, when the music stops and the environment is truly quiet, after the SNR exceeds 3dB, the system will gradually reduce the volume after a longer release time (e.g., 2 seconds) to avoid the discomfort of sudden silence.

[0096] Step 206: Determine the real-time signal-to-noise ratio based on the ambient sound signal and the target sound signal.

[0097] It should be noted that the specific implementation of step 206 can be referred to the above embodiments, and will not be repeated here.

[0098] Step 207: Obtain the hearing threshold curve uploaded by the user.

[0099] The following methods are used to obtain the hearing threshold curve uploaded by the user:

[0100] Method 1: The headphones use their own speakers and follow the calibration process, or perform a pure tone test according to the prompts on the mobile phone screen. Users need to complete the test in a quiet environment. At each typical frequency point (such as 250Hz, 500Hz, 1kHz, 2kHz, 4kHz, 6kHz, 8kHz, etc.), the system will automatically record the threshold and generate a hearing threshold curve.

[0101] Method 2: Users can directly import existing air conduction hearing threshold curve data into the App. This data may come from professional hearing testing institutions or other channels. As long as the data format meets the App's requirements, it can be successfully imported.

[0102] It should be noted that the hearing threshold curve, also known as the air conduction hearing threshold curve, is a user's personal hearing data. It shows the minimum sound pressure level HL(f) that the user's ear can just barely hear at several typical frequencies. The higher the value at each frequency, the more severe the hearing loss in that frequency band, and the higher the compensation gain required.

[0103] Step 208: Based on the air conduction hearing values ​​corresponding to each frequency point in the hearing threshold curve, generate a band-limited gain adjustment table. The air conduction hearing values ​​are used to represent the minimum sound pressure level that the user can perceive.

[0104] Specifically, if the air conduction hearing value at any frequency point is less than a first preset threshold, then the gain compensation value at that frequency point can be determined to be zero.

[0105] The first preset threshold can be 20dB.

[0106] Specifically, the compensation strategy can be divided into three ranges based on the degree of hearing loss (HL value). If it is mild loss (HL<20dB), the compensation ratio is 1:1 (no compensation is required), and the gain compensation formula is G(f)=0, which means that the hearing is close to normal and no additional gain is required.

[0107] Optionally, if the air conduction hearing value at any frequency point is greater than the second preset threshold, then based on the first preset adjustment rule, the gain compensation value at any frequency point can be determined, and the second preset threshold is greater than the first preset threshold.

[0108] The first preset threshold can be 50dB.

[0109] If the air conduction hearing value corresponding to any frequency point is greater than the second preset threshold, it is considered as severe hearing loss (HL>50dB). The compensation ratio is 1:3 (strong compensation but limited amplitude), and the gain compensation formula is: G(f)=15+(HL(f)-50) / 3. First, a fixed compensation of 15dB is applied (maximum compensation for moderate hearing loss), and then 1 / 3 is applied to the portion exceeding 50dB to prevent excessive gain from causing discomfort.

[0110] For example, the air conduction hearing value corresponding to any frequency point can be subtracted from the second preset threshold to obtain a first difference. Then, the first ratio of the first difference to the first preset ratio can be determined. After that, the sum of the first ratio and the target preset value can be used as the gain compensation value corresponding to any frequency point.

[0111] The first preset ratio can be 3. The target preset value can be 15dB.

[0112] Optionally, if the air conduction hearing value at any frequency point is not less than the first preset threshold and not greater than the second preset threshold, then the gain compensation value at any frequency point can be determined based on the second preset adjustment rule.

[0113] Specifically, if the air conduction hearing value corresponding to any frequency point is not less than the first preset threshold and not greater than the second preset threshold, it indicates that the frequency point is in the range: 20dB≤HL≤50dB, the compensation ratio is 1:2, and the gain compensation formula is G(f)=(HL(f)-20) / 2. HL(f) is the air conduction hearing value corresponding to any frequency point f, compensating for half of the hearing loss to avoid excessive amplification leading to dynamic range compression.

[0114] For example, the air conduction hearing value corresponding to any frequency point can be subtracted from the first preset threshold to obtain a second difference. The second difference and the second ratio of the second preset ratio can be used as the gain compensation value corresponding to any frequency point, where the second preset ratio is less than the first preset ratio. The second preset ratio can be 2.

[0115] Then, a band-limited gain adjustment table can be generated based on the gain compensation value corresponding to each frequency point.

[0116] Specifically, the user's HL value can be substituted into the segmented compression formula to calculate the compensation gain G(f) for each frequency point, and a continuous band-limited gain curve (such as linear interpolation or spline interpolation) can be generated through an interpolation algorithm.

[0117] Step 209: Adjust the sound gain of the device based on the band-limited gain adjustment table, the reference threshold, and the real-time signal-to-noise ratio.

[0118] The initial gain can be determined first based on the reference threshold and the real-time signal-to-noise ratio. Then, the initial gain can be compensated by combining the band-limited gain adjustment table, thereby adjusting the sound gain of the device.

[0119] Specifically, the audio signal is first converted to the frequency domain by Fourier transform, and the total gain after superimposing G(f) and G_auto is applied to each frequency band. Then, the inverse Fourier transform is performed to convert it back to the time domain, and the compensated audio is output.

[0120] For example, assuming the user's right ear has an HL of 55dB at 4kHz and the current automatic gain G_auto = 6dB, and the range is defined as 55dB > 50dB → severe hearing loss, the hearing compensation gain is calculated as: G(4k) = 15 + (55-50) / 3 ≈ 15 + 1.67 ≈ 16.7dB

[0121] Total gain: G_total = G_auto + G(4k) = 6dB + 16.7dB ≈ 22.7dB. If the predicted output sound pressure level exceeds 95dB, the gain will be cut to a safe range using soft limiting.

[0122] Optionally, the current battery level of the device can be obtained first. If the current battery level is less than a preset battery threshold, the scene classification model can be stopped. Based on the Infinite Impulse Response (IIR) filter path, the third sound power corresponding to the ambient sound signal and the fourth sound power corresponding to the target sound signal can be identified. The difference between the fourth sound power and the third sound power can be used as the real-time signal-to-noise ratio. Based on the real-time signal-to-noise ratio, the sound gain of the device can be adjusted.

[0123] The preset battery threshold can be 20%, or it can be 15%, 10%, or other values, which are not limited here.

[0124] Optionally, an IIR bandpass filter can be used to extract the characteristic frequency band of ambient noise (e.g., 200Hz-5kHz), and the root mean square (RMS) calculation can be performed on the filtered signal to obtain the ambient sound power (third sound power). An IIR bandpass filter can then be designed to enhance the characteristic frequency band of the target signal (e.g., speech) (e.g., 800Hz-3kHz). The target sound power (fourth sound power) can also be obtained through RMS calculation.

[0125] It's worth noting that when the device's battery level drops below a preset threshold, the system automatically disables the high-energy-consuming scene classification model. Instead, it quickly calculates the power difference between the environment and the target sound using an IIR filter to obtain the real-time signal-to-noise ratio, and then dynamically adjusts the sound gain. This mechanism significantly reduces audio processing power consumption and extends device battery life while still maintaining adaptive volume adjustment based on the fundamental signal-to-noise ratio logic. This ensures that users can still enjoy a basically clear listening experience when battery power is low, achieving a balance between energy efficiency and the availability of core audio functions.

[0126] Step 210: Obtain the parameters to be uploaded, which include at least one of the following: sound gain adjustment information, acoustic features, and keyword trigger count.

[0127] It should be noted that the terminal can collect anonymous statistical data such as sound gain adjustment information, acoustic features (e.g., MFCC mean / variance), and keyword trigger counts as parameters to be uploaded, without uploading the original audio to protect privacy. The sound gain adjustment information records parameters during the adaptive gain process (e.g., dynamic changes in SNR and Tlow / Thigh) to optimize threshold strategies. The acoustic features extract the statistical characteristics (mean and variance) of MFCC (Mel-frequency cepstral coefficients) to reflect the environmental acoustic properties.

[0128] Step 211: Send the parameters to be uploaded to the cloud so that the cloud can iteratively update the model parameters and reference thresholds of the scene classification model. The scene classification model is used to identify the scene in which the device is located.

[0129] Specifically, the cloud iteratively optimizes the parameters of the scene classification model (such as a lightweight CNN) based on weakly labeled data uploaded from a large number of terminals, and updates the scene threshold table. The cloud generates encrypted differential packets (AES-256 encryption) and sends them to the terminals.

[0130] Specifically, the cloud can utilize implicit feedback reported by the terminal (such as gain adjustment behavior) as weak labels, combined with a small amount of manually labeled data to train the model. Only the changed parts of the model (differential packets) are updated, reducing transmission volume and terminal storage pressure. Based on global data analysis, threshold parameters such as Tlow and Thigh are dynamically adjusted to improve adaptability in different scenarios.

[0131] Optionally, the differential packets can be protected using the AES-256 symmetric encryption algorithm, with the key dynamically generated and periodically changed. Before updating, the MD5 / SHA-256 hash value is verified; if verification fails, the system rolls back to the backup model. The terminal only downloads the changed portion of the model, significantly reducing transmission traffic (typically only 5%-10% of the original model size).

[0132] Step 212: Receive the update data packet returned by the cloud, and update the model parameters and reference thresholds based on the update data packet.

[0133] Specifically, the terminal can verify the validity of the update package and use a dual-model backup mechanism to ensure automatic rollback in case of update failure.

[0134] In summary, by acquiring dual signals through external and in-ear microphones, a scene reference threshold is first matched based on ambient sound. Then, the threshold is dynamically updated based on the user's movement state (moving / lying still). Simultaneously, the real-time signal-to-noise ratio is calculated using the dual signals, and a band-limited gain adjustment table is generated by incorporating the user's uploaded hearing threshold curve. Finally, the sound gain is precisely adjusted based on the adjustment table, thresholds, and signal-to-noise ratio. Furthermore, parameters can be collected and uploaded to the cloud for model iteration, achieving edge-cloud collaborative optimization. The beneficial effects are that multi-dimensional perception (environment, movement, hearing) makes gain adjustment more closely aligned with scene and individual needs; dynamic thresholds adapt to different movement states; personalized curves compensate for hearing differences; and edge-cloud collaborative continuous model evolution improves the accuracy, comfort, and scene adaptability of the audio experience. It also allows device functionality to be continuously optimized with use through a data closed loop.

[0135] To facilitate better implementation of the sound gain control method of this disclosure, this disclosure also provides a sound gain control device based on the above-described sound gain control method. The meanings of the terms used are the same as in the sound gain control method described above, and specific implementation details can be found in the descriptions of the method embodiments.

[0136] Please see Figure 3 , Figure 3 This is a schematic diagram of the structure of a sound gain control device 300 provided in an embodiment of this disclosure. The sound gain control device 300 includes:

[0137] The first acquisition module 310 is used to acquire ambient sound signals collected by an external microphone and target sound signals collected by an in-ear microphone.

[0138] The second acquisition module 320 is used to acquire a reference threshold corresponding to the scene where the device is located based on the ambient sound signal;

[0139] The first determining module 330 is used to determine the real-time signal-to-noise ratio based on the ambient sound signal and the target sound signal;

[0140] The adjustment module 340 is used to adjust the sound gain of the device based on the reference threshold and the real-time signal-to-noise ratio.

[0141] Optionally, the second acquisition module 320 is specifically used for:

[0142] Feature extraction is performed on the environmental sound signal to obtain multidimensional Mel-frequency cepstral coefficient features;

[0143] A pre-trained scene classification model is used to classify the Mel frequency cepstral coefficient features to obtain the scene in which the device is located.

[0144] Based on the preset mapping relationship between scenarios and scenario threshold combinations, the scenario threshold combination corresponding to the scenario where the device is located is used as the reference threshold.

[0145] Optionally, the first determining module 330 is specifically used for:

[0146] The ambient sound signal is processed into frames based on a preset frame length and frame shift.

[0147] The environmental sound signal after frame segmentation is subjected to a first weighted filtering process to obtain the first sound power corresponding to each frame of audio signal. The first weighting is a frequency weighting that simulates the hearing characteristics of the human ear.

[0148] The power of the target sound signal is calculated to obtain the second sound power;

[0149] The difference between the second sound power and the first sound power is used as the real-time signal-to-noise ratio.

[0150] Optionally, the first determining module 330 is specifically used for:

[0151] Based on the preset fast trajectory time constant and slow trajectory time constant, dual trajectory energy tracking processing is performed on the first sound power corresponding to each frame of audio signal to obtain the sound power difference corresponding to each frame of audio signal.

[0152] Determine the impulse noise threshold included in the reference threshold;

[0153] If the difference in sound power corresponding to the audio signal is greater than or equal to the impulse noise threshold, the audio signal is determined to be impulse noise; otherwise, the audio signal is not impulse noise.

[0154] If the audio signal is impulse noise, the audio signal is discarded.

[0155] Optionally, the reference threshold includes an upper gain threshold, a low gain threshold, and a high gain threshold. The adjustment module 340 is specifically used for:

[0156] If the real-time signal-to-noise ratio remains below the low gain threshold for a first preset time, the sound gain of the device is increased according to a preset step size, and the sound gain of the device does not exceed the upper limit of the gain value.

[0157] If the real-time signal-to-noise ratio remains higher than the high gain threshold for a second preset time, the sound gain of the device is reduced according to the preset step size, wherein the first preset time is less than the second preset time.

[0158] Optionally, the second acquisition module 320 is also used for:

[0159] Detect the user's movement status;

[0160] When the user's motion state is in a moving state, the reference threshold is updated to a first threshold, where the first threshold is higher than the reference threshold;

[0161] When the user's movement state is a resting state, the reference threshold is updated to a second threshold, which is lower than the reference threshold.

[0162] Optionally, adjustment module 340 is used for:

[0163] Obtain the hearing threshold curve uploaded by the user;

[0164] Based on the air conduction hearing value corresponding to each frequency point in the hearing threshold curve, a band-limited gain adjustment table is generated, wherein the air conduction hearing value is used to represent the minimum sound pressure level that the user can perceive.

[0165] The sound gain of the device is adjusted based on the band-limited gain adjustment table, the reference threshold, and the real-time signal-to-noise ratio.

[0166] Optionally, adjustment module 340 is used for:

[0167] If the air conduction hearing value corresponding to any frequency point is less than the first preset threshold, then the gain compensation value corresponding to any frequency point is determined to be zero.

[0168] If the air conduction hearing value corresponding to any frequency point is greater than the second preset threshold, then based on the first preset adjustment rule, the gain compensation value corresponding to any frequency point is determined, and the second preset threshold is greater than the first preset threshold.

[0169] If the air conduction hearing value corresponding to any frequency point is not less than the first preset threshold and not greater than the second preset threshold, then the gain compensation value corresponding to any frequency point is determined based on the second preset adjustment rule.

[0170] The band-limited gain adjustment table is generated based on the gain compensation value corresponding to each frequency point.

[0171] Optionally, adjustment module 340 is used for:

[0172] Subtract the air conduction hearing value corresponding to any frequency point from the second preset threshold to obtain the first difference;

[0173] Determine a first ratio between the first difference and the first preset ratio;

[0174] The sum of the first ratio and the target preset value is used as the gain compensation value corresponding to any frequency point.

[0175] Optionally, adjustment module 340 is used for:

[0176] Subtract the air conduction hearing value corresponding to any frequency point from the first preset threshold to obtain the second difference;

[0177] The second ratio of the second difference to the second preset ratio is used as the gain compensation value corresponding to any frequency point, wherein the second preset ratio is less than the first preset ratio.

[0178] Optionally, the device is also used for:

[0179] Obtain the current battery level of the device;

[0180] If the current battery level is less than a preset battery level threshold, the scenario classification model will be stopped.

[0181] Based on the infinite impulse response (IIR) filter path, the third sound power corresponding to the ambient sound signal and the fourth sound power corresponding to the target sound signal are identified.

[0182] The difference between the fourth sound power and the third sound power is used as the real-time signal-to-noise ratio;

[0183] The sound gain of the device is adjusted based on the real-time signal-to-noise ratio.

[0184] Optionally, the adjustment module is also used for:

[0185] Obtain the parameters to be uploaded, wherein the parameters to be uploaded include at least one of the following: sound gain adjustment information, acoustic features, and keyword trigger count;

[0186] The parameters to be uploaded are sent to the cloud so that the cloud can iteratively update the model parameters of the scene classification model and the reference threshold. The scene classification model is used to identify the scene in which the device is located.

[0187] The system receives the update data packet returned by the cloud and updates the model parameters and the reference threshold based on the update data packet.

[0188] In this embodiment, the system first acquires ambient sound signals collected by an external microphone and target sound signals collected by an in-ear microphone. Then, based on the ambient sound signals, a reference threshold corresponding to the scene in which the device is located is obtained. Next, based on the ambient sound signals and the target sound signals, the real-time signal-to-noise ratio (SNR) is determined. Finally, based on the reference threshold and the real-time SNR, the device's sound gain is adjusted. By acquiring signals from both the external and in-ear microphones, combined with the collaborative processing of scene reference thresholds and real-time SNR, multi-dimensional auditory experience optimization can be achieved. The dual-microphone array can accurately distinguish between ambient noise and user speech. The external microphone captures ambient sound field characteristics to match the reference threshold for the corresponding scene, while the in-ear microphone focuses on near-field sounds from the user, improving the accuracy of target signal pickup. Dynamically adjusting the gain based on the real-time SNR can ensure speech clarity while avoiding noise amplification. For example, when ambient noise suddenly increases, the system automatically enhances noise reduction and appropriately increases speech gain based on the scene threshold. Furthermore, the near-field pickup characteristics of the in-ear microphone can reduce echo and wind noise interference. This solution utilizes scene classification to achieve adaptive adjustment in different environments, and improves the accuracy of signal-to-noise ratio calculation through dual-microphone signal processing. Ultimately, it brings users a clearer, more comfortable, and power-optimized auditory experience, making it particularly suitable for scenarios such as smart headphones and hearing aids that require real-time audio enhancement.

[0189] In addition, this disclosure also provides an electronic device, such as Figure 4 As shown, it illustrates a schematic diagram of the structure of the electronic device involved in this disclosure, specifically:

[0190] The electronic device may include components such as a processor 401 with one or more processing cores, a memory 402 with one or more computer-readable storage media, a power supply 403, and an input unit 404. Those skilled in the art will understand that... Figure 4 The electronic device structure shown does not constitute a limitation on the electronic device and may include more or fewer components than shown, or combine certain components, or have different component arrangements. Wherein:

[0191] The processor 401 is the control center of the electronic device. It connects various parts of the electronic device via various interfaces and lines. By running or executing software programs and / or modules stored in the memory 402, and by calling data stored in the memory 402, it performs various functions and processes data, thereby providing overall monitoring of the electronic device. Optionally, the processor 401 may include one or more processing cores; preferably, the processor 401 may integrate an application processor and a modem processor, wherein the application processor mainly handles the operating system, user interface, and applications, and the modem processor mainly handles wireless communication. It is understood that the modem processor may not be integrated into the processor 401.

[0192] The memory 402 can be used to store software programs and modules. The processor 401 executes various functional applications and data processing by running the software programs and modules stored in the memory 402. The memory 402 may mainly include a program storage area and a data storage area. The program storage area may store the operating system, application programs required for at least one function (such as sound playback function, image playback function, etc.), etc.; the data storage area may store data created according to the use of the electronic device, etc. In addition, the memory 402 may include high-speed random access memory, and may also include non-volatile memory, such as at least one disk storage device, flash memory device, or other volatile solid-state storage device. Accordingly, the memory 402 may also include a memory controller to provide the processor 401 with access to the memory 402.

[0193] The electronic device also includes a power supply 403 that supplies power to the various components. Preferably, the power supply 403 can be logically connected to the processor 401 through a power management system, thereby enabling functions such as charging, discharging, and power consumption management through the power management system. The power supply 403 may also include one or more DC or AC power supplies, recharging systems, power equipment debugging circuits, power converters or inverters, power status indicators, and other arbitrary components.

[0194] The electronic device may also include an input unit 404, which can be used to receive input digital or character information, and generate keyboard, mouse, joystick, optical or trackball signal inputs related to user settings and function control.

[0195] Although not shown, the electronic device may also include a display unit, etc., which will not be described in detail here. Specifically, in this embodiment, the processor 401 in the electronic device loads the executable files corresponding to the processes of one or more application programs into the memory 402 according to the following instructions, and the processor 401 runs the application programs stored in the memory 402, thereby implementing the steps in any of the sound gain control methods provided in the embodiments of this disclosure.

[0196] For details on the implementation of each of the above operations, please refer to the previous examples, which will not be repeated here.

[0197] In this embodiment, the system first acquires ambient sound signals collected by an external microphone and target sound signals collected by an in-ear microphone. Then, based on the ambient sound signals, a reference threshold corresponding to the scene in which the device is located is obtained. Next, based on the ambient sound signals and the target sound signals, the real-time signal-to-noise ratio (SNR) is determined. Finally, based on the reference threshold and the real-time SNR, the device's sound gain is adjusted. By acquiring signals from both the external and in-ear microphones, combined with the collaborative processing of scene reference thresholds and real-time SNR, multi-dimensional auditory experience optimization can be achieved. The dual-microphone array can accurately distinguish between ambient noise and user speech. The external microphone captures ambient sound field characteristics to match the reference threshold for the corresponding scene, while the in-ear microphone focuses on near-field sounds from the user, improving the accuracy of target signal pickup. Dynamically adjusting the gain based on the real-time SNR can ensure speech clarity while avoiding noise amplification. For example, when ambient noise suddenly increases, the system automatically enhances noise reduction and appropriately increases speech gain based on the scene threshold. Furthermore, the near-field pickup characteristics of the in-ear microphone can reduce echo and wind noise interference. This solution utilizes scene classification to achieve adaptive adjustment in different environments, and improves the accuracy of signal-to-noise ratio calculation through dual-microphone signal processing. Ultimately, it brings users a clearer, more comfortable, and power-optimized auditory experience, making it particularly suitable for scenarios such as smart headphones and hearing aids that require real-time audio enhancement.

[0198] Those skilled in the art will understand that all or part of the steps in the various methods of the above embodiments can be performed by instructions, or by instructions controlling related hardware. These instructions can be stored in a computer-readable storage medium and loaded and executed by a processor.

[0199] To this end, the present disclosure provides a computer-readable storage medium storing a computer program that can be loaded by a processor to perform the steps of any of the sound gain control methods provided in the present disclosure.

[0200] For details on the implementation of each of the above operations, please refer to the previous examples, which will not be repeated here.

[0201] The computer-readable storage medium may include: read-only memory (ROM), random access memory (RAM), disk or optical disk, etc.

[0202] Since the instructions stored in the computer-readable storage medium can execute the steps of any of the sound gain control methods provided in this disclosure, the beneficial effects that any of the sound gain control methods provided in this disclosure can achieve can be realized, as detailed in the preceding embodiments, and will not be repeated here.

[0203] The above provides a detailed description of a sound gain control method, apparatus, electronic device, and computer-readable storage medium provided in this disclosure. Specific examples have been used to illustrate the principles and implementation methods of the present invention. The descriptions of the above embodiments are only for the purpose of helping to understand the method and core ideas of the present invention. At the same time, those skilled in the art will recognize that, based on the ideas of the present invention, there will be changes in the specific implementation methods and application scope. Therefore, the content of this specification should not be construed as a limitation of the present invention.

Claims

1. A sound gain control method, characterized in that, include: Acquire ambient sound signals collected via an external microphone, and target sound signals collected via an in-ear microphone; Based on the ambient sound signal, obtain a reference threshold corresponding to the scene where the device is located; The real-time signal-to-noise ratio is determined based on the ambient sound signal and the target sound signal; The sound gain of the device is adjusted based on the reference threshold and the real-time signal-to-noise ratio.

2. The method according to claim 1, characterized in that, The step of obtaining a reference threshold corresponding to the scene where the device is located based on the ambient sound signal includes: Feature extraction is performed on the environmental sound signal to obtain multidimensional Mel-frequency cepstral coefficient features; A pre-trained scene classification model is used to classify the Mel frequency cepstral coefficient features to obtain the scene in which the device is located. Based on the preset mapping relationship between scenarios and scenario threshold combinations, the scenario threshold combination corresponding to the scenario where the device is located is used as the reference threshold.

3. The method according to claim 1, characterized in that, Determining the real-time signal-to-noise ratio based on the ambient sound signal and the target sound signal includes: The ambient sound signal is processed into frames based on a preset frame length and frame shift. The environmental sound signal after frame segmentation is subjected to a first weighted filtering process to obtain the first sound power corresponding to each frame of audio signal. The first weighting is a frequency weighting that simulates the hearing characteristics of the human ear. The power of the target sound signal is calculated to obtain the second sound power; The difference between the second sound power and the first sound power is used as the real-time signal-to-noise ratio.

4. The method according to claim 3, characterized in that, After performing a first weighted filtering process on the frame-segmented ambient sound signal to obtain the first sound power corresponding to each frame of audio signal, the method further includes: Based on the preset fast trajectory time constant and slow trajectory time constant, dual trajectory energy tracking processing is performed on the first sound power corresponding to each frame of audio signal to obtain the sound power difference corresponding to each frame of audio signal. Determine the impulse noise threshold included in the reference threshold; If the difference in sound power corresponding to the audio signal is greater than or equal to the impulse noise threshold, the audio signal is determined to be impulse noise; otherwise, the audio signal is not impulse noise. If the audio signal is impulse noise, the audio signal is discarded.

5. The method according to claim 1, wherein The reference thresholds include an upper gain threshold, a low gain threshold, and a high gain threshold. Adjusting the sound gain of the device based on the reference thresholds and the real-time signal-to-noise ratio includes: If the real-time signal-to-noise ratio remains below the low gain threshold for a first preset time, the sound gain of the device is increased according to a preset step size, and the sound gain of the device does not exceed the upper limit of the gain value. If the real-time signal-to-noise ratio remains higher than the high gain threshold for a second preset time, the sound gain of the device is reduced according to the preset step size, wherein the first preset time is less than the second preset time.

6. The method according to claim 1, characterized in that, After obtaining the reference threshold corresponding to the scene where the device is located based on the ambient sound signal, the method further includes: Detect the user's movement status; When the user's motion state is in a moving state, the reference threshold is updated to a first threshold, where the first threshold is higher than the reference threshold; When the user's movement state is a resting state, the reference threshold is updated to a second threshold, which is lower than the reference threshold.

7. The method according to claim 1, characterized in that, The adjustment of the sound gain of the device based on the reference threshold and the real-time signal-to-noise ratio includes: Obtain the hearing threshold curve uploaded by the user; Based on the air conduction hearing value corresponding to each frequency point in the hearing threshold curve, a band-limited gain adjustment table is generated, wherein the air conduction hearing value is used to represent the minimum sound pressure level that the user can perceive. The sound gain of the device is adjusted based on the band-limited gain adjustment table, the reference threshold, and the real-time signal-to-noise ratio.

8. The method according to claim 7, characterized in that, The step of generating a band-limited gain adjustment table based on the air conduction hearing values ​​corresponding to each frequency point in the hearing threshold curve includes: If the air conduction hearing value corresponding to any frequency point is less than the first preset threshold, then the gain compensation value corresponding to any frequency point is determined to be zero. If the air conduction hearing value corresponding to any frequency point is greater than the second preset threshold, then based on the first preset adjustment rule, the gain compensation value corresponding to any frequency point is determined, and the second preset threshold is greater than the first preset threshold. If the air conduction hearing value at any frequency point is not less than the first preset threshold and not greater than the second preset threshold, then the gain compensation value at any frequency point is determined based on the second preset adjustment rule. The band-limited gain adjustment table is generated based on the gain compensation value corresponding to each frequency point.

9. The method according to claim 8, characterized in that, The step of determining the gain compensation value corresponding to any frequency point based on the first preset adjustment rule includes: Subtract the air conduction hearing value corresponding to any frequency point from the second preset threshold to obtain the first difference; Determine a first ratio between the first difference and the first preset ratio; The sum of the first ratio and the target preset value is used as the gain compensation value corresponding to any frequency point.

10. The method according to claim 9, characterized in that, The step of determining the gain compensation value corresponding to any frequency point based on the second preset adjustment rule includes: Subtract the air conduction hearing value corresponding to any frequency point from the first preset threshold to obtain the second difference; The second ratio of the second difference to the second preset ratio is used as the gain compensation value corresponding to any frequency point, wherein the second preset ratio is less than the first preset ratio.

11. The method according to claim 1, characterized in that, Also includes: Obtain the current battery level of the device; If the current battery level is less than a preset battery level threshold, the scenario classification model will be stopped. Based on the infinite impulse response (IIR) filter path, the third sound power corresponding to the ambient sound signal and the fourth sound power corresponding to the target sound signal are identified. The difference between the fourth sound power and the third sound power is used as the real-time signal-to-noise ratio; The sound gain of the device is adjusted based on the real-time signal-to-noise ratio.

12. The method according to claim 1, characterized in that, After adjusting the sound gain of the device based on the reference threshold and the real-time signal-to-noise ratio, the method further includes: Obtain the parameters to be uploaded, wherein the parameters to be uploaded include at least one of the following: sound gain adjustment information, acoustic features, and keyword trigger count; The parameters to be uploaded are sent to the cloud so that the cloud can iteratively update the model parameters of the scene classification model and the reference threshold. The scene classification model is used to identify the scene in which the device is located. The system receives the update data packet returned by the cloud and updates the model parameters and the reference threshold based on the update data packet.

13. A sound gain control device, characterized in that, include: The first acquisition module is used to acquire ambient sound signals collected by an external microphone and target sound signals collected by an in-ear microphone. The second acquisition module is used to acquire a reference threshold corresponding to the scene where the device is located based on the ambient sound signal; The first determining module is used to determine the real-time signal-to-noise ratio based on the ambient sound signal and the target sound signal; An adjustment module is used to adjust the sound gain of the device based on the reference threshold and the real-time signal-to-noise ratio.

14. An electronic device, characterized in that, The method includes a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor, when executing the computer program, implements the steps of the method as described in any one of claims 1-12.

15. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program that, when executed by a processor, implements the steps of the method as described in any one of claims 1-12.