Volume adjusting method and device and computer equipment
By analyzing ambient audio data, the system automatically adjusts the volume, solving the problems of inconvenience, lag, and inaccuracy in manually adjusting the volume. This achieves precise volume adaptation and improves the user experience.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- SHENZHEN CHOUMEI CULTURAL BROADCASTING CO LTD
- Filing Date
- 2026-01-30
- Publication Date
- 2026-05-01
AI Technical Summary
Users need to manually adjust the volume to adapt to different environments, which leads to inconvenience, lag, and inaccuracy in adjustment, affecting the user experience, especially in mobile scenarios where the current activity needs to be interrupted to adjust the volume.
By analyzing the ambient audio data of the device, noise characteristics are extracted, the noise index is calculated, the volume adjustment strategy is determined, and the device volume is automatically adjusted. This includes identifying the noise level, volume adjustment coefficient, and baseline noise index, and making precise adjustments based on acoustic scene and user preference data.
It enables automatic and precise adjustment of device volume, improves user experience, avoids the inconvenience and lag of manual adjustment, and ensures that the volume is within a safe and comfortable range to adapt to environmental changes.
Smart Images

Figure CN121967595A_ABST
Abstract
Description
Volume adjustment methods, devices, and computer equipment Technical Field
[0001] This specification relates to the field of computer technology, and in particular to a volume adjustment method, apparatus, and computer device. Background Technology
[0002] Smartphones and tablets have become an indispensable part of users' work and life.
[0003] Taking smartphones as an example, smartphones not only have voice calls but also many entertainment functions, such as watching videos and listening to music. In related technologies, when users watch videos or listen to music on their smartphones, they need to manually adjust the volume according to the ambient noise level as the external sound changes, which affects the user experience. Summary of the Invention
[0004] This specification provides a volume adjustment method, apparatus, and computer device to enable the device to automatically adjust the volume according to the ambient noise level, thereby improving the user experience.
[0005] This specification provides a volume adjustment method, comprising: extracting a first ambient noise feature based on first ambient audio data of the device; calculating a first ambient noise index based on the first ambient noise feature; determining a first volume adjustment strategy based on a first noise level to which the first ambient noise index belongs; and adjusting the device volume according to the first volume adjustment strategy.
[0006] In some embodiments, the first environmental noise feature includes time-domain features and frequency-domain features. A basic noise index can be calculated based on the time-domain features; noise complexity and energy of a specified frequency band can be calculated based on the frequency-domain features; and the basic noise index, noise complexity, and energy of the specified frequency band can be weighted and summed to obtain the first environmental noise index.
[0007] In some embodiments, the time-domain features include the root mean square of the sampled values, and the frequency-domain features include the power spectral density; the fundamental noise index can be calculated based on the root mean square; the energy, spectral bandwidth, and spectral centroid of a specified frequency band can be calculated based on the power spectral density; and the noise complexity can be calculated based on the spectral bandwidth and the spectral centroid.
[0008] In some embodiments, a first noise level corresponding to the first environmental noise index can be identified in a noise level set; the noise level set includes multiple noise levels, each noise level corresponding to a matching condition; the first environmental index satisfies the matching condition corresponding to the first noise level.
[0009] In some embodiments, the first volume adjustment strategy includes a volume adjustment coefficient corresponding to the first noise level and a reference noise index. A compensation volume can be calculated based on the first ambient noise index, the volume adjustment coefficient, and the reference noise index; a target volume can be calculated based on the compensation volume; and the current volume of the device can be adjusted to the target volume.
[0010] In some embodiments, the adjustment step size can be determined based on the compensation volume and the preset adjustment time; the current volume of the device can be adjusted to the target volume according to the adjustment step size.
[0011] In some embodiments, a volume constraint condition matching the first noise level can be obtained; if the target volume does not meet the volume constraint condition, the target volume can be updated according to the volume constraint condition; the current volume of the device can be adjusted to the updated target volume.
[0012] In some embodiments, an acoustic scene corresponding to the first ambient noise feature is identified in an acoustic scene set; the acoustic scene set includes multiple acoustic scenes, each acoustic scene has a matching condition, and the first ambient noise feature satisfies the matching condition corresponding to the acoustic scene; the device sound effect can be adjusted according to the acoustic scene.
[0013] In some embodiments, a second ambient noise feature can be extracted based on the second ambient audio data of the device; a second ambient noise index can be calculated based on the second ambient noise feature; the second ambient noise index and the first ambient noise index can be compared to determine the adjustment effect; if the adjustment effect does not meet the preset conditions and the second noise level to which the second ambient noise index belongs is the same as the first noise level to which the first ambient noise index belongs, the device volume can be fine-tuned; if the adjustment effect does not meet the preset conditions and the second noise level to which the second ambient noise index belongs is different from the first noise level to which the first ambient noise index belongs, a second volume adjustment strategy can be determined based on the second noise level, and the device volume can be adjusted based on the second volume adjustment strategy.
[0014] In some embodiments, volume adjustment behavior data corresponding to multiple noise levels can be obtained; the volume adjustment behavior data corresponding to each noise level can be input into the model to obtain user preference data for that noise level; and the volume adjustment strategy corresponding to that noise level can be updated based on the user preference data for each noise level.
[0015] This specification also provides a volume adjustment device, comprising: an extraction unit for extracting a first ambient noise feature based on first ambient audio data of the device; a calculation unit for calculating a first ambient noise index based on the first ambient noise feature; a determination unit for determining a first volume adjustment strategy based on a first noise level to which the first ambient noise index belongs; and an adjustment unit for adjusting the device volume according to the first volume adjustment strategy.
[0016] This specification also provides a computer device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the volume adjustment method described above.
[0017] This specification also provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the volume adjustment method described above.
[0018] This specification also provides a computer program product, which includes a computer program that, when executed by a processor, implements the volume adjustment method described above.
[0019] The technical solution of this specification embodiment can extract a first ambient noise feature based on the first ambient audio data of the device; calculate a first ambient noise index based on the first ambient noise feature; determine a first volume adjustment strategy based on the first noise level to which the first ambient noise index belongs; and adjust the device volume based on the first volume adjustment strategy. Therefore, this specification embodiment can automatically adjust the volume by analyzing ambient noise, thereby improving the user experience. Attached Figure Description
[0020] To more clearly illustrate the technical solutions in the embodiments or prior art of this specification, the drawings used in the description of the embodiments or prior art will be briefly introduced below. The drawings described below are only some embodiments recorded in this specification. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0021] Figure 1 is a flowchart illustrating the volume adjustment method in an embodiment of this specification; Figure 2 is a structural diagram illustrating the volume adjustment device in an embodiment of this specification. Detailed Implementation
[0022] The technical solutions in the embodiments of this specification will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this specification, and not all embodiments. The specific embodiments described herein are only used to explain this disclosure, and not to limit this disclosure. All other embodiments obtained by those skilled in the art based on the described embodiments of this disclosure are within the scope of protection of this disclosure. In addition, relational terms such as "first" and "second" are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. In addition, the acquisition, transmission, storage, use, and processing of data in the technical solutions of this application comply with the relevant provisions of national laws and regulations. In addition, in the embodiments of this application, certain software, components, models, and other existing solutions in the industry may be mentioned. These should be considered as exemplary, and their purpose is only to illustrate the feasibility of implementing the technical solutions of this application, but it does not mean that the applicant has used or necessarily used such solutions.
[0023] The aforementioned technologies require users to manually adjust the volume, which presents the following problems: 1. Inconvenience of manual adjustment. Users need to frequently adjust the volume manually to adapt to different environments. For example, when moving from a quiet environment to a noisy one, the volume needs to be manually increased. Conversely, when moving from a noisy environment to a quiet one, the volume needs to be manually decreased. Frequent manual adjustments reduce the user experience, especially in mobile scenarios where users need to interrupt their current activity to adjust the volume. 2. Lag in adjustment. Users often only realize the need to adjust the volume after the ambient noise changes, resulting in a time lag that affects the listening experience. 3. Inaccurate adjustment. When manually adjusting the volume, users find it difficult to accurately judge the ambient noise level, leading to inaccurate volume adjustment. For example, problems may arise such as the volume being too low to hear clearly, or the volume being too high to disturb others or damage hearing.
[0024] Therefore, this specification provides a volume adjustment method to automatically adjust the device volume by analyzing ambient noise, thereby solving problems such as inconvenience, lag, and inaccuracy of manual adjustment in related technologies and improving user experience.
[0025] The volume adjustment method can be applied to a device. The device includes a portable terminal device. For example, the device includes, but is not limited to, smartphones, tablet computers, portable computers, and smart wearable devices. The device has a sound acquisition component. The sound acquisition component may, for example, include a microphone for acquiring audio data. The device may also have a speaker. The speaker is used to play sound.
[0026] Please refer to Figure 1. The volume adjustment method may include the following steps.
[0027] Step 11: Extract the first ambient noise feature based on the first ambient audio data of the device; Step 12: Calculate the first ambient noise index based on the first ambient noise feature; Step 13: Determine the first volume adjustment strategy based on the first noise level to which the first ambient noise index belongs; Step 14: Adjust the device volume according to the first volume adjustment strategy.
[0028] The technical solution of the embodiments in this specification can extract a first ambient noise feature based on the first ambient audio data of the device; calculate a first ambient noise index based on the first ambient noise feature; determine a first volume adjustment strategy based on the first noise level to which the first ambient noise index belongs; and adjust the device volume based on the first volume adjustment strategy. Therefore, the embodiments in this specification can automatically adjust the volume by analyzing ambient noise, thereby improving the user experience.
[0029] In some embodiments, the first ambient audio data is ambient audio data collected by the device for adjusting the device volume. The device can collect the first ambient audio data via a microphone. As an example, the first ambient audio data may include audio components of the environment in which the device is located, but exclude audio components of the sound emitted by the device itself. This eliminates interference from the device's own sound, thereby facilitating the accurate extraction of ambient noise characteristics and the accurate calculation of the ambient noise index. For example, when the device's speaker is muted (no sound is playing), the device can directly use the raw audio data collected by the microphone as the first ambient audio data. Alternatively, when the device's speaker is playing sound, the device can remove its own audio components from the raw audio data collected by the microphone to obtain the first ambient audio data. The device can employ acoustic echo cancellation technology (AEC, Acoustic Echo Cancel) to remove its own audio components from the raw audio data. As another example, the first ambient audio data may include audio components of the environment in which the device is located, or it may include audio components of the sound emitted by the device itself. This reduces computational load and improves the response speed of device volume adjustment.
[0030] In some embodiments, the device can collect first ambient audio data in real time. Alternatively, the device can collect first ambient audio data at set intervals. Alternatively, the device can collect first ambient audio data when it detects a change in its own geographical location. This balances the timeliness of volume adjustment with the consumption of computing resources.
[0031] The device can acquire first ambient audio data according to a preset sampling rate and a preset bit depth. The preset sampling rate can be, for example, 44.1 kHz, and the preset bit depth can be, for example, 16 bits. The length of the first ambient audio data can be a preset length. The preset length ensures both frequency resolution and real-time performance. For example, the first ambient audio data includes 1024 sampling points. At a sampling rate of 44.1 kHz, the length of the first ambient audio data can be 23.2 ms.
[0032] In some embodiments, the device can directly extract the first ambient noise features based on the first ambient audio data. Alternatively, the device can also perform denoising processing on the first ambient audio data; the first ambient noise features can then be extracted based on the denoised first ambient audio data. For example, the device can use a high-pass filter to remove low-frequency noise from the first ambient audio data, and a low-pass filter to remove high-frequency noise from the first ambient audio data. The cutoff frequency of the high-pass filter can be, for example, 80Hz, and the cutoff frequency of the low-pass filter can be, for example, 8000Hz.
[0033] In some embodiments, a first ambient noise feature is used to represent the noise characteristics of the environment in which the terminal is located. The first ambient noise feature may include time-domain features and frequency-domain features. The device can calculate a fundamental noise index based on the time-domain features; calculate noise complexity and energy of a specified frequency band based on the frequency-domain features; and calculate a first ambient noise index based on the fundamental noise index, noise complexity, and energy of the specified frequency band. The first ambient noise index is used to represent the intensity of ambient noise. By combining time-domain and frequency-domain features to calculate the first ambient noise index, the accuracy of ambient noise detection can be improved.
[0034] The time-domain characteristics include at least one of the following: the root mean square of the first ambient audio data, the peak value of the first ambient audio data, and the dynamic range value of the first ambient audio data. The device can then use the following formula: Calculate the root mean square (RMS). RMS represents the root mean square, N represents the number of sampling points for the first environmental audio data, and X... i This represents the value of the i-th sampling point. The root mean square (RMS) is used to represent the average energy level of the first ambient audio data. The device can be configured according to the formula Peak = max(|X1|, ..., |X...). i |、...、|X N |) Calculate the peak value of the first ambient audio data. Peak represents the maximum amplitude of the first ambient audio data. The device can calculate the dynamic range value based on the root mean square and peak value. For example, the device can calculate it using the formula... DR stands for Dynamic Range. Dynamic Range is used to represent the degree of dynamic change.
[0035] Frequency domain characteristics can include power spectral density. The device can obtain the power spectral density by performing a Fourier transform on the first ambient audio data. For example, the device can calculate the power spectral density using the formula PSD = |FFT(X)|². FFT stands for Fast Fourier Transform (FFT). PSD represents the power spectral density, and X represents the first ambient audio data. The first ambient audio data can include the values of N sample points, X... i This represents the value at the i-th sampling point, where 1 ≤ i ≤ N. The power spectral density (PSD) can include the energy corresponding to K frequency bin indices. j Let K represent the energy of the j-th frequency bin index, where 1 ≤ j ≤ K.
[0036] The device can calculate the energy of one or more key frequency bands based on the power spectral density. Each key frequency band can include a band formed by a start frequency value and a cutoff frequency value. For each key frequency band, the device can calculate the frequency bin index corresponding to the start frequency value in the power spectral density as the start index, based on the number of sampling points and the sampling rate; it can also calculate the frequency bin index corresponding to the cutoff frequency value in the power spectral density as the cutoff index, based on the number of sampling points and the sampling rate; the energy of each frequency bin index between the start and cutoff indices in the power spectral density can be accumulated to obtain the energy of the key frequency band. For example, the device can use a formula... Calculate the starting index. Indicates the starting index. This represents the initial frequency value, and N represents the number of sampling points. This indicates the sampling rate. The device uses the formula... Calculate the cutoff index. Indicates the cutoff index. This represents the cutoff frequency value, and N represents the number of sampling points. This indicates the sampling rate. Key frequency bands include low frequency (Low Freq), mid frequency (Mid Freq), and high frequency (High Freq). The start and cutoff frequencies for the low frequency band are 80Hz and 250Hz, respectively; for the mid frequency band, they are 250Hz and 2000Hz; and for the high frequency band, they are 2000Hz and 8000Hz. The low frequency band corresponds to the fundamental frequency range of human voice, the mid frequency band to the resonance range of human voice, and the high frequency band to the ambient noise range.
[0037] The device can also calculate the total energy, spectral centroid, and spectral bandwidth of the first ambient audio data based on the power spectral density. The device can sum the energies of each frequency band index in the power spectral density to obtain the total energy of the first ambient audio data. For example, the device can use a formula... Calculate the total energy of the first ambient audio data. The device can use the formula... Calculate the spectral centroid. SC represents the spectral centroid, and K represents the number of frequency bin indices in the power spectral density. This represents the total energy of the first ambient audio data. This represents the energy of the j-th frequency bin index. This represents the energy of the l-th frequency bin index. This represents the frequency value corresponding to the l-th frequency bin index. N represents the number of sampling points. This indicates the sampling rate. The spectral centroid is used to represent the location of concentrated spectral energy. The device can be determined using the formula... Calculate the spectral bandwidth. Spectral bandwidth represents the degree of dispersion of spectral energy. The device can determine the spectral rolloff based on the power spectral density. The spectral rolloff can be a frequency value. The device can calculate the index of a specified frequency bin corresponding to the spectral rolloff. Then, the energy below the specified frequency bin index in the power spectral density accounts for 85% of the total energy. For example, the device can use the formula... Calculate the index of the specified frequency bin corresponding to the spectral roll-off point. SR represents the spectral roll-off point, and SR represents the specified frequency bin index. The energy below the specified frequency bin index is obtained by accumulating the energy of each frequency bin index in the power spectral density that is less than the specified frequency bin index. .
[0038] The equipment can calculate the fundamental noise level (NL) based on the root mean square (RMS) formula. For example, the equipment can use the formula NL = 20 × log0 10(RMS) calculates the basic noise index. RMS represents the root mean square of the first ambient audio data. The device can calculate the noise complexity based on the spectral bandwidth and spectral centroid. For example, the device can calculate the noise complexity using the formula NC=SB / SC. NC represents the noise complexity, SB represents the spectral bandwidth, and SC represents the spectral centroid. The device can obtain the first ambient noise index by weighted summing of the basic noise index, noise complexity, and energy of a specified frequency band. The energy of the specified frequency band can include the energy of the high-frequency band. EN=α×NL+β×NC+γ×HFE. EN represents the first ambient noise index, NL represents the basic noise index, NC represents the noise complexity, HFE represents the energy of the high-frequency band, α represents the weighting coefficient of the basic noise index (e.g., 0.6), β represents the weighting coefficient of the noise complexity (e.g., 0.3), and γ represents the weighting coefficient of the high-frequency energy (e.g., 0.1).
[0039] In some embodiments, the device provides a noise level set. The noise level set includes one or more noise levels. Each noise level corresponds to a noisy environment. Each noise level corresponds to a matching condition, which may include a noise intensity range. Different noise levels have different noise intensity ranges. For example, the noise level set may include noise level A, noise level B, noise level C, and noise level D. Noise level A corresponds to a quiet environment, noise level B corresponds to a normal environment, noise level C corresponds to a noisy environment, and noise level D corresponds to an extremely noisy environment. The matching condition for noise level A may be: noise index < 40 dB. The matching condition for noise level B may be: 40 dB ≤ noise index < 60 dB. The matching condition for noise level C may be: 60 dB ≤ noise index < 80 dB. The matching condition for noise level D may be: noise index ≥ 80 dB. The device can identify a noise level in the noise level set that corresponds to a first environmental noise index as the first noise level. The first environmental index satisfies the matching condition corresponding to the first noise level.
[0040] In some embodiments, noise levels are grouped, and each noise level corresponds to a volume adjustment coefficient. The higher the noise intensity of a noise level, the larger the corresponding volume adjustment coefficient; the lower the noise intensity of a noise level, the smaller the corresponding volume adjustment coefficient. For example, the volume adjustment coefficient for noise level A is 0.8, for noise level B it is 1.0, for noise level C it is 1.2, and for noise level D it is 1.5.
[0041] The device can obtain the volume adjustment coefficient corresponding to a first noise level. The first volume adjustment strategy may include the volume adjustment coefficient corresponding to the first noise level. The device can calculate the compensation volume based on a first ambient noise index, the volume adjustment coefficient corresponding to the first noise level, and a reference noise index. For example, the device can subtract the reference noise index from the first ambient noise index to obtain the difference; it can then multiply the difference by the volume adjustment coefficient corresponding to the first noise level to obtain the compensation volume. The device can calculate the target volume based on the compensation volume; it can then adjust the device's current volume to the target volume. For example, the device can add the compensation volume to the current volume to obtain the target volume.
[0042] A baseline noise level can be used as a reference for calculating the compensated volume. Optionally, the baseline noise level can be the same for different noise levels. Alternatively, the baseline noise level can be different for different noise levels. In a noise level set, each noise level corresponds to a baseline noise level. The device can obtain the baseline noise level corresponding to a first noise level. A first volume adjustment strategy can include the baseline noise level corresponding to the first noise level. Therefore, the device can calculate the compensated volume based on the first ambient noise level, the volume adjustment coefficient corresponding to the first noise level, and the baseline noise level corresponding to the first noise level; it can calculate the target volume based on the compensated volume; and it can adjust the device's current volume to the target volume. This allows for differentiated and precise volume compensation for different noise levels, further improving the user experience.
[0043] In some embodiments, the device can determine the adjustment step size based on the compensation volume and a preset adjustment time; the current volume of the device can be adjusted to the target volume according to the adjustment step size. The preset adjustment time can be, for example, 2 seconds. The adjustment step size can be decibels (dB), or it can be a proportion (e.g., percentage) of the total volume of the device.
[0044] For example, the device can divide the compensated volume by a preset adjustment time to obtain the adjustment step size. For example, the device can perform a volume adjustment every certain time interval (e.g., 100 milliseconds), and the magnitude of the volume adjustment can be the adjustment step size.
[0045] This allows for a smooth transition in device volume, avoiding sudden volume changes and further enhancing the user experience.
[0046] In some embodiments, the device can acquire volume constraints; if the target volume does not meet the volume constraints, the target volume can be updated according to the volume constraints, and the device's current volume can be adjusted to the updated target volume; if the target volume meets the volume constraints, the target volume can be kept unchanged, and the device's current volume can be adjusted to the target volume. This ensures that the adjusted volume is within a safe and comfortable range, preventing the volume from being too loud or too soft, thereby satisfying the user's personalized needs while further enhancing the auditory experience and safety.
[0047] Volume constraints can include a lower volume limit and a higher volume limit. The device can determine whether the target volume is between the lower and upper volume limits; if so, the target volume satisfies the volume constraint; if the target volume is lower than the lower limit or higher than the upper volume limit, the target volume does not satisfy the volume constraint. If the target volume is lower than the lower limit, the device can use the lower limit as the updated target volume. If the target volume is higher than the upper volume limit, the device can use the upper volume limit as the updated target volume.
[0048] Optionally, the volume constraints corresponding to different noise levels can be the same. Alternatively, the volume constraints corresponding to different noise levels can also be different. In the noise level set, each noise level corresponds to a volume constraint. The device can obtain the volume constraint corresponding to the first noise level. The first volume adjustment strategy can include the volume constraint corresponding to the first noise level. Thus, the device can calculate the compensation volume based on the first ambient noise index, the volume adjustment coefficient corresponding to the first noise level, and the reference noise index; calculate the target volume based on the compensation volume; determine whether the target volume meets the volume constraint corresponding to the first noise level; if the target volume does not meet the volume constraint corresponding to the first noise level, the target volume can be updated based on the volume constraint corresponding to the first noise level, and the device's current volume can be adjusted to the updated target volume; if the target volume meets the volume constraint corresponding to the first noise level, the target volume can be kept unchanged, and the device's current volume can be adjusted to the target volume. In this way, differentiated safety and comfort ranges can be configured for different noise levels, further improving the user experience.
[0049] Optionally, in the noise level set, each noise level corresponds to a volume adjustment coefficient, a reference noise index, and a volume constraint condition. The device can obtain the volume adjustment coefficient corresponding to the first noise level, the reference noise index corresponding to the first noise level, and the volume constraint condition corresponding to the first noise level. The first volume adjustment strategy may include the volume adjustment coefficient corresponding to the first noise level, the reference noise index corresponding to the first noise level, and the volume constraint condition corresponding to the first noise level. Therefore, the device can calculate the compensation volume based on the first ambient noise index, the volume adjustment coefficient corresponding to the first noise level, and the reference noise index corresponding to the first noise level; calculate the target volume based on the compensation volume; determine whether the target volume meets the volume constraint condition corresponding to the first noise level; if the target volume does not meet the volume constraint condition corresponding to the first noise level, the target volume can be updated according to the volume constraint condition corresponding to the first noise level, and the device's current volume can be adjusted to the updated target volume; if the target volume meets the volume constraint condition corresponding to the first noise level, the target volume can be kept unchanged, and the device's current volume can be adjusted to the target volume.
[0050] In some embodiments, the device can identify the acoustic environment in which it is located based on a first ambient noise characteristic.
[0051] The device provides an acoustic scene set. The acoustic scene set may include one or more acoustic scenes. These acoustic scenes may include indoor office areas, city parks, and areas along main traffic arteries. Each acoustic scene corresponds to matching conditions. Matching conditions may include time-domain conditions and frequency-domain conditions. Time-domain conditions may include a dynamic range value interval. The dynamic range value interval may include a first value interval and a second value interval. The first value interval corresponds to steady-state noise. Steady-state noise has a smooth sound, such as drizzle or air conditioning. The second value interval corresponds to impulsive noise. Impulsive noise is sudden and abrupt, such as door slamming, collisions, glass breaking, explosions, or car horns. The length of the second value interval is greater than the length of the first value interval. Frequency-domain conditions may include a spectral roll-off point value interval. The spectral roll-off point value interval may include a third and a fourth value interval. The third value interval corresponds to high-frequency noise. High-frequency noise has energy concentrated at high frequencies, such as cicada chirping, alarms, car horns, or metallic friction sounds. The fourth value interval corresponds to low-frequency noise. Low-frequency noise energy is concentrated in the low frequencies, and can include sounds such as air conditioners and thunder.
[0052] The device can identify the acoustic scene corresponding to a first ambient noise feature within an acoustic scene set. The first ambient noise feature satisfies the matching conditions corresponding to the acoustic scene. The matching conditions can include time-domain conditions and frequency-domain conditions. The time-domain conditions can include a dynamic range value interval, which may include a first value interval and a second value interval. The frequency-domain conditions can include a spectral roll-off point value interval, which may include a third value interval and a fourth value interval. The first ambient noise feature can include a dynamic range value and a spectral roll-off point. The dynamic range value and the spectral roll-off point satisfy the time-domain conditions and the frequency-domain conditions, respectively. For example, the dynamic range value may be located within the dynamic range value interval in the time-domain conditions, and the spectral roll-off point may be located within the spectral roll-off point value interval in the frequency-domain conditions. In practical applications, the device can select the dynamic range value interval to which the dynamic range value belongs from the first and second value intervals; it can select the spectral roll-off point value interval to which the spectral roll-off point belongs from the third and fourth value intervals; and it can select the corresponding acoustic scene from the acoustic scene set based on the selected dynamic range value interval and the selected spectral roll-off point value interval.
[0053] Optionally, in the acoustic scene set, the matching conditions for each acoustic scene can also include environment type. Environment types include urban buildings, parks, sports fields, etc. The device can obtain the user's environment type based on its own geographic location data; it can identify the acoustic scene corresponding to both the first ambient noise feature and the user's environment type in the acoustic scene set. The first ambient noise feature and the user's environment type satisfy the matching conditions corresponding to the acoustic scene. For example, the dynamic range value and spectral roll-off point in the first ambient noise feature satisfy the time domain condition and the frequency domain condition, respectively, and the user's environment type is consistent with the environment type in the matching conditions. Thus, combining the environment type can improve the accuracy of acoustic scene recognition.
[0054] The device can adjust its sound effects based on acoustic scenarios. Within a set of acoustic scenarios, each scenario corresponds to a sound effect adjustment strategy. The device can then acquire this strategy and adjust its sound effects accordingly. Sound effect adjustment strategies can include boosting low and mid frequencies, and boosting high and mid frequencies. Boosting low and mid frequencies and boosting high and mid frequencies can be achieved by adjusting the equalizer. For example, the equalizer can be adjusted to boost low and mid frequencies as follows: appropriately attenuate within the 60Hz-250Hz range to reduce muddiness from noise; moderately boost within the 1kHz-4kHz range to enhance the clarity of vocals and main melodies; and slightly boost above 8kHz to increase airiness and detail. Similarly, the equalizer can be adjusted to boost high and mid frequencies as follows: moderately boost within the 100Hz-300Hz range to enhance the rhythm and drum impact; maintain fullness within the 400Hz-800Hz range to avoid thinning the sound; and carefully fine-tune within the 2kHz-5kHz range.
[0055] Within the acoustic scene set, the sound effect adjustment strategy for each acoustic scene is adapted to the frequency domain conditions corresponding to that scene. For example, if the frequency domain condition for an acoustic scene is in the third value range, the sound effect adjustment strategy for that scene could be to boost the mid-low frequencies, thereby enhancing clarity and detail, improving sound quality, and enhancing the user experience. As another example, if the frequency domain condition for an acoustic scene is in the fourth value range, the sound effect adjustment strategy for that scene could be to boost the mid-high frequencies, thus stabilizing the foundation and enhancing the user experience.
[0056] The sound effect adjustment strategy can be an equalizer adjustment strategy. The device can adjust the equalizer according to the equalizer adjustment strategy, thereby achieving the adjustment of the device's sound effects.
[0057] In some embodiments, after adjusting the device's current volume to a target volume, the device can also collect second ambient audio data via a microphone; second ambient noise features can be extracted based on the second ambient audio data; and a second ambient noise index can be calculated based on the second ambient noise features. The process of the device extracting the second ambient noise features based on the second ambient audio data is similar to the process of extracting the first ambient noise features based on the first ambient audio data, and the two can be explained by comparison. The process of the device calculating the second ambient noise index based on the second ambient noise features is similar to the process of calculating the first ambient noise index based on the first ambient noise features, and the two can be explained by comparison.
[0058] The device can compare a second ambient noise index with a first ambient noise index to determine the adjustment effect; it can also determine whether the adjustment effect meets preset conditions. If the adjustment effect does not meet the preset conditions, the device can further obtain the second noise level corresponding to the second ambient noise index; it can then compare the second noise level with the first noise level. If the second noise level is the same as the first noise level, it indicates that the noise environment in which the device is located has not changed, and the device volume can be fine-tuned. If the second noise level is different from the first noise level, it indicates that the noise environment in which the device is located has changed, and a second volume adjustment strategy can be determined based on the second noise level, allowing for further adjustment of the device volume. This achieves more precise and adaptive volume adjustment, thereby improving the user experience.
[0059] The device can subtract a second ambient noise index from a first ambient noise index. The adjustment effect includes the difference between the second and first ambient noise indices. Preset conditions include a threshold. The device can determine whether the difference is less than or equal to the threshold; if so, it determines that the adjustment effect meets the preset conditions; if not, it determines that the adjustment effect does not meet the preset conditions.
[0060] The device can identify the noise level corresponding to the second ambient noise index within a noise level set, and use this as the second noise level. The second ambient noise index satisfies the matching condition corresponding to the second noise level.
[0061] The process by which the device determines a second volume adjustment strategy based on a second noise level is similar to the process by which it determines a first volume adjustment strategy based on a first noise level; the two can be explained by comparison. Similarly, the process by which the device adjusts its volume based on the second volume adjustment strategy is similar to the process by which it adjusts its volume based on the first volume adjustment strategy; the two can be explained by comparison.
[0062] The device can be fine-tuned in any way. For example, the device can increase the default volume. Or, the device can decrease the default volume. The default volume can be in decibels (dB), or it can be a percentage of the device's total volume.
[0063] Optionally, the device can multiply the difference between the second and first ambient noise indices by a preset fine-tuning coefficient to obtain a fine-tuning amount; it can also add the fine-tuning amount to the target volume to achieve fine-tuning of the device's volume. The fine-tuning amount can be positive, negative, or zero. When the second ambient noise index is greater than the first ambient noise index, the fine-tuning amount is positive. When the second ambient noise index is equal to the first ambient noise index, the fine-tuning amount is zero. When the second ambient noise index is less than the first ambient noise index, the fine-tuning amount is negative.
[0064] Of course, if the adjustment effect does not meet the preset conditions, the device can further acquire the acoustic scene corresponding to the second ambient noise feature; it can then compare the acoustic scene corresponding to the second ambient noise feature with the acoustic scene corresponding to the first ambient noise feature. If the acoustic scene corresponding to the second ambient noise feature is the same as the acoustic scene corresponding to the first ambient noise feature, it means that the user's acoustic scene has not changed, and the volume can be fine-tuned. If the acoustic scene corresponding to the second ambient noise feature is different from the acoustic scene corresponding to the first ambient noise feature, it means that the user's acoustic scene has changed, and a second volume adjustment strategy can be determined based on the second noise level to which the second ambient noise index belongs. The device volume can then be adjusted again based on the second volume adjustment strategy. This achieves more precise and adaptive volume adjustment, improving the user experience. The process of the device acquiring the acoustic scene corresponding to the second ambient noise feature is similar to the process of acquiring the acoustic scene corresponding to the first ambient noise feature, and the two can be explained by comparison.
[0065] Alternatively, if the adjustment effect does not meet the preset conditions, the device can further acquire the geographic location data corresponding to the second environmental audio data; it can acquire the geographic location data corresponding to the first environmental audio data; and it can calculate the geographic distance based on the geographic location data corresponding to the second and first environmental audio data. The geographic location data corresponding to the second environmental audio data represents the geographic location where the device collected the second environmental audio data. The geographic location data corresponding to the first environmental audio data represents the geographic location where the device collected the first environmental audio data. If the geographic distance is less than or equal to a distance threshold, it indicates that the user's location has not changed significantly, and the volume can be fine-tuned. If the geographic distance is greater than the distance threshold, it indicates that the user's location has changed significantly, and a second volume adjustment strategy can be determined based on the second noise level to which the second environmental noise index belongs. The device volume can then be adjusted again based on the second volume adjustment strategy. This achieves more precise and adaptive volume adjustment, improving the user experience.
[0066] The device has a positioning component. Through this positioning component, the device can acquire geographic location data.
[0067] In some embodiments, the device can also acquire volume adjustment behavior data corresponding to multiple noise levels; input the volume adjustment behavior data corresponding to each noise level into the model to obtain user preference data for that noise level; and update the volume adjustment strategy corresponding to that noise level based on the user preference data for each noise level.
[0068] After automatic volume adjustment, if the user is not satisfied with the automatically adjusted volume, they can manually adjust the volume again. For this purpose, the device can collect user volume adjustment behavior data. This data can include the noise level during manual adjustment, the volume before adjustment, and the volume after adjustment. The device can collect multiple volume adjustment behavior data sets and classify them to obtain multiple datasets. These datasets correspond to noise level sets. Each noise level in the noise level set can correspond to one dataset. Each dataset can include one or more volume adjustment behavior data sets. Different volume adjustment behavior data sets within the same dataset can contain the same noise level. For each noise level in the noise level set, the device can input the dataset corresponding to that noise level into a model to obtain user preference data for that noise level. The model can include a machine learning model, such as a neural network model.
[0069] Optionally, user preference data may include preference types. Preference types reflect the user's preferred volume level. The device can update the volume adjustment strategy corresponding to each noise level based on the preference type. For example, preference types may include high volume, low volume, and moderate volume. The volume adjustment strategy may include a volume adjustment coefficient. If the preferred volume type for a noise level is high volume, the device can increase the volume adjustment coefficient corresponding to that noise level. If the preferred volume type for a noise level is low volume, the device can decrease the volume adjustment coefficient corresponding to that noise level. If the preferred volume type for a noise level is moderate volume, the device can keep the volume adjustment coefficient unchanged.
[0070] Optionally, user preference data may include preference type and preference magnitude. Preference type reflects the user's preferred volume direction, and preference magnitude reflects the intensity of the user's preferred volume. The device can update the volume adjustment strategy corresponding to each noise level based on the preference type and preference magnitude. For example, preference type may include high volume, low volume, moderate volume, etc. Preference magnitude can be the magnitude of change in the volume adjustment coefficient. The volume adjustment strategy may include the volume adjustment coefficient. If the preferred volume type for a noise level is high volume, the device can add the volume adjustment coefficient corresponding to that noise level to the preference magnitude, thereby increasing the volume adjustment coefficient corresponding to that noise level. If the preferred volume type for a noise level is low volume, the device can subtract the volume adjustment coefficient corresponding to that noise level from the preference magnitude, thereby decreasing the volume adjustment coefficient corresponding to that noise level. If the preferred volume type for a noise level is moderate volume, the device can keep the volume adjustment coefficient corresponding to that noise level unchanged.
[0071] The device can acquire volume adjustment behavior data corresponding to multiple noise levels at regular intervals (e.g., 1 month) to update the volume adjustment strategy for the multiple noise levels.
[0072] Therefore, based on user feedback, the volume adjustment strategy can be continuously optimized in a quantitative and personalized manner.
[0073] Please refer to Figure 2. This specification also provides a volume adjustment device, comprising: an extraction unit 21, configured to extract a first ambient noise feature based on first ambient audio data of the device; a calculation unit 22, configured to calculate a first ambient noise index based on the first ambient noise feature; a determination unit 23, configured to determine a first volume adjustment strategy based on a first noise level to which the first ambient noise index belongs; and an adjustment unit 24, configured to adjust the device volume according to the first volume adjustment strategy.
[0074] In some embodiments, the first environmental noise feature includes time-domain features and frequency-domain features. The calculation unit 22 is further configured to calculate a basic noise index based on the time-domain features; calculate noise complexity and energy of a specified frequency band based on the frequency-domain features; and weightedly sum the basic noise index, noise complexity, and energy of the specified frequency band to obtain the first environmental noise index.
[0075] In some embodiments, the time-domain features include the root mean square of the sampled values, and the frequency-domain features include the power spectral density. The calculation unit 22 is further configured to calculate the fundamental noise figure based on the root mean square; calculate the energy, spectral bandwidth, and spectral centroid of a specified frequency band based on the power spectral density; and calculate the noise complexity based on the spectral bandwidth and the spectral centroid.
[0076] In some embodiments, the volume adjustment device may further include a first identification unit. The first identification unit is used to identify a first noise level corresponding to the first ambient noise index in a noise level set; the noise level set includes multiple noise levels, each noise level corresponding to a matching condition; the first ambient noise index satisfies the matching condition corresponding to the first noise level.
[0077] In some embodiments, the first volume adjustment strategy includes a volume adjustment coefficient corresponding to a first noise level and a reference noise index. The adjustment unit 24 is further configured to calculate a compensation volume based on the first ambient noise index, the volume adjustment coefficient, and the reference noise index; calculate a target volume based on the compensation volume; and adjust the current volume of the device to the target volume.
[0078] In some embodiments, the adjustment unit 24 is further configured to determine an adjustment step size based on the compensation volume and a preset adjustment time; and adjust the current volume of the device to the target volume according to the adjustment step size.
[0079] In some embodiments, the volume adjustment device further includes a first updating unit. The first updating unit is configured to acquire volume constraints matching the first noise level; if the target volume does not meet the volume constraints, the target volume is updated according to the volume constraints. The adjustment unit 24 is further configured to adjust the current volume of the device to the updated target volume.
[0080] In some embodiments, the volume adjustment device may further include a second identification unit. The second identification unit is used to identify the acoustic scene corresponding to the first ambient noise feature in an acoustic scene set; the acoustic scene set includes multiple acoustic scenes, each acoustic scene has a corresponding matching condition, and the first ambient noise feature satisfies the matching condition corresponding to the acoustic scene; and adjusts the device sound effect according to the acoustic scene.
[0081] In some embodiments, the volume adjustment device may further include a fine-tuning unit. The fine-tuning unit is configured to extract a second ambient noise feature based on the second ambient audio data of the device; calculate a second ambient noise index based on the second ambient noise feature; compare the second ambient noise index with the first ambient noise index to determine the adjustment effect; and fine-tune the device volume if the adjustment effect does not meet a preset condition and the second noise level to which the second ambient noise index belongs is the same as the first noise level to which the first ambient noise index belongs. The adjustment unit 24 is further configured to, if the adjustment effect does not meet a preset condition and the second noise level to which the second ambient noise index belongs is different from the first noise level to which the first ambient noise index belongs, determine a second volume adjustment strategy based on the second noise level, and adjust the device volume according to the second volume adjustment strategy.
[0082] In some embodiments, the volume adjustment device may further include a second updating unit. The second updating unit is configured to acquire volume adjustment behavior data corresponding to multiple noise levels; input the volume adjustment behavior data corresponding to each noise level into a model to obtain user preference data for that noise level; and update the volume adjustment strategy corresponding to that noise level based on the user preference data for each noise level.
[0083] The volume adjustment device of this embodiment can extract a first ambient noise feature based on the device's first ambient audio data; calculate a first ambient noise index based on the first ambient noise feature; determine a first volume adjustment strategy based on the first noise level to which the first ambient noise index belongs; and adjust the device volume according to the first volume adjustment strategy. Therefore, this embodiment can automatically adjust the volume by analyzing ambient noise, thereby improving the user experience.
[0084] This specification also provides a computer device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the above-described volume adjustment method.
[0085] This specification also provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the volume adjustment method described above.
[0086] This specification also provides a computer program product, which includes a computer program that, when executed by a processor, implements the volume adjustment method described above.
[0087] Those skilled in the art will understand that this specification can be provided as a method, system, or computer program product. Therefore, this specification may take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware. Furthermore, this specification may take the form of a computer program product embodied on one or more computer-usable storage media (including, but not limited to, disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0088] This specification is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments thereof. It should be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. The computer may be a personal computer, laptop computer, cellular phone, camera phone, smartphone, personal digital assistant, media player, navigation device, email device, game console, tablet computer, wearable device, or any combination of these devices.
[0089] The functional units in the embodiments of this specification can be integrated into one processing unit, or each functional unit can exist physically separately, or two or more functional units can be integrated into one processing unit.
[0090] Those skilled in the art will understand that the descriptions of the various embodiments in this specification have different focuses, and parts not described in detail in a certain embodiment can be referred to in the relevant descriptions of other embodiments. Furthermore, it is understood that those skilled in the art, after reading this specification, can conceive of any combination of some or all of the embodiments listed in this specification without creative effort, and such combinations are also within the scope of disclosure and protection of this specification.
[0091] Although this specification has been described through embodiments, those skilled in the art will understand that the above embodiments are merely illustrative of the core ideas of this specification. Those skilled in the art will appreciate that many variations and modifications are possible with this specification. It is intended that the appended claims encompass these variations and modifications without departing from the spirit of this specification.
Claims
1. A volume adjustment method, characterized in that, include: Based on the device's first ambient audio data, extract the first ambient noise features; Based on the first environmental noise characteristics, calculate the first environmental noise index; based on the first noise level to which the first environmental noise index belongs, determine the first volume adjustment strategy; and adjust the device volume according to the first volume adjustment strategy.
2. The method according to claim 1, characterized in that, The first environmental noise feature includes time-domain features and frequency-domain features. The calculation of the first environmental noise index includes: calculating the basic noise index based on the time-domain features; calculating the noise complexity and the energy of a specified frequency band based on the frequency-domain features; and weighted summing the basic noise index, noise complexity, and energy of the specified frequency band to obtain the first environmental noise index.
3. The method according to claim 2, characterized in that, The time-domain feature includes the root mean square of the sampled values, and the frequency-domain feature includes the power spectral density. Calculating the fundamental noise index includes: calculating the fundamental noise index based on the root mean square. Calculating the noise complexity and the energy of a specified frequency band includes: calculating the energy, spectral bandwidth, and spectral centroid of the specified frequency band based on the power spectral density; and calculating the noise complexity based on the spectral bandwidth and the spectral centroid.
4. The method according to claim 1, characterized in that, The method further includes: identifying a first noise level corresponding to the first environmental noise index in a noise level set; the noise level set includes multiple noise levels, each noise level corresponding to a matching condition; the first environmental index satisfies the matching condition corresponding to the first noise level.
5. The method according to claim 1, characterized in that, The first volume adjustment strategy includes a volume adjustment coefficient corresponding to the first noise level and a reference noise index; The method of adjusting the device volume includes: calculating a compensation volume based on the first ambient noise index, the volume adjustment coefficient, and the reference noise index; calculating a target volume based on the compensation volume; and adjusting the current volume of the device to the target volume.
6. The method according to claim 5, characterized in that, Adjusting the current volume of the device to the target volume includes: determining an adjustment step size based on the compensation volume and a preset adjustment time; and adjusting the current volume of the device to the target volume according to the adjustment step size.
7. The method according to claim 5, characterized in that, Also includes: Obtain a volume constraint that matches the first noise level; if the target volume does not meet the volume constraint, update the target volume according to the volume constraint; adjusting the current volume of the device to the target volume includes: adjusting the current volume of the device to the updated target volume.
8. The method according to claim 5, characterized in that, Also includes: Identify the acoustic scene corresponding to the first environmental noise feature in the acoustic scene set; the acoustic scene set includes multiple acoustic scenes, each acoustic scene has a matching condition, and the first environmental noise feature satisfies the matching condition corresponding to the acoustic scene. Adjust the device's sound effects according to the described acoustic scenario.
9. The method according to claim 1, characterized in that, Also includes: Based on the second ambient audio data of the device, extract the second ambient noise features; Based on the second environmental noise characteristics, calculate the second environmental noise index; compare the second environmental noise index with the first environmental noise index to determine the adjustment effect; if the adjustment effect does not meet the preset conditions, and the second noise level to which the second environmental noise index belongs is the same as the first noise level to which the first environmental noise index belongs, fine-tune the device volume.
10. The method according to claim 9, characterized in that, Also includes: If the adjustment effect does not meet the preset conditions, and the second noise level to which the second environmental noise index belongs is different from the first noise level to which the first environmental noise index belongs, a second volume adjustment strategy is determined according to the second noise level, and the device volume is adjusted according to the second volume adjustment strategy.
11. The method according to claim 1, characterized in that, Also includes: Acquire volume adjustment behavior data corresponding to multiple noise levels; input the volume adjustment behavior data corresponding to each noise level into the model to obtain user preference data for that noise level; Based on user preference data for each noise level, update the volume adjustment strategy corresponding to that noise level.
12. A volume control device, characterized in that, include: The extraction unit is used to extract first ambient noise features based on the first ambient audio data of the device; The calculation unit is used to calculate the first environmental noise index based on the first environmental noise characteristics. The determining unit is configured to determine a first volume adjustment strategy based on the first noise level to which the first environmental noise index belongs; The adjustment unit is used to adjust the device volume according to the first volume adjustment strategy.
13. A computer device, comprising a memory, a processor, and computer program instructions stored in the memory, characterized in that, The processor executes the computer program instructions to implement the steps of the method according to any one of claims 1 to 11.
14. A computer-readable storage medium having computer program instructions stored thereon, characterized in that, When the computer program instructions are executed by a processor, they implement the steps of the method according to any one of claims 1 to 11.