A method and system for automatic volume adjustment of a loudspeaker
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-05-14
- Publication Date
- 2026-08-14
AI Technical Summary
该方案的缺陷在于不检测外部环境噪音,在多元场景(如交通噪音、对话、音乐等)下难以适应或者响应滞后,无法实时根据实际环境检测环境噪音并结合用户听觉偏好和播放内容类型进行音量的自适应调整
(1)本案通过提取环境噪音的多维特征,不仅识别噪音强度,还能准确分类噪音类型,并区分瞬时突变型噪音与缓慢变化型噪音,采取差异化的增益补偿与PID参数调节策略。对突变的强噪音能触发快速前馈补偿实现瞬间保护,对缓慢变化的噪音则执行平滑渐进式动态调节,从根本上避免了因噪音短暂波动导致的音量忽大忽小现象,兼顾了响应的迅捷性与听觉的舒适性;
Smart Images

Figure CN122569876A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of loudspeaker technology, and in particular to a method and system for automatic adaptive volume adjustment of loudspeakers. Background Technology
[0002] In existing terminal products (such as mobile phones, tablets, laptops, wireless speakers, etc.), the volume adjustment of the speakers is mostly manual. That is, users actively set the volume through buttons, touch, or remote control. However, this manual adjustment method has obvious limitations, especially when switching audio, users need to manually adjust the volume. It lacks an intelligent volume adjustment mechanism, resulting in a poor user experience and easily causing the volume to be too loud or too soft.
[0003] To address the aforementioned issues, patent document CN111416909A discloses a volume adaptive adjustment method, system, storage medium, and mobile terminal. This method collects audio of a preset duration at preset intervals, identifies the audio's speech data using big data methods, compares the volume of the audio data with the difference between the volume of the speech data, and adjusts the speech volume based on this difference to suit the user's comfort level. However, this solution lacks the ability to detect external environmental noise, making it difficult to adapt to or responding sluggishly in diverse scenarios (such as traffic noise, conversations, and music). It cannot detect environmental noise in real time and adaptively adjust the volume based on user listening preferences and the type of content being played. Summary of the Invention
[0004] In order to overcome the shortcomings of the prior art, the technical problem to be solved by the present invention is to propose a speaker adaptive volume automatic adjustment method and system, which can detect ambient noise in real time and combine user listening preferences and playback content type to automatically and smoothly adjust the output speaker volume to suit the scene, improve the user's listening experience, and is applicable to a variety of complex scenarios.
[0005] To achieve this objective, the present invention adopts the following technical solution: The present invention provides a speaker adaptive volume automatic adjustment method, comprising the following steps: S00: Ambient sound signals and audio signals played by the speaker are collected separately by a multi-microphone array, wherein the multi-microphone array includes at least two microphone arrays, and an adaptive filtering algorithm is used to separate the pure ambient noise signal and the speaker playback signal from the collected mixed signal. S10: Extract the multidimensional features of the environmental noise signal, identify the noise type and noise intensity, extract the acoustic features of the currently playing audio, and identify the content type of the playing content; S20: Analyze the spatial characteristic parameters of the signals collected by the multi-microphone array, construct a spatial soundscape map centered on the speaker, and dynamically divide the surrounding physical space into at least one listening area and at least one noise source area based on the spatial soundscape map; S30: Combining the noise type, noise intensity, content type, and area division information in the spatial soundscape map, and obtaining the user's historical preferred volume data in this scenario, calculate the optimal volume gain value and spatial orientation compensation parameters using a PID algorithm; S40: Based on the optimal volume gain value and spatial orientation compensation parameters, adjust the overall gain of the speaker audio signal, apply differentiated frequency band equalization and sound field rendering to different spatial directions, obtain the adjusted audio signal, and play it.
[0006] It also includes a speaker adaptive volume automatic adjustment system for implementing the speaker adaptive volume automatic adjustment method described above, comprising the following modules: A multi-microphone array is used to collect ambient sound signals and audio signals played by the speaker itself; An adaptive filtering and separation module is used to separate the clean ambient noise signal and the speaker playback signal from the acquired mixed signal; The noise feature extraction and classification module is used to extract multidimensional features of the environmental noise signal and identify the noise type and noise intensity. The spatial soundscape construction module is used to analyze the spatial characteristic parameters of the signals collected by the multi-microphone array, construct a spatial soundscape map centered on the speaker, and dynamically divide the surrounding physical space into at least one listening area and at least one noise source area based on the spatial soundscape map. The playback content recognition module is used to extract the acoustic features of the currently playing audio and identify the content type of the playback content; The user preference storage and update module is used to record and update users' historical preference volume data under different noise scenarios, different content types and different spatial partitions. It is also used to receive preference parameters manually set by users through mobile applications, calibrate users' historical preference volume data, and use the updated data for subsequent volume decisions. The volume and spatial compensation decision module is used to combine the noise type, noise intensity, content type, spatial area division information and user preference volume data to calculate the optimal volume gain value and spatial orientation compensation parameters through a PID algorithm. The gain and spatial equalization adjustment module is used to adjust the overall gain of the speaker audio signal according to the optimal volume gain value and spatial orientation compensation parameters, and to apply differentiated frequency band equalization and sound field rendering to different spatial directions to obtain the adjusted audio signal. An audio playback module is used to play the adjusted audio signal; The change detection and wake-up module is used to monitor changes in ambient noise or playback content in real time. When a sudden change is detected, it quickly wakes up each module to enter the working state, with a response time of no more than 100 milliseconds. The low-power management module is used to control relevant modules to enter a low-power mode when the speaker is in standby mode or when there is no significant change in the environment. The hardware self-diagnostic module is used to compare the features of the separated speaker playback signal with the original audio signal, calculate the distortion and frequency response offset, and output a warning signal when an anomaly is detected.
[0007] The beneficial effects of this invention are as follows: (1) This case extracts multidimensional features of environmental noise, which can not only identify noise intensity, but also accurately classify noise types and distinguish between instantaneous and sudden noise and slow-changing noise. Differentiated gain compensation and PID parameter adjustment strategies are adopted. For sudden strong noise, fast feedforward compensation can be triggered to achieve instant protection, while for slow-changing noise, smooth and gradual dynamic adjustment is performed, which fundamentally avoids the phenomenon of fluctuating volume caused by short-term noise fluctuations, and takes into account both the speed of response and the comfort of hearing. (2) When adjusting the volume, the content type of the currently playing audio is taken into account. For speech content, the gain of the speech segment is given priority. For music content, the balance of high and low frequencies is taken into account. For film and television sound effects, differentiated constraints are imposed. This breaks the logic of the traditional solution that arbitrarily adjusts the volume based solely on the noise intensity, and significantly improves the speech clarity and sound quality balance in different content scenarios. (3) This case constructs a spatial soundscape map centered on the loudspeaker and dynamically divides the core listening area, the dominant noise source area, and the spatial reflection area, realizing an upgrade from one-dimensional intensity perception to three-dimensional spatial dynamic perception with spectrum. Based on this spatial functional zoning, beamforming is used to actively gather sound energy in the listening area to form a bright area and generate an acoustic dark area in the noise source area. The masked frequency band is directionally compensated to protect the sound quality of the listening area in a "fence-like" manner, rather than simply eliminating the noise source. It has the ability to understand the environmental intent that traditional beamforming does not have. At the same time, the sound field is pre-compensated by modeling the reflection area, which effectively improves the working stability and listening experience in complex real spaces. (4) This case integrates multi-source heterogeneous data such as noise multidimensional characteristics, content type, spatial partition information and user historical preferences through PID algorithm to make unified decisions, generate the optimal volume gain value and spatial orientation compensation parameters such as beam focusing and zero trap control, so that the overall volume adjustment and spatial orientation sound field rendering are highly coordinated, the control process is smooth and stable without overshoot, and realizes personalized and imperceptible intelligent volume following. (5) A hierarchical adaptive filtering architecture is adopted, which combines a fast-converging least mean square algorithm and a fine separation module based on deep neural networks. The confidence of the separation results is evaluated, which effectively suppresses nonlinear echo and reverberation and ensures high confidence in the acquisition of environmental noise signals. At the same time, with the low power management and burst wake-up mechanism, as well as the speaker hardware self-diagnosis function, the reliability, real-time performance and long-term stable operation of the system are guaranteed. Attached Figure Description
[0008] Figure 1 This is a flowchart illustrating a speaker adaptive volume automatic adjustment method provided in a specific embodiment of the present invention. Figure 2 This is a schematic diagram of a speaker adaptive volume automatic adjustment system provided in a specific embodiment of the present invention; Figure 3 This is a system architecture block diagram of a speaker adaptive volume automatic adjustment system provided in a specific embodiment of the present invention. Detailed Implementation
[0009] The technical solution of the present invention will be further described below with reference to the accompanying drawings and specific embodiments.
[0010] To enable real-time and accurate detection of ambient noise, and combined with user auditory preferences and playback content types, this invention provides a speaker adaptive volume automatic adjustment method and system to improve the user's listening experience in diverse and complex environments. The core concept is to construct a closed-loop system of perception-separation-modeling-decision-rendering, deeply integrating spatial auditory perception with content, environment, and user habits. Specifically, it integrates and synergizes four aspects: multi-channel adaptive acoustic front-end separation, multi-dimensional environmental perception and spatial cognition, closed-loop decision-making integrating content and user, and spatial directional compensation and sound field rendering. This achieves intelligent, spatially directional, and personalized volume adjustment, significantly improving the listening experience in complex noisy environments.
[0011] Example 1: A method for automatic adaptive volume adjustment of a loudspeaker, which is described in the following five steps: physical signal separation, acoustic feature extraction, spatial cognition construction, multi-dimensional information fusion decision-making, and directional sound field control. S00: Multi-channel adaptive acoustic signal front-end separation, which uses a multi-microphone array to collect ambient sound signals and audio signals played by the speaker itself. The multi-microphone array includes at least two microphone arrays. Preferably, a dual-microphone array is used as an example. The first microphone is used to collect the near-field audio signal played by the speaker (i.e., the speaker playback signal) and can be set near the speaker's sound outlet. The second microphone is used to collect the spatial mixed signal containing ambient noise and playback reflections (i.e., the ambient noise signal) and can be set on the side of the device housing away from the sound outlet. Since the first microphone is also affected by ambient noise, the speaker playback signal can also directly obtain the digital reference signal of the played audio from the audio digital link. This digital reference signal and the near-field audio signal collected by the first microphone are used together as the reference input of the subsequent adaptive filter to eliminate the error introduced by the speaker's nonlinear distortion.
[0012] An adaptive filtering algorithm (such as the LMS algorithm) is used to separate the clean ambient noise signal and the speaker playback signal from the acquired mixed signal. A multi-microphone array can utilize spatial information, combined with an adaptive filter whose step size is dynamically adjusted based on the real-time analysis of the ambient noise characteristics and the type of playback content, to separate the speaker playback signal and the ambient noise signal in real time and dynamically. For example, using the near-field audio signal acquired by the first microphone as a reference signal, the mixed signal acquired by the second microphone is filtered to eliminate the playback signal component and output a clean ambient noise signal. The adaptive filter can adopt a hierarchical filtering architecture. The first-stage filter uses a fast-converging least mean square algorithm or a normalized least mean square algorithm for initial separation. The second-stage filter uses a deep neural network-based filtering model for fine separation of non-steady-state noise and spectral overlap noise. The initially separated noise signal is input to a lightweight deep neural network post-processing module to suppress residual nonlinear echoes and reverberation, outputting a high-confidence final ambient noise signal. Simultaneously, the confidence level of the final ambient noise signal is evaluated in real time, and a degradation processing strategy is triggered when the confidence level falls below a preset threshold.
[0013] After separating the clean environmental noise signal and the speaker playback signal, feature mining processing can be performed on these two clean signals. Specifically, in step S10, multi-dimensional features of the environmental noise signal are extracted to identify the noise type and intensity. This involves multi-dimensional analysis of the clean environmental noise signal, including extracting intensity and spectral features. The spectral features are then input into a pre-trained support vector machine or convolutional neural network model to identify the noise type, which includes at least one of traffic noise, human voice noise, and wind noise. Simultaneously, the instantaneous and slow changes in noise intensity are analyzed, and differentiated volume compensation coefficients are set for different noise types. This allows not only the noise intensity to be measured but also its nature to be identified, enabling different decisions based on the noise's characteristics. For example, in the face of sudden high-decibel alarms (strong, requiring immediate yielding) and continuously increasing traffic noise (strong, requiring gradual volume increase), this is significantly different from the logic of traditional solutions that only identify volume levels for response.
[0014] The above process identifies the noise type and intensity. However, in real-world scenarios, noise is not static but constantly changing. Furthermore, even if the peak volume of instantaneously abruptly changing noise and slowly changing noise are the same, their impact on the human ear is completely different. Therefore, after identifying the noise type and intensity, the process also includes: determining whether the noise is instantaneously abruptly changing or slowly changing based on the noise change trend parameter; employing rapid gain compensation for instantaneously abruptly changing noise and smooth, gradual dynamic adjustment for slowly changing noise; and feeding the noise change trend as input into the PID algorithm. Thus, when instantaneously abruptly changing noise is identified, the derivative term (the derivative term in the proportional-integral-derivative algorithm) can be temporarily increased to suppress volume overshoot; when dealing with slowly changing noise, the integral term is emphasized to ensure long-term accurate tracking.
[0015] Specifically, for transient, abrupt noises (such as those caused by closing doors, honking horns, or impacts, including intermittent noises when they recur), traditional solutions analyze the noise before reacting, thus lengthening the entire response chain. This solution, however, immediately identifies the abrupt change, triggering feedforward compensation for an extremely fast response and instantaneous protection, rapidly increasing the volume before the PID controller completes steady-state calculations. It also quickly restores volume after the noise disappears. For slowly changing noises, traditional solutions tend to rapidly increase the volume when noise levels rise and rapidly decrease it when noise levels fall. This solution employs a smooth, gradual dynamic adjustment, allowing the volume to change smoothly with the environment, avoiding sudden fluctuations in volume due to brief noise intensity changes. This mechanism, which adjusts volume based on noise trend parameters, makes volume control more intelligent, ensuring that no sudden protection needs are missed while avoiding overreaction during stable changes.
[0016] It also extracts the acoustic features of the currently playing audio to identify the content type. Specifically, it extracts and identifies features of the currently playing audio content, including: extracting the Mel-frequency cepstral coefficients and fundamental frequency features; matching the extracted features with a preset audio type feature library to identify the playback content type, which includes music, speech, and film / television sound effects. This step addresses the poor user experience caused by arbitrarily adjusting volume based solely on noise level or type, ignoring the characteristics of the content being played by the speaker. For example, for speech content, sufficient gain is prioritized when adjusting volume; for music content, a balance between high and low frequencies is considered. This provides a basis for accurately adopting different adjustment strategies (setting different volume adjustment benchmarks and dynamic range constraints) for different content types.
[0017] Preferably, when the ambient noise intensity and the type of content being played remain stable within a preset time, the microphone sampling rate and the operating frequency for content recognition are reduced; when a sudden change in ambient noise or a change in the type of content being played are detected, the system resumes full-function operation within 100 milliseconds.
[0018] The above describes volume adjustment based on the characteristics of the noise source itself (such as type, intensity, and trend of change). However, in real-world scenarios, the relative positions between the noise source and the person hearing the noise are dynamic. In step S20: the spatial characteristic parameters of the signals collected by the multi-microphone array are analyzed to construct a spatial soundscape map centered on the speaker. Based on this spatial soundscape map, the surrounding physical space is dynamically divided into at least one listening area and at least one noise source area. By performing spatial spectrum analysis and other techniques on the multi-microphone array signals, the three-dimensional coordinates of the main sound sources in the environment can be estimated, constructing a dynamic sound source distribution map. Then, combined with data such as device posture, the "listening area" and "noise source area" are intelligently delineated, providing precise spatial coordinates for directional audio processing.
[0019] The steps involved in constructing a spatial soundscape map include: S21: Using the phase difference, time difference, and intensity difference information of the same sound source signal received by the multi-microphone array, calculate the azimuth and elevation angles of the sound source relative to the speaker. This step is a conversion from physical signal to spatial coordinates. By utilizing the inherent time difference, phase difference, and intensity difference when the multi-microphone array receives the same signal, the azimuth and elevation angles of any sound source relative to the speaker can be accurately calculated. By assigning spatial attributes to the noise, the direction of noise transmission can be identified, providing the necessary geometric coordinates for subsequent directional analysis and compensation.
[0020] S22: Blind source separation is performed on the environmental noise signal to obtain multiple independent noise components, and their spatial orientation and energy changes are tracked separately. Noise in the real environment is often a mixture of multiple sounds. This blind source separation step allows for the extraction of multiple independent noise components from the mixed signal, and their spatial orientation and energy changes are tracked separately. This simultaneous processing of multiple noise sources of different natures provides a physical basis for delineating "noise source areas."
[0021] S23: By integrating the azimuth, elevation, energy changes, and spectrum information, a three-dimensional spatial soundscape map containing sound source location and distribution is generated. Steps S21 and S22 above generated scattered location and energy data for multiple sound sources, but this is insufficient to support intelligent decision-making. Therefore, S23 integrates this data with spectrum information (corresponding to the noise type in S10) to generate a real-time updated three-dimensional spatial soundscape map, comprehensively depicting the type, location, intensity, and dynamic changes of surrounding sound sources, completing the transition from environmental perception (steps S10 and S20) to subsequent fusion decision-making (steps S30 and S40).
[0022] In this way, the three sub-steps described above enhance the perception of environmental noise signals from one-dimensional intensity to a dynamic multi-dimensional space with added spectrum, laying the geometric and semantic foundation for subsequent precise and imperceptible directional compensation.
[0023] In the above scheme, the dynamic division of the surrounding physical space into at least one listening zone and at least one noise source zone based on the spatial soundscape map includes: By detecting the user's location or user-defined listening points, the core listening area is located; thus, a service direction for subsequent sound field rendering and beamforming is determined. By identifying the main incident directions of environmental noise, the dominant noise source area is calibrated; this prioritizes the multiple independent noise sources separated in S22. Computational and acoustic resources can be concentrated to perform beam suppression in the direction of the dominant noise source, while background noise from other directions is only processed conventionally, achieving the most efficient use of acoustic resources. By analyzing early reflected sound, the spatial reflection boundary is estimated, and the spatial reflection zone is divided. In this way, the reflection boundary can be estimated by analyzing early reflected sound, which can distinguish between direct sound and reflected sound, and avoid misjudging reflected noise as an independent noise source. When performing beamforming, the reflection boundary can be included in the calculation, and reflection interference can be canceled through pre-filtering or other methods, or even beneficial reflected sound can be used to enhance the spatial sense of the listening area.
[0024] Based on the location and attributes of sound objects in the spatial soundscape, the boundaries and priorities of each area are dynamically adjusted; based on the real-time location and attribute changes of sound objects, the areas are dynamically shrunk, expanded, merged, or reassigned priorities, ensuring that the functional zoning is optimal in any emergency scenario and avoiding adjustment errors caused by scene switching.
[0025] Building upon the above, further fusion decision-making and spatial orientation compensation are performed based on the PID algorithm. In step S30: combining the noise type, noise intensity, content type, and regional division information in the spatial soundscape map, and obtaining the user's historical preferred volume data in this scenario, the optimal volume gain value and spatial orientation compensation parameters are calculated using the PID algorithm. The optimal volume gain value determines the overall volume adjustment, while the spatial orientation compensation parameters determine how the sound frequency band and sound field change in different directions. Based on these parameters, differentiated equalization and rendering are applied to different spatial directions, and the adjusted audio is output. This completes the closed loop from physical perception to intelligent control. In step S40: based on the optimal volume gain value and spatial orientation compensation parameters, the overall gain of the speaker audio signal is adjusted, and differentiated frequency band equalization and sound field rendering are applied to different spatial directions to obtain the adjusted audio signal, which is then played. Thus, through the PID control algorithm, the volume adjustment action becomes smooth, stable, and without overshoot, avoiding the discomfort caused by abrupt changes. Spatial directional compensation focuses on sound processing: at the physical or algorithmic level, it actively enhances the sound energy directed towards the listening area or creates suppression in the direction of noise sources, thereby fundamentally improving the signal-to-noise ratio and listening experience.
[0026] In step S30, obtaining the user's historical preferred volume data in this scenario specifically includes reading the user's historical volume adjustment records from the memory. These historical volume adjustment records are associated with different environmental noise information and playback content types. The user's listening habits are mined through big data analysis algorithms to generate personalized volume adjustment benchmark values and adjustment sensitivity parameters. Users can manually set preference parameters and calibrate the personalized model through a mobile application.
[0027] In step S30, the spatial orientation compensation parameters include beam focusing parameters and null control parameters; these two parameters are key parameters in the beamforming scheme. The beam focusing parameters maximize the gain of the target sound source, while the null control parameters maximize the attenuation of interfering sound sources (noise sources). Based on this, the differential frequency band equalization and sound field rendering of the speaker audio signal in different spatial directions includes: S31: By adjusting the phase and time delay of different units in the loudspeaker array through beamforming, the sound wave energy forms a bright area with concentrated energy in the listening area and an acoustic dark area with weakened energy in the direction of the noise source area. In this way, by forming a bright area with concentrated energy in the listening area and a dark area with weakened energy in the noise source area, an acoustic isolation zone that is difficult for noise to enter can be directly created for the listening area to prevent environmental noise from interfering with the listening area.
[0028] S32: Based on the spectral characteristics of the ambient noise and the direction of the noise source area, directional gain compensation is performed only towards the listening area for specific frequency bands masked by the ambient noise; directional gain compensation is also performed for key frequency bands masked by the ambient noise, targeting the listening area. This improves speech clarity and the perceptibility of musical details.
[0029] It should be noted that in existing technologies, using microphone array beamforming for passive noise source identification and localization, and combining it with active noise cancellation systems to achieve targeted suppression, is a relatively mature technology. However, it often exists as a standalone module. The logic of this traditional approach is to detect the noise source, calculate its location, and create a null area (dark zone) at that location to suppress the noise. Its characteristics are a single beam target and static triggering rules; that is, once the noise location is identified, a fixed pattern of suppression is applied to that location. Of course, it does not consider factors such as the content being played by the speaker, the user, or the scene.
[0030] In this solution, spatial directional compensation focuses and enhances sound in the listening area while suppressing and reducing noise in the noise source area. Simultaneously, it locks onto the user service within the listening area, ensuring optimal listening experience for those within that area. Compared to traditional beamforming solutions, spatial directional compensation in this case is the final execution stage of the system decision-making process. The control parameters for beamforming come from an intelligent decision-making layer that integrates environmental, content, and user preferences. Furthermore, it is based on a three-dimensional spatial soundscape map, incorporating multi-dimensional information such as the listening area, noise source area, reflection area, noise type, content type, and user historical preferences. This will be further discussed from the following three perspectives: (1) First, a fence-like spatial orientation compensation logic is established. The core listening area is finely delineated through S20. The processing logic of steps S31 and S32 is to protect the listening area and establish a safety barrier to isolate environmental noise, rather than eliminating the noise source. For example, when noise is generated right next to the user (such as when people walking together are talking), the system may judge that this belongs to the user's social scene, and thus only perform slight suppression or no processing. This processing strategy based on spatial functional zoning is an intention understanding capability that beamforming schemes in traditional solutions for noise suppression obviously do not have.
[0031] (2) Next, this solution performs differentiated execution based on content and preferences through steps S10 to S40, taking into account the scenario, content, and user preferences when making decisions. Based on this, steps S41 and S42 are then utilized. For example, for voice content, the frequency band compensation in S42 will accurately enhance the frequency bands related to voice clarity; for movies, the beam focusing in S41 will pay more attention to maintaining the breadth of the sound field and the sense of immersion.
[0032] (3) This case also utilizes the reflection zone to optimize the beam, that is, the spatial reflection zone is specially divided through S20 to model the spatial sound scene map close to the real environment. When S41 performs beamforming, its algorithm will include these reflection boundaries in the calculation, perform pre-compensation or utilization, so as to improve the stability and effect of working in actual application scenarios. In contrast, beamforming in traditional solutions usually assumes a free sound field, which cannot effectively deal with reflections from walls and other obstacles, and may even cause new interference due to reflections.
[0033] It should be added that, based on the understanding of the acoustic space in step S20, the changes in the environment or the user's position are further monitored in real time, and the spatial soundscape map and area division are updated; when the spatial position of the user or noise is detected to change, the spatial orientation compensation parameters are recalculated and the sound field is adjusted for a smooth transition.
[0034] Further steps include step S50: a self-diagnostic step, which compares the separated speaker playback signal with the original input audio signal to calculate the distortion and frequency response difference; establishes a speaker status model based on long-term comparison results, and issues an early warning when abnormal features are detected.
[0035] Example 2: A speaker adaptive volume automatic adjustment system, used to implement the speaker adaptive volume automatic adjustment method described above, includes the following modules in the order of sensing, analysis and processing, fusion decision-making, adjustment execution, and other safeguards: (1) Perception process A multi-microphone array is used to collect ambient sound signals and audio signals played by the speaker itself; An adaptive filtering and separation module is used to separate the pure ambient noise signal and the speaker playback signal from the acquired mixed signal; such as Figure 2 As shown, it can be integrated into a microcontroller.
[0036] The playback content recognition module is used to extract the acoustic features of the currently playing audio and identify the content type of the playback content; such as... Figure 2 As shown, it can be integrated into Type Recognition.
[0037] (2) Analysis and processing The noise feature extraction and classification module is used to extract multidimensional features of the environmental noise signal and identify the noise type and intensity; for example... Figure 2 As shown, it can be integrated into Noise Analysis.
[0038] A spatial soundscape construction module is used to analyze the spatial characteristic parameters of the signals collected by the multi-microphone array, construct a spatial soundscape map centered on the speaker, and dynamically divide the surrounding physical space into at least one listening zone and at least one noise source zone based on the spatial soundscape map; such as Figure 2 As shown, it can be integrated into the Construction of Spatial Soundscape, and the spatial soundscape construction module includes: The azimuth calculation unit is used to calculate the azimuth and elevation angles of the sound source by utilizing the phase difference, time difference, and intensity difference of the signals received by the multi-microphone array. A blind source separation unit is used to separate the environmental noise signal into multiple independent noise components and track their spatial trajectories; The map generation and partitioning unit is used to integrate sound source location, energy and spectrum information to generate a three-dimensional spatial soundscape map, and dynamically divide the core listening area, dominant noise source area, spatial reflection boundary area and interactive perception area.
[0039] The user preference storage and update module records and updates users' historical preferred volume data under different noise scenarios, content types, and spatial partitions. This scenario-integrated preference record represents a composite preference across noise scenarios, content types, and spatial partitions. It also receives preference parameters manually set by users through a mobile application, calibrates the user's historical preferred volume data, and uses the updated data for subsequent volume decisions. Figure 2 As shown, it can be integrated into User Storaga.
[0040] (3) Integrated decision-making The volume and spatial compensation decision module is used to combine the noise type, noise intensity, content type, spatial area division information, and user preference volume data to calculate the optimal volume gain value and spatial orientation compensation parameters using a PID algorithm; for example... Figure 2 As shown, it can be integrated into a microcontroller.
[0041] (4) Execution of adjustment instructions Gain and spatial equalization adjustment modules, such as Figure 2 As shown, it can be integrated into a microcontroller to adjust the overall gain of the speaker audio signal based on the optimal volume gain value and spatial orientation compensation parameters, and to apply differentiated frequency band equalization and sound field rendering to different spatial directions to obtain the adjusted audio signal; the gain and spatial equalization adjustment module includes: A beam controller is used to adjust the phase and time delay of each unit in the loudspeaker array to form a bright acoustic area in the listening area and an acoustic dark area in the noise source area. A directional equalizer is used to provide directional gain compensation for the masked frequency bands only in the direction of the listening area, based on the noise spectrum masking characteristics.
[0042] An audio playback module is used to play the adjusted audio signal; such as Figure 2 As shown, it can be integrated into the Speaker.
[0043] (5) Other functional modules The change detection and wake-up module is used to monitor changes in ambient noise or playback content in real time. When a sudden change is detected, it quickly wakes up each module to enter the working state, with a response time of no more than 100 milliseconds. The low-power management module is used to control relevant modules to enter a low-power mode when the speaker is in standby mode or when there is no significant change in the environment. The hardware self-diagnostic module is used to compare the features of the separated speaker playback signal with the original audio signal, calculate the distortion and frequency response offset, and output a warning signal when an anomaly is detected.
[0044] This invention has been described through preferred embodiments. Those skilled in the art will understand that various changes or equivalent substitutions can be made to these features and embodiments without departing from the spirit and scope of the invention. This invention is not limited to the specific embodiments disclosed herein; other embodiments falling within the scope of the claims are also within the protection scope of this invention.
Claims
1. A method for automatic adaptive volume adjustment of a speaker, characterized in that, Includes the following steps: S00: Ambient sound signals and audio signals played by the speaker are collected separately by a multi-microphone array, wherein the multi-microphone array includes at least two microphone arrays, and an adaptive filtering algorithm is used to separate the pure ambient noise signal and the speaker playback signal from the collected mixed signal. S10: Extract the multidimensional features of the environmental noise signal, identify the noise type and noise intensity, extract the acoustic features of the currently playing audio, and identify the content type of the playing content; S20: Analyze the spatial characteristic parameters of the signals collected by the multi-microphone array, construct a spatial soundscape map centered on the speaker, and dynamically divide the surrounding physical space into at least one listening area and at least one noise source area based on the spatial soundscape map; S30: Combining the noise type, noise intensity, content type, and area division information in the spatial soundscape map, and obtaining the user's historical preferred volume data in this scenario, calculate the optimal volume gain value and spatial orientation compensation parameters using a PID algorithm; S40: Based on the optimal volume gain value and spatial orientation compensation parameters, adjust the overall gain of the speaker audio signal, apply differentiated frequency band equalization and sound field rendering to different spatial directions, obtain the adjusted audio signal, and play it.
2. The speaker adaptive volume automatic adjustment method according to claim 1, characterized in that, In step S10, after identifying the noise type and noise intensity, the method further includes: determining whether the noise is a sudden change or a slow change based on the noise change trend parameter; using fast gain compensation for sudden change noise and smooth, gradual dynamic adjustment for slow change noise; and feeding the noise change trend as an input into the PID algorithm.
3. The speaker adaptive volume automatic adjustment method according to claim 1, characterized in that: In step S20, the step of analyzing the spatial characteristic parameters of the signals acquired by the multi-microphone array and constructing a spatial soundscape map includes: S21: Calculate the azimuth and elevation angles of the sound source relative to the loudspeaker using the phase difference, time difference, and intensity difference information of the same sound source signal received by the multi-microphone array; S22: Perform blind source separation on the environmental noise signal to obtain multiple independent noise components, and track their spatial orientation and energy changes respectively; S23: Based on the azimuth, elevation, energy changes and spectrum information, generate a three-dimensional spatial soundscape map that includes sound source localization and distribution.
4. The speaker adaptive volume automatic adjustment method according to claim 1, characterized in that, In step S20, the dynamic division of the surrounding physical space into at least one listening zone and at least one noise source zone based on the spatial soundscape map includes: The core listening area is located by detecting the user's location or user-defined listening points; By identifying the main incident directions of environmental noise, the dominant noise source area can be pinpointed; By analyzing early reflected sound, the spatial reflection boundary is estimated, and the spatial reflection zone is delineated. Based on the location and attributes of sound objects in the spatial soundscape, the boundaries and priorities of each area are dynamically adjusted.
5. The speaker adaptive volume automatic adjustment method according to claim 1, characterized in that, In step S20, environmental changes or user location changes are monitored in real time, and the spatial soundscape map and area division are updated accordingly. When a change in the spatial location of a user or noise is detected, the spatial orientation compensation parameters are recalculated to smoothly adjust the sound field.
6. The speaker adaptive volume automatic adjustment method according to claim 1, characterized in that, In step S30, the spatial orientation compensation parameters include beam focusing parameters and null control parameters; The differential frequency band equalization and sound field rendering of the speaker audio signal in different spatial directions includes: S31: By adjusting the phase and time delay of different units in the loudspeaker array through beamforming, the sound wave energy forms a bright area where energy is concentrated in the listening area and an acoustic dark area where energy is weakened in the direction of the noise source area. S32: Based on the spectral characteristics of the ambient noise and the direction of the noise source area, perform directional gain compensation only towards the direction of the listening area for specific frequency bands masked by the ambient noise.
7. The speaker adaptive volume automatic adjustment method according to claim 1, characterized in that, It also includes step S50: a self-diagnosis step, which compares the separated speaker playback signal with the original input audio signal, calculates the distortion and frequency response difference; establishes a speaker status model based on long-term comparison results, and issues an early warning when abnormal features are detected.
8. A speaker adaptive volume automatic adjustment system, used to implement the speaker adaptive volume automatic adjustment method as described in claims 1-7, characterized in that, Includes the following modules: A multi-microphone array is used to collect ambient sound signals and audio signals played by the speaker itself; An adaptive filtering and separation module is used to separate the clean ambient noise signal and the speaker playback signal from the acquired mixed signal; The noise feature extraction and classification module is used to extract multidimensional features of the environmental noise signal and identify the noise type and noise intensity. The spatial soundscape construction module is used to analyze the spatial characteristic parameters of the signals collected by the multi-microphone array, construct a spatial soundscape map centered on the speaker, and dynamically divide the surrounding physical space into at least one listening area and at least one noise source area based on the spatial soundscape map. The playback content recognition module is used to extract the acoustic features of the currently playing audio and identify the content type of the playback content; The user preference storage and update module is used to record and update users' historical preference volume data under different noise scenarios, different content types and different spatial partitions. It is also used to receive preference parameters manually set by users through mobile applications, calibrate users' historical preference volume data, and use the updated data for subsequent volume decisions. The volume and spatial compensation decision module is used to combine the noise type, noise intensity, content type, spatial area division information and user preference volume data to calculate the optimal volume gain value and spatial orientation compensation parameters through a PID algorithm. The gain and spatial equalization adjustment module is used to adjust the overall gain of the speaker audio signal according to the optimal volume gain value and spatial orientation compensation parameters, and to apply differentiated frequency band equalization and sound field rendering to different spatial directions to obtain the adjusted audio signal. An audio playback module is used to play the adjusted audio signal; The change detection and wake-up module is used to monitor changes in ambient noise or playback content in real time. When a sudden change is detected, it quickly wakes up each module to enter the working state, with a response time of no more than 100 milliseconds. The low-power management module is used to control relevant modules to enter a low-power mode when the speaker is in standby mode or when there is no significant change in the environment. The hardware self-diagnostic module is used to compare the features of the separated speaker playback signal with the original audio signal, calculate the distortion and frequency response offset, and output a warning signal when an anomaly is detected.
9. A speaker adaptive volume automatic adjustment system according to claim 8, characterized in that, The spatial soundscape construction module includes: The azimuth calculation unit is used to calculate the azimuth and elevation angles of the sound source by utilizing the phase difference, time difference, and intensity difference of the signals received by the multi-microphone array. A blind source separation unit is used to separate the environmental noise signal into multiple independent noise components and track their spatial trajectories; The map generation and partitioning unit is used to integrate sound source location, energy and spectrum information to generate a three-dimensional spatial soundscape map, and dynamically divide the core listening area, dominant noise source area, spatial reflection boundary area and interactive perception area.
10. A speaker adaptive volume automatic adjustment system according to claim 8, characterized in that, The gain and spatial equalization adjustment module includes: A beam controller is used to adjust the phase and time delay of each unit in the loudspeaker array to form a bright acoustic area in the listening area and an acoustic dark area in the noise source area. A directional equalizer is used to provide directional gain compensation for the masked frequency bands only in the direction of the listening area, based on the noise spectrum masking characteristics.
Citation Information
Patent Citations
Volume adaptive adjustment method and system, storage medium and mobile terminal
CN111416909A