Self-adaptive environment sound box control system and method based on artificial intelligence
By setting up a weak echo cavity and a ball source swing model in the speaker, combined with spectrum analysis and volume adjustment, the noise and echo problems of the speaker in a noisy environment are solved, and the sound quality and user experience are improved.
Patent Information
- Application Number
- CN202510869714.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-26
- Publication Date
- 2025-09-16
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
Existing speakers have difficulty adaptively eliminating noise and echoes in noisy environments, affecting sound quality and user experience.
A weak echo chamber is set up in the speaker, and the ambient noise is collected using a recording device. The noise is separated through spectrum analysis and adaptive filters. The sound effect is adjusted by combining the spherical source swing and boundary element model, and the volume is dynamically adjusted to eliminate echoes and reduce noise.
It achieves adaptive noise and echo elimination in complex environments, improves sound quality, reduces the user's need to turn up the volume, reduces fatigue from long-term listening, expands the usage scenarios of speakers, and improves the level of intelligence.
Smart Images

Figure CN120652810A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of speaker control, and in particular to an artificial intelligence-based adaptive environmental speaker control system and method. Background Art
[0002] A speaker is a device that converts electrical signals into sound. A typical speaker structure includes a housing, diaphragm, voice coil, sound source, and tuner. When the speaker is powered on, the voice coil wrapped around the diaphragm is forced to move in a magnetic field, driving the sound source to produce a swinging or pulsating sound. This sound is then amplified by the diaphragm and reproduced as sound. Speakers with adaptive functions can automatically adjust audio output based on changes in the surrounding environment, dynamically optimizing sound parameters and providing an optimal listening experience.
[0003] Most environments contain noise of varying waveforms, such as in a noisy living room or outdoors. This high volume and chaotic noise can interfere with the speaker's audio playback, reducing the user's listening experience in a noisy environment. Most environmentally adaptive speakers only reduce the impact of noise by adjusting the volume, but cannot achieve adaptive noise cancellation.
[0004] In addition, sound waves will produce echoes when they touch environmental obstacles. The echo standing waves mix with the original audio to produce reverberation. Reverberation will increase the volume heard by users at a fixed position and produce noise, affecting the user's listening experience. Due to the complexity of environmental echoes and the uncertainty of user positions, conventional reverberation smoothing methods cannot adapt to complex environments, and it is more difficult to control the sound effects. Summary of the Invention
[0005] The purpose of the present invention is to provide an adaptive ambient sound box control system and method based on artificial intelligence to solve the problems raised in the above background technology.
[0006] In order to solve the above technical problems, the present invention provides the following technical solutions: an adaptive environmental speaker control system based on artificial intelligence, comprising: a noise acquisition module, a spectrum analysis module, an echo test module, a field strength modeling module and a sound effect adjustment module; The noise collection module is used to set up a weak echo chamber in the speaker, which is soundproofed from the outside world, and play standard audio at a fixed volume in the weak echo chamber. The standard audio is white noise or a linear frequency modulation signal. Before the audio is played, a microphone recording device is used to obtain ambient noise, mix the ambient noise into the audio in the weak echo chamber, and collect the audio in the weak echo chamber as mixed audio. The mixed audio spectrum is uploaded to the speaker processing chip; The spectrum analysis module is used to process the mixed audio spectrum through short-time Fourier transform, compare the standard audio with the mixed audio, separate the noise components using spectral subtraction, output the noise spectrum, obtain the audio to be played from the speaker, divide the audio to be played into time segments, subtract the original noise from the spectrum of each segment of the audio to be played, and output the processed audio to the voice coil magnetic field; The echo test module is used to establish a field strength polar coordinate system with the speaker as the center based on the volume and the spectrum of the output audio, collect the reverberation audio of the audio echo and the ambient noise with a recording device, use an adaptive filter to eliminate the direct sound, and then use the noise spectrum to clean the noise to obtain the echo spectrum, and play the echo spectrum in the echo chamber; The field intensity modeling module is used to calculate the swing speed, angle and pulse interval of the ball source based on the frequency spectrum in the echo chamber, so that the sound emitted by the ball source is consistent with the frequency spectrum in the echo chamber. A micro servo motor is used to drive the metal ball inside the speaker to swing, simulating the movement of the sound field scatterer. The swing frequency and swing amplitude are monitored in real time by an accelerometer. The convolution reverberation algorithm is used to superimpose the swing of the ball source on the audio playback. The ball source is regarded as a secondary sound source to calculate the sound pressure distribution at the field point, and a BEM boundary element model is established. The sound effect adjustment module is used to obtain the user position coordinates through UW or infrared sensors, map them to the sound pressure model, calculate the masking threshold according to the noise power spectrum, and dynamically adjust the main speaker EQ based on the volume set by the user to make the sound pressure at the user position equal to the sound pressure at the reference position with the set volume.
[0007] Furthermore, the noise collection module includes: an echo chamber unit and a noise recording unit; The echo chamber unit uses an inner wall of sound-absorbing material to suppress internal sound wave reflection, and the cavity volume meets the Helmholtz resonance condition to ensure non-destructive testing of sound; The noise recording unit is composed of a spherical array of omnidirectional microphones arranged in a speaker housing and is used to collect spatial echo signals.
[0008] Furthermore, the spectrum analysis module includes: a mixed extraction unit and an audio processing unit; The mixing extraction unit is used to select full-band white noise or linear frequency modulation signal to mix standard audio and noise; The audio processing unit is used to separate the noise components and process the original audio at the same time so that the environmental noise is included in the audio composition.
[0009] Furthermore, the echo test module includes: a noise suppression unit, a speaker coordinate unit and an echo spectrum unit; The noise suppression unit is used to subtract the noise power spectrum and the direct sound spectrum from the reverberant audio to retain the pure echo component; The speaker coordinate unit is used to determine the echo angle by estimating the direction of arrival and to construct a polar coordinate sound field model with the speaker as the center; The echo spectrum unit is used to reduce the amplitude of the echo spectrum and play the echo spectrum in the echo chamber.
[0010] Furthermore, the field intensity modeling module includes: a field point sound pressure unit, a user positioning unit and a spherical source reverberation unit; The field point sound pressure unit is used to calculate the sound characteristics of the ball source according to the echo spectrum, so that the sound of the ball source is consistent with the echo; The user positioning unit is used to obtain the user's position coordinates and optimize the user's directional hearing sense; The spherical source reverberation unit is used to adjust the reverberation decay time to match the environmental characteristics and establish a BEM boundary element model.
[0011] Furthermore, the sound effect adjustment module includes: a user interaction unit and an audio modulation unit; The user interaction unit is used to obtain the user's playback settings for the original audio, initialize the audio volume, and calculate the sound pressure at the reference position; The audio modulation unit is used to dynamically adjust the playback volume of the main speaker to ensure that the output sound pressure is consistent with the set sound pressure.
[0012] An artificial intelligence-based adaptive ambient speaker control method comprises the following steps: Step S1. A weak echo chamber is set up in the speaker box to be soundproofed from the outside world. A spherical sound source is set up in the weak echo chamber. Standard audio is played at a fixed volume in the weak echo chamber. External ambient noise is introduced using a recording device to extract the spectrum of the mixed audio in the weak echo chamber. Step S2. Process the mixed audio using a short-time Fourier transform, separate the noise component from the mixed audio using spectral subtraction, output the noise spectrum, inversely transform the noise spectrum to obtain the original noise, subtract the original noise from the spectrum of the audio to be played on the speaker to obtain the output audio, and control the speaker to play the output audio; Step S3. Using a recording device to capture the reverberation of the audio echo and ambient noise, the original noise and direct sound are removed from the reverberation audio to obtain the echo spectrum. A micro-servo motor is used to drive the metal ball inside the speaker to swing, simulating the motion of a sound field scatterer. The swing frequency and amplitude of the sphere sound source are adjusted so that the sound pressure emitted by the sphere source at a fixed distance is consistent with the sound pressure of the echo spectrum. Step S4. Using an accelerometer to monitor the swing frequency and amplitude of the spherical sound source in real time, playing the output audio in a weak echo chamber, obtaining the sound pressure distribution of the output audio based on the sound pressure sampling data in the weak echo chamber, superimposing the swing audio of the spherical sound source on the output audio using a convolution reverberation algorithm, and establishing a boundary element model centered on the speaker; Step S5. Obtain the user position coordinates and map them to the boundary element model to obtain the actual sound pressure at the user's position, and dynamically adjust the main speaker volume based on the volume set by the user so that the actual sound pressure at the user's position is equal to the sound pressure at the reference position with the set volume.
[0013] Furthermore, step S1 includes: Step S11. A weak echo chamber is provided inside the speaker. The weak echo chamber uses an inner wall made of sound-absorbing material to suppress internal sound wave reflections. The chamber volume satisfies the Helmholtz resonance condition. An omnidirectional microphone spherical array is arranged on the speaker housing. The omnidirectional microphone spherical array has two modes: capturing ambient noise and amplifying sound. Step S12: Play standard audio in the weak echo chamber. The standard audio is white noise or linear frequency modulation signal. The playback volume is controlled at 30-40dB to avoid interfering with the main speaker output. Step S13: At regular intervals, external environmental noise is introduced into the weak echo chamber, mixed with the standard audio to obtain mixed audio, and the spectrum of the mixed audio is obtained through a spectrum analysis chip.
[0014] Furthermore, step S2 includes: Step S21. Process the mixed audio spectrum by short-time Fourier transform, subtract the mixed audio spectrum from the standard audio spectrum to obtain a noise spectrum, and inversely transform the noise spectrum to obtain the original noise; Step S22. Acquire the audio to be played from the speaker in real time, subtract the original noise from the spectrum of the audio to be played, obtain the output audio, make the ambient noise participate in the audio composition, output the processed audio to the voice coil magnetic field, and control the speaker to play the output audio.
[0015] Furthermore, step S3 includes: Step S31. After the speaker plays audio, reverberant audio is acquired through an omnidirectional microphone spherical array. The reverberant audio includes: direct sound from the speaker, ambient noise, and echo. The direct sound in the reverberant audio is eliminated using an adaptive filter. The ambient noise is then removed from the reverberant audio using spectral subtraction to obtain an echo spectrum. Step S32. Use a micro servo motor to drive the metal ball inside the speaker to swing, forming a spherical sound source, collect the sound pressure of the spherical sound source and the sound pressure of the echo spectrum at a fixed position in the weak echo cavity, adjust the swing frequency and swing amplitude of the spherical sound source, and make the sound pressure of the spherical sound source equal to the sound pressure of the echo spectrum.
[0016] Furthermore, step S4 includes: Step S41. Use an accelerometer to monitor the swing frequency and swing amplitude of the spherical sound source in real time, and substitute the swing frequency and swing amplitude of the spherical sound source into the sound source transfer function: ; Where H(A,f) is the sound source transfer function, which represents the change in sound pressure of the spherical sound source when the swing frequency is f and the swing amplitude is A. f is the swing frequency of the metal ball, A is the swing amplitude, d is the average displacement of the metal ball, c is the speed of sound in air, and the sinc function satisfies sinc(x)=sin(π·x) / (π·x); Step S42: Play the output audio in a weak echo chamber and sample the sound pressure at different distances. The output audio is synthesized audio, and the sound pressure is related to the set volume amplitude, frequency, and distance. The user-set volume amplitude and output audio frequency are known quantities. Based on the sampling results, a function G(r) is fitted between the output audio sound pressure and distance, where r represents the distance between the user and the sound source. Step S43: Use the convolution reverberation algorithm to superimpose the spherical sound source's swing audio and output audio, and establish a boundary element model with the speaker as the center: ; Where p(r) represents the distance between the user and the center of the sound source, G(r) is the function of the output audio sound pressure and distance, k is the angle constant, and e is the base of the natural logarithm.
[0017] Furthermore, step S5 includes: Step S51. Obtain the user's position coordinates through the UW or infrared sensor, determine the distance between the user and the speaker, substitute it into the boundary element model, obtain the sound pressure at the user's location, and simultaneously obtain the user's playback settings for the original audio and initialize the audio volume; Step S52. Dynamically adjust the main speaker volume based on the volume set by the user so that the main speaker volume satisfies: ; Where p is the sound pressure at the user's location after adjusting the main speaker volume, p0 is the sound pressure at the user's location before adjusting the main speaker volume, Gp is the sound pressure increment caused by adjusting the main speaker volume, Q1 and Q2 represent the power of the ambient noise and the power of the output audio, respectively. Adjust the volume of the main speaker so that the value of p is equal to the sound pressure of the set volume at a reference position, which is preset by the user.
[0018] Compared with the prior art, the present invention has the following beneficial effects: The present invention sets a weak echo chamber in the speaker, uses a recording device to introduce ambient noise into the weak echo chamber, collects the audio in the weak echo chamber and provides a carrier for spectrum analysis, thereby obtaining a noise power spectrum, thereby separating the sound from the background noise. It can filter out the background noise, present the dynamic range of the music more clearly, and reduce the need for users to turn up the volume.
[0019] The present invention can establish a field intensity polar coordinate system with the speaker as the center, collect audio echo spectra, scale down the echoes and play them in the echo chamber, use the speaker swing ball to model the field intensity in the echo chamber, determine the reverberation swing speed and reverberation pulse speed of the ball source, realize adaptive adjustment of the sound effect, reduce fatigue after long-term listening, expand the use scenarios of the speaker, improve the intelligence level of the speaker, and enhance the sound quality and user experience.
[0020] The present invention can determine the sound pressure attenuation according to the noise power spectrum and modulate the volume of the original audio so that the superimposed sound pressure at the user's location minus the sound pressure attenuation is equal to the sound pressure at the user's original volume setting, thereby reducing the impact of reverberation on the user's listening to audio, improving the performance of the speaker in complex environments, maintaining a comfortable sound pressure level, and reducing hearing damage caused by echo. BRIEF DESCRIPTION OF THE DRAWINGS
[0021] The accompanying drawings are used to provide a further understanding of the present invention and constitute a part of the specification. Together with the embodiments of the present invention, they are used to explain the present invention and do not constitute a limitation of the present invention. In the accompanying drawings: Figure 1 This is a schematic structural diagram of an artificial intelligence-based adaptive ambient speaker control system of the present invention; Figure 2 This is a schematic diagram of the steps of an artificial intelligence-based adaptive environmental speaker control method of the present invention. DETAILED DESCRIPTION
[0022] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.
[0023] See also Figure 1 , the present invention provides a technical solution: an adaptive environmental speaker control system based on artificial intelligence, comprising: a noise acquisition module, a spectrum analysis module, an echo test module, a field strength modeling module and a sound effect adjustment module; The noise collection module is used to set up a weak echo chamber in the speaker, which is soundproofed from the outside world, and play standard audio at a fixed volume in the weak echo chamber. The standard audio is white noise or a linear frequency modulation signal. Before the audio is played, a microphone recording device is used to obtain ambient noise, mix the ambient noise into the audio in the weak echo chamber, and collect the audio in the weak echo chamber as mixed audio. The mixed audio spectrum is uploaded to the speaker processing chip; The noise collection module includes: an echo chamber unit and a noise recording unit; The echo chamber unit uses an inner wall of sound-absorbing material to suppress internal sound wave reflection, and the cavity volume meets the Helmholtz resonance condition to ensure non-destructive testing of sound; The noise recording unit is composed of a spherical array of omnidirectional microphones arranged in a speaker housing and is used to collect spatial echo signals.
[0024] The spectrum analysis module is used to process the mixed audio spectrum through short-time Fourier transform, compare the standard audio with the mixed audio, separate the noise components using spectral subtraction, output the noise spectrum, obtain the audio to be played from the speaker, divide the audio to be played into time segments, subtract the original noise from the spectrum of each segment of the audio to be played, and output the processed audio to the voice coil magnetic field; The spectrum analysis module includes: a mixed extraction unit and an audio processing unit; The mixing extraction unit is used to select full-band white noise or linear frequency modulation signal to mix standard audio and noise; The audio processing unit is used to separate the noise components and process the original audio at the same time so that the environmental noise is included in the audio composition.
[0025] The echo test module is used to establish a field strength polar coordinate system with the speaker as the center based on the volume and the spectrum of the output audio, collect the reverberation audio of the audio echo and the ambient noise with a recording device, use an adaptive filter to eliminate the direct sound, and then use the noise spectrum to clean the noise to obtain the echo spectrum, and play the echo spectrum in the echo chamber; The echo test module includes: a noise suppression unit, a speaker coordinate unit and an echo spectrum unit; The noise suppression unit is used to subtract the noise power spectrum and the direct sound spectrum from the reverberant audio to retain the pure echo component; The speaker coordinate unit is used to determine the echo angle by estimating the direction of arrival and to construct a polar coordinate sound field model with the speaker as the center; The echo spectrum unit is used to reduce the amplitude of the echo spectrum and play the echo spectrum in the echo chamber.
[0026] The field intensity modeling module is used to calculate the swing speed, angle and pulse interval of the ball source based on the frequency spectrum in the echo chamber, so that the sound emitted by the ball source is consistent with the frequency spectrum in the echo chamber. A micro servo motor is used to drive the metal ball inside the speaker to swing, simulating the movement of the sound field scatterer. The swing frequency and swing amplitude are monitored in real time by an accelerometer. The convolution reverberation algorithm is used to superimpose the swing of the ball source on the audio playback. The ball source is regarded as a secondary sound source to calculate the sound pressure distribution at the field point, and a BEM boundary element model is established. The field intensity modeling module includes: a field point sound pressure unit, a user positioning unit and a spherical source reverberation unit; The field point sound pressure unit is used to calculate the sound characteristics of the ball source according to the echo spectrum, so that the sound of the ball source is consistent with the echo; The user positioning unit is used to obtain the user's position coordinates and optimize the user's directional hearing sense; The spherical source reverberation unit is used to adjust the reverberation decay time to match the environmental characteristics and establish a BEM boundary element model.
[0027] The sound effect adjustment module is used to obtain the user position coordinates through UW or infrared sensors, map them to the sound pressure model, calculate the masking threshold according to the noise power spectrum, and dynamically adjust the main speaker EQ based on the volume set by the user to make the sound pressure at the user position equal to the sound pressure at the reference position with the set volume.
[0028] The sound effect adjustment module includes: a user interaction unit and an audio modulation unit; The user interaction unit is used to obtain the user's playback settings for the original audio, initialize the audio volume, and calculate the sound pressure at the reference position; The audio modulation unit is used to dynamically adjust the playback volume of the main speaker to ensure that the output sound pressure is consistent with the set sound pressure.
[0029] like Figure 2 As shown, an artificial intelligence-based adaptive ambient speaker control method includes the following steps: Step S1. A weak echo chamber is set up in the speaker box to be soundproofed from the outside world. A spherical sound source is set up in the weak echo chamber. Standard audio is played at a fixed volume in the weak echo chamber. External ambient noise is introduced using a recording device to extract the spectrum of the mixed audio in the weak echo chamber. Step S1 includes: Step S11. A weak echo chamber is provided inside the speaker. The weak echo chamber uses an inner wall made of sound-absorbing material to suppress internal sound wave reflections. The chamber volume satisfies the Helmholtz resonance condition. An omnidirectional microphone spherical array is arranged on the speaker housing. The omnidirectional microphone spherical array has two modes: capturing ambient noise and amplifying sound. Step S12: Play standard audio in the weak echo chamber. The standard audio is white noise or linear frequency modulation signal. The playback volume is controlled at 30-40dB to avoid interfering with the main speaker output. Step S13: At regular intervals, external environmental noise is introduced into the weak echo chamber, mixed with the standard audio to obtain mixed audio, and the spectrum of the mixed audio is obtained through a spectrum analysis chip.
[0030] Step S2. Process the mixed audio using a short-time Fourier transform, separate the noise component from the mixed audio using spectral subtraction, output the noise spectrum, inversely transform the noise spectrum to obtain the original noise, subtract the original noise from the spectrum of the audio to be played on the speaker to obtain the output audio, and control the speaker to play the output audio; Step S2 includes: Step S21. Process the mixed audio spectrum by short-time Fourier transform, subtract the mixed audio spectrum from the standard audio spectrum to obtain a noise spectrum, and inversely transform the noise spectrum to obtain the original noise; Step S22. Acquire the audio to be played from the speaker in real time, subtract the original noise from the spectrum of the audio to be played, obtain the output audio, make the ambient noise participate in the audio composition, output the processed audio to the voice coil magnetic field, and control the speaker to play the output audio.
[0031] Step S3. Using a recording device to capture the reverberation of the audio echo and ambient noise, the original noise and direct sound are removed from the reverberation audio to obtain the echo spectrum. A micro-servo motor is used to drive the metal ball inside the speaker to swing, simulating the motion of a sound field scatterer. The swing frequency and amplitude of the sphere sound source are adjusted so that the sound pressure emitted by the sphere source at a fixed distance is consistent with the sound pressure of the echo spectrum. Step S3 includes: Step S31. After the speaker plays audio, reverberant audio is acquired through an omnidirectional microphone spherical array. The reverberant audio includes: direct sound from the speaker, ambient noise, and echo. The direct sound in the reverberant audio is eliminated using an adaptive filter. The ambient noise is then removed from the reverberant audio using spectral subtraction to obtain an echo spectrum. Step S32. Use a micro servo motor to drive the metal ball inside the speaker to swing, forming a spherical sound source, collect the sound pressure of the spherical sound source and the sound pressure of the echo spectrum at a fixed position in the weak echo cavity, adjust the swing frequency and swing amplitude of the spherical sound source, and make the sound pressure of the spherical sound source equal to the sound pressure of the echo spectrum.
[0032] Step S4. Using an accelerometer to monitor the swing frequency and amplitude of the spherical sound source in real time, playing the output audio in a weak echo chamber, obtaining the sound pressure distribution of the output audio based on the sound pressure sampling data in the weak echo chamber, superimposing the swing audio of the spherical sound source on the output audio using a convolution reverberation algorithm, and establishing a boundary element model centered on the speaker; Step S4 includes: Step S41. Use an accelerometer to monitor the swing frequency and swing amplitude of the spherical sound source in real time, and substitute the swing frequency and swing amplitude of the spherical sound source into the sound source transfer function: ; Where H(A,f) is the sound source transfer function, which represents the change in sound pressure of the spherical sound source when the swing frequency is f and the swing amplitude is A. f is the swing frequency of the metal ball, A is the swing amplitude, d is the average displacement of the metal ball, c is the speed of sound in air, and the sinc function satisfies sinc(x)=sin(π·x) / (π·x); Step S42: Play the output audio in a weak echo chamber and sample the sound pressure at different distances. The output audio is synthesized audio, and the sound pressure is related to the set volume amplitude, frequency, and distance. The user-set volume amplitude and output audio frequency are known quantities. Based on the sampling results, a function G(r) is fitted between the output audio sound pressure and distance, where r represents the distance between the user and the sound source. Step S43: Use the convolution reverberation algorithm to superimpose the spherical sound source's swing audio and output audio, and establish a boundary element model with the speaker as the center: ; Where p(r) represents the distance between the user and the center of the sound source, G(r) is the function of the output audio sound pressure and distance, k is the angle constant, and e is the base of the natural logarithm.
[0033] Step S5. Obtain the user position coordinates and map them to the boundary element model to obtain the actual sound pressure at the user's position, and dynamically adjust the main speaker volume based on the volume set by the user so that the actual sound pressure at the user's position is equal to the sound pressure at the reference position with the set volume.
[0034] Step S5 includes: Step S51. Obtain the user's position coordinates through the UW or infrared sensor, determine the distance between the user and the speaker, substitute it into the boundary element model, obtain the sound pressure at the user's location, and simultaneously obtain the user's playback settings for the original audio and initialize the audio volume; Step S52. Dynamically adjust the main speaker volume based on the volume set by the user so that the main speaker volume satisfies: ; Where p is the sound pressure at the user's location after adjusting the main speaker volume, p0 is the sound pressure at the user's location before adjusting the main speaker volume, Gp is the sound pressure increment caused by adjusting the main speaker volume, Q1 and Q2 represent the power of the ambient noise and the power of the output audio, respectively. Adjust the volume of the main speaker so that the value of p is equal to the sound pressure of the set volume at a reference position, which is preset by the user.
[0035] Example: A spherical sound source vibrates at a frequency of 5 Hz and an amplitude of 0.2 m. The sound pressure variation H(5, 0.2) is calculated, and the function between the output audio sound pressure and distance is fitted as G(r) = 100 - 0.001 e r ,The distance between the user and the speaker is 10m, and the angle constant is 0. Then the actual sound pressure at the user's location is p0=4.6+25·H(5,0.2). Adjust the speaker volume to make the actual sound pressure consistent with the set sound pressure.
[0036] It should be noted that, in this document, relational terms such as first and second, etc., are used only to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the terms "comprises," "comprising," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that includes a list of elements includes not only those elements but also other elements not explicitly listed, or elements inherent to such process, method, article, or apparatus.
[0037] Finally, it should be noted that the above descriptions are merely preferred embodiments of the present invention and are not intended to limit the present invention. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art will be able to modify the technical solutions described in the aforementioned embodiments or substitute equivalents for some of the technical features. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of the present invention shall be included within the scope of protection of the present invention.
Claims
1. An adaptive environmental speaker control method based on artificial intelligence, characterized in that: The method comprises the following steps: Step S1. A weak echo chamber is set up in the speaker box to be soundproofed from the outside world. A spherical sound source is set up in the weak echo chamber. Standard audio is played at a fixed volume in the weak echo chamber. External ambient noise is introduced using a recording device to extract the spectrum of the mixed audio in the weak echo chamber. Step S2. Process the mixed audio using a short-time Fourier transform, separate the noise component from the mixed audio using spectral subtraction, output the noise spectrum, inversely transform the noise spectrum to obtain the original noise, subtract the original noise from the spectrum of the audio to be played on the speaker to obtain the output audio, and control the speaker to play the output audio; Step S3. Using a recording device to capture the reverberation of the audio echo and ambient noise, the original noise and direct sound are removed from the reverberation audio to obtain the echo spectrum. A micro-servo motor is used to drive the metal ball inside the speaker to swing, simulating the motion of a sound field scatterer. The swing frequency and amplitude of the sphere sound source are adjusted so that the sound pressure emitted by the sphere source at a fixed distance is consistent with the sound pressure of the echo spectrum. Step S4. Using an accelerometer to monitor the swing frequency and amplitude of the spherical sound source in real time, playing the output audio in a weak echo chamber, obtaining the sound pressure distribution of the output audio based on the sound pressure sampling data in the weak echo chamber, superimposing the swing audio of the spherical sound source on the output audio using a convolution reverberation algorithm, and establishing a boundary element model centered on the speaker; Step S5. Obtain the user position coordinates and map them to the boundary element model to obtain the actual sound pressure at the user's position, and dynamically adjust the main speaker volume based on the volume set by the user so that the actual sound pressure at the user's position is equal to the sound pressure at the reference position with the set volume.
2. The method for controlling an adaptive ambient sound box based on artificial intelligence according to claim 1, wherein: Step S1 includes: Step S11. A weak echo chamber is provided inside the speaker. The weak echo chamber uses an inner wall made of sound-absorbing material to suppress internal sound wave reflections. The chamber volume satisfies the Helmholtz resonance condition. An omnidirectional microphone spherical array is arranged on the speaker housing. The omnidirectional microphone spherical array has two modes: capturing ambient noise and amplifying sound. Step S12: Play standard audio in the weak echo chamber. The standard audio is white noise or linear frequency modulation signal. The playback volume is controlled at 30-40dB to avoid interfering with the main speaker output. Step S13: At regular intervals, external environmental noise is introduced into the weak echo chamber, mixed with the standard audio to obtain mixed audio, and the spectrum of the mixed audio is obtained through a spectrum analysis chip.
3. The method for controlling an adaptive ambient sound box based on artificial intelligence according to claim 2, wherein: Step S2 includes: Step S21. Process the mixed audio spectrum by short-time Fourier transform, subtract the mixed audio spectrum from the standard audio spectrum to obtain a noise spectrum, and inversely transform the noise spectrum to obtain the original noise; Step S22: Acquire the audio to be played from the speaker in real time, subtract the original noise from the spectrum of the audio to be played, obtain the output audio, include the ambient noise in the audio composition, output the processed audio to the voice coil magnetic field, and control the speaker to play the output audio; Step S3 includes: Step S31. After the speaker plays audio, reverberant audio is acquired through an omnidirectional microphone spherical array. The reverberant audio includes: direct sound from the speaker, ambient noise, and echo. The direct sound in the reverberant audio is eliminated using an adaptive filter. The ambient noise is then removed from the reverberant audio using spectral subtraction to obtain an echo spectrum. Step S32. Use a micro servo motor to drive the metal ball inside the speaker to swing, forming a spherical sound source, collect the sound pressure of the spherical sound source and the sound pressure of the echo spectrum at a fixed position in the weak echo cavity, adjust the swing frequency and swing amplitude of the spherical sound source, and make the sound pressure of the spherical sound source equal to the sound pressure of the echo spectrum.
4. The method for controlling an adaptive ambient sound box based on artificial intelligence according to claim 3, wherein: Step S4 include: Step S41. Use an accelerometer to monitor the swing frequency and swing amplitude of the spherical sound source in real time, and substitute the swing frequency and swing amplitude of the spherical sound source into the sound source transfer function: ; Where H(A,f) is the sound source transfer function, which represents the change in sound pressure of the spherical sound source when the swing frequency is f and the swing amplitude is A. f is the swing frequency of the metal ball, A is the swing amplitude, d is the average displacement of the metal ball, c is the speed of sound in air, and the sinc function satisfies sinc(x)=sin(π·x) / (π·x); Step S42: Play the output audio in a weak echo chamber and sample the sound pressure at different distances. The output audio is synthesized audio, and the sound pressure is related to the set volume amplitude, frequency, and distance. The user-set volume amplitude and output audio frequency are known quantities. Based on the sampling results, a function G(r) is fitted between the output audio sound pressure and distance, where r represents the distance between the user and the sound source. Step S43: Use the convolution reverberation algorithm to superimpose the spherical sound source's swing audio and output audio, and establish a boundary element model with the speaker as the center: ; Where p(r) represents the distance between the user and the center of the sound source, G(r) is the function of the output audio sound pressure and distance, k is the angle constant, and e is the base of the natural logarithm.
5. The method for controlling an adaptive ambient sound box based on artificial intelligence according to claim 4, characterized in that: Step S5 includes: Step S51. Obtain the user's position coordinates through the UW or infrared sensor, determine the distance between the user and the speaker, substitute it into the boundary element model, obtain the sound pressure at the user's location, and simultaneously obtain the user's playback settings for the original audio and initialize the audio volume; Step S52. Dynamically adjust the main speaker volume based on the volume set by the user so that the main speaker volume satisfies: ; Where p is the sound pressure at the user's location after adjusting the main speaker volume, p0 is the sound pressure at the user's location before adjusting the main speaker volume, Gp is the sound pressure increment caused by adjusting the main speaker volume, Q1 and Q2 represent the power of the ambient noise and the power of the output audio, respectively. Adjust the volume of the main speaker so that the value of p is equal to the sound pressure of the set volume at a reference position, which is preset by the user.
6. An adaptive ambient speaker control system based on artificial intelligence, characterized in that: The system includes the following modules: noise acquisition module, spectrum analysis module, echo test module, field strength modeling module and sound effect adjustment module; The noise collection module is used to set up a weak echo chamber in the speaker, which is soundproofed from the outside world, and play standard audio at a fixed volume in the weak echo chamber. The standard audio is white noise or a linear frequency modulation signal. Before the audio is played, a microphone recording device is used to obtain ambient noise, mix the ambient noise into the audio in the weak echo chamber, and collect the audio in the weak echo chamber as mixed audio. The mixed audio spectrum is uploaded to the speaker processing chip; The spectrum analysis module is used to process the mixed audio spectrum through short-time Fourier transform, compare the standard audio with the mixed audio, separate the noise components using spectral subtraction, output the noise spectrum, obtain the audio to be played from the speaker, divide the audio to be played into time segments, subtract the original noise from the spectrum of each segment of the audio to be played, and output the processed audio to the voice coil magnetic field; The echo test module is used to establish a field strength polar coordinate system with the speaker as the center based on the volume and the spectrum of the output audio, collect the reverberation audio of the audio echo and the ambient noise with a recording device, use an adaptive filter to eliminate the direct sound, and then use the noise spectrum to clean the noise to obtain the echo spectrum, and play the echo spectrum in the echo chamber; The field intensity modeling module is used to calculate the swing speed, angle and pulse interval of the ball source based on the frequency spectrum in the echo chamber, so that the sound emitted by the ball source is consistent with the frequency spectrum in the echo chamber. A micro servo motor is used to drive the metal ball inside the speaker to swing, simulating the movement of the sound field scatterer. The swing frequency and swing amplitude are monitored in real time by an accelerometer. The convolution reverberation algorithm is used to superimpose the swing of the ball source on the audio playback. The ball source is regarded as a secondary sound source to calculate the sound pressure distribution at the field point, and a BEM boundary element model is established. The sound effect adjustment module is used to obtain the user position coordinates through UW or infrared sensors, map them to the sound pressure model, calculate the masking threshold according to the noise power spectrum, and dynamically adjust the main speaker EQ based on the volume set by the user to make the sound pressure at the user position equal to the sound pressure at the reference position with the set volume.
7. The artificial intelligence-based adaptive ambient sound box control system according to claim 6, characterized in that: The noise collection module includes: an echo chamber unit and a noise recording unit; The echo chamber unit uses an inner wall of sound-absorbing material to suppress internal sound wave reflection, and the cavity volume meets the Helmholtz resonance condition to ensure non-destructive testing of sound; The noise recording unit is composed of a spherical array of omnidirectional microphones arranged in a speaker housing, and is used to collect spatial echo signals; The spectrum analysis module includes: a mixed extraction unit and an audio processing unit; The mixing extraction unit is used to select full-band white noise or linear frequency modulation signal to mix standard audio and noise; The audio processing unit is used to separate the noise components and process the original audio at the same time so that the environmental noise is included in the audio composition.
8. The artificial intelligence-based adaptive ambient sound box control system according to claim 7, characterized in that: The echo test module includes: a noise suppression unit, a speaker coordinate unit and an echo spectrum unit; The noise suppression unit is used to subtract the noise power spectrum and the direct sound spectrum from the reverberant audio to retain the pure echo component; The speaker coordinate unit is used to determine the echo angle by estimating the direction of arrival and to construct a polar coordinate sound field model with the speaker as the center; The echo spectrum unit is used to reduce the amplitude of the echo spectrum and play the echo spectrum in the echo chamber.
9. The artificial intelligence-based adaptive ambient sound box control system according to claim 8, characterized in that: The field intensity modeling module includes: a field point sound pressure unit, a user positioning unit and a spherical source reverberation unit; The field point sound pressure unit is used to calculate the sound characteristics of the ball source according to the echo spectrum, so that the sound of the ball source is consistent with the echo; The user positioning unit is used to obtain the user's position coordinates and optimize the user's directional hearing sense; The spherical source reverberation unit is used to adjust the reverberation decay time to match the environmental characteristics and establish a BEM boundary element model.
10. The artificial intelligence-based adaptive ambient sound box control system according to claim 9, characterized in that: The sound effect adjustment module includes: a user interaction unit and an audio modulation unit; The user interaction unit is used to obtain the user's playback settings for the original audio, initialize the audio volume, and calculate the sound pressure at the reference position; The audio modulation unit is used to dynamically adjust the playback volume of the main speaker to ensure that the output sound pressure is consistent with the set sound pressure.