Sound box configuration optimization method, system and device and storage medium
By controlling the speaker to play test sounds and collecting sound wave data, the speaker configuration is optimized, solving the problems of low efficiency and insufficient accuracy of audio device configuration parameters in existing technologies, and achieving uniform audio quality and a personalized listening experience.
Patent Information
- Application Number
- CN202511278273.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-08
- Publication Date
- 2025-11-18
AI Technical Summary
Existing audio calibration techniques are inefficient and inaccurate in determining audio device configuration parameters, making it difficult to achieve uniform audio quality across different regions.
By controlling the speaker to play multiple sets of test sounds, sound wave data is collected, acoustic test results are determined based on the sound wave data, and recommended adjustment parameters for the speaker are generated. The speaker configuration is then optimized in combination with spatial and environmental characteristics.
It improves the efficiency and accuracy of determining speaker configuration parameters, ensures the uniformity and consistency of audio quality in space, and provides a personalized listening experience.
Smart Images

Figure CN120980435A_ABST
Abstract
Description
Technical Field
[0001] This specification relates to the field of audio technology, and in particular to a speaker configuration optimization method, system, device, and storage medium. Background Technology
[0002] When configuring audio equipment in a space, it's sometimes difficult to effectively calibrate audio playback quality, such as inconsistent audio quality across different areas that is hard to standardize. Adjusting the physical position or angle of the audio playback equipment and performing audio calibration is an effective solution. Traditional audio calibration techniques, such as Audyssey or Dirac Live, primarily use room equalization (REQ) algorithms for audio calibration. This algorithm analyzes room acoustic defects by playing test tones at the listening position and uses digital signal processing (DSP) to adjust the device's acoustic parameters (such as frequency and gain) to compensate for these defects, thereby optimizing the listening experience at specific locations.
[0003] However, the aforementioned audio calibration methods are generally based on a crucial assumption: that the physical placement and angle of the audio playback device are already reasonably optimized. Existing methods often employ an exhaustive approach, trying every possible method to determine the optimal placement and angle. This method is extremely inefficient and its accuracy needs improvement.
[0004] Therefore, it is desirable to provide a speaker configuration optimization method to improve the efficiency and accuracy of determining audio device configuration parameters. Summary of the Invention
[0005] This specification provides one or more embodiments of a speaker configuration optimization method. The method includes: controlling a speaker to play multiple sets of test sounds, and controlling an audio acquisition device to receive the multiple sets of test sounds; acquiring multiple sets of sound wave data corresponding to the multiple sets of test sounds from the audio acquisition device; determining acoustic test results based on the multiple sets of sound wave data; and generating recommended adjustment parameters for the speaker in response to the acoustic test results not meeting preset conditions.
[0006] Optionally, the method further includes: determining at least one preset acoustic parameter based on spatial and environmental characteristics; and determining the plurality of test sounds based on the at least one preset acoustic parameter.
[0007] Optionally, the audio acquisition device receives the multiple sets of test sounds at multiple listening points; determining the acoustic test result based on the multiple sets of sound wave data includes: determining the acoustic feature fit degree corresponding to the combination of each listening point and preset acoustic parameters based on the multiple sets of sound wave data; determining the comprehensive fit degree based on the acoustic feature fit degree, and determining the comprehensive fit degree as the acoustic test result.
[0008] Optionally, the step of generating recommended adjustment parameters for the speaker in response to the acoustic test results not meeting preset conditions includes: generating a sound field map based on the multiple sets of sound wave data; determining the speaker's distribution data based on the sound field map; and generating the recommended adjustment parameters based on the distribution data and overall fit.
[0009] This specification also provides a speaker configuration optimization system through one or more embodiments. The system includes: a control module configured to control a speaker to play multiple sets of test sounds and to control an audio acquisition device to receive the multiple sets of test sounds; an acquisition module configured to acquire multiple sets of sound wave data corresponding to the multiple sets of test sounds from the audio acquisition device; a determination module configured to determine acoustic test results based on the multiple sets of sound wave data; and a parameter generation module configured to generate recommended adjustment parameters for the speaker in response to the acoustic test results not meeting preset conditions.
[0010] Optionally, the control module is further configured to: determine at least one preset acoustic parameter based on spatial and environmental characteristics; and determine the multiple sets of test sounds based on the at least one preset acoustic parameter.
[0011] Optionally, the determining module is further configured to: determine the acoustic feature fit degree corresponding to each combination of the listening point and the preset acoustic parameters based on the multiple sets of sound wave data; determine the comprehensive fit degree based on the acoustic feature fit degree; and determine the comprehensive fit degree as the acoustic test result.
[0012] Optionally, the generation module is further configured to: generate a sound field map based on the multiple sets of sound wave data; determine the distribution data of the speaker based on the sound field map; and generate the recommended adjustment parameters based on the distribution data and the overall fit.
[0013] One or more embodiments of this specification also provide a speaker configuration optimization apparatus, the apparatus including a processor for executing the speaker configuration optimization method described in any of the above embodiments.
[0014] This specification also provides a computer-readable storage medium that stores computer instructions. When a computer reads the computer instructions in the storage medium, the computer executes the speaker configuration optimization method described in any of the above embodiments. Attached Figure Description
[0015] This specification will be further described by way of exemplary embodiments, which will be described in detail with reference to the accompanying drawings. These embodiments are not limiting; in these embodiments, the same reference numerals denote the same structures, wherein: Figure 1These are exemplary module diagrams of a speaker configuration optimization system according to some embodiments of this specification; Figure 2 This is an exemplary flowchart of a speaker configuration optimization method according to some embodiments of this specification; Figure 3 This is an exemplary flowchart illustrating the determination of recommended adjustment parameters according to some embodiments of this specification. Detailed Implementation
[0016] To more clearly illustrate the technical solutions of the embodiments in this specification, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are merely some examples or embodiments of this specification. For those skilled in the art, these drawings can be applied to other similar scenarios without creative effort. Unless obvious from the context or otherwise specified, the same reference numerals in the drawings represent the same structures or operations.
[0017] It should be understood that the terms “system,” “device,” “unit,” and / or “module” used herein are one way to distinguish different components, elements, parts, sections, or assemblies at different levels. However, if other terms can achieve the same purpose, they may be replaced by other expressions.
[0018] As indicated in this specification and claims, unless the context clearly indicates otherwise, the words "a," "an," "an," and / or "the" do not specifically refer to the singular and may also include the plural. Generally speaking, the terms "comprising" and "including" only indicate the inclusion of expressly identified steps and elements, which do not constitute an exclusive list, and the method or apparatus may also include other steps or elements.
[0019] Flowcharts are used in this specification to illustrate the operations performed by the system according to embodiments of this specification. It should be understood that the preceding or following operations are not necessarily performed in exact order. Instead, the steps can be processed in reverse order or simultaneously. Furthermore, other operations can be added to these processes, or one or more steps can be removed from them.
[0020] Figure 1 These are exemplary module diagrams of a speaker configuration optimization system according to some embodiments of this specification. Figure 1 As shown, system 100 includes a control module 110, an acquisition module 120, a determination module 130, and a parameter generation module 140.
[0021] The control module 110 can be configured to control the speaker to play multiple sets of test sounds, and to control the audio acquisition device to receive multiple sets of test sounds.
[0022] In some embodiments, the control module is further configured to: determine at least one preset acoustic parameter based on spatial and environmental characteristics; and determine multiple sets of test sounds based on the at least one preset acoustic parameter.
[0023] The acquisition module 120 can be configured to acquire multiple sets of sound wave data corresponding to multiple sets of test sounds from the audio acquisition device.
[0024] The determination module 130 can be configured to determine acoustic test results based on multiple sets of acoustic wave data.
[0025] In some embodiments, the determining module is further configured to: determine the acoustic feature fit degree corresponding to the combination of each listening point and preset acoustic parameters based on multiple sets of sound wave data; determine the comprehensive fit degree based on the acoustic feature fit degree; and determine the comprehensive fit degree as the acoustic test result.
[0026] The parameter generation module 140 can be configured to generate recommended adjustment parameters for the speaker in response to acoustic test results not meeting preset conditions.
[0027] In some embodiments, the parameter generation module 140 is further configured to: generate a sound field map based on multiple sets of sound wave data; determine the speaker distribution data based on the sound field map; and generate recommended adjustment parameters based on the distribution data and overall fit.
[0028] For more information on test sound, audio acquisition equipment, spatial characteristics, environmental characteristics, preset acoustic parameters, sound wave data, acoustic test results, listening point, acoustic feature fit, overall fit, recommended adjustment parameters, preset conditions, sound field diagrams, and distribution data, please refer to [link to relevant documentation]. Figure 2 , Figure 3 And its related descriptions.
[0029] It should be noted that the above description of the audio device layout system and its modules is for convenience only and should not be construed as limiting this specification to the embodiments described. It is understood that those skilled in the art, after understanding the principles of this system, may arbitrarily combine the various modules or construct subsystems connected to other modules without departing from these principles. For example, the modules may share a single storage module, or each module may have its own independent storage module. Such modifications are all within the scope of this specification.
[0030] Figure 2 This is an exemplary flowchart illustrating a speaker configuration optimization method according to some embodiments of this specification. Figure 2 As shown, process 200 may include steps 210, 220, 230, and 240. In some embodiments, process 200 may be executed by a processor.
[0031] Step 210: Control the speaker to play multiple sets of test sounds, and control the audio acquisition device to receive multiple sets of test sounds.
[0032] Test sounds are audio signals used to measure the acoustic properties of a room. In some embodiments, test sounds are specific audio signals with known acoustic characteristics. For example, test sounds include sinusoidal sweep signals, powder noise, pulse signals, etc.
[0033] In some embodiments, the processor generates test sounds based on specific testing requirements and objectives. For example, the processor can determine the corresponding test sounds based on testing experience, according to specific testing requirements and objectives. Testing requirements and objectives may include: testing whether the playback sound is muffled, unclear, or whether the sound is balanced across different locations, etc.
[0034] In some embodiments, the processor determines at least one preset acoustic parameter based on spatial and environmental characteristics, and determines multiple sets of test sounds based on the at least one preset acoustic parameter.
[0035] Spatial characteristics refer to the attributes of a region's space in terms of its physical layout. For example, spatial characteristics include spatial dimensions, spatial shape, and building materials. Spatial shape can include circular domes, irregular polygons, and long, narrow corridors.
[0036] Environmental characteristics are the environmentally relevant attributes of a regional space. For example, environmental characteristics include environmental noise, environmental temperature, and environmental humidity.
[0037] Preset acoustic parameters refer to the configuration parameters of the devices used to generate test sounds. For example, preset acoustic parameters include resonant frequency compensation parameters and reverberation time compression parameters of the audio signal processing module.
[0038] In some embodiments, the preset acoustic parameters can be obtained by adjusting the parameter configuration of the audio signal processing module through an audio processing program. The audio signal processing module includes a digital equalizer (EQ) module, a resonance compensator, a sound pressure level controller, etc. For example, the adjustment logic of the digital EQ module for adjusting the sound pressure level of the equalizer frequency band is as follows: the gain value corresponding to different frequency bands (such as 31Hz, 62Hz, 125Hz, ..., 16kHz) is set through the audio processing program, and the audio processing program is run to adjust the digital EQ module.
[0039] In some embodiments, the processor acquires spatial features via an image sensor built into the speaker. The image sensor may include a camera, an ultrasonic sensor array, an optical sensor, etc. The processor acquires environmental features by monitoring with temperature and humidity sensors, and analyzes the frequency distribution of background noise in the environment by measuring the background noise spectrum to obtain the environmental noise level.
[0040] In some embodiments, the processor determines at least one preset acoustic parameter based on spatial and environmental characteristics using a second preset table. The second preset table includes the correspondence between different spatial and environmental characteristics and different preset acoustic parameters. For example, in the second preset table, when the sound wave wavelength is close to the room size, which easily generates standing waves and resonant frequency shifts, resulting in significant non-uniformity in the low-frequency response, the corresponding preset acoustic parameters may include: resonant frequency compensation, reverberation time compression, and low-frequency attenuation slope adjustment.
[0041] Resonance frequency compensation attenuates specific low-frequency standing wave frequencies (such as 63Hz, 125Hz) of a sound. For example, attenuating a bass signal by -3dB to -5dB to suppress peaks. Reverberation time compression refers to a slight attenuation of the mid-to-high frequency (500Hz-4kHz) range of a sound, such as attenuating by -2dB. Low-frequency attenuation slope adjustment refers to using a low-cut filter (HPF) to raise the cutoff frequency of a sound. For example, adjusting the cutoff frequency from 80Hz to 100Hz.
[0042] The second preset table can be constructed based on the preset acoustic parameters of historical data when the playback effect was good, as well as the corresponding spatial and environmental characteristics at that time.
[0043] In some embodiments, preset acoustic parameters can also be determined in other ways, such as by analyzing spatial and environmental features using a machine learning model to adaptively generate preset acoustic parameters.
[0044] In some embodiments, the processor determines multiple sets of test sounds based on at least one preset acoustic parameter.
[0045] For example, the processor can determine multiple sets of test sounds based on the type of preset acoustic parameters. If the preset acoustic parameters include resonant frequency compensation, the test sound is powder noise; if the preset acoustic parameters include reverberation time compression, the test sound is a pulse signal; if the preset acoustic parameters include frequency response adjustment, the test sound is a sinusoidal sweep signal.
[0046] In some embodiments, the processor may also determine the test sound by other methods based on preset acoustic parameters, such as dynamically generating the corresponding test sound based on the type and value of the preset acoustic parameters.
[0047] Some embodiments in this specification combine spatial and environmental characteristics to determine preset acoustic parameters, making the test sound more closely match the needs of actual scenarios and avoiding errors in general testing; the test signal (e.g., frequency range, sound pressure level) is dynamically adjusted according to the real-time environment to ensure the accuracy of sound wave data acquisition and provide a more reliable basis for speaker adjustment.
[0048] In some embodiments, the number of preset acoustic parameters is related to the number of speakers.
[0049] In some embodiments, the processor dynamically adjusts the number of preset acoustic parameters based on the number of speakers. The more speakers there are, the longer the fitting evaluation time becomes. To shorten the overall evaluation time, the processor can select fewer preset acoustic parameters. For example, the processor can choose to test only the high-frequency portion of the preset acoustic parameters. The processor can have a built-in preset lookup table that directly matches the corresponding preset acoustic parameter set based on the number of detected speakers. The preset lookup table can be set based on historical experience or simulation experiments.
[0050] In some embodiments, the correlation between the number of preset acoustic parameters and the number of speakers can also be established in other ways. For example, using a machine learning model, when the number of speakers is large, the model can automatically identify and focus only on the core acoustic parameters that have the greatest impact on sound quality. This can further optimize processing efficiency while ensuring calibration effectiveness.
[0051] In some embodiments of this specification, the system can automatically select to use a simple or complex set of acoustic parameters (e.g., the number of frequency bands) based on the estimated complexity of the problem, thereby optimizing processing efficiency and avoiding unnecessary computational overhead while ensuring calibration effectiveness.
[0052] Audio acquisition equipment refers to devices used to collect and test sound. For example, audio acquisition equipment can include mobile phones, microphones, recording devices, etc.
[0053] In some embodiments, the audio acquisition device may also be a professional microphone array or other acoustic measurement device.
[0054] In some embodiments, the processor controls the speaker to play multiple preset test sounds. For example, the processor can control the speaker to play multiple preset pink noise, sine sweep signal, and pulse signal sequentially in a certain playback order. Alternatively, the processor can randomly play multiple test sounds.
[0055] In some embodiments, the processor can control the speaker to play test sounds in various ways. For example, it can control the playback device to play sounds via wireless signals, wired connections, or by sending voice commands.
[0056] In some embodiments, while multiple sets of test sounds are played, the processor controls an audio acquisition device to receive the played test sounds. The multiple sets of test sounds acquired by the audio acquisition device can be electrical signals.
[0057] Step 220: Obtain multiple sets of sound wave data corresponding to multiple sets of test sounds from the audio acquisition device.
[0058] Sound wave data refers to raw audio data captured and recorded by audio acquisition devices. Sound wave data can be unprocessed or partially processed raw audio data. For example, sound wave data can be a set of audio signals with amplitude values varying over time, presented in the form of sound waves.
[0059] In some embodiments, after receiving multiple sets of test sounds, the audio acquisition device can generate multiple sets of sound wave data by amplifying the signals and performing analog-to-digital conversion. In some embodiments, the processor can acquire multiple sets of test sounds from the audio acquisition device, and then amplify and perform analog-to-digital conversion on the multiple sets of test sounds to obtain multiple sets of sound wave data.
[0060] Multiple sets of acoustic data correspond to multiple sets of test sounds. For example, when the received test sounds include powder noise, a sinusoidal sweep signal, and a pulse signal, the multiple sets of acoustic data include acoustic data corresponding to the powder noise, acoustic data corresponding to the sinusoidal sweep signal, and acoustic data corresponding to the pulse signal. In some embodiments, each set of acoustic data can be labeled with its corresponding test sound type and stored.
[0061] In some embodiments, the processor may acquire sound wave data by means of high-speed sampling directly connected to a memory of the audio acquisition device, or by acquiring it in real time from the audio acquisition device via network transmission.
[0062] Step 230: Determine the acoustic test results based on multiple sets of sound wave data.
[0063] Acoustic test results refer to data used to evaluate the playback effect of test sounds within a space. Acoustic test results include sound pressure level deviations and the presence of standing waves, among other things, for multiple sets of test sound data. Good acoustic test results indicate that the configuration parameters (such as placement position and angle) of each speaker in the current space meet the requirements.
[0064] In some embodiments, the processor determines the acoustic test results by preprocessing multiple sets of acoustic wave data. Preprocessing methods may include short-time Fourier transform (STFT), wavelet transform, etc.
[0065] In some embodiments, for the acoustic wave data corresponding to powder noise, the processor can compare the measured spectrum corresponding to the acoustic wave data with the target frequency response curve (such as the Harman target curve), calculate the sound pressure level deviation (dB) for three key frequency bands of 200Hz, 1kHz, and 5kHz, and determine the acoustic test results based on the sound pressure level deviation. Sound pressure level refers to the pressure change caused by sound waves propagating in the air. Sound pressure level deviation refers to the difference between the sound pressure and a reference sound pressure value (such as the first deviation).
[0066] For example, if the measured value in the 200Hz band is greater than the sum of the target value and the first deviation on the target frequency response curve, the acoustic test result is determined to be excessively strong in the low frequencies. If the measured value in the 1kHz band is less than the difference between the target value and the first deviation on the target frequency response curve, the acoustic test result is determined to be a mid-frequency dip. If the measured value in the 5kHz band is less than the difference between the target value and the first deviation on the target frequency response curve, the acoustic test result is determined to be insufficient in the high frequencies.
[0067] The first deviation refers to the maximum permissible deviation between the measured value on the measured spectrum and the corresponding value on the target frequency response curve, which is determined through empirical preset, such as 3dB.
[0068] Excessive low frequencies refer to excessive energy in the low-frequency range, resulting in a muddy, unclear, and indistinct sound. Mid-frequency depression refers to a significant lack of energy in the mid-frequency range, resulting in a hollow, weak, and unimpactful sound. Insufficient high frequencies refer to insufficient energy in the high-frequency range, resulting in reduced clarity, a muffled sound, and impaired spatial and stereo perception.
[0069] In some embodiments, for a sinusoidal sweep frequency signal, the processor performs peak detection on the acoustic wave data of the sinusoidal sweep frequency signal. If the sound pressure level at a certain frequency point is higher than or equal to the sound pressure level of the adjacent frequency band by a second deviation, and the peak bandwidth is less than 1 / 6 octave, then the acoustic test result is determined to be that a standing wave exists.
[0070] In some embodiments, the processor may also employ other acoustic analysis methods to determine the acoustic test results. Examples include cross-correlation analysis and transfer function measurement. The processor may also analyze other frequency bands or acoustic characteristics to obtain more comprehensive acoustic test results.
[0071] In some embodiments, the audio acquisition device receives multiple sets of test sounds at multiple listening points. Determining the acoustic test result based on the multiple sets of sound wave data includes: determining the acoustic feature fit degree corresponding to the combination of each listening point and preset acoustic parameters based on the multiple sets of sound wave data; determining the comprehensive fit degree based on the acoustic feature fit degree; and determining the comprehensive fit degree as the acoustic test result.
[0072] A listening point is a sampling point selected within a space for collecting audio data. For example, a listening point could be a location where users frequently linger, or the center point of each grid after the space has been divided into grids.
[0073] In some embodiments, the selection of listening points can be based on methods other than user dwell location and grid division. For example, it can be determined by random sampling or by heatmap analysis of user dwell location.
[0074] In some embodiments, multiple listening points are determined based on the time characteristics of the current time period.
[0075] Time characteristics refer to the time attributes corresponding to the current time period. For example, time characteristics can be daytime or nighttime. Other examples include work hours, leisure and entertainment hours, mealtimes, and sleep hours.
[0076] In some embodiments, the processor can determine the listening points based on the time characteristics of the current period. For example, if the current period is daytime, the user's location may mainly be concentrated in the sofa area of the living room, so the listening points can be mainly selected near the sofa to optimize the sound experience in that area. If the current period is nighttime, the user may mainly be concentrated in bed in the bedroom, and the listening points can be selected accordingly near the bed. As another example, if the current period is the user's work period, leisure and entertainment period, mealtime, or sleep period, the processor can determine the areas where the user spends the most time during the corresponding period and set the listening points in the corresponding areas.
[0077] In some embodiments, the processor automatically identifies the current time characteristics based on the usage data of smart home devices or the user's daily routine, and dynamically adjusts the layout of listening points most suitable for that time period accordingly. In some embodiments, the determination of listening points can be based not only on time characteristics, but also on the user's location information (e.g., via Wi-Fi positioning, Bluetooth beacons, etc.) and the user's device usage habits at different times for more refined judgment.
[0078] Some embodiments in this specification, by associating different time periods and listening positions, allow the system to automatically switch and optimize to the corresponding best sound field for users in different scenarios, providing a truly personalized and intelligent listening experience.
[0079] In some embodiments, the audio acquisition device can be configured at multiple preset listening points in a room. For example, three listening points can be set in the sofa area, two in the dining table area, and three in the bedroom to cover the user's main listening needs. The processor can also control a mobile robot equipped with the audio acquisition device to automatically receive sound at the multiple preset listening points.
[0080] In some embodiments, the processor determines the acoustic feature fit degree corresponding to the combination of each listening point and preset acoustic parameters based on multiple sets of sound wave data.
[0081] Acoustic fit refers to the degree to which an audio playback device fits the listening point. For example, acoustic fit can be measured by acoustic error; the larger the acoustic error, the lower the fit. Acoustic error refers to the deviation between the actual measured sound response and the ideal or target response.
[0082] In some embodiments, for each listening point and a combination of preset acoustic parameters, the processor can determine the type of multiple sets of acoustic wave data and determine the acoustic feature fit based on the type of the multiple sets of acoustic wave data. For example, if the acoustic wave data is a sinusoidal sweep signal, the processor obtains the actual frequency response curve through Fast Fourier Transform (FFT) analysis. Then, the processor calculates the root mean square error between this frequency response curve and the target frequency response curve (e.g., a Harman curve) as the audio playback error. As another example, if the acoustic wave data is a pulse signal, the processor calculates the actual reverberation time at the listening point and uses the absolute difference between the actual reverberation time and the preset reverberation time as the reverberation time error. The preset reverberation time can be set based on experience or requirements.
[0083] In some embodiments, the processor determines the acoustic feature fit based on audio playback error and reverberation duration error. For example, the acoustic feature fit can be expressed as [audio playback error, reverberation duration error].
[0084] In some embodiments, acoustic feature fit may also include errors in other acoustic parameters, such as signal-to-noise ratio and distortion. Furthermore, the processor may determine acoustic feature fit through other means, such as fuzzy logic or neural networks.
[0085] In some embodiments, the processor determines the overall fit based on the acoustic feature fit and uses the overall fit as the acoustic test result.
[0086] Overall fit refers to the degree to which an audio playback device is compatible with the entire playback space. Similarly, overall fit can also be measured by acoustic error.
[0087] In some embodiments, the processor can calculate the mean error of multiple listening points and use this mean error as the overall fit. The overall fit can be [mean audio playback error, mean reverberation duration error].
[0088] In some embodiments, the determination of overall fit may also include weighted average, maximum value selection, etc., and may be personalized in combination with user preferences.
[0089] Some embodiments in this specification improve the accuracy and rationality of determining the fit of the audio playback device in the playback space by setting multiple listening points and determining the acoustic feature fit of each listening point based on the sound wave data corresponding to the combination of multiple listening points and preset acoustic parameters. Furthermore, setting multiple listening points provides a data foundation for subsequently determining a playback effect that better meets the user's actual listening needs, avoiding the uncertainty caused by subjective listening perception and ensuring the scientific nature and reliability of the optimization process.
[0090] In some embodiments, the processor determines the overall fit based on the weights corresponding to different listening points and the acoustic feature fit; wherein, the weights corresponding to different listening points are different.
[0091] To more accurately reflect the user's listening experience in real-world scenarios, the system assigns different weights to different listening points. In some embodiments, the processor determines the corresponding weight based on the frequency with which the user stays in different locations. For example, if the user typically spends the most time in the sofa area, the listening point in that area can be assigned a higher weight, such as 0.6. This weighting ensures that sound field optimization prioritizes the listening experience in the user's most frequently occupied location. For less important listening areas, such as the dining table area, the listening points can be assigned lower weights, such as 0.3 or 0.1.
[0092] In some embodiments, the processor determines weights based on the frequency with which the user stays at different locations, combined with factors such as the spatial layout of each location, the placement of the speakers, and the distance from the listening point to the speakers. Listening points that are closer to the speakers and have stronger direct sound can be assigned higher weights.
[0093] In some embodiments, the processor may also determine weights based on other factors. For example, the sound absorption or propagation characteristics of building materials at different locations, or the type of audio that users at different locations need to play. Another example is the use of a machine learning-based adaptive weight allocation method, which dynamically adjusts the weights of each listening point by analyzing user behavior patterns, room heat distribution data, and other factors.
[0094] In some embodiments, the processor determines the overall fit by weighted summation based on the weights of different listening points and the acoustic feature fit. For example, the processor can perform weighted summation on the audio playback error and reverberation duration error of each listening point, and then determine the overall fit based on the weighted overall audio playback error and overall reverberation duration error.
[0095] For example, the overall fit can be expressed as [overall audio playback error, overall reverberation duration error].
[0096] Some embodiments in this specification introduce weights to make sound field optimization more targeted and user-friendly. The system can prioritize the listening effect at the user's most frequently located position while also considering other less important positions. Furthermore, factors such as the spatial layout, building materials, and distance of each listening point are taken into account when determining the weights, thereby achieving optimal user experience and sound field calibration that meets actual requirements in complex environments with multiple listening points.
[0097] Some embodiments in this specification provide an objective and comprehensive standard for evaluating the physical placement of speakers by generating sound field maps and quantifying sound field uniformity. This effectively identifies poor placements that may perform well in the optimal position but cause a significant decline in the experience in other positions, thereby ensuring the overall balance and consistency of the sound field and improving the user experience of the entire listening area.
[0098] Step 240: In response to the acoustic test results not meeting the preset conditions, generate recommended adjustment parameters for the speaker.
[0099] Preset conditions refer to the conditions used to determine whether the acoustic test results meet the requirements. For example, preset conditions could be that the test results do not have problems such as excessively strong low frequencies, insufficient high frequencies, or standing waves.
[0100] In some embodiments, the processor can determine whether there are problems such as excessively strong low frequencies, insufficient high frequencies, and standing waves in the acoustic test results. If none of these exist, the acoustic test results are determined to meet the preset conditions.
[0101] In some embodiments, the preset conditions may further include overall fit satisfaction with fit conditions. Fit conditions include a playback error threshold and a reverberation error threshold. The playback error threshold refers to the condition that the audio playback error needs to meet in the preset overall fit, and the reverberation error threshold refers to the condition that the reverberation duration error needs to meet in the preset overall fit.
[0102] In some embodiments, when the audio playback error in the overall fit exceeds the playback error threshold, or the reverberation duration error exceeds the reverberation error threshold, the processor can determine that the acoustic test results do not meet the preset conditions.
[0103] Recommended adjustment parameters refer to the optimal values for the speaker configuration after adjustment. For example, recommended adjustment parameters include the optimal values for the installation position, orientation angle, and height of each speaker after adjustment. Other examples include speaker EQ settings and phase settings.
[0104] In some embodiments, when the acoustic test results do not meet preset conditions, the processor can retrieve recommended adjustment parameters from a first preset table based on the acoustic test results. The first preset table contains various acoustic test results and corresponding reference adjustment parameters.
[0105] For example, if the low frequencies are too strong in the first preset table, the corresponding reference adjustment parameter could be to move the speaker 10-20cm away from the wall; if the high frequencies are insufficient, the corresponding reference adjustment parameter could be to tilt the speaker down by 5°; if there are standing waves, the corresponding reference adjustment parameter could be to raise the speaker by 10-20cm.
[0106] The first preset table may also include reference adjustment parameters corresponding to different audio playback errors and reverberation duration errors.
[0107] The first preset table can be determined through statistical analysis of historical speaker configuration data and audio calibration data, or generated through simulation experiments. In some embodiments, the first preset table may also include more detailed acoustic test results and their corresponding more precise reference adjustment parameters.
[0108] In some embodiments, the processor can determine speaker distribution data based on the sound field map; and generate recommended adjustment parameters based on the distribution data and overall fit. More details can be found in [link to relevant documentation]. Figure 3 And its related descriptions.
[0109] In some embodiments, the processor may also generate recommended adjustment parameters in other ways. For example, more refined or personalized adjustment recommendations may be automatically generated based on the relationship between acoustic test results and preset conditions using expert systems, machine learning models, or other intelligent algorithms.
[0110] In some embodiments, the preset condition also includes that the sound field difference between multiple listening points is less than a preset difference threshold.
[0111] The preset difference threshold refers to the threshold condition that the sound field difference needs to meet, which can be determined manually based on experience or actual needs.
[0112] Sound field difference refers to the difference in sound emitted by a speaker at different listening points. For example, sound field difference can include the difference between the maximum and minimum sound pressure level, the difference in frequency response between different listening points, or the difference in reverberation time between different listening points. Sound field difference can be used to measure sound pressure level uniformity, frequency response consistency, and so on.
[0113] In some embodiments, the processor generates a sound field map based on multiple sets of sound wave data, and then determines the sound field differences based on the sound field map.
[0114] A sound field diagram is a graph that visualizes the acoustic parameters of a listening area. For example, a sound field diagram can be a two-dimensional or three-dimensional heatmap, where color intensity or height represents the magnitude of acoustic parameters. Acoustic parameters can include sound pressure level, reverberation time, and frequency response.
[0115] In some embodiments, a sound field map can be generated by processing multiple sets of sound wave data collected from multiple listening points. For example, for sound pressure level (SPL), the processor can perform spectral analysis on the sound wave data received at each listening point using Fast Fourier Transform (FFT) to extract SPL information for a specific frequency band (e.g., the mid-to-high frequency band 500Hz-8kHz). Finally, interpolation algorithms such as bilinear interpolation and Kriging interpolation are used to smooth and visualize the discrete SPL data, generating a two-dimensional or three-dimensional sound field map. In this sound field map, different colors or color levels represent different SPL intensities, which can intuitively show the SPL distribution across the entire listening area.
[0116] In some embodiments, the processor can continuously acquire multiple sets of acoustic wave data in a short period of time and process the sound pressure level data using time series analysis methods to eliminate the influence of instantaneous noise and improve the stability and accuracy of the sound field map.
[0117] In some embodiments, the processor can use various methods to determine the sound field difference based on the sound field map. For example, the processor can analyze the sound field map, obtain the sound parameters (such as sound pressure level) corresponding to each listening point in the sound field map, and determine the sound field difference as the difference between the maximum and minimum values of the obtained sound parameters. As another example, the processor can divide the sound field map into multiple sub-regions, calculate the average sound parameters within each sub-region, and then determine the maximum difference between the average sound parameters corresponding to each sub-region as the sound field difference.
[0118] In some embodiments, when the sound field difference is greater than or equal to a preset difference threshold, the processor can determine that the acoustic test results do not meet the preset conditions.
[0119] Some embodiments in this specification measure the uniformity of the sound field across the entire listening area by determining the sound field difference. This can more effectively assess the rationality of the physical placement of the speakers and effectively identify poor placements that, even if the effect is acceptable in the optimal position, will cause a serious decline in the experience in other positions, thus ensuring the overall balance and consistency of the sound field.
[0120] Some embodiments in this manual utilize automated testing of sound playback and sound wave data acquisition to accurately pinpoint speaker placement issues without requiring human experience. This significantly reduces the setup threshold and time cost, improving the efficiency and accuracy of speaker configuration. Based on acoustic test results, dynamic adjustment parameters such as position and orientation are generated, allowing for targeted optimization of sound field uniformity and frequency response, thus significantly enhancing audio playback quality.
[0121] Figure 3 This is an exemplary flowchart illustrating the determination of recommended adjustment parameters according to some embodiments of this specification. Figure 3 As shown, process 300 includes steps 310, 320 and 330.
[0122] In some embodiments, in response to the acoustic test results not meeting preset conditions, the processor generating recommended adjustment parameters for the speaker further includes: generating a sound field map based on multiple sets of sound wave data; determining the speaker distribution data based on the sound field map; and generating recommended adjustment parameters based on the distribution data and overall fit.
[0123] Step 310: Generate a sound field map based on multiple sets of sound wave data.
[0124] For details regarding acoustic wave data and generating sound field maps from multiple sets of acoustic wave data, please refer to the relevant descriptions in steps 210 and 240.
[0125] Step 320: Determine the speaker distribution data based on the sound field diagram.
[0126] Distribution data refers to the spatial distribution of speakers' positions and orientation angles. Distribution data can be represented by coordinates and angle values.
[0127] In some embodiments, the processor determines the peak points of acoustic parameters based on the sound field map, and determines distribution data based on the peak points.
[0128] In some embodiments, the processor identifies multiple peak points of acoustic parameters (such as sound pressure level) from the sound field map and determines the spatial location corresponding to the multiple peak points as the speaker's position. For example, by identifying the darkest color or the largest value region in the sound field map as the peak point, this point is regarded as the approximate position coordinates of the speaker.
[0129] In some embodiments, the processor determines the gradient direction of the acoustic parameters centered on the peak point, determines the direction in which the acoustic parameters increase the fastest (such as the direction in which the sound pressure level increases the fastest), and then determines the opposite direction of that direction as the speaker's orientation angle.
[0130] Step 330: Generate recommended adjustment parameters based on distribution data and overall fit.
[0131] In some embodiments, the processor first determines the main error types based on the overall fit, and then determines recommended adjustment parameters by combining the speaker distribution data. For example, the processor first identifies which error items in the overall fit (such as audio playback error, reverberation duration error) exceed a preset threshold. More details about the recommended adjustment parameters can be found in the relevant description of step 240.
[0132] If the main audio playback error exceeds a preset threshold, the processor determines the recommended adjustment parameters corresponding to the speaker's position adjustment. Then, based on the distribution data, the processor adjusts the speaker's position by a first-step length to obtain the recommended adjustment parameters. The first-step length can be preset. In some embodiments, the processor can also randomly generate multiple preset first-step lengths, and then evaluate the playback effect of the speaker placement after adjusting the preset first-step lengths through simulation, determining the preset first-step length with the best playback effect as the first-step length.
[0133] If the main issue is that the reverberation duration error exceeds a preset threshold, the processor determines the recommended adjustment parameters corresponding to the speaker's orientation angle. Then, based on the distributed data, the processor adjusts the speaker's orientation angle with a second compensation to obtain the recommended adjustment parameters. The method for determining the second step length is similar to that of the first step length.
[0134] In some embodiments, when both the audio playback error and the reverberation duration error exceed a preset threshold, the processor simultaneously adjusts the speaker position by a first step length and the orientation angle by a second step length based on the distribution data to obtain recommended adjustment parameters.
[0135] In some embodiments, the processor can determine recommended adjustment parameters iteratively based on distributed data. Each iteration adjusts the speaker position and orientation angle by a first step and a second step, respectively, based on the distributed data, and then evaluates the playback effect of the adjusted parameters. If the playback effect is unsatisfactory, the processor performs another first step and a second step adjustment based on the previous iteration until the playback effect meets the requirements. The iteration then stops, and the adjustment parameters from the last iteration are determined as the recommended adjustment parameters. The process of determining whether the playback effect meets the requirements can be performed using audio configuration simulation software, etc.
[0136] In some embodiments, the processor generates recommended adjustment parameters based on distribution data and overall fitness by: processing the distribution data and overall fitness through a parameter recommendation model to determine the recommended adjustment parameters.
[0137] Parameter recommendation models can include any model that can be used to determine the parameters to be adjusted for recommendations, such as machine learning models and neural network models. Examples include machine learning models, deep neural network models, and convolutional neural network models acquired through training.
[0138] In some embodiments, the input to the parameter recommendation model may include distribution data and overall fitness score, or it may be an input vector generated based on the distribution data and overall fitness score. The processor can vectorize the distribution data and overall fitness score to generate the input vector. Vectorization methods include using deep learning models, embedding layers, etc., which will not be elaborated here.
[0139] In some embodiments, the parameter recommendation model can be obtained through training. Training samples include multiple sets of sample distribution data and overall sample fitness scores, or multiple vectors constructed based on the sample distribution data and overall sample fitness scores. Training labels are the sample adjustment parameters corresponding to the training samples.
[0140] In some embodiments, training samples and training labels are obtained based on historical data. The processor selects the recorded speaker position and orientation angle from the historical data as sample adjustment parameters when the speaker position and orientation angle are adjusted to achieve the user's expected playback effect. The historical speaker distribution data before adjustment and the historical comprehensive fit are used as training samples.
[0141] In some embodiments, the training methods for the parameter recommendation model include, but are not limited to, gradient descent and directional propagation algorithms.
[0142] Some embodiments in this specification introduce a parameter recommendation model to process distributed data and overall fit, determining recommended adjustment parameters. This leverages the powerful data processing capabilities of machine learning models to comprehensively analyze and learn the deeper data information reflected by the overall fit, thereby predicting and outputting a theoretically better, but potentially never-before-tested, precise location. This overcomes the limitations of traditional discrete step-length search, enabling the system to find a solution superior to all measured points, thus obtaining more accurate target physical speaker configuration parameters and significantly improving the accuracy and effectiveness of speaker configuration optimization.
[0143] Some embodiments in this specification generate recommended adjustment parameters by combining comprehensive fit and speaker distribution data. Based on acoustic test results, they can specifically generate adjustment parameters for speaker position and orientation angle, realizing intelligent and automated optimization of speaker physical placement. This effectively solves the problems of low efficiency and subjectivity in traditional manual adjustment, and significantly improves sound field optimization efficiency and user listening experience.
[0144] Some embodiments of this specification also provide a speaker configuration optimization apparatus, including a processor, which is used to execute the speaker configuration optimization method described in any of the above embodiments.
[0145] Some embodiments of this specification also provide a computer-readable storage medium, which, when read by a computer from computer instructions in the storage medium, enables the computer to execute the speaker configuration optimization method described in any of the above embodiments.
[0146] The basic concepts have been described above. Obviously, for those skilled in the art, the detailed disclosure above is merely illustrative and does not constitute a limitation of this specification. Although not explicitly stated herein, those skilled in the art may make various modifications, improvements, and corrections to this specification. Such modifications, improvements, and corrections are suggested in this specification and therefore remain within the spirit and scope of the exemplary embodiments described herein.
[0147] Furthermore, this specification uses specific terms to describe embodiments thereof. For example, "an embodiment," "one embodiment," and / or "some embodiments" refer to a particular feature, structure, or characteristic associated with at least one embodiment of this specification. Therefore, it should be emphasized and noted that references to "an embodiment," "one embodiment," or "an alternative embodiment" in different locations throughout this specification do not necessarily refer to the same embodiment. Moreover, certain features, structures, or characteristics in one or more embodiments of this specification can be appropriately combined.
[0148] Furthermore, unless expressly stated in the claims, the order of processing elements and sequences, the use of numbers and letters, or other names described in this specification are not intended to limit the order of the processes and methods described herein. Although various examples have been discussed in the foregoing disclosure of some embodiments of the invention that are currently considered useful, it should be understood that such details are for illustrative purposes only, and the appended claims are not limited to the disclosed embodiments; rather, the claims are intended to cover all modifications and equivalent combinations that conform to the spirit and scope of the embodiments described herein. For example, while the system components described above can be implemented using hardware devices, they can also be implemented solely using software solutions, such as installing the described system on existing servers or mobile devices.
[0149] Similarly, it should be noted that, in order to simplify the description disclosed herein and thus aid in the understanding of one or more embodiments of the invention, the foregoing description of embodiments in this specification may sometimes combine multiple features into a single embodiment, drawing, or description thereof. However, this method of disclosure does not imply that the subject matter of this specification requires more features than those mentioned in the claims. In fact, the embodiments contain fewer features than all the features of a single embodiment disclosed above.
[0150] In some embodiments, numbers describing the quantity of components and attributes are used. It should be understood that such numbers used in the description of embodiments are modified in some examples with the terms "approximately," "approximately," or "generally." Unless otherwise stated, "approximately," "approximately," or "generally" indicates that the numbers are allowed to vary by ±20%. Accordingly, in some embodiments, the numerical parameters used in the specification and claims are approximate values, which may be changed depending on the characteristics required by individual embodiments. In some embodiments, numerical parameters should take into account specified significant digits and employ a general method of digit reservation. Although the numerical ranges and parameters used to confirm their breadth of range in some embodiments of this specification are approximate values, in specific embodiments, such values are set as precisely as feasible.
[0151] For each patent, patent application, patent application publication, and other material such as articles, books, specifications, publications, and documents referenced in this specification, the entire contents of which are incorporated herein by reference. This excludes historical application documents that are inconsistent with or conflict with the content of this specification, as well as documents that limit the broadest scope of the claims in this specification (currently or subsequently appended to this specification). It should be noted that in the event of any inconsistency or conflict between the descriptions, definitions, and / or terminology used in the supplementary materials to this specification and the content of this specification, the descriptions, definitions, and / or terminology used in this specification shall prevail.
[0152] Finally, it should be understood that the embodiments described in this specification are merely illustrative of the principles of the embodiments described herein. Other variations may also fall within the scope of this specification. Therefore, alternative configurations of the embodiments described herein are intended to be illustrative rather than limiting, and should be considered consistent with the teachings of this specification. Accordingly, the embodiments described herein are not limited to those explicitly introduced and described herein.
Claims
1. A speaker configuration optimization method, characterized in that, include: Control the speaker to play multiple sets of test sounds, and control the audio acquisition device to receive the multiple sets of test sounds; Acquire multiple sets of sound wave data corresponding to the multiple sets of test sounds from the audio acquisition device; Based on the multiple sets of sound wave data, the acoustic test results are determined; In response to the acoustic test results not meeting the preset conditions, recommended adjustment parameters for the speaker are generated.
2. The method according to claim 1, characterized in that, The method further includes: Based on spatial and environmental characteristics, at least one preset acoustic parameter is determined; The plurality of test sounds are determined based on the at least one preset acoustic parameter.
3. The method according to claim 1, characterized in that, The audio acquisition device receives the multiple sets of test sounds at multiple listening points; determining the acoustic test results based on the multiple sets of sound wave data includes: Based on the multiple sets of sound wave data, determine the acoustic feature fit degree corresponding to each combination of the listening point and the preset acoustic parameters; Based on the acoustic feature fit, a comprehensive fit is determined, and the comprehensive fit is used as the acoustic test result.
4. The method according to claim 1, characterized in that, The recommended adjustment parameters for the speaker generated in response to the acoustic test results not meeting the preset conditions include: A sound field map is generated based on the multiple sets of sound wave data; The distribution data of the speaker is determined based on the sound field diagram; The recommended adjustment parameters are generated based on the distribution data and overall fit.
5. A speaker configuration optimization system, characterized in that, include: The control module is configured to control the speaker to play multiple sets of test sounds, and to control the audio acquisition device to receive the multiple sets of test sounds; The acquisition module is configured to acquire multiple sets of sound wave data corresponding to the multiple sets of test sounds from the audio acquisition device; The determination module is configured to determine the acoustic test results based on the multiple sets of acoustic wave data; The parameter generation module is configured to generate recommended adjustment parameters for the speaker in response to the acoustic test results not meeting preset conditions.
6. The system according to claim 5, characterized in that, The control module is further configured to: Based on spatial and environmental characteristics, at least one preset acoustic parameter is determined; The plurality of test sounds are determined based on the at least one preset acoustic parameter.
7. The system according to claim 5, characterized in that, The determining module is further configured to: Based on the multiple sets of sound wave data, determine the acoustic feature fit degree corresponding to each combination of the listening point and the preset acoustic parameters; Based on the acoustic feature fit, a comprehensive fit is determined, and the comprehensive fit is used as the acoustic test result.
8. The system according to claim 5, characterized in that, The parameter generation module is further configured to: A sound field map is generated based on the multiple sets of sound wave data; The distribution data of the speaker is determined based on the sound field diagram; The recommended adjustment parameters are generated based on the distribution data and overall fit.
9. A speaker configuration optimization device, characterized in that, Includes a processor for executing the speaker configuration optimization method according to claims 1-4.
10. A computer-readable storage medium storing computer instructions, wherein when a computer reads the computer instructions in the storage medium, the computer executes the speaker configuration optimization method as described in claims 1-4.
Citation Information
Cited By
Single-chip double-atmosphere lamp music rhythm control method and system and electronic equipment
CN121968418A