Intelligent sound field calibration method and system

By collecting frequency sweep signals from multiple microphone locations in the sound field, constructing an acoustic feature database and generating a sound field topology map, identifying howling frequency points, and optimizing the frequency response curve by combining user preference data, the problem of insufficient perception of sound field changes in existing sound field calibration methods is solved, and precise sound quality adjustment and personalized audio optimization are achieved.

CN121665172APending Publication Date: 2026-03-13SHENZHEN SOUNDFIT TECH CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-11-28
Publication Date
2026-03-13

AI Technical Summary

Technical Problem

Existing sound field calibration methods are mainly based on audio signals collected by microphones at fixed locations. They lack the ability to model spatial differences with multiple microphones and cannot accurately perceive changes in the sound field at different locations, resulting in limited improvement in sound quality.

Method used

By sending frequency sweep signals within a preset frequency range to multiple microphone acquisition positions in the sound field, an acoustic feature database is constructed, the relative positional relationship between microphones and echo signal characteristics are extracted, a sound field spatial topology map is generated and acoustic zones are divided, spectral analysis is combined to identify howling frequency points, the frequency response curve is adjusted according to user historical preference data, and iterative optimization is carried out through user feedback.

Benefits of technology

It achieves fine perception and modeling of the sound field spatial structure, accurately locates the howling frequency area, provides personalized audio optimization strategies, and improves the user's subjective listening experience and the system's adaptability in dynamic scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121665172A_ABST
    Figure CN121665172A_ABST
Patent Text Reader

Abstract

The invention relates to an intelligent sound field calibration method and system, and the method comprises the steps: transmitting sweep frequency signals in a preset frequency range to a plurality of microphone collection positions in a sound field, collecting echo signals through each microphone, and constructing an acoustic feature database based on the echo signals; extracting a relative position relationship between microphones and corresponding echo signal features from the acoustic feature database, constructing a sound field spatial topological graph and dividing acoustic partitions based on the relative position relationship and the echo signal features, and determining response weights of the partitions; based on the acoustic characteristic database, a frequency spectrum analysis algorithm is adopted to identify and mark howling frequency points, and a preliminary frequency response curve is generated in combination with the howling frequency points and the response weights of the acoustic partitions; adjusting the initial frequency response curve according to historical preference data of the user to generate a target frequency response curve; and applying the target frequency response curve to the audio output path to output the audio signal. The method has the effect of improving the accuracy of sound field calibration.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical field of audio processing, and in particular to an intelligent sound field calibration method and system. Background Technology

[0002] Currently, with the widespread adoption of smart speakers, home theaters, immersive conferencing equipment, and in-vehicle entertainment systems, users are placing higher demands on the environmental adaptability and personalized adjustment capabilities of audio playback. To enhance the overall listening experience, systems often need to dynamically calibrate the sound field based on the usage scenario, speaker layout, and user preferences to achieve effects such as frequency response curve optimization, suppression of feedback frequencies, and multi-zone response equalization.

[0003] Current technologies primarily rely on microphones at fixed locations to sample test sounds emitted by the audio system, obtain the initial frequency response, and then perform simple corrections using preset frequency response curves. While this approach has some effectiveness, it lacks the ability to model spatial differences with multiple microphones and cannot address issues such as reflections, obstructions, or multi-source reverberation in complex rooms.

[0004] The existing technical solutions mentioned above have the following drawbacks: the existing sound field calibration methods are mainly based on collecting audio signals from microphones at fixed positions and correcting the frequency response through static filtering parameters. They lack fine modeling capabilities, especially in multi-user, multi-terminal or complex indoor environments. They cannot accurately perceive changes in the sound field at different locations and cannot accurately match specific environments, resulting in limited improvement in sound quality. Therefore, there is room for improvement. Summary of the Invention

[0005] To improve the accuracy of sound field calibration, this application provides an intelligent sound field calibration method and system.

[0006] The above-mentioned objective of this application is achieved through the following technical solution: A smart sound field calibration method, the smart sound field calibration method comprising: A frequency sweep signal within a preset frequency range is sent to multiple microphone acquisition positions in the sound field, and each microphone acquires the echo signal, and an acoustic feature database is constructed based on the echo signal; The relative positional relationships between microphones and their corresponding echo signal features are extracted from the acoustic feature database. Based on the relative positional relationships and echo signal features, a sound field spatial topology map is constructed and acoustic partitions are divided, and the response weight of each partition is determined. Based on the acoustic feature database, a spectrum analysis algorithm is used to identify and mark the howling frequency points, and a preliminary frequency response curve is generated by combining the howling frequency points and the response weights of each acoustic zone. The initial frequency response curve is adjusted based on the user's historical preference data to generate the target frequency response curve; The target frequency response curve is applied to the audio output path to output an audio signal, and the corresponding frequency response adjustment result is displayed through the interface program. The system collects user feedback on the current audio signal. If the feedback is unsatisfactory, it iteratively optimizes the preliminary frequency response curve based on the feedback to generate a new target frequency response curve to replace the original target frequency response curve.

[0007] By employing the above technical solution, a frequency sweep signal within a preset frequency range is sent to multiple microphone acquisition positions in the sound field. Each microphone collects echo signals, and an acoustic feature database is constructed. This allows for a comprehensive acquisition of the acoustic reflection and attenuation characteristics of each region in the space, thereby enhancing the system's ability to perceive the spatial structure of the sound field. By extracting the relative positional relationships between microphones and echo signal characteristics, a sound field spatial topology map is constructed and acoustic zones are divided. This enables the modeling of the spatial distribution characteristics within the sound field, thereby achieving the identification of differences in sound wave propagation effects at different locations. By performing spectral analysis based on the acoustic feature database to identify and mark howling frequency points, and combining the response weights of each acoustic zone, a preliminary frequency response is generated. The frequency response curve can accurately locate high-risk frequency areas that may cause feedback, thereby improving the stability and accuracy of frequency response adjustment. By adjusting the initial frequency response curve based on users' historical preference data and generating a target frequency response curve, personalized audio optimization strategies can be implemented, thereby enhancing the user's subjective listening experience. By applying the target frequency response curve to the audio output path and displaying the adjustment results, the user's perception and participation can be improved, thereby building an interactive audio optimization feedback loop. By collecting user feedback and iteratively optimizing the initial frequency response curve, the system can continuously follow up on users' subjective feelings and adaptively optimize parameters, thereby improving the system's adaptability and learning ability in dynamic scenarios.

[0008] In one example, this application can be further configured as follows: sending a frequency sweep signal within a preset frequency range to multiple microphone acquisition positions in the sound field, having each microphone acquire echo signals, and constructing an acoustic feature database based on the echo signals, specifically includes: A frequency sweep signal is generated within a preset frequency range, and the frequency sweep signal is synchronously sent to multiple microphone acquisition positions deployed in the sound field. Each microphone receives the echo signal at the corresponding location, extracts features from the echo signal to obtain acoustic feature parameters, and summarizes the acoustic feature parameters to construct an acoustic feature database. The acoustic feature parameters include frequency domain response features, energy attenuation parameters, and reflection delay information.

[0009] By adopting the above technical solution, by generating a sweep frequency signal within a preset frequency range and synchronously sending it to multiple microphone acquisition positions, the acoustic response of different spatial points can be acquired under a unified timing sequence, thereby avoiding data errors caused by timing inconsistencies. By receiving echo signals from each microphone and extracting frequency domain response characteristics, energy attenuation parameters, and reflection delay information, the acoustic transmission characteristics of different spatial locations can be accurately reflected, thus providing a reliable data foundation for subsequent sound field modeling and frequency response optimization.

[0010] In one example, this application can be further configured as follows: Based on the relative positional relationship and echo signal characteristics, constructing a sound field spatial topology map and dividing it into acoustic partitions, and determining the response weight of each partition, specifically includes: Based on the relative positional relationship of each microphone and the propagation delay features extracted from the corresponding echo signal features, the spatial geometric relationship between each microphone is calculated, and the sound field spatial topology map is generated based on the geometric relationship. Based on the density of microphone distribution, the spatial region they are located in, and the sound wave reflection characteristics in the aforementioned spatial topology diagram, the sound field is divided into multiple acoustic zones. Based on each acoustic zone, and taking into account the sound wave attenuation characteristics and the user's primary listening position, the sound wave contribution of each acoustic zone at the user's primary listening position is comprehensively evaluated, and a response weight is assigned to each acoustic zone according to the sound wave contribution.

[0011] By adopting the above technical solution, spatial geometric relationships are calculated based on the relative positional relationship and propagation delay characteristics between microphones to generate a sound field topology map. This enables the establishment of a spatial structure mapping that conforms to the actual layout, thus laying the foundation for refined area division. By dividing acoustic zones according to the microphone distribution density, spatial region, and reflection characteristics in the topology map, different acoustic behavior regions can be identified and adapted, thereby improving the regional adaptability of the overall frequency response adjustment. By combining sound wave attenuation characteristics and the user's primary listening position to assign response weights to zones, the main listening area can be ensured to obtain the best calibration effect, thereby enhancing the personalized listening experience.

[0012] In one example, this application can be further configured as follows: based on the acoustic feature database, a spectrum analysis algorithm is used to identify and mark howling frequency points, and a preliminary frequency response curve is generated by combining the howling frequency points and the response weights of each acoustic zone, specifically including: Perform spectral analysis on the echo spectrum in the acoustic feature database, extract the amplitude response features and frequency persistence features of each frequency point, and determine the energy value of the frequency point based on the amplitude response features; The frequency persistence characteristics are compared with a preset frequency stability threshold, and the energy value of the frequency point is compared with a preset energy threshold. Frequency points that simultaneously meet the judgment conditions are selected as the target howling frequency point set. The target howling frequency set is weighted and fused with the response weights corresponding to each acoustic zone to generate energy adjustment parameters for each frequency point, and the preliminary frequency response curve is constructed based on the energy adjustment parameters.

[0013] By employing the above technical solutions, spectral analysis is performed on the echo spectrum in the acoustic feature database to extract the amplitude response and frequency persistence features of frequency points and determine the energy value. This allows for the construction of a frequency-dimensional response intensity map, thereby enabling accurate identification of abnormal frequency points. By jointly filtering the extracted features with frequency stability thresholds and energy thresholds, a set of target howling frequency points that meet the conditions can be extracted, effectively distinguishing high-risk frequency points from normal frequencies and thus improving the accuracy of howling identification. By weighted fusion of howling frequency points with the response weights of acoustic zones, energy adjustment parameters are generated and a preliminary frequency response curve is constructed. This enables a frequency response adjustment strategy that combines acoustic spatial characteristics, thereby improving the targeting and effectiveness of frequency response adjustment.

[0014] In one example, this application can be further configured as follows: adjusting the preliminary frequency response curve based on user historical preference data to generate the target frequency response curve specifically includes: Extract user's historical listening records, manual adjustment behavior and scene label information from the user's historical preference data to construct a user preference feature set, wherein the user preference feature set includes frequency band preference, loudness selection and filter adjustment parameters; The preliminary frequency response curve is compared with the user preference feature set, and the preliminary frequency response curve is adjusted based on a preset loudness compensation model to generate the target frequency response curve.

[0015] By adopting the above technical solution, a set of preference features can be constructed by extracting users' historical listening records, manual adjustment behavior, and scene tag information. This can fully restore the user's audio preference characteristics, thus providing real data basis for personalized adjustment. By comparing the preliminary frequency response curve with the user preference feature set and adjusting it based on the preset loudness compensation model, the output effect of each frequency band can be adjusted according to the user's real needs, thereby improving the matching degree between the frequency response curve and the user's listening experience and enhancing the user's subjective experience.

[0016] In one example, this application can be further configured as follows: applying the target frequency response curve to the audio output path to output an audio signal, and displaying the corresponding frequency response adjustment result through a user interface program, specifically includes: Based on the target frequency response curve, filter gain control parameters corresponding to each frequency band are generated, and the control parameters are loaded into the digital filtering module in the audio output path; Based on the filter control parameters generated from the target frequency response curve, the frequency response of the output audio signal is adjusted, and the target frequency response curve is displayed in real time in the form of a frequency-gain relationship graph through the interface program, with the adjustment range of each frequency band marked.

[0017] By adopting the above technical solution, and generating filter gain control parameters corresponding to each frequency band based on the target frequency response curve, and loading them into the digital filtering module in the audio output path, frequency response adjustment can be quickly achieved in the playback link, thereby avoiding delay and inconsistency issues. By performing frequency response adjustment based on filter control parameters and displaying the frequency response curve in real time, the user's visual understanding of the audio adjustment results and the degree of adjustment participation can be improved, thereby realizing a user-friendly calibration process.

[0018] In one example, this application can be further configured such that: the iterative optimization of the preliminary frequency response curve based on the feedback information specifically includes: The feedback information is parsed into a frequency response adjustment factor, and a corresponding set of frequency response optimization parameters is generated. The parsing process includes weighted analysis of frequency band tendencies and frequency response difference fitting of the preference tilt direction. Based on the frequency response optimization parameter set, the initial frequency response curve is adjusted to generate a new target frequency response curve and replace the original target frequency response curve to update the response characteristics in the audio output path. The adjustment operation includes gain fine-tuning, filter parameter reconstruction, or local suppression enhancement.

[0019] By adopting the above technical solution, user feedback information is parsed into frequency response adjustment factors, and a set of frequency response optimization parameters is generated through weighted analysis and fitting with frequency response differences. This enables a structured understanding of user feedback, thereby improving the intelligence of feedback processing. By performing gain fine-tuning, filter parameter reconstruction, or local suppression enhancement based on the optimized parameter set, and updating the frequency response curve to adapt to the feedback results, dynamic response and adaptive optimization can be achieved, thereby continuously improving user satisfaction and auditory consistency.

[0020] The second objective of this invention is achieved through the following technical solution: A smart sound field calibration system, characterized in that the smart sound field calibration system comprises: The acoustic playback and acquisition module is used to send a frequency sweep signal within a preset frequency range to multiple microphone acquisition positions in the sound field, and each microphone acquires the echo signal and constructs an acoustic feature database based on the echo signal. The topology modeling module is used to extract the relative positional relationships between microphones and the corresponding echo signal features from the acoustic feature database, and based on the relative positional relationships and echo signal features, construct a sound field space topology map and divide it into acoustic partitions, and determine the response weight of each partition. The frequency response calculation module is used to identify and mark howling frequency points based on the acoustic feature database using a spectrum analysis algorithm, and generate a preliminary frequency response curve by combining the howling frequency points and the response weights of each acoustic zone; The preference matching module is used to adjust the preliminary frequency response curve based on the user's historical preference data to generate the target frequency response curve; The output display module is used to apply the target frequency response curve to the audio output path to output an audio signal, and to display the corresponding frequency response adjustment result through the interface program. The feedback optimization module is used to collect user feedback information on the current audio signal. If the feedback result is unsatisfactory, the initial frequency response curve is iteratively optimized based on the feedback information to generate a new target frequency response curve to replace the original target frequency response curve.

[0021] By employing the above technical solution, a frequency sweep signal within a preset frequency range is sent to multiple microphone acquisition positions in the sound field. Each microphone collects echo signals, and an acoustic feature database is constructed. This allows for a comprehensive acquisition of the acoustic reflection and attenuation characteristics of each region in the space, thereby enhancing the system's ability to perceive the spatial structure of the sound field. By extracting the relative positional relationships between microphones and echo signal characteristics, a sound field spatial topology map is constructed and acoustic zones are divided. This enables the modeling of the spatial distribution characteristics within the sound field, thereby achieving the identification of differences in sound wave propagation effects at different locations. By performing spectral analysis based on the acoustic feature database to identify and mark howling frequency points, and combining the response weights of each acoustic zone, a preliminary frequency response is generated. The frequency response curve can accurately locate high-risk frequency areas that may cause feedback, thereby improving the stability and accuracy of frequency response adjustment. By adjusting the initial frequency response curve based on users' historical preference data and generating a target frequency response curve, personalized audio optimization strategies can be implemented, thereby enhancing the user's subjective listening experience. By applying the target frequency response curve to the audio output path and displaying the adjustment results, the user's perception and participation can be improved, thereby building an interactive audio optimization feedback loop. By collecting user feedback and iteratively optimizing the initial frequency response curve, the system can continuously follow up on users' subjective feelings and adaptively optimize parameters, thereby improving the system's adaptability and learning ability in dynamic scenarios.

[0022] In summary, this application includes the following beneficial technical effects: 1. By sending frequency sweep signals within a preset frequency range to multiple microphone acquisition positions in the sound field, each microphone collects echo signals, and an acoustic feature database is constructed. This allows for a comprehensive acquisition of the acoustic reflection and attenuation characteristics of each region in the space, thereby enhancing the system's ability to perceive the spatial structure of the sound field. By extracting the relative positional relationships between microphones and echo signal characteristics, a sound field spatial topology map is constructed and acoustic zones are divided. This enables the modeling of the spatial distribution characteristics within the sound field, thereby achieving the identification of differences in sound wave propagation effects at different locations. By performing spectral analysis based on the acoustic feature database to identify and mark howling frequency points, and combining the response weights of each acoustic zone to generate preliminary frequency response curves, high-risk frequency areas that may cause howling can be accurately located, thereby improving the stability and accuracy of frequency response adjustment. 2. By adjusting the initial frequency response curve based on users' historical preference data and generating a target frequency response curve, personalized audio optimization strategies can be implemented, thereby enhancing users' subjective listening experience. By applying the target frequency response curve to the audio output path and displaying the adjustment results, user perception and participation can be improved, thus building an interactive audio optimization feedback loop. By collecting user feedback and iteratively optimizing the initial frequency response curve, the system can continuously follow up on users' subjective feelings and adaptively optimize parameters, thereby improving the system's adaptability and learning ability in dynamic scenarios. Attached Figure Description

[0023] Figure 1 This is a flowchart of an intelligent sound field calibration method according to an embodiment of this application; Figure 2 This is a flowchart illustrating the implementation of step S10 in an embodiment of an intelligent sound field calibration method of this application. Figure 3 This is a flowchart illustrating the implementation of step S20 in an embodiment of an intelligent sound field calibration method of this application. Figure 4 This is a flowchart illustrating the implementation of step S30 in an embodiment of an intelligent sound field calibration method of this application. Figure 5 This is a flowchart illustrating the implementation of step S40 in an embodiment of an intelligent sound field calibration method of this application. Figure 6 This is a flowchart illustrating the implementation of step S50 in an embodiment of an intelligent sound field calibration method of this application. Figure 7 This is a flowchart illustrating the implementation of step S60 in an embodiment of an intelligent sound field calibration method of this application. Figure 8 This is a schematic diagram of a smart sound field calibration system according to one embodiment of this application. Detailed Implementation

[0024] The present application will be further described in detail below with reference to the accompanying drawings.

[0025] In one embodiment, such as Figure 1 As shown, this application discloses an intelligent sound field calibration method, which specifically includes the following steps: S10: Send a frequency sweep signal within a preset frequency range to multiple microphone acquisition positions in the sound field, and have each microphone acquire the echo signal and construct an acoustic feature database based on the echo signal.

[0026] Specifically, when transmitting the frequency sweep signal, a linearly increasing sine wave signal is emitted at fixed intervals within the range of 20Hz to 20kHz. The duration of each frequency point is controlled to be 20 milliseconds. Multiple microphone sampling points are evenly distributed in the sound field. Each microphone records the echo signal response at its location. Subsequently, the amplitude attenuation, arrival delay, and frequency domain response characteristics of the echo signal are analyzed and stored in the data record associated with the corresponding microphone number. For example, after setting up 9 microphones in an 8×8 meter conference room and transmitting a frequency sweep signal of 20Hz to 20kHz, 9 sets of echo response curves are obtained and used to generate an acoustic feature database. Each record contains the transmission frequency point, echo amplitude, arrival time, and phase information to support subsequent spatial topology reconstruction and frequency response tuning.

[0027] S20: Extract the relative positional relationship between microphones and the corresponding echo signal features from the acoustic feature database. Based on the relative positional relationship and echo signal features, construct a sound field spatial topology map and divide the acoustic partitions, and determine the response weight of each partition.

[0028] Specifically, the relative geometric distance between each microphone is calculated based on the echo delay difference, and a two-dimensional or three-dimensional coordinate system is reconstructed using the triangulation method. Then, the frequency response mean, peak frequency point, reverberation time, and other characteristics of the echo signal are used as the basis for zoning evaluation to construct a topological structure diagram of the sound field space. In the diagram, nodes represent microphone positions, and edge weights represent signal similarity and spatial adjacency. Furthermore, the K-means algorithm is used to divide the sound field into multiple acoustic zones. For example, in the aforementioned conference room, the nine microphones are clustered into three regions, representing the areas near the walls, the central open area, and the corner areas, respectively. Each region is assigned a response weight based on the echo response energy and howling frequency sensitivity. For example, the weight of the central region is set to 0.5, and the weight of the corner area is set to 0.2, to reflect the reference priority of the region for frequency response calibration.

[0029] S30: Based on the acoustic feature database, the spectrum analysis algorithm is used to identify and mark the howling frequency points, and a preliminary frequency response curve is generated by combining the howling frequency points and the response weights of each acoustic zone.

[0030] Specifically, a Fast Fourier Transform (FFT) is used to perform spectral analysis on the echo signals of each microphone, and high-amplitude narrowband peaks are detected to mark possible howling frequencies. Then, the weighting coefficients of each acoustic zone are combined to perform zone-weighted accumulation of all detected howling frequencies to generate multi-band response curves. Frequency points in high-weight regions are given higher correction priority. For example, two strong howling points at 1.2kHz and 2.5kHz are detected in the central region, and a secondary strong howling point at 4.8kHz is detected in the edge region. Finally, the initial frequency response curve reduces the gain by more than 3dB at 1.2kHz and 2.5kHz, while it only reduces it by 1.5dB at 4.8kHz, so as to ensure that the optimization focus is on the howling sensitive points in the core region.

[0031] S40: Adjust the preliminary frequency response curve based on the user's historical preference data to generate the target frequency response curve.

[0032] Specifically, the system retrieves the historical preference dataset stored in the user's account, performs feature fitting analysis on the frequency response adjustment model, and maps the user's preference parameters for high-frequency enhancement, low-frequency smoothing, or vocal clarity to the shape of the frequency response curve. For example, if a user has repeatedly enhanced the vocal frequency range of 2kHz to 4kHz and reduced the gain of sharp audio ranges above 8kHz in the past, the current preliminary frequency response curve will be further increased by 1.5dB in the 2kHz to 4kHz range and further decreased by 2dB in the range above 8kHz, thereby generating a target frequency response curve that better matches the individual's subjective listening experience, improving user satisfaction and listening comfort.

[0033] S50: Apply the target frequency response curve to the audio output path to output the audio signal, and display the corresponding frequency response adjustment results through the interface program.

[0034] Specifically, the target frequency response curve parameters are loaded into the output path to a digital audio processor, such as a DSP or a custom filter chain, and the gain parameters and phase delay of the filters are adjusted in real time to output the calibrated audio signal. At the same time, the frequency response curve is presented in a visual graphical form in the user interface, such as a curve with frequency on the horizontal axis and gain on the vertical axis displayed on a webpage or APP. The adjusted howling frequencies, user-preferred enhancement bands, and system-suggested optimization areas are clearly marked. Users can click to view the gain value of each segment and the reason for the adjustment, which makes it easier to understand the calibration process and further fine-tune manually.

[0035] S60: Collect user feedback information on the current audio signal. If the feedback result is unsatisfactory, iteratively optimize the preliminary frequency response curve based on the feedback information and generate a new target frequency response curve to replace the original target frequency response curve.

[0036] Specifically, after users listen to the audio signal, their feedback ratings, text opinions, or quick tag selections on the interface, such as "high frequencies are harsh," "low frequencies are insufficient," and "obvious echo," are collected. These subjective feedback contents are converted into quantifiable feature tags, and then combined with machine learning models such as decision tree algorithms to evaluate the frequency ranges most likely to cause dissatisfaction. Targeted gain fine-tuning is then performed on the initial frequency response curve. For example, if a user marks "high frequencies are harsh" multiple times, the model will automatically identify the 8kHz~12kHz range as the problem frequency range, reduce this frequency range by 3dB, generate a new target frequency response curve, and display the optimized area and next listening suggestion on the interface, guiding the user to complete the feedback-optimization closed loop.

[0037] By employing the above technical solution, a frequency sweep signal within a preset frequency range is sent to multiple microphone acquisition positions in the sound field. Each microphone collects echo signals, and an acoustic feature database is constructed. This allows for a comprehensive acquisition of the acoustic reflection and attenuation characteristics of each region in the space, thereby enhancing the system's ability to perceive the spatial structure of the sound field. By extracting the relative positional relationships between microphones and echo signal characteristics, a sound field spatial topology map is constructed and acoustic zones are divided. This enables the modeling of the spatial distribution characteristics within the sound field, thereby achieving the identification of differences in sound wave propagation effects at different locations. By performing spectral analysis based on the acoustic feature database to identify and mark howling frequency points, and combining the response weights of each acoustic zone, a preliminary frequency response is generated. The frequency response curve can accurately locate high-risk frequency areas that may cause feedback, thereby improving the stability and accuracy of frequency response adjustment. By adjusting the initial frequency response curve based on users' historical preference data and generating a target frequency response curve, personalized audio optimization strategies can be implemented, thereby enhancing the user's subjective listening experience. By applying the target frequency response curve to the audio output path and displaying the adjustment results, the user's perception and participation can be improved, thereby building an interactive audio optimization feedback loop. By collecting user feedback and iteratively optimizing the initial frequency response curve, the system can continuously follow up on users' subjective feelings and adaptively optimize parameters, thereby improving the system's adaptability and learning ability in dynamic scenarios.

[0038] In one embodiment, such as Figure 2 As shown, in step S10, a frequency sweep signal within a preset frequency range is sent to multiple microphone acquisition positions in the sound field. Each microphone acquires the echo signal, and an acoustic feature database is constructed based on the echo signal. Specifically, this includes: S11: Generate a sweep frequency signal within a preset frequency range and synchronously send the sweep frequency signal to multiple microphone acquisition positions deployed in the sound field.

[0039] Specifically, the starting and ending frequencies are set according to the frequency coverage required by the target sound field. For example, a sweep frequency signal sequence is generated in a linear or logarithmic manner within the range of 20Hz to 20kHz. At the same time, an appropriate duration is set to ensure that each frequency point has sufficient excitation time. Then, the sweep frequency signal is output to the sound field through a loudspeaker. Multiple microphone sampling points are distributed in advance according to the structural characteristics of the sound field. For example, they can be set at the front wall, rear wall, side wall, and center position. Each microphone node receives the excitation signal from the loudspeaker and its echo signal at that point synchronously with a unified clock. This ensures that each sampling node receives the reflection and response performance of the sweep frequency output at the same moment in its spatial position. For example, in a 20-square-meter conference room environment, six microphone sampling points are set up to record the response around the sound source location and the boundary wall to obtain the initial acquisition of spatial multidimensional acoustic data.

[0040] S12: Each microphone receives the echo signal at the corresponding location, extracts features from the echo signal to obtain acoustic feature parameters, and summarizes the acoustic feature parameters to build an acoustic feature database. The acoustic feature parameters include frequency domain response features, energy attenuation parameters, and reflection delay information.

[0041] Specifically, the echo signals received by each microphone are first converted to the frequency domain using a Fast Fourier Transform algorithm, and the amplitude response of different frequency components is extracted to form frequency domain response characteristics. Simultaneously, the original power of the swept signal is compared with the power of the received signal at the current sampling point, and the energy attenuation coefficient at different frequencies is calculated and normalized and represented in the database. For example, at a frequency of 1kHz, if the input power is 85dB and the power measured at the receiving point is 75dB, then the attenuation at that frequency is recorded as 10dB. Furthermore, echo signals are extracted using correlation function analysis and peak detection methods. The time delay difference between the primary reflection and multiple reflections in the sound path is recorded as reflection time delay information. This information helps to analyze the room's reverberation level and obstacle distribution characteristics. Finally, the above features are summarized according to the spatial coordinates corresponding to each sampling point to construct an acoustic feature database with a location index. For example, the database records information such as "microphone position A-frequency domain response: [X1, X2, ...], energy attenuation: [Y1, Y2, ...], reflection time delay: [Z1, Z2, ...]", etc., providing a data foundation for subsequent calibration parameter inference and sound field optimization strategies.

[0042] In one embodiment, such as Figure 3 As shown, in step S20, based on the relative positional relationship and echo signal characteristics, a sound field spatial topology map is constructed and acoustic partitions are divided, and the response weight of each partition is determined. Specifically, this includes: S21: Based on the relative positional relationship of each microphone and the propagation delay features extracted from the corresponding echo signal features, calculate the spatial geometric relationship between each microphone, and generate a sound field spatial topology map based on the geometric relationship.

[0043] Specifically, based on the relative positional relationship of each microphone and the propagation delay features extracted from its echo signals, a triangulation algorithm is executed to determine the relative distance relationship between the microphones. At the same time, spatial coordinate mapping is combined to project it onto a three-dimensional spatial model, thereby constructing an initial geometric map representing the distribution state of the microphones. On this basis, a graph structure optimization algorithm based on connectivity and minimum hop count is introduced to abstract the topology structure and generate a sound field spatial topology map covering all microphone nodes. Each side of this topology map represents the structural relationship between two microphones that can perceive the sound propagation path. For example, in a typical home theater scenario, seven microphones may be deployed on the front wall, back wall, and ceiling, respectively. The constructed topology map will automatically reflect the distance differences and relative layout features between each microphone.

[0044] S22: Based on the density of microphone distribution, the spatial region, and the sound wave reflection characteristics in the spatial topology diagram, the sound field is divided into multiple acoustic zones.

[0045] Specifically, based on the distribution density of each microphone node in the constructed spatial topology map and the boundary conditions of its spatial region, combined with the characteristics of sound wave reflection and absorption on different surfaces, a partitioning algorithm is executed to cluster the sound field region. The density-based DBSCAN clustering method is preferentially used to divide the space into several acoustic sub-regions, where the microphones in each sub-region have strong acoustic propagation similarity. Then, the boundary regions are corrected and compensated by combining sound wave multipath reflection simulation to ensure the coherence and effectiveness of acoustic partitioning. For example, in a conference room scenario, microphones near windows may be automatically classified into high-reflection acoustic partitions based on the reflection characteristics of glass walls, while areas near sound-absorbing materials form low-reflection acoustic partitions.

[0046] S23: Based on each acoustic zone, combined with the sound wave attenuation characteristics and the user's primary listening position, comprehensively evaluate the sound wave contribution of each acoustic zone at the user's primary listening position, and assign response weights to each acoustic zone according to the sound wave contribution.

[0047] Specifically, for each acoustic zone, the sound pressure level propagating from the microphone in that zone to the main listening position is calculated based on a preset sound wave attenuation model. The number of reflections and refraction angles in the propagation path are corrected by combining the coordinate information of the user's main listening position. Furthermore, the sound wave contribution index is calculated based on the sound wave arrival time difference, sound intensity amplitude, and frequency band consistency. On this basis, a normalized weighting strategy is used to assign corresponding response weights to each acoustic zone. Zones with higher response weights are given priority in the sound field adjustment strategy during subsequent calibration. For example, in a home theater scenario, if the user's main listening position is located in the center of the sofa, the zones directly in front and on the left and right sides are assigned higher response weights because of their higher direct sound contribution. Furthermore, the user's primary listening position can be determined in various ways. For example, during actual deployment, the user's most frequently used seating area, the main sofa position, or the presentation area can be preset as the primary listening position. Alternatively, the primary listening position can be automatically determined by combining the location information reported by the user's terminal device, the user's manual marking on the interface, or the spatial node with the longest device connection time in historical listening behavior. In some scenarios equipped with infrared sensors or cameras, the user's current activity area can also be obtained as the primary listening reference point through visual recognition or human body sensing technology. For example, when the system detects that the user spends more than 90% of their time listening in a corner of a room, it can automatically use the acoustic sub-zone corresponding to that area as the primary listening position for subsequent weight calculation and response curve optimization, thereby achieving personalized sound field matching that varies from person to person and from location to location.

[0048] In one embodiment, such as Figure 4 As shown, in step S30, a spectrum analysis algorithm is used to identify and mark the howling frequency points, and a preliminary frequency response curve is generated by combining the howling frequency points and the response weights of each acoustic zone. Specifically, this includes: S31: Perform spectral analysis on the echo spectrum in the acoustic feature database, extract the amplitude response features and frequency persistence features of each frequency point, and determine the energy value of the frequency point based on the amplitude response features.

[0049] Specifically, by performing a short-time Fourier transform on the echo signal collected by the user equipment in a specific spatial scene, the spectral intensity distribution map of each frequency point in the range of 0Hz to 20kHz is obtained. Then, the amplitude change trend of each frequency point in different time windows is calculated to extract amplitude response features. At the same time, it is recorded whether the frequency point continues to appear in several consecutive frames to form frequency persistence features. Based on this, the energy value of each frequency point is estimated by numerically normalizing the amplitude response features and combining them with the energy model curve in the historical sampling data. For example, in a certain audio calibration scenario, it is found that the amplitude response of the 3200Hz frequency point always maintains a high value and is stable and continuous. This frequency point can be identified as an energy concentration point and its corresponding energy value is recorded for subsequent processing.

[0050] S32: Compare the frequency persistence characteristics with the preset frequency stability threshold, and compare the energy value of the frequency point with the preset energy threshold to select the frequency points that simultaneously meet the judgment conditions as the target howling frequency point set.

[0051] Specifically, the frequency stability threshold is set to appear continuously for no less than 5 frames, and the energy threshold is the frequency energy range between -10dB and -5dB. In the spectrum analysis results, it is determined whether the persistence characteristics of each frequency point are higher than the stability threshold and whether its energy value is within the valid range. If both conditions are met, the frequency point is included in the target howling frequency point set. For example, in actual testing, it was found that after playing back voice in a room, sharp howling phenomena repeatedly occurred at 4800Hz and 5200Hz. After analysis, it was found that the above two frequency points appeared continuously for more than 8 frames and the energy values ​​were -6.3dB and -5.1dB, respectively, which met the preset threshold standard. Therefore, they were identified as target howling frequency points.

[0052] S33: Based on the target howling frequency set and the response weights corresponding to each acoustic zone, perform weighted fusion to generate energy adjustment parameters for each frequency point, and construct a preliminary frequency response curve based on the energy adjustment parameters.

[0053] Specifically, firstly, based on the preset sound field partitioning model in the acoustic space, the echo response intensity corresponding to the target howling frequency point in each partition is converted into a response weight. Then, the energy value of each target frequency point is weighted and summed using this weight as a weighting coefficient to determine the energy attenuation value or gain coefficient of the frequency point that needs to be adjusted. Finally, the adjustment parameters of all frequency points are combined into a frequency response correction curve, which serves as the initial basis for subsequent sound field compensation. For example, in a test in a conference room, the howling frequency was detected to be concentrated between 4000Hz and 6000Hz, with the first half of the space mainly reflecting. After weighted processing, a frequency response curve with obvious low-pass suppression characteristics is generated, which will be used for spectrum adjustment in the dynamic EQ processing stage.

[0054] In one embodiment, such as Figure 5 As shown, in step S40, the preliminary frequency response curve is adjusted based on the user's historical preference data to generate the target frequency response curve, specifically including: S41: Extract user's historical listening records, manual adjustment behavior and scene label information from user historical preference data to construct a user preference feature set, which includes frequency band preference, loudness selection and filter adjustment parameters.

[0055] Specifically, historical preference data is structured and analyzed to extract behavioral characteristics related to sound field adjustment during past use. These characteristics include the types of audio content listened to by the user at different times, the manual equalizer adjustments performed on specific tracks, and the label information of the environment. Based on this, a set of user preference features is generated. Parameters such as the gain amplitude preferred by the user in the low-frequency region, the compression behavior in the high-frequency region, and the adjustment method of the mid-frequency filter are mapped to three types of structured information: frequency band preference, loudness selection, and filter adjustment parameters. For example, if a user often listens to classical music at night and repeatedly increases the gain values ​​of the 250Hz and 1kHz bands, while frequently reducing the overall loudness in quiet environments, this is recorded as a composite feature of enhanced mid-low frequency preference and reduced loudness preference, so as to facilitate subsequent personalized sound field adaptation.

[0056] S42: Compare the preliminary frequency response curve with the user preference feature set, adjust the preliminary frequency response curve based on the preset loudness compensation model, and generate the target frequency response curve.

[0057] Specifically, the gain values ​​of each frequency band in the preliminary frequency response curve are compared one by one with the user preference feature set to identify frequency ranges with significant differences and determine the part of the curve that needs to be corrected. Then, a preset loudness compensation model is called to perform gain or attenuation adjustment on the identified target frequency band. The loudness compensation model uses a curve superposition method to combine the user's psychological loudness curve at common volume levels to reshape the current frequency response curve. For example, when the user prefers to obtain a higher energy perception in the 500Hz to 2kHz range, but there is a significant dip in this range in the preliminary frequency response curve, the model will make reference adjustments based on the ISO226 equal loudness curve, appropriately increase the gain value of this frequency band, and control the smoothness of the change between adjacent frequency points, finally generating a target frequency response curve that is consistent with the user's actual listening expectation. Furthermore, the construction of the loudness compensation model includes modeling and training based on users' historical audio playback behavior and psychoacoustic standards. In the training phase, subjective rating data of multiple users on multiple sets of frequency response curves in different listening environments are first collected. The correlation mapping relationship between frequency response curves and psychological loudness perception is established by combining the actual gain of each frequency band with the ISO 226 equal loudness curve. Then, a shallow neural network is used as the basic model structure, and user preference features (including volume usage habits, frequency band gain preferences, scene labels, etc.) are used as input variables. The loudness sensitivity scores of users in each frequency band are used as training labels. Iterative training is carried out to optimize the model parameters, so that the model can automatically output a set of gain adjustment amounts that match the user's subjective loudness preferences based on the input frequency response curve and user characteristics. For example, during the training process, it was found that some users prefer mid-frequency enhancement in conference mode and pay more attention to low-frequency energy expression in cinema mode. The model can then form a differentiated compensation weight strategy and solidify it into a set of preset parameters that can be called. Thus, in the running phase, combined with the initial frequency response curve input, the target curve that matches the user's listening perception is quickly output for dynamic compensation adjustment.

[0058] In one embodiment, such as Figure 6 As shown, in step S50, the target frequency response curve is applied to the audio output path to output the audio signal, and the corresponding frequency response adjustment result is displayed through the interface program. Specifically, this includes: S51: Based on the target frequency response curve, generate filter gain control parameters corresponding to each frequency band, and load the control parameters into the digital filtering module in the audio output path.

[0059] Specifically, based on the constructed target frequency response curve, the complete audible frequency band is segmented and matched with corresponding gain adjustment requirements. By analyzing the target gain value of each frequency point, a set of digital filter gain control parameters corresponding to it is generated. This set of control parameters is then written into a configuration file and loaded into the digital filter module in the audio output path through audio processing instructions. This allows the subsequent audio signal to apply the corresponding frequency response calibration strategy in real time on the output path. For example, when the target frequency response curve is set to increase the gain by +3dB at the 2kHz frequency band, the generated filter control parameters will include the +3dB gain factor corresponding to the 2kHz frequency point. This parameter will be applied to the filter to amplify the signal in that frequency band, thereby achieving the purpose of frequency response compensation.

[0060] S52: Based on the filter control parameters generated from the target frequency response curve, the frequency response of the output audio signal is adjusted, and the target frequency response curve is displayed in real time in the form of a frequency-gain relationship graph through the interface program, and the adjustment range of each frequency band is marked.

[0061] Specifically, after loading the generated filter control parameters, frequency band separation and gain transformation operations are performed on the output audio signal. Each frequency band signal is input to the corresponding filter channel for amplitude adjustment. The adjustment results are simultaneously fed back to the graphical interface program during the real-time construction of the frequency-gain relationship graph by the graphics rendering engine. The interface displays an interactive curve view with frequency on the horizontal axis and gain on the vertical axis. The gain adjustment amplitude of each key frequency band is marked on the curve with text or color labels. For example, if the attenuation is set to -4dB in the 800Hz to 1.2kHz range, the interface curve will show a concave trend in this range and be marked with the value "-4dB" to improve the user's visual understanding of the frequency response adjustment strategy and the efficiency of calibration confirmation.

[0062] In one embodiment, such as Figure 7 As shown, in step S60, the preliminary frequency response curve is iteratively optimized based on the feedback information, specifically including: S61: The feedback information is parsed into frequency response adjustment factors, and the corresponding frequency response optimization parameter set is generated. The parsing process includes weighted analysis of frequency band tendencies and frequency response difference fitting of the preference tilt direction.

[0063] Specifically, the feedback information includes subjective evaluation data generated by users during the playback of reference audio segments and acoustic feedback parameters collected from the device. By calculating the distribution weights of feedback preference values ​​on different frequency bands and constructing frequency band response score vectors, a weighted analysis of frequency band tendencies is performed to determine the response characteristic trends presented by user preferences on each frequency band. Furthermore, based on the direction of this trend, a frequency response difference model is constructed between the original frequency response curve and the target response. The key change intervals are extracted using linear fitting or polynomial fitting methods. Finally, a set of frequency response optimization parameters is generated, which includes gain adjustment coefficients, filter shift intervals, and weighted target frequency band indices. For example, when users express dissatisfaction with the excessive brightness of the mid-frequency sound in their feedback, the generated optimization parameters will include a gain reduction instruction for the 1kHz to 3kHz frequency band and a suppression transition strategy within adjacent bandwidths.

[0064] S62: Based on the frequency response optimization parameter set, perform adjustment operations on the initial frequency response curve to generate a new target frequency response curve and replace the original target frequency response curve to update the response characteristics in the audio output path. The adjustment operations include gain fine-tuning, filter parameter reconstruction, or local suppression enhancement.

[0065] Specifically, during the adjustment operation, the adjustment coefficients of each frequency band in the frequency response optimization parameter set are applied to the corresponding preliminary frequency response curve data points. In the filter parameter reconstruction stage, an appropriate filter type, such as a Butterworth filter, Chebyshev filter, or adaptive IIR filter, is selected according to the target characteristics of the frequency band. Based on the optimization parameters, the cutoff frequency, passband width, and stopband attenuation are recalculated. On this basis, a new target frequency response curve is constructed to cover the original curve and the replacement is completed in the audio signal processing path to ensure that the response characteristics of the subsequent output signal are consistent with the optimized frequency response characteristics. For example, if it is detected that the user expects an enhanced low-frequency bass during a certain adjustment process, a gain boost can be performed in the 50Hz to 120Hz range and a band-stop filter structure can be constructed at 300Hz to avoid low-frequency muddiness, thereby generating an output frequency response profile with stronger layering and clarity.

[0066] It should be understood that the sequence number of each step in the above embodiments does not imply the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of this application.

[0067] In one embodiment, an intelligent sound field calibration system is provided, which corresponds one-to-one with the intelligent sound field calibration method described in the above embodiments. For example... Figure 8 As shown, this intelligent sound field calibration system includes an acoustic acquisition module, a topology modeling module, a frequency response calculation module, a preference matching module, an output display module, and a feedback optimization module. Detailed descriptions of each functional module are as follows: The acoustic playback and acquisition module is used to send a frequency sweep signal within a preset frequency range to multiple microphone acquisition positions in the sound field, and each microphone acquires the echo signal and constructs an acoustic feature database based on the echo signal; The topology modeling module is used to extract the relative positional relationships between microphones and the corresponding echo signal features from the acoustic feature database. Based on the relative positional relationships and echo signal features, it constructs a sound field spatial topology map and divides the acoustic partitions, and determines the response weight of each partition. The frequency response calculation module is used to identify and mark howling frequency points based on the acoustic feature database and the spectrum analysis algorithm, and generate a preliminary frequency response curve by combining the howling frequency points and the response weights of each acoustic zone. The preference matching module is used to adjust the initial frequency response curve based on the user's historical preference data and generate the target frequency response curve; The output display module is used to apply the target frequency response curve to the audio output path and output the audio signal, and display the corresponding frequency response adjustment results through the interface program. The feedback optimization module is used to collect user feedback information on the current audio signal. If the feedback result is unsatisfactory, the initial frequency response curve is iteratively optimized based on the feedback information to generate a new target frequency response curve to replace the original target frequency response curve.

[0068] Optional acoustic playback and acquisition modules include: The frequency sweep signal generation submodule is used to generate a frequency sweep signal within a preset frequency range and synchronously send the frequency sweep signal to multiple microphone acquisition positions deployed in the sound field. The acoustic parameter extraction submodule is used to receive echo signals from corresponding locations at each microphone, extract features from the echo signals to obtain acoustic feature parameters, and summarize the acoustic feature parameters to build an acoustic feature database. The acoustic feature parameters include frequency domain response features, energy attenuation parameters, and reflection delay information.

[0069] Optionally, the topology modeling module includes: The geometric relationship calculation submodule is used to calculate the spatial geometric relationship between each microphone based on the relative positional relationship of each microphone and the propagation delay features extracted from the corresponding echo signal features, and to generate a sound field spatial topology map based on the geometric relationship. The partitioning submodule is used to divide the sound field into multiple acoustic partitions based on the density of microphone distribution, the spatial region, and the sound wave reflection characteristics in the spatial topology diagram. The weight allocation submodule is used to comprehensively evaluate the sound wave contribution of each acoustic zone at the user's primary listening position based on each acoustic zone, combined with the sound wave attenuation characteristics and the user's primary listening position, and to allocate response weights to each acoustic zone according to the sound wave contribution.

[0070] Optionally, the frequency response calculation module includes: The spectrum analysis submodule is used to perform spectrum analysis on the echo spectrum in the acoustic feature database, extract the amplitude response characteristics and frequency duration characteristics of each frequency point, and determine the energy value of the frequency point based on the amplitude response characteristics. The frequency point filtering submodule is used to compare the frequency persistence characteristics with the preset frequency stability threshold, and at the same time compare the energy value of the frequency point with the preset energy threshold, and filter out the frequency points that meet the judgment conditions as the target howling frequency point set. The frequency response construction submodule is used to perform weighted fusion based on the target howling frequency point set and the response weights corresponding to each acoustic zone, generate energy adjustment parameters for each frequency point, and construct a preliminary frequency response curve based on the energy adjustment parameters.

[0071] Optional, the preference matching module includes: The user feature extraction submodule is used to extract user historical listening records, manual adjustment behavior and scene label information from user historical preference data, and construct a user preference feature set, which includes frequency band preference, loudness selection and filter adjustment parameters. The personalized adjustment submodule is used to compare the preliminary frequency response curve with the user preference feature set, adjust the preliminary frequency response curve based on the preset loudness compensation model, and generate the target frequency response curve.

[0072] Optional, the output display module includes: The filter parameter generation submodule is used to generate filter gain control parameters corresponding to each frequency band based on the target frequency response curve, and load the control parameters into the digital filtering module in the audio output path; The frequency response display submodule is used to adjust the frequency response of the output audio signal based on the filter control parameters generated by the target frequency response curve. The target frequency response curve is displayed in real time in the form of a frequency-gain relationship graph through the interface program, and the adjustment range of each frequency band is marked.

[0073] Optional, the feedback optimization module includes: The feedback analysis submodule is used to analyze the feedback information into frequency response adjustment factors and generate the corresponding frequency response optimization parameter set. The analysis process includes weighted analysis of frequency band tendencies and frequency response difference fitting for the preference tilt direction. The iterative optimization submodule is used to perform adjustment operations on the initial frequency response curve based on the frequency response optimization parameter set, generate a new target frequency response curve and replace the original target frequency response curve to update the response characteristics in the audio output path. The adjustment operations include gain fine-tuning, filter parameter reconstruction or local suppression enhancement.

[0074] For specific limitations regarding an intelligent sound field calibration system, please refer to the limitations of an intelligent sound field calibration method described above, which will not be repeated here. Each module in the aforementioned intelligent sound field calibration system can be implemented entirely or partially through software, hardware, or a combination thereof. These modules can be embedded in or independent of the processor in a computer device in hardware form, or stored in the memory of a computer device in software form, so that the processor can call and execute the corresponding operations of each module.

[0075] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the above-described division of functional units and modules is used as an example. In practical applications, the above functions can be assigned to different functional units and modules as needed, that is, the internal structure of the system can be divided into different functional units or modules to complete all or part of the functions described above.

[0076] The above-described embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit them. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of this application, and should all be included within the protection scope of this application.

Claims

1. A smart sound field calibration method, characterized in that, The intelligent sound field calibration method includes: A frequency sweep signal within a preset frequency range is sent to multiple microphone acquisition positions in the sound field, and each microphone acquires the echo signal, and an acoustic feature database is constructed based on the echo signal; The relative positional relationships between microphones and their corresponding echo signal features are extracted from the acoustic feature database. Based on the relative positional relationships and echo signal features, a sound field spatial topology map is constructed and acoustic partitions are divided, and the response weight of each partition is determined. Based on the acoustic feature database, a spectrum analysis algorithm is used to identify and mark the howling frequency points, and a preliminary frequency response curve is generated by combining the howling frequency points and the response weights of each acoustic zone. The initial frequency response curve is adjusted based on the user's historical preference data to generate the target frequency response curve; The target frequency response curve is applied to the audio output path to output an audio signal, and the corresponding frequency response adjustment result is displayed through the interface program. The system collects user feedback on the current audio signal. If the feedback is unsatisfactory, it iteratively optimizes the preliminary frequency response curve based on the feedback to generate a new target frequency response curve to replace the original target frequency response curve.

2. The intelligent sound field calibration method according to claim 1, characterized in that, The process of sending a frequency sweep signal within a preset frequency range to multiple microphone acquisition positions in the sound field, having each microphone acquire echo signals, and constructing an acoustic feature database based on the echo signals specifically includes: A frequency sweep signal is generated within a preset frequency range, and the frequency sweep signal is synchronously sent to multiple microphone acquisition positions deployed in the sound field. Each microphone receives the echo signal at the corresponding location, extracts features from the echo signal to obtain acoustic feature parameters, and summarizes the acoustic feature parameters to construct an acoustic feature database. The acoustic feature parameters include frequency domain response features, energy attenuation parameters, and reflection delay information.

3. The intelligent sound field calibration method according to claim 1, characterized in that, Based on the relative positional relationships and echo signal characteristics, the sound field spatial topology map is constructed and acoustic partitions are divided, and the response weight of each partition is determined, specifically including: Based on the relative positional relationship of each microphone and the propagation delay features extracted from the corresponding echo signal features, the spatial geometric relationship between each microphone is calculated, and the sound field spatial topology map is generated based on the geometric relationship. Based on the density of microphone distribution, the spatial region they are located in, and the sound wave reflection characteristics in the aforementioned spatial topology diagram, the sound field is divided into multiple acoustic zones. Based on each acoustic zone, and taking into account the sound wave attenuation characteristics and the user's primary listening position, the sound wave contribution of each acoustic zone at the user's primary listening position is comprehensively evaluated, and a response weight is assigned to each acoustic zone according to the sound wave contribution.

4. The intelligent sound field calibration method according to claim 1, characterized in that, Based on the acoustic feature database, a spectrum analysis algorithm is used to identify and mark howling frequency points. A preliminary frequency response curve is generated by combining the howling frequency points with the response weights of each acoustic zone. Specifically, this includes: Perform spectral analysis on the echo spectrum in the acoustic feature database, extract the amplitude response features and frequency persistence features of each frequency point, and determine the energy value of the frequency point based on the amplitude response features; The frequency persistence characteristics are compared with a preset frequency stability threshold, and the energy value of the frequency point is compared with a preset energy threshold. Frequency points that simultaneously meet the judgment conditions are selected as the target howling frequency point set. The target howling frequency set is weighted and fused with the response weights corresponding to each acoustic zone to generate energy adjustment parameters for each frequency point, and the preliminary frequency response curve is constructed based on the energy adjustment parameters.

5. The intelligent sound field calibration method according to claim 1, characterized in that, The step of adjusting the preliminary frequency response curve based on user historical preference data to generate the target frequency response curve specifically includes: Extract user's historical listening records, manual adjustment behavior and scene label information from the user's historical preference data to construct a user preference feature set, wherein the user preference feature set includes frequency band preference, loudness selection and filter adjustment parameters; The preliminary frequency response curve is compared with the user preference feature set, and the preliminary frequency response curve is adjusted based on a preset loudness compensation model to generate the target frequency response curve.

6. The intelligent sound field calibration method according to claim 1, characterized in that, The step of applying the target frequency response curve to the audio output path and outputting the audio signal, and displaying the corresponding frequency response adjustment result through the interface program, specifically includes: Based on the target frequency response curve, filter gain control parameters corresponding to each frequency band are generated, and the control parameters are loaded into the digital filtering module in the audio output path; Based on the filter control parameters generated from the target frequency response curve, the frequency response of the output audio signal is adjusted, and the target frequency response curve is displayed in real time in the form of a frequency-gain relationship graph through the interface program, with the adjustment range of each frequency band marked.

7. The intelligent sound field calibration method according to claim 1, characterized in that, The iterative optimization of the preliminary frequency response curve based on the feedback information specifically includes: The feedback information is parsed into a frequency response adjustment factor, and a corresponding set of frequency response optimization parameters is generated. The parsing process includes weighted analysis of frequency band tendencies and frequency response difference fitting of the preference tilt direction. Based on the frequency response optimization parameter set, the initial frequency response curve is adjusted to generate a new target frequency response curve and replace the original target frequency response curve to update the response characteristics in the audio output path. The adjustment operation includes gain fine-tuning, filter parameter reconstruction, or local suppression enhancement.

8. An intelligent sound field calibration system, characterized in that, The intelligent sound field calibration system includes: The acoustic playback and acquisition module is used to send a frequency sweep signal within a preset frequency range to multiple microphone acquisition positions in the sound field, and each microphone acquires the echo signal and constructs an acoustic feature database based on the echo signal. The topology modeling module is used to extract the relative positional relationships between microphones and the corresponding echo signal features from the acoustic feature database, and based on the relative positional relationships and echo signal features, construct a sound field space topology map and divide it into acoustic partitions, and determine the response weight of each partition. The frequency response calculation module is used to identify and mark howling frequency points based on the acoustic feature database using a spectrum analysis algorithm, and generate a preliminary frequency response curve by combining the howling frequency points and the response weights of each acoustic zone; The preference matching module is used to adjust the preliminary frequency response curve based on the user's historical preference data to generate the target frequency response curve; The output display module is used to apply the target frequency response curve to the audio output path to output an audio signal, and to display the corresponding frequency response adjustment result through the interface program. The feedback optimization module is used to collect user feedback information on the current audio signal. If the feedback result is unsatisfactory, the initial frequency response curve is iteratively optimized based on the feedback information to generate a new target frequency response curve to replace the original target frequency response curve.

Citation Information

Cited By

  • Room sound calibration method based on hearing tendency control and related equipment

    CN122054066A