Sound effect optimization method and device, equipment and storage medium
By acquiring the relative positional relationship between the speaker and the user and analyzing the room impulse response, a real-time acoustic spatial model is constructed. This solves the problems of sound image positioning shift and audio clarity degradation in multi-speaker systems when the device moves or the listener's position changes, thus achieving a stable immersive listening experience.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- LINKPLAY TECHNOLOGY INC NANJING
- Filing Date
- 2025-12-16
- Publication Date
- 2026-05-08
AI Technical Summary
Existing multi-speaker systems cannot detect changes in spatial relationships when the equipment is moved or the listener's position changes, resulting in sound image positioning deviation, decreased audio clarity, and inconsistent listening experience. They also lack the ability to analyze room impulse response in real time.
By acquiring data on the relative positional relationship between the speaker and the user, and combining this with room impulse response analysis, acoustic space modeling is performed to construct a real-time spatial model. Channel mapping weights are calculated, and equalizer curves, delay compensation, and phase parameters are adjusted in real time.
It achieves adaptive calculation and allocation of channel assignment, and can optimize the channel assignment scheme in real time according to the user's displacement, improve the accuracy of sound image positioning, the clarity of human voice dialogue and the consistency of low frequency response, and maintain a stable and immersive listening experience.
Smart Images

Figure CN122002183A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of electronic digital data processing technology, and in particular to a sound effect optimization method, apparatus, device, and storage medium. Background Technology
[0002] Currently, most mainstream multi-speaker systems employ a fixed channel allocation mode, meaning that audio rendering is performed based on preset roles for left and right channels, center channel, and low-frequency effect channels during device initialization. The system typically relies on static configuration parameters (such as speaker distance and channel roles) for equalization (EQ) and delay compensation, and only supports manual user calibration, failing to respond to device displacement or user movement during use. This static sound field optimization method has significant limitations: when speakers are moved or the listener's position changes, the system cannot perceive changes in spatial relationships, leading to problems such as sound image localization shift, decreased clarity of dialogue, and overlapping or missing low-frequency responses. Furthermore, traditional methods lack real-time analysis capabilities of room impulse response, making it difficult to adapt to the acoustic propagation characteristics in complex environments, ultimately affecting the consistency and stability of the listening experience. Summary of the Invention
[0003] This invention provides a sound effect optimization method, apparatus, device, and storage medium to solve the problems in the prior art where the static sound field configuration cannot adapt to device displacement and listener movement, resulting in sound image positioning deviation, decreased audio clarity, and inconsistent listening experience.
[0004] The first aspect of this invention provides a sound effect optimization method, comprising: acquiring relative positional relationship data between each speaker and the user; performing acoustic spatial modeling based on the relative positional relationship and combining it with the analysis of room impulse response to obtain a real-time spatial model including geometric attributes and acoustic propagation characteristics; calculating and allocating channel mapping weights based on the real-time spatial model to generate a channel allocation scheme that dynamically changes with the user's position; and adjusting the equalizer curve, delay compensation, and phase parameters of each speaker in real time according to the channel allocation scheme.
[0005] In one feasible implementation, acquiring the relative positional relationship data between each speaker and the user includes: initially calculating the first relative positional data between the devices based on the wireless communication signal strength and flight time between the main speaker and each subordinate speaker; acquiring visual scene information through a smart camera associated with the system, and using a computer vision algorithm to identify the user's position and outline to obtain the second relative positional data between the user and the main speaker; inputting the first relative positional data and the second relative positional data into a Kalman filter model for data fusion and trajectory prediction, and outputting the dynamic relative positional relationship between each speaker and the user.
[0006] In one feasible implementation, the step of performing spatial modeling processing based on the relative positional relationship data to obtain a real-time spatial model describing the spatial distribution relationship between the speaker and the user includes: constructing an initial spatial network describing the geometric positional relationship between the speaker and the user based on the relative positional relationship; playing a known test signal through at least one speaker and receiving it through at least one microphone based on the initial spatial network; extracting the impulse response of the room by calculating the cross-correlation function between the test signal and the received signal; analyzing the impulse response, identifying and quantifying the energy and delay of direct sound, early reflection sound, and reverberation sound, and mapping them to the initial spatial network to generate a real-time spatial model that combines geometric properties and acoustic propagation characteristics.
[0007] In one feasible implementation, the step of constructing an initial spatial network describing the geometric positional relationship between the speakers and the user based on the relative positional relationship includes: obtaining the absolute positional coordinates of each speaker and the user based on the relative positional relationship; connecting the position points of each speaker and the user to construct a spatial triangular mesh describing the sound wave propagation path; labeling the device type attribute of each node in the spatial triangular mesh, and calculating the straight-line distance between each adjacent node to complete the construction of the initial spatial network.
[0008] In one feasible implementation, based on the initial spatial network, a known test signal is played through at least one speaker and received by at least one microphone. The impulse response of the room is extracted by calculating the cross-correlation function between the test signal and the received signal. This includes: dynamically selecting one or more optimal speakers as test signal transmission sources according to the topology and node density of the initial spatial network, and generating a composite excitation signal combining an exponentially swept frequency signal and a pseudo-random sequence for playback; synchronously acquiring reverberation signals reflected by the room through a microphone array distributed in the user's typical listening area, and using an adaptive filter based on minimum mean square error to perform echo cancellation and background noise suppression on the acquired signals; performing a fast cross-correlation operation on the preprocessed received signal and the original excitation signal, and using a frequency-domain weighted Wiener deconvolution algorithm to separate and extract the high signal-to-noise ratio impulse response that only reflects the acoustic transmission characteristics of the room.
[0009] In one feasible implementation, the analysis of the impulse response, identifying and quantifying the energy and delay of direct sound, early reflections, and reverberation, and mapping them to the initial spatial network to generate a real-time spatial model with both geometric properties and acoustic propagation characteristics, includes: In the impulse response, segmenting the time-domain intervals of direct sound, early reflections, and reverberation based on a dynamic threshold detection algorithm, and quantifying and recording the arrival time, sound pressure level, and spectral characteristics of the acoustic components within each interval; inputting the quantized early reflection sequence into a clustering analysis model to identify the main reflection clusters, and reconstructing the corresponding virtual reflectors in the initial spatial network based on the arrival time difference and azimuth information using a ray tracing method; fusing the decay time parameter of the reverberation field with the reconstructed virtual reflector structure, labeling the initial spatial network with acoustic properties, and generating a real-time spatial model containing geometric topology and acoustic propagation paths.
[0010] In one feasible implementation, the calculation and allocation of channel mapping weights based on the real-time spatial model to generate a channel allocation scheme that dynamically changes with the user's position includes: calculating the binaural acoustic transfer function from each speaker to the user's position based on the geometric relationship and acoustic propagation characteristics between each speaker and the user in the real-time spatial model; using the binaural acoustic transfer function as the sound field reconstruction target, decomposing the standard multi-channel signal based on the principle of sound image localization using a vector basis amplitude translation algorithm, and calculating the initial weight coefficients corresponding to each speaker; combining the user's real-time movement data, and according to the constraint relationship of the binaural acoustic transfer function on sound image localization, applying a psychoacoustic model to dynamically optimize and smooth the transition of the initial weights to generate a channel allocation scheme.
[0011] In one feasible implementation, the step of calculating the binaural acoustic transfer function from each speaker to the user's location based on the geometric relationship and acoustic propagation characteristics between the user and each speaker in the real-time spatial model includes: extracting the direct sound propagation path from each speaker to the user's ears in the real-time spatial model, and calculating the length difference and azimuth angle of each path; combining the room reflection parameters obtained from impulse response analysis to simulate and calculate the superposition effect of early reflected sound and reverberation sound on each propagation path; and generating a binaural acoustic transfer function that includes spatial filtering effects by integrating the amplitude, time delay, and phase relationship of direct sound and reflected sound based on the principle of acoustic interference.
[0012] In one feasible implementation, the step of using the binaural acoustic transfer function as the sound field reconstruction target, decomposing the standard multichannel signal based on the sound image localization principle and employing a vector basis amplitude translation algorithm to calculate the initial weight coefficients corresponding to each speaker includes: mapping each channel in the standard multichannel signal to a virtual sound source plane to generate a target sound field distribution containing the directional information of each virtual sound source; using the target sound field distribution as the desired sound image and the binaural acoustic transfer function as the acoustic constraint for sound field reconstruction, constructing a sound field reconstruction mathematical model based on the vector basis amplitude translation algorithm; solving the sound field reconstruction mathematical model, calculating the reconstruction contribution of each speaker to each virtual sound source in the target sound field under the acoustic constraint, thereby obtaining the initial weight coefficients of each speaker for each channel signal.
[0013] In one feasible implementation, the step of combining the user's real-time movement data and, based on the constraint relationship of the binaural acoustic transfer function on sound image localization, applying a psychoacoustic model to dynamically optimize and smooth the initial weights to generate a channel allocation scheme includes: calculating the vector deviation between the current sound image position and the target sound image position based on real-time acquired user head position and orientation data, combined with the binaural acoustic transfer function; inputting the vector deviation into the psychoacoustic model to generate dynamic compensation coefficients for the initial weights of each speaker; and applying the dynamic compensation coefficients to the initial weight coefficients using a first-order inertial smoothing algorithm to generate a channel allocation scheme.
[0014] In one feasible implementation, the real-time adjustment of equalizer curves, delay compensation, and phase parameters for each speaker according to the channel allocation scheme includes: generating a personalized target equalization curve for each speaker based on the weighted role of each speaker in the channel allocation scheme and its acoustic propagation path in the real-time spatial model, wherein near-field vocal speakers enhance mid-frequency clarity and far-field surround speakers expand high-frequency spatial sense; calculating microsecond-level delay compensation values based on the time synchronization of sound wave propagation to the user's position and the length difference of the acoustic path of each speaker, and applying group delay calibration to all speakers; and calculating and applying an anti-phase compensation filter by analyzing the impulse response coherence of each speaker in the cross-band to eliminate phase cancellation caused by multi-speaker interference, thereby completing the collaborative optimization of acoustic parameters.
[0015] In one feasible implementation, the step of using the time synchronization of sound wave propagation to the user's location as a reference, calculating microsecond-level delay compensation values based on the length difference of the acoustic paths of each speaker, and applying group delay calibration to all speakers includes: selecting the speaker with the longest acoustic path to the user's location as a reference, calculating the sound wave propagation time difference of other speakers relative to this reference; converting the sound wave propagation time difference into digital delay parameters with sampling point precision, and configuring corresponding fractional delay filters; and achieving sub-sampling precision delay calibration through multi-phase filtering interpolation technology to ensure that the sound waves emitted by all speakers are synchronized at the user's location.
[0016] A second aspect of the present invention provides a sound effect optimization device, comprising: an acquisition module for acquiring relative positional relationship data between each speaker and the user; a modeling module for performing acoustic spatial modeling based on the relative positional relationship and combined with the analysis of room impulse response to obtain a real-time spatial model including geometric attributes and acoustic propagation characteristics; a generation module for calculating and allocating channel mapping weights based on the real-time spatial model to generate a channel allocation scheme that dynamically changes with the user's position; and an adjustment module for performing real-time adjustments to the equalizer curves, delay compensation, and phase parameters of each speaker according to the channel allocation scheme.
[0017] In one feasible implementation, the acquisition module is specifically used to: preliminarily calculate the first relative position data between the devices by using the wireless communication signal strength and flight time between the main speaker and each subordinate speaker; collect visual scene information by using the smart camera associated with the system, and use computer vision algorithms to identify the user's position and outline to obtain the second relative position data between the user and the main speaker; input the first relative position data and the second relative position data into a Kalman filter model for data fusion and trajectory prediction, and output the dynamic relative position relationship between each speaker and the user.
[0018] In one feasible implementation, the modeling module includes: a construction unit for constructing an initial spatial network describing the geometric positional relationship between the speaker and the user based on the relative positional relationship; an extraction unit for extracting the impulse response of the room by playing a known test signal through at least one speaker and receiving it through at least one microphone based on the initial spatial network, and by calculating the cross-correlation function between the test signal and the received signal; and a first generation unit for analyzing the impulse response, identifying and quantifying the energy and delay of direct sound, early reflections, and reverberation, and mapping them to the initial spatial network to generate a real-time spatial model that combines geometric properties and acoustic propagation characteristics.
[0019] In one feasible implementation, the construction unit is specifically used to: obtain the absolute position coordinates of each speaker and user based on the relative position relationship; connect the position points of each speaker and user to construct a spatial triangular mesh describing the sound wave propagation path; label the device type attribute of each node in the spatial triangular mesh, and calculate the straight-line distance between each adjacent node to complete the construction of the initial spatial network.
[0020] In one feasible implementation, the extraction unit is specifically used to: dynamically select one or more optimal speakers as test signal transmission sources based on the topology and node density of the initial spatial network, and generate a composite excitation signal combining an exponentially swept frequency signal and a pseudo-random sequence for playback; synchronously acquire reverberation signals reflected by the room through a microphone array distributed in the user's typical listening area, and use an adaptive filter based on minimum mean square error to perform echo cancellation and background noise suppression on the acquired signals; perform fast cross-correlation calculation on the preprocessed received signal and the original excitation signal, and use a frequency-domain weighted Wiener deconvolution algorithm to separate and extract the high signal-to-noise ratio impulse response that only reflects the acoustic transmission characteristics of the room.
[0021] In one feasible implementation, the first generation unit is specifically used to: segment the time domain intervals of direct sound, early reflection sound, and reverberation sound in the impulse response based on a dynamic threshold detection algorithm, and quantize and record the arrival time, sound pressure level, and spectral characteristics of the acoustic components in each interval; input the quantized early reflection sound sequence into a clustering analysis model to identify the main reflection sound clusters, and reconstruct the corresponding virtual reflector in the initial spatial network based on the arrival time difference and azimuth information using a ray tracing method; fuse the decay time parameter of the reverberation sound field with the reconstructed virtual reflector structure, label the acoustic attributes of the initial spatial network, and generate a real-time spatial model containing geometric topology and acoustic propagation paths.
[0022] In one feasible implementation, the generation module includes: a first calculation unit, used to calculate the binaural acoustic transfer function from each speaker to the user's position based on the geometric relationship and acoustic propagation characteristics between each speaker and the user in the real-time spatial model; a second calculation unit, used to decompose the standard multichannel signal using the binaural acoustic transfer function as the sound field reconstruction target, based on the principle of sound image localization, and using a vector basis amplitude translation algorithm to calculate the initial weight coefficients corresponding to each speaker; and a second generation unit, used to combine the user's real-time movement data and, according to the constraint relationship of the binaural acoustic transfer function on sound image localization, apply a psychoacoustic model to dynamically optimize and smooth the transition of the initial weights to generate a channel allocation scheme.
[0023] In one feasible implementation, the first computing unit is specifically used to: extract the direct sound propagation paths from each speaker to the user's ears in the real-time spatial model, and calculate the length difference and azimuth angle of each path; combine the room reflection parameters obtained from impulse response analysis to simulate and calculate the superposition effect of early reflected sound and reverberation sound on each propagation path; and based on the principle of acoustic interference, synthesize the amplitude, time delay, and phase relationship of direct sound and reflected sound to generate a binaural acoustic transfer function that includes spatial filtering effects.
[0024] In one feasible implementation, the second computing unit is specifically used to: map each channel in the standard multi-channel signal to a virtual sound source plane to generate a target sound field distribution containing the orientation information of each virtual sound source; construct a sound field reconstruction mathematical model based on the vector basis amplitude translation algorithm, using the target sound field distribution as the desired sound image and the binaural acoustic transfer function as the acoustic constraint for sound field reconstruction; solve the sound field reconstruction mathematical model, calculate the reconstruction contribution of each speaker to each virtual sound source in the target sound field under the acoustic constraint, and thus obtain the initial weighting coefficient of each speaker to each channel signal.
[0025] In one feasible implementation, the second generation unit is specifically used to: calculate the vector deviation between the current sound image position and the target sound image position based on the real-time acquired user head position and orientation data, combined with the binaural acoustic transfer function; input the vector deviation into the psychoacoustic model to generate dynamic compensation coefficients for the initial weights of each speaker; and apply the dynamic compensation coefficients to the initial weight coefficients using a first-order inertial smoothing algorithm to generate a channel allocation scheme.
[0026] In one feasible implementation, the adjustment module includes: a third generation unit, used to generate a personalized target equalization curve for each speaker based on the weighted role of each speaker in the channel allocation scheme and its acoustic propagation path in the real-time spatial model, wherein near-field vocal speakers enhance mid-frequency clarity and far-field surround speakers expand high-frequency spatial sense; a processing unit, used to calculate microsecond-level delay compensation values based on the time synchronization of sound wave propagation to the user's position and the length difference of the acoustic path of each speaker, and apply group delay calibration to all speakers; and an optimization unit, used to calculate and apply an anti-phase compensation filter by analyzing the impulse response coherence of each speaker in the cross-band, to eliminate phase cancellation caused by multi-speaker interference, and to complete the collaborative optimization of acoustic parameters.
[0027] In one feasible implementation, the processing unit is specifically used to: select the speaker with the longest acoustic path to the user's location as a reference, calculate the sound wave propagation time difference of other speakers relative to the reference; convert the sound wave propagation time difference into a digital delay parameter with sampling point precision, and configure a corresponding fractional delay filter; and achieve sub-sampling precision delay calibration through multi-phase filtering interpolation technology to ensure that the sound waves emitted by all speakers are synchronized at the user's location.
[0028] A third aspect of the present invention provides an electronic device, comprising: a memory and at least one processor, wherein the memory stores instructions; the at least one processor invokes the instructions in the memory to cause the electronic device to perform the above-described sound effect optimization method.
[0029] A fourth aspect of the present invention provides a computer-readable storage medium storing instructions that, when executed on a computer, cause the computer to perform the above-described sound effect optimization method.
[0030] The technical solution provided by this invention involves acquiring relative positional relationship data between each speaker and the user; based on the relative positional relationship, and combined with the analysis of room impulse response, performing acoustic spatial modeling to obtain a real-time spatial model containing geometric attributes and acoustic propagation characteristics; calculating and allocating channel mapping weights based on the real-time spatial model to generate a channel allocation scheme that dynamically changes with the user's position; and adjusting the equalizer curve, delay compensation, and phase parameters of each speaker in real time according to the channel allocation scheme.
[0031] In this embodiment of the invention, by acquiring the relative positional relationship between the speakers and the user in real time and combining it with room impulse response analysis, a dynamic spatial model integrating geometric properties and acoustic propagation characteristics is constructed, realizing adaptive calculation and allocation of channel mapping weights. The system can optimize the channel allocation scheme in real time according to the user's displacement and automatically adjust the equalizer, delay, and phase parameters of each speaker, thereby effectively overcoming the limitations of the traditional fixed channel mode, significantly improving the accuracy of sound image positioning, the clarity of human voice dialogue, and the consistency of low-frequency response, and ultimately maintaining a stable and immersive listening experience even when the user moves or the device position changes. Attached Figure Description
[0032] Figure 1 This is a schematic diagram of one embodiment of the sound effect optimization method in this invention; Figure 2 This is a schematic diagram of another embodiment of the sound effect optimization method in this invention; Figure 3 This is a schematic diagram of one embodiment of the sound effect optimization device in this invention; Figure 4This is a schematic diagram of another embodiment of the sound effect optimization device in this invention; Figure 5 This is a schematic diagram of one embodiment of the electronic device in this invention. Detailed Implementation
[0033] This invention provides a sound effect optimization method, apparatus, device, and storage medium that can automatically adjust speaker parameters according to the user's location and the room's acoustic characteristics, thereby maintaining accurate sound field positioning and optimal listening effect as the user moves.
[0034] The terms "first," "second," "third," "fourth," etc. (if present) in the specification, claims, and accompanying drawings of this invention are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" or "having" and any variations thereof are intended to cover a non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.
[0035] It is understood that the executing entity of this invention can be a sound effect optimization device, a terminal, or a server; no specific limitation is made here. This embodiment of the invention will be described using a server as an example.
[0036] For ease of understanding, the specific process of the embodiments of the present invention is described below. Please refer to [link / reference]. Figure 1 One embodiment of the sound effect optimization method in this invention includes: 101. Obtain the relative positional relationship data between each speaker and the user; The master device periodically transmits UWB ultra-wideband signals to each slave speaker (e.g., portable speaker, subwoofer), and accurately measures the relative distance and orientation between the devices by calculating the signal's time of flight and angle of arrival, forming a first location dataset. Simultaneously, the system invokes a smart TV camera linked to the master device, using a deep learning-based object detection algorithm to capture the user's silhouette in real time, and determines the relative coordinates between the user and the master device using triangulation, forming a second location dataset. Finally, these two datasets are input into an extended Kalman filter for data fusion and trajectory prediction, filtering out measurement noise and predicting the user's movement trend, ultimately outputting a real-time "speaker-user" relative position relationship matrix.
[0037] 102. Based on the relative positional relationship and combined with the analysis of the room impulse response, acoustic space modeling is performed to obtain a real-time spatial model that includes geometric properties and acoustic propagation characteristics. Using the main device as the origin, the relative position matrix is converted into absolute coordinates, and an initial spatial geometric grid containing all speakers and user nodes is constructed using the Delaunay triangulation algorithm. A speaker in the central position is dynamically selected to play a composite test signal combining exponential frequency sweep and maximum length sequence, which is received by the microphone built into the user's smartphone. The impulse response of the room is extracted by performing generalized cross-correlation calculation and frequency domain Wiener deconvolution processing on the transmitted and received signals. Subsequently, the impulse response is analyzed, and the direct sound, early reflection sound (5ms-80ms), and reverberation sound (after 80ms) are separated using the Schroeder integral method. Their energy, delay, and spectral characteristics are quantified, and finally these acoustic parameters are mapped as attributes to the corresponding nodes and links of the initial geometric grid to generate a real-time spatial model that combines physical structure and sound wave propagation characteristics.
[0038] 103. Calculate and allocate channel mapping weights based on a real-time spatial model to generate a channel allocation scheme that dynamically changes with the user's position; The acoustic transfer function from each physical speaker to the user's ears is extracted from the real-time spatial model. This acoustic transfer function fully characterizes the sound wave propagation characteristics, including room reflection effects. Each channel of the standard multi-channel signal is positioned as a virtual sound source, and a hybrid algorithm combining high-order Ambisonics encoding and vector basis amplitude shift is used for sound field decomposition. First, the virtual sound source is encoded into a high-order B-Format signal to capture its spatial orientation information. Then, in the playback array composed of physical speakers, an optimal reconstruction unit containing 3 physical speakers is established for each virtual sound source. Using the acoustic transfer function as a constraint, the initial weight coefficients required for each physical speaker to reconstruct the virtual sound source are calculated by solving a weighted least squares optimization problem, ensuring that the sound field reconstruction error is minimized at the user's position. Then, a dynamic perception optimization mechanism is introduced to acquire user head orientation data in real time and dynamically adjust the initial weights according to the priority effect psychoacoustic model. When the user's head turns, the weight priority of the speaker group directly in front is increased, and the weight change is subjected to first-order inertial smoothing. Finally, a dynamic channel allocation scheme is generated that can adapt to the acoustic characteristics of the room and maintain sound image stability as the user moves.
[0039] 104. Based on the channel allocation scheme, perform real-time adjustments to the equalizer curves, delay compensation, and phase parameters of each speaker.
[0040] Based on the role and positioning of each speaker in the channel allocation scheme, a preset equalizer template library is used for personalized configuration. Specifically, for the near-field speakers responsible for the main output of human voices, a "voice enhancement" curve is applied, boosting the frequency band by 3-6dB in the 1-3kHz range to improve voice clarity, while implementing high-pass filtering below 200Hz to eliminate low-frequency muddiness caused by near-field effects. For the far-field speakers responsible for ambient sounds, a "spatial expansion" curve is used, moderately boosting the frequency band above 8kHz by 2-4dB to enhance the sense of airiness, and slightly attenuating the frequency band in the 500-800Hz range to reduce acoustic coloration. Taking the user's position as the acoustic focus, the required delay compensation value for each speaker is calculated based on the sound wave propagation path length in the real-time spatial model. Specifically, the speaker farthest from the user is first determined as the reference delay, and then the path difference of other speakers is converted into microsecond-level delay parameters to ensure that all sound waves arrive at the listening position simultaneously. Phase coherence calibration is initiated by analyzing the impulse response of each speaker in the crossband (especially the 80-500Hz transition region), detecting and calculating the phase difference between them, and using an all-pass filter chain to perform 0-180 degree phase rotation compensation on specific speakers to effectively eliminate the phase cancellation phenomenon caused by position asymmetry. Finally, the entire set of parameters is synchronously sent to each speaker for execution through a low-latency communication protocol.
[0041] In this embodiment of the invention, by acquiring the relative positional relationship between the speakers and the user in real time and combining it with room impulse response analysis, a dynamic spatial model integrating geometric properties and acoustic propagation characteristics is constructed, realizing adaptive calculation and allocation of channel mapping weights. The system can optimize the channel allocation scheme in real time according to the user's displacement and automatically adjust the equalizer, delay, and phase parameters of each speaker, thereby effectively overcoming the limitations of the traditional fixed channel mode, significantly improving the accuracy of sound image positioning, the clarity of human voice dialogue, and the consistency of low-frequency response, and ultimately maintaining a stable and immersive listening experience even when the user moves or the device position changes.
[0042] Please see Figure 2 Another embodiment of the sound effect optimization method in this invention includes: 201. Obtain the relative positional relationship data between each speaker and the user; The first relative position data between the devices is initially calculated by using the wireless communication signal strength and flight time between the main speaker and each slave speaker. Visual scene information is collected by the smart camera associated with the system, and the user's position and outline are identified by computer vision algorithms to obtain the second relative position data between the user and the main speaker. The first and second relative position data are input into the Kalman filter model for data fusion and trajectory prediction, and the dynamic relative position relationship between each speaker and the user is output.
[0043] The main speaker acts as the positioning and coordination center, periodically transmitting pulse signals with nanosecond-level timestamps to all slave speakers via its UWB module. Each slave speaker immediately returns an acknowledgment frame upon receiving the signal. The main speaker calculates the round-trip time of the signal and, combined with the angle of arrival measurement, calculates the straight-line distance to each slave speaker based on the Time-of-Flight (TOF) principle. Simultaneously, cross-validation is performed using an RSSI signal attenuation model to form first relative position data containing distance values and confidence levels. Further, a camera linked to the main speaker is activated via HDMI-CEC or an IP network, initiating a real-time human pose detection algorithm based on improved YOLOv5 to capture scene images. Specifically, background subtraction and skin tone modeling are used to initially locate the user area, and then the OpenPose keypoint detection model is applied to identify the user's head and shoulder contour features. Finally, using the principle of stereo vision—utilizing the camera's known intrinsic parameters such as focal length and sensor size, combined with prior knowledge of the user's pixel coordinates in the image and actual human body dimensions—the straight-line distance and horizontal azimuth angle between the user and the main speaker are calculated, forming the second relative position data. A three-dimensional Cartesian coordinate system with the main speaker as the origin is established, and the first and second relative position data are uniformly transformed into this coordinate system. Subsequently, these two sets of time-domain aligned data are input into an extended Kalman filter with a uniform motion model: the filter uses the user's and speaker's positions and velocities as state variables, and wireless ranging data and visual positioning data as observation variables. In the prediction step, it estimates the current state based on the previous state, and in the update step, it calculates the Kalman gain using the current observation data and optimally corrects the predicted state. This process not only effectively filters out multipath interference from wireless signals and random errors in visual detection, but also achieves short-term prediction of the user's movement trends (such as uniform walking and turning) through the integrity of the state vector, ultimately outputting a dynamic relative position relationship matrix containing three-dimensional coordinates, movement speed, and confidence intervals.
[0044] 202. Based on the relative positional relationship, construct an initial spatial network describing the geometric positional relationship between the speaker and the user; Based on the relative positional relationship, the absolute coordinates of each speaker and the user in a three-dimensional Cartesian coordinate system are obtained; the position points of each speaker and the user are connected to construct a spatial triangular mesh describing the sound wave propagation path; the device type attribute of each node is marked in the spatial triangular mesh, and the straight-line distance between each adjacent node is calculated to complete the construction of the initial spatial network.
[0045] In a three-dimensional Cartesian coordinate system established with the acoustic center of the main speaker as the origin, the absolute coordinates of the speaker and the user, already placed in this unified coordinate system, are directly read. A constrained Delaunay triangulation algorithm is used to construct a spatial triangular mesh. Specifically, using all speaker and user position points as discrete vertex sets, the Bowyer-Watson algorithm is used to iteratively generate the triangular mesh, and specific constraints are applied to ensure that each user position point is contained within at least one triangle. At the same time, the physical connections between the main speaker and each subordinate speaker are forced to be retained as mesh edges, thus forming a triangular mesh structure that conforms to the actual sound wave propagation path. Each node in the mesh is labeled with device type attributes (including main speaker, front subordinate speaker, rear surround speaker, subwoofer, user, etc.), and each edge is assigned a sound wave propagation attribute. Finally, the straight-line distance between all adjacent nodes is calculated using the Euclidean distance formula in three-dimensional space, constructing an initial spatial network containing complete topological relationships and geometric metrics.
[0046] 203. Based on the initial spatial network, a known test signal is played through at least one speaker and received by at least one microphone. The impulse response of the room is extracted by calculating the cross-correlation function between the test signal and the received signal. Based on the initial spatial network topology and node density, one or more optimal speakers are dynamically selected as test signal sources, and a composite excitation signal combining an exponentially swept frequency signal and a pseudo-random sequence is generated for playback. Reverberation signals reflected from the room are synchronously acquired using a microphone array distributed across typical listening areas of the user, and an adaptive filter based on minimum mean square error is used to cancel echoes and suppress background noise in the acquired signals. A fast cross-correlation operation is performed on the preprocessed received signal and the original excitation signal, and a frequency-domain weighted Wiener deconvolution algorithm is used to separate and extract the high signal-to-noise ratio impulse response that reflects only the acoustic transmission characteristics of the room.
[0047] Based on the initial spatial network, by analyzing the distribution density of network nodes, when a high-density node cluster is detected in the user area, a multi-speaker cooperative transmission mode is adopted, and the three speakers with the greatest spatial differences are selected as synchronous transmission sources; if the nodes in the user area are sparse, the single speaker closest to the user and distributed in an isosceles triangle with other speakers is selected as the main transmission source.
[0048] During the signal generation phase, the digital signal processor in the main device executes a composite excitation signal synthesis algorithm. Specifically, the processor first synchronously generates two independent fundamental signals in the time domain: an exponentially swept frequency signal covering the audible frequency band, whose frequency changes continuously with time according to an exponential law; and a pseudo-random sequence with good autocorrelation characteristics. Subsequently, the system linearly superimposes the two signals according to preset amplitude weights using a digital mixer, and embeds several specific waveforms at key frequency points as marker signals for system calibration. The composite signal undergoes pre-emphasis filtering to optimize the high-frequency signal-to-noise ratio, and then a peak limiter prevents signal clipping, ultimately generating a test signal containing rich frequency components and time-domain characteristics, which is sent to the selected transmitting speaker for playback via a digital audio interface.
[0049] In the signal acquisition and preprocessing stage, a microphone array deployed in the user's listening area synchronously captures the sound field signal. This microphone array employs a specific geometric configuration, and its output signal is first subjected to sampling rate unification and time-domain alignment. Subsequently, a normalized least mean square adaptive filter based on multiple reference channels is initiated: using the original test signals played from each transmitter as reference input, an adaptive algorithm is used to construct a transfer function model to predict the direct sound component. This predicted value is then subtracted from the mixed received signal to eliminate acoustic echoes. A background noise spectrum estimation model is established by monitoring quiet segments of the signal, and an improved spectral subtraction method is used to suppress noise in the remaining signal. The noise suppression factor is dynamically adjusted according to the signal-to-noise ratio (SNR), ultimately outputting a reverberant signal that preserves room reflection characteristics and significantly improves the SNR.
[0050] In the impulse response extraction stage, a fast cross-correlation operation is first performed on the preprocessed received signal and the time-delay-calibrated original excitation signal: by converting the two signals to the frequency domain, performing conjugate multiplication, and then inversely transforming them to the time domain, an initial impulse response estimate is obtained. Subsequently, a frequency-domain weighted Wiener deconvolution is used for precise extraction: a transfer function model incorporating the frequency response characteristics of the loudspeaker-microphone system is constructed, and a weighted Wiener filter based on signal-to-noise ratio (SNR) estimation is designed in the frequency domain. The frequency domain weighting function of this filter is dynamically adjusted according to the signal quality of each frequency band, applying strong suppression to low SNR bands and maintaining weak processing to high SNR bands. By performing deconvolution on the initial impulse response using this filter, the inherent response of the device and the room's acoustic transmission characteristics are effectively separated, ultimately outputting a high-fidelity impulse response containing only the room's acoustic features.
[0051] 204. Analyze the impulse response, identify and quantify the energy and delay of direct sound, early reflection sound and reverberation sound, and map them to the initial spatial network to generate a real-time spatial model that combines geometric properties and acoustic propagation characteristics. In the impulse response, a dynamic threshold detection algorithm is used to segment the time domain intervals of direct sound, early reflections, and reverberation, and the arrival time, sound pressure level, and spectral characteristics of the acoustic components in each interval are quantified and recorded. The quantized early reflection sequence is input into a clustering analysis model to identify the main reflection clusters, and the corresponding virtual reflectors are reconstructed in the initial spatial network based on the arrival time difference and azimuth information using the ray tracing method. The decay time parameter of the reverberation field is fused with the reconstructed virtual reflector structure, and the initial spatial network is labeled with acoustic attributes to generate a real-time spatial model that includes geometric topology and acoustic propagation paths.
[0052] Time-domain analysis was performed on the extracted high signal-to-noise ratio impulse response: An adaptive dynamic threshold detection algorithm was adopted, using a preset attenuation ratio of the maximum amplitude of the impulse response as the initial detection threshold (for example, the preset attenuation ratio could correspond to attenuating the peak amplitude to 10% of its original value (i.e., -20dB) as the initial detection threshold). Zero-crossing points and extreme points were detected by sliding window to segment the time-domain boundaries of direct sound, early reflection sound (e.g., the interval 5-80ms after the direct sound) and reverberation sound (e.g., the interval from 80ms to attenuation to -60dB). For each segmented interval, the arrival time of the direct sound, the relative time delay of each component of the early reflection sound, and the attenuation envelope of the reverberation sound were quantized and recorded. At the same time, the sound pressure level distribution and spectral centroid of each acoustic component in a specific critical frequency band (e.g., using a 1 / 3 octave bandwidth for analysis) were extracted by short-time Fourier transform. Next, the quantified parameters of early reflected sound (e.g., time delay, azimuth, energy) are input into a DBSCAN-based clustering analysis model. Density clustering identifies the main reflected sound groups forming clusters in the time delay-azimuth space, with each group representing an important reflection interface. For each reflected sound cluster, a ray tracing algorithm is initiated in the initial spatial network: using the corresponding speaker as the sound source and the arrival direction of the reflected sound as a constraint, the sound wave propagation path is calculated using the mirror method, and the position, size, and orientation of the virtual reflector are reconstructed in space. Finally, the decay time parameters of the reverberant sound field, the virtual reflector structure of the early reflected sound, and the propagation path of the direct sound are integrated into the initial spatial network. Each triangular mesh facet in the network is labeled with its corresponding acoustic reflection properties, and each edge is assigned a time delay and attenuation coefficient for sound wave propagation. Ultimately, a parameterized real-time spatial model that includes both physical geometry and topology and fully describes the sound wave propagation mechanism is generated.
[0053] When analyzing the impulse response, an adaptive impulse response component segmentation algorithm is used, specifically including: Dynamic threshold calculation formula:
[0054] The formula for adaptive correction of time-domain boundaries is as follows:
[0055] in, Let be the dynamic detection threshold at time t; This represents the maximum amplitude of the impulse response; This is a balance factor between global and local features, with a value of [0,1]. The decay rate is exponential and is related to the room reverberation time RT60, with the following specific relationship: t0 is the arrival time of the direct sound; The average signal amplitude within the sliding window; is the time-domain boundary of the c-type component, where c=1 represents the boundary between direct sound and early reflections, and c=2 represents the boundary between early reflections and reverberation. This represents the cumulative energy from t0 to t; The threshold for the energy percentage of class C components; This represents the total energy of the impulse response.
[0056] 205. Based on the geometric relationship and acoustic propagation characteristics between each speaker and the user in the real-time spatial model, calculate the binaural acoustic transfer function from each speaker to the user's position; The direct sound propagation paths from each speaker to the user's ears in the real-time spatial model are extracted, and the length difference and azimuth of each path are calculated. The room reflection parameters obtained by impulse response analysis are combined to simulate the superposition effect of early reflected sound and reverberation sound on each propagation path. Based on the principle of acoustic interference, the amplitude, time delay and phase relationship of direct sound and reflected sound are integrated to generate a binaural acoustic transfer function that includes spatial filtering effect.
[0057] The three-dimensional direct sound path from each speaker to the user's left and right ears is extracted from the real-time spatial model. Vector operations are used to calculate the binaural path length difference (for calculating the binaural time difference (ITD)) and the horizontal azimuth and vertical elevation angles relative to the user's head direction (for calculating the binaural sound level difference (ILD)). Subsequently, combined with room reflection parameters obtained from impulse response analysis, the propagation path of the main early reflections during the initial reflection phase (e.g., the first 50ms) is simulated using the mirror source method, including the time delay, attenuation, and direction information of each reflection reaching both ears. Simultaneously, the energy distribution and coherence characteristics of the reverberant sound field at both ears are estimated based on a statistical acoustic model. Finally, based on the principle of sound wave superposition, the transfer functions of the direct sound, early reflections, and reverberation are vector-synthesized in the frequency domain, fully considering the comb filtering effect caused by sound wave interference, head diffraction effect, and shoulder reflection influence. This results in a high-fidelity binaural acoustic transfer function dataset containing complete room acoustic characteristics and spatial positioning cues.
[0058] 206. Using the binaural acoustic transfer function as the sound field reconstruction target, based on the sound image localization principle, the standard multi-channel signal is decomposed using the vector basis amplitude translation algorithm to calculate the initial weight coefficients corresponding to each speaker. Each channel in the standard multi-channel signal is mapped to a virtual sound source plane to generate a target sound field distribution containing the orientation information of each virtual sound source. Using the target sound field distribution as the desired sound image and the binaural acoustic transfer function as the acoustic constraint for sound field reconstruction, a sound field reconstruction mathematical model based on the vector basis amplitude translation algorithm is constructed. The sound field reconstruction mathematical model is solved to calculate the reconstruction contribution of each speaker to each virtual sound source in the target sound field under acoustic constraints, thereby obtaining the initial weighting coefficient of each speaker to each channel signal.
[0059] The input multi-channel audio signal (e.g., the left, center, right, left surround, and right surround channels of a 5.1 channel system) is mapped onto a spherical sound field space centered on the listening position, generating a set of virtual sound sources with fixed azimuth angles. Using the target acoustic transfer function generated by these virtual sound sources at both ears as the ideal response, and combining this with the acoustic transfer function generated by the actual speaker layout at both ears as physical constraints, a sound field reconstruction optimization problem based on vector basis amplitude translation is constructed. The objective function of this problem aims to minimize the mean square error between the ideal binaural response and the actual reproduced binaural response, while the constraints ensure a reasonable distribution of sound pressure energy at the listening position. By solving this constrained optimization problem, the optimal contribution weight of each physical speaker to each virtual sound source is calculated, forming an initial weight coefficient matrix. This weight coefficient matrix reflects the amplitude and phase adjustment required by each speaker when reconstructing the target sound field using an existing speaker array in a real acoustic environment.
[0060] 207. Combining the user's real-time movement data and the constraint relationship of the binaural acoustic transfer function on the sound image localization, the psychoacoustic model is applied to dynamically optimize and smooth the initial weights to generate a channel allocation scheme. Based on real-time acquired user head position and orientation data, combined with binaural acoustic transfer function, the vector deviation between the current sound image position and the target sound image position is calculated; the vector deviation is input into the psychoacoustic model to generate dynamic compensation coefficients for the initial weights of each speaker; a first-order inertial smoothing algorithm is used to apply the dynamic compensation coefficients to the initial weight coefficients to generate a channel allocation scheme.
[0061] By acquiring real-time head position and Euler angle data in 3D space using a head tracker, a sound image offset perception model is established using binaural acoustic transfer functions. This model calculates the azimuth and distance deviations between the sound image position generated by the actual speaker layout and the target virtual sound source position in a spherical coordinate system. This vector deviation is then input into a psychoacoustic model incorporating priority effects and auditory masking effects. This model dynamically generates weight compensation coefficients based on the magnitude and direction of the deviation—when a rapid head turn is detected, the weight of the speakers directly in front is strengthened while the contribution of the speakers to the sides and rear is attenuated; when the sound image distance deviation exceeds the auditory perception threshold, distance attenuation compensation is initiated. Finally, a first-order inertial smoother with an adaptive time constant is used to process the compensated weight coefficients: the smoothing time is automatically shortened (e.g., 100ms) during rapid head movement to ensure response speed, and the smoothing time is extended (e.g., 500ms) during stationary or slow movement to ensure a natural listening experience. The final output is a real-time channel allocation scheme that can quickly track user movement while maintaining a stable and smooth sound image.
[0062] When applying psychoacoustic models for dynamic optimization, a psychoacoustic dynamic weight compensation algorithm is used, which specifically includes: The formula for calculating the compensation coefficient is:
[0063]
[0064] The formula for the direction sensitivity function is:
[0065] The adaptive inertial smoothing formula is:
[0066]
[0067] in, Let be the dynamic compensation coefficient for the i-th speaker; This represents the acoustic image vector deviation between the current acoustic image position and the target acoustic image position. Let be the unit vector pointing to the i-th speaker; For acoustic image vector deviation With the unit vector pointing to the i-th speaker The angle between them; This is the compensation intensity factor; The angle of the user's head orientation (with the user's direct frontal direction as the reference for 0°); Let be the azimuth angle of the i-th speaker relative to the front of the user; Here is the direction sensitivity function; The weight coefficient for the i-th speaker at the n-th step; Let be the initial weighting coefficient for the i-th speaker; The time constant for adaptive inertial smoothing; and These represent the maximum and minimum values of the time constant; The magnitude of the angular velocity of the user's head movement; This is the velocity sensitivity constant, used for adjustment. Follow The rate of change.
[0068] 208. Based on the channel allocation scheme, perform real-time adjustments to the equalizer curves, delay compensation, and phase parameters of each speaker.
[0069] Based on the weighted roles of each speaker in the dynamic channel allocation scheme and the acoustic propagation paths in the real-time spatial model, a personalized target equalization curve is generated for each speaker. Near-field vocal speakers are given enhanced mid-frequency clarity, while far-field surround speakers expand high-frequency spatiality. Specifically, the target frequency response characteristics are determined according to the weight ratio of each speaker in the channel allocation; high-weight speakers use flat response curves, while low-weight speakers use attenuation compensation curves. Combining the acoustic propagation characteristics in the real-time spatial model, high-frequency roll-off compensation is applied to speakers in the near-field propagation path, and low-frequency cutoff optimization is performed on speakers in the far-field propagation path. Based on the room resonance modes obtained from impulse response analysis, targeted notch filters are set in the target equalization curves to suppress acoustic coloration in specific frequency bands.
[0070] Using the time synchronization of sound wave propagation to the user's location as a benchmark, microsecond-level delay compensation values are calculated based on the length difference of the acoustic paths of each speaker, and group delay calibration is applied to all speakers. Specifically, the speaker with the longest acoustic path to the user's location is selected as the benchmark reference, and the sound wave propagation time difference of other speakers relative to this benchmark is calculated. The sound wave propagation time difference is converted into a digital delay parameter with sampling point precision, and a corresponding fractional delay filter is configured. Sub-sampling precision delay calibration is achieved through multi-phase filtering interpolation technology to ensure that the sound waves emitted by all speakers are synchronized at the user's location.
[0071] By analyzing the impulse response coherence of each speaker in the crossover frequency band, an anti-phase compensation filter is calculated and applied to eliminate phase cancellation caused by multi-speaker interference, thus achieving synergistic optimization of acoustic parameters. Specifically, the impulse response of each speaker at the user's position is extracted, and the mutual coherence function between speakers in the main interference frequency band (e.g., 80-500Hz) is calculated. The main interference modes causing phase cancellation are identified by eigenvalue decomposition, and the corresponding anti-phase compensation filter coefficients are calculated. An all-pass filter chain is designed using a least-squares optimization algorithm to apply precise phase rotation compensation to specific speakers, reconstructing a coherent sound wavefront.
[0072] In this embodiment of the invention, a real-time spatial model integrating geometric and acoustic characteristics is constructed, and based on binaural acoustic transfer function and dynamically optimized channel allocation, accurate sound field reconstruction of a multi-speaker system is achieved. The system can automatically generate personalized equalization curves, microsecond-level delay compensation, and phase-inverting filters for each speaker according to user movement and spatial acoustic characteristics, thereby effectively improving the accuracy of sound image positioning, voice clarity, and frequency response consistency in complex indoor environments, eliminating multi-speaker interference, and ultimately providing users with an immersive, stable, and high-quality listening experience that adapts to their real-time position.
[0073] The sound effect optimization method in the embodiments of the present invention has been described above. The sound effect optimization device in the embodiments of the present invention is described below. Please refer to [link / reference]. Figure 3 One embodiment of the sound effect optimization device in this invention includes: The acquisition module 301 is used to acquire the relative positional relationship data between each speaker and the user; Modeling module 302 is used to perform acoustic space modeling based on relative positional relationships and combined with the analysis of room impulse response, to obtain a real-time spatial model that includes geometric properties and acoustic propagation characteristics. The generation module 303 is used to calculate and allocate channel mapping weights based on a real-time spatial model, and generate a channel allocation scheme that dynamically changes with the user's position. The adjustment module 304 is used to adjust the equalizer curve, delay compensation and phase parameters of each speaker in real time according to the channel allocation scheme.
[0074] In this embodiment of the invention, by acquiring the relative positional relationship between the speakers and the user in real time and combining it with room impulse response analysis, a dynamic spatial model integrating geometric properties and acoustic propagation characteristics is constructed, realizing adaptive calculation and allocation of channel mapping weights. The system can optimize the channel allocation scheme in real time according to the user's displacement and automatically adjust the equalizer, delay, and phase parameters of each speaker, thereby effectively overcoming the limitations of the traditional fixed channel mode, significantly improving the accuracy of sound image positioning, the clarity of human voice dialogue, and the consistency of low-frequency response, and ultimately maintaining a stable and immersive listening experience even when the user moves or the device position changes.
[0075] Please see Figure 4 Another embodiment of the sound effect optimization device in this invention includes: The acquisition module 301 is used to acquire the relative positional relationship data between each speaker and the user; Modeling module 302 is used to perform acoustic space modeling based on relative positional relationships and combined with the analysis of room impulse response, to obtain a real-time spatial model that includes geometric properties and acoustic propagation characteristics. The generation module 303 is used to calculate and allocate channel mapping weights based on a real-time spatial model, and generate a channel allocation scheme that dynamically changes with the user's position. The adjustment module 304 is used to adjust the equalizer curve, delay compensation and phase parameters of each speaker in real time according to the channel allocation scheme.
[0076] Optionally, the acquisition module 301 can be specifically used for: The first relative position data between the devices is initially calculated by using the wireless communication signal strength and flight time between the main speaker and each slave speaker. Visual scene information is collected by the smart camera associated with the system, and the user's position and outline are identified by computer vision algorithms to obtain the second relative position data between the user and the main speaker. The first and second relative position data are input into the Kalman filter model for data fusion and trajectory prediction, and the dynamic relative position relationship between each speaker and the user is output.
[0077] Optionally, modeling module 302 includes: Building unit 3021 is used to construct an initial spatial network describing the geometric positional relationship between the speaker and the user based on the relative positional relationship; Extraction unit 3022 is used to extract the impulse response of the room by playing a known test signal through at least one speaker and receiving it by at least one microphone based on an initial spatial network, and by calculating the cross-correlation function between the test signal and the received signal. The first generation unit 3023 is used to analyze the impulse response, identify and quantify the energy and delay of direct sound, early reflection sound and reverberation sound, and map them to the initial spatial network to generate a real-time spatial model that combines geometric properties and acoustic propagation characteristics.
[0078] Optionally, building unit 3021 can be specifically used for: Based on the relative positional relationship, the absolute position coordinates of each speaker and user are obtained; the position points of each speaker and user are connected to construct a spatial triangular mesh describing the sound wave propagation path; the device type attribute of each node is labeled in the spatial triangular mesh, and the straight-line distance between each adjacent node is calculated to complete the construction of the initial spatial network.
[0079] Optionally, the extraction unit 3022 can be specifically used for: Based on the initial spatial network topology and node density, one or more optimal speakers are dynamically selected as test signal sources, and a composite excitation signal combining an exponentially swept frequency signal and a pseudo-random sequence is generated for playback. Reverberation signals reflected from the room are synchronously acquired using a microphone array distributed across typical listening areas of the user, and an adaptive filter based on minimum mean square error is used to cancel echoes and suppress background noise in the acquired signals. A fast cross-correlation operation is performed on the preprocessed received signal and the original excitation signal, and a frequency-domain weighted Wiener deconvolution algorithm is used to separate and extract the high signal-to-noise ratio impulse response that reflects only the acoustic transmission characteristics of the room.
[0080] Optionally, the first generating unit can be specifically used for: In the impulse response, a dynamic threshold detection algorithm is used to segment the time domain intervals of direct sound, early reflections, and reverberation, and the arrival time, sound pressure level, and spectral characteristics of the acoustic components in each interval are quantified and recorded. The quantized early reflection sequence is input into a clustering analysis model to identify the main reflection clusters, and the corresponding virtual reflectors are reconstructed in the initial spatial network based on the arrival time difference and azimuth information using the ray tracing method. The decay time parameter of the reverberation field is fused with the reconstructed virtual reflector structure, and the initial spatial network is labeled with acoustic attributes to generate a real-time spatial model that includes geometric topology and acoustic propagation paths.
[0081] Optionally, the generation module 303 includes: The first computing unit 3031 is used to calculate the binaural acoustic transfer function from each speaker to the user's position based on the geometric relationship and acoustic propagation characteristics between each speaker and the user in the real-time spatial model. The second calculation unit 3032 is used to decompose the standard multi-channel signal based on the sound field reconstruction target with the binaural acoustic transfer function and the sound image localization principle, and to calculate the initial weight coefficients corresponding to each speaker by using the vector basis amplitude translation algorithm. The second generation unit 3033 is used to combine the user's real-time movement data and, based on the constraint relationship of the binaural acoustic transfer function on the sound image localization, apply a psychoacoustic model to dynamically optimize and smooth the initial weights to generate a channel allocation scheme.
[0082] Optionally, the first computing unit 3031 can be specifically used for: The direct sound propagation paths from each speaker to the user's ears in the real-time spatial model are extracted, and the length difference and azimuth of each path are calculated. The room reflection parameters obtained by impulse response analysis are combined to simulate the superposition effect of early reflected sound and reverberation sound on each propagation path. Based on the principle of acoustic interference, the amplitude, time delay and phase relationship of direct sound and reflected sound are integrated to generate a binaural acoustic transfer function that includes spatial filtering effect.
[0083] Optionally, the second computing unit 3032 can be specifically used for: Each channel in the standard multi-channel signal is mapped to a virtual sound source plane to generate a target sound field distribution containing the orientation information of each virtual sound source. Using the target sound field distribution as the desired sound image and the binaural acoustic transfer function as the acoustic constraint for sound field reconstruction, a sound field reconstruction mathematical model based on the vector basis amplitude translation algorithm is constructed. The sound field reconstruction mathematical model is solved to calculate the reconstruction contribution of each speaker to each virtual sound source in the target sound field under acoustic constraints, thereby obtaining the initial weighting coefficient of each speaker to each channel signal.
[0084] Optionally, the second generation unit 3033 can be specifically used for: Based on real-time acquired user head position and orientation data, combined with binaural acoustic transfer function, the vector deviation between the current sound image position and the target sound image position is calculated; the vector deviation is input into the psychoacoustic model to generate dynamic compensation coefficients for the initial weights of each speaker; a first-order inertial smoothing algorithm is used to apply the dynamic compensation coefficients to the initial weight coefficients to generate a channel allocation scheme.
[0085] Optionally, adjustment module 304 includes: The third generation unit 3041 is used to generate a personalized target equalization curve for each speaker based on the weight role of each speaker in the channel allocation scheme and the acoustic propagation path in its real-time spatial model. The near-field vocal speakers enhance mid-frequency clarity, while the far-field surround speakers expand the high-frequency spatial sense. The processing unit 3042 is used to calculate the microsecond-level delay compensation value based on the time synchronization of sound wave propagation to the user's position and the length difference of the acoustic path of each speaker, and to apply group delay calibration to all speakers. The optimization unit 3043 is used to calculate and apply an anti-phase compensation filter by analyzing the impulse response coherence of each speaker in the cross frequency band, thereby eliminating phase cancellation caused by multi-speaker interference and completing the collaborative optimization of acoustic parameters.
[0086] Optionally, the processing unit 3042 is specifically used for: The speaker with the longest acoustic path to the user's location is selected as the reference, and the sound wave propagation time difference of other speakers relative to this reference is calculated. The sound wave propagation time difference is converted into a digital delay parameter with sampling point precision, and a corresponding fractional delay filter is configured. Sub-sampling precision delay calibration is achieved through multi-phase filtering interpolation technology to ensure that the sound waves emitted by all speakers are synchronized at the user's location.
[0087] In this embodiment of the invention, a real-time spatial model integrating geometric and acoustic characteristics is constructed, and based on binaural acoustic transfer function and dynamically optimized channel allocation, accurate sound field reconstruction of a multi-speaker system is achieved. The system can automatically generate personalized equalization curves, microsecond-level delay compensation, and phase-inverting filters for each speaker according to user movement and spatial acoustic characteristics, thereby effectively improving the accuracy of sound image positioning, voice clarity, and frequency response consistency in complex indoor environments, eliminating multi-speaker interference, and ultimately providing users with an immersive, stable, and high-quality listening experience that adapts to their real-time position.
[0088] above Figure 3 and Figure 4 The sound effect optimization device in the embodiments of the present invention will be described in detail from the perspective of modular functional entities. The electronic device in the embodiments of the present invention will be described in detail from the perspective of hardware processing.
[0089] See Figure 5 As shown, the electronic device includes a processor 500 and a memory 501. The memory 501 stores machine-executable instructions that can be executed by the processor 500. The processor 500 executes the machine-executable instructions to implement the above-mentioned sound effect optimization method.
[0090] Furthermore, Figure 5 The electronic device shown also includes a bus 502 and a communication interface 503. The processor 500, the communication interface 503 and the memory 501 are connected via the bus 502.
[0091] The memory 501 may include high-speed random access memory (RAM) and may also include non-volatile memory, such as at least one disk storage device. Communication between this system network element and at least one other network element is achieved through at least one communication interface 503 (which can be wired or wireless), such as the Internet, wide area network, local area network, metropolitan area network, etc. The bus 502 may be an ISA bus, PCI bus, or EISA bus, etc. The bus can be divided into address bus, data bus, control bus, etc. For ease of representation, Figure 5 The symbol is represented by a single double-headed arrow, but this does not mean that there is only one bus or one type of bus.
[0092] The processor 500 may be an integrated circuit chip with signal processing capabilities. In implementation, each step of the above method can be completed by the integrated logic circuitry in the hardware of the processor 500 or by instructions in software form. The processor 500 may be a general-purpose processor, including a central processing unit (CPU), a network processor (NP), etc.; it may also be a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components. It can implement or execute the methods, steps, and logic block diagrams disclosed in the embodiments of this disclosure. The general-purpose processor may be a microprocessor or any conventional processor. The steps of the methods disclosed in the embodiments of this disclosure can be directly embodied in the execution of a hardware decoding processor, or executed by a combination of hardware and software modules in the decoding processor. The software modules may reside in random access memory, flash memory, read-only memory, programmable read-only memory, electrically erasable programmable memory, registers, or other mature storage media in the art. The storage medium is located in memory 501. The processor 500 reads the information in memory 501 and, in conjunction with its hardware, completes the method steps of the aforementioned embodiment.
[0093] The present invention also provides an electronic device, the computer device including a memory and a processor, wherein the memory stores computer-readable instructions, and when the computer-readable instructions are executed by the processor, the processor performs the steps of the sound effect optimization method in the above embodiments. The present invention also provides a computer-readable storage medium, which can be a non-volatile computer-readable storage medium or a volatile computer-readable storage medium, wherein the computer-readable storage medium stores instructions, and when the instructions are executed on a computer, the computer performs the steps of the sound effect optimization method.
[0094] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working processes of the systems, devices, and units described above can be referred to the corresponding processes in the foregoing method embodiments, and will not be repeated here.
[0095] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0096] The above-described embodiments are only used to illustrate the technical solutions of the present invention, and are not intended to limit it. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A sound effect optimization method, characterized in that, The sound effect optimization method includes: Obtain the relative positional relationship data between each speaker and the user; Based on the relative positional relationship, acoustic space modeling is performed in conjunction with the analysis of the room impulse response to obtain a real-time spatial model that includes geometric properties and acoustic propagation characteristics. Based on the real-time spatial model, the channel mapping weights are calculated and allocated to generate a channel allocation scheme that dynamically changes with the user's position. Based on the channel allocation scheme, the equalizer curve, delay compensation, and phase parameters of each speaker are adjusted in real time.
2. The sound effect optimization method according to claim 1, characterized in that, The acquisition of relative positional relationship data between each speaker and the user includes: The first relative position data between the devices is initially calculated by using the wireless communication signal strength and flight time between the main speaker and each slave speaker. Visual scene information is collected by the smart camera associated with the system, and computer vision algorithms are used to identify the user's position and outline to obtain the second relative position data between the user and the main speaker. The first relative position data and the second relative position data are input into the Kalman filter model for data fusion and trajectory prediction, and the dynamic relative position relationship between each speaker and the user is output.
3. The sound effect optimization method according to claim 1, characterized in that, The step of performing spatial modeling processing based on the relative positional relationship data to obtain a real-time spatial model describing the spatial distribution relationship between the speaker and the user includes: Based on the aforementioned relative positional relationship, an initial spatial network describing the geometric positional relationship between the speaker and the user is constructed; Based on the initial spatial network, a known test signal is played through at least one speaker and received by at least one microphone. The impulse response of the room is extracted by calculating the cross-correlation function between the test signal and the received signal. The impulse response is analyzed to identify and quantify the energy and delay of direct sound, early reflections, and reverberation, and then mapped to the initial spatial network to generate a real-time spatial model that combines geometric properties and acoustic propagation characteristics.
4. The sound effect optimization method according to claim 3, characterized in that, The initial spatial network describing the geometrical positional relationship between the speaker and the user, based on the aforementioned relative positional relationship, includes: Based on the relative positional relationship, obtain the absolute position coordinates of each speaker and the user; Connect each speaker and user location point to construct a spatial triangular mesh that describes the sound wave propagation path; The device type attributes of each node are labeled in the spatial triangular mesh, and the straight-line distance between each adjacent node is calculated to complete the construction of the initial spatial network.
5. The sound effect optimization method according to claim 3, characterized in that, Based on the initial spatial network, a known test signal is played through at least one speaker and received by at least one microphone. The impulse response of the room is extracted by calculating the cross-correlation function between the test signal and the received signal, including: Based on the topology and node density of the initial spatial network, one or more optimal speakers are dynamically selected as test signal transmission sources, and a composite excitation signal combining an exponential sweep frequency signal and a pseudo-random sequence is generated for playback. The reverberation signal reflected by the room is synchronously acquired by a microphone array distributed in the user's typical listening area, and the acquired signal is echo-cancelled and background noise suppressed by an adaptive filter based on minimum mean square error. A fast cross-correlation operation is performed on the preprocessed received signal and the original excitation signal, and a frequency-domain weighted Wiener deconvolution algorithm is used to separate and extract the high signal-to-noise ratio impulse response that only reflects the acoustic transmission characteristics of the room.
6. The sound effect optimization method according to claim 3, characterized in that, The analysis of the impulse response identifies and quantifies the energy and delay of direct sound, early reflections, and reverberation, and maps them to the initial spatial network to generate a real-time spatial model that combines geometric properties and acoustic propagation characteristics, including: In the impulse response, the time domain intervals of direct sound, early reflection sound and reverberation sound are segmented based on the dynamic threshold detection algorithm, and the arrival time, sound pressure level and spectral characteristics of the acoustic components in each interval are quantitatively recorded. The quantized early reflected sound sequence is input into the cluster analysis model to identify the main reflected sound clusters, and the corresponding virtual reflective surfaces are reconstructed in the initial spatial network based on the arrival time difference and azimuth information using the ray tracing method. The decay time parameter of the reverberant sound field is fused with the reconstructed virtual reflective surface structure, and the initial spatial network is labeled with acoustic properties to generate a real-time spatial model that includes geometric topology and acoustic propagation path.
7. The sound effect optimization method according to claim 1, characterized in that, The calculation and allocation of channel mapping weights based on the real-time spatial model to generate a channel allocation scheme that dynamically changes with the user's position includes: Based on the geometric relationship and acoustic propagation characteristics between each speaker and the user in the real-time spatial model, the binaural acoustic transfer function from each speaker to the user's position is calculated. Using the binaural acoustic transfer function as the sound field reconstruction target, based on the sound image localization principle, the standard multi-channel signal is decomposed using the vector basis amplitude translation algorithm to calculate the initial weight coefficients corresponding to each speaker. By combining the user's real-time movement data and based on the constraints of the binaural acoustic transfer function on sound image localization, a psychoacoustic model is applied to dynamically optimize and smooth the initial weights, thereby generating a channel allocation scheme.
8. The sound effect optimization method according to claim 7, characterized in that, The step of calculating the binaural acoustic transfer function from each speaker to the user's location based on the geometric relationship and acoustic propagation characteristics between the user and each speaker in the real-time spatial model includes: Extract the direct sound propagation path from each speaker to the user's ears in the real-time spatial model, and calculate the length difference and azimuth angle of each path; Based on the room reflection parameters obtained from impulse response analysis, the superposition effect of early reflected sound and reverberation sound on each propagation path is simulated and calculated. Based on the principle of acoustic wave interference, the amplitude, time delay, and phase relationship between direct sound and reflected sound are integrated to generate a binaural acoustic transfer function that includes spatial filtering effects.
9. The sound effect optimization method according to claim 7, characterized in that, The method uses the binaural acoustic transfer function as the sound field reconstruction target, and based on the principle of sound image localization, employs a vector basis amplitude shift algorithm to decompose the standard multi-channel signal, calculating the initial weighting coefficients corresponding to each speaker, including: Each channel in a standard multi-channel signal is mapped onto a virtual sound source plane to generate a target sound field distribution containing the directional information of each virtual sound source. Using the target sound field distribution as the desired sound image and the binaural acoustic transfer function as the acoustic constraint for sound field reconstruction, a mathematical model for sound field reconstruction based on the vector basis amplitude translation algorithm is constructed. Solve the mathematical model for sound field reconstruction, calculate the reconstruction contribution of each speaker to each virtual sound source in the target sound field under the acoustic constraints, and thus obtain the initial weighting coefficient of each speaker to each channel signal.
10. The sound effect optimization method according to claim 7, characterized in that, The method combines real-time user movement data and, based on the constraints of the binaural acoustic transfer function on sound image localization, applies a psychoacoustic model to dynamically optimize and smooth the initial weights, generating a channel allocation scheme, including: Based on real-time acquired user head position and orientation data, combined with the binaural acoustic transfer function, the vector deviation between the current acoustic image position and the target acoustic image position is calculated; The vector deviation is input into the psychoacoustic model to generate dynamic compensation coefficients for the initial weights of each speaker; A first-order inertial smoothing algorithm is used to apply the dynamic compensation coefficient to the initial weight coefficient to generate a channel allocation scheme.
11. The sound effect optimization method according to claim 1, characterized in that, The step of adjusting the equalizer curve, delay compensation, and phase parameters of each speaker in real time according to the channel allocation scheme includes: Based on the weighted roles of each speaker in the channel allocation scheme and the acoustic propagation path in its real-time spatial model, a personalized target equalization curve is generated for each speaker, wherein the near-field vocal speakers enhance mid-frequency clarity and the far-field surround speakers expand the high-frequency spatial sense. Based on the time synchronization of sound wave propagation to the user's location, microsecond-level delay compensation values are calculated according to the length difference of the acoustic path of each speaker, and group delay calibration is applied to all speakers; By analyzing the impulse response coherence of each speaker in the crossover frequency band, an anti-phase compensation filter is calculated and applied to eliminate phase cancellation caused by multi-speaker interference, thus achieving synergistic optimization of acoustic parameters.
12. The sound effect optimization method according to claim 11, characterized in that, The method, based on the time synchronization of sound wave propagation to the user's location, calculates microsecond-level delay compensation values according to the length difference of the acoustic path of each speaker, and applies group delay calibration to all speakers, including: The speaker with the longest acoustic path to the user's location is selected as the reference, and the sound wave propagation time difference of other speakers relative to this reference is calculated. The sound wave propagation time difference is converted into a digital delay parameter with sampling point precision, and a corresponding fractional delay filter is configured. Multi-phase filtering interpolation technology is used to achieve sub-sampling precision delay calibration, ensuring that the sound waves emitted by all speakers are synchronized at the user's location.
13. A sound effect optimization device, characterized in that, The sound effect optimization device includes: The acquisition module is used to acquire data on the relative positional relationship between each speaker and the user; The modeling module is used to perform acoustic space modeling based on the relative positional relationship and the analysis of the room impulse response, so as to obtain a real-time spatial model that includes geometric properties and acoustic propagation characteristics. The generation module is used to calculate and allocate channel mapping weights based on the real-time spatial model, and generate a channel allocation scheme that dynamically changes with the user's position. The adjustment module is used to adjust the equalizer curve, delay compensation and phase parameters of each speaker in real time according to the channel allocation scheme.
14. An electronic device, characterized in that, The electronic device includes: a memory and at least one processor, wherein the memory stores instructions; The at least one processor invokes the instructions in the memory to cause the electronic device to perform the sound optimization method as described in any one of claims 1-12.
15. A computer-readable storage medium storing instructions thereon, characterized in that, When the instruction is executed by the processor, it implements the sound effect optimization method as described in any one of claims 1-12.