Audio detection-based sound control method, device, equipment, medium and program product
Patent Information
- Application Number
- CN202610837082.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-06-10
- Publication Date
- 2026-08-18
AI Technical Summary
[0004]本发明提供一种基于音频检测的音响控制方法、装置、设备、介质及程序产品,本发明解决了现有技术中传递函数建模精度不足导致声场控制误差累积的技术问题,保障了听音者对声源空间方向的准确感知
[0015] In one of the solutions provided by the aforementioned audio control methods, devices, equipment, media, and program products based on audio detection, sound pressure response signals are synchronously acquired at various measurement points in both bright and dark areas using a microphone array. Frequency domain deconvolution modeling is then performed in conjunction with a frequency sweep excitation signal. This accurately constructs the sound pressure transfer function matrix for the bright area, the sound pressure transfer function matrix for the dark area, and the sound intensity transfer function pair at the center point of the bright area within the frequency domain. This solves the technical problem of insufficient transfer function modeling accuracy leading to the accumulation of sound field control errors in existing technologies. This invention incorporates the pressure matching error term and the flatness deviation term into a unified joint objective function. Simultaneously, it introduces three types of hard constraints: a lower bound constraint on the sound intensity direction, an upper bound constraint on the sound pressure amplitude in the bright area, and a sound pressure suppression constraint in the dark area. Within a single optimization framework, it simultaneously considers the accuracy of sound pressure reconstruction in the bright area, the spatial consistency of sound pressure in the bright area, and the sound pressure suppression effect in the dark area. The introduction of the sound intensity direction constraint ensures that the direction of the synthesized sound energy propagation in the bright area is precisely aligned with the target location, effectively eliminating the virtual sound image offset problem caused by neglecting the directionality of sound intensity in existing technologies, and ensuring the listener's accurate perception of the spatial direction of the sound source. This invention transforms a non-convex optimization problem into a sequential convex subproblem through a continuous convex approximation strategy and solves iteratively using the interior point method. While ensuring the convergence of the solution, the optimal loudspeaker weight vector is directly converted into the time-domain filter coefficients of each channel through inverse transformation and loaded into the DSP processor for real-time driving.
Smart Images

Figure CN122602038A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of audio detection technology, and in particular to a sound control method, device, equipment, medium, and program product based on audio detection. Background Technology
[0002] Due to the wide range of applications of personal audio systems in scenarios such as multi-zone directional playback, private calls, and zoned audio, existing audio control technologies use weighted processing of the speaker array output signal to create differences in sound pressure distribution within the target space, thereby achieving zoned control effects such as enhancing sound in bright areas and suppressing sound in dark areas.
[0003] However, existing technologies only aim to minimize the sound pressure reconstruction error in the bright area, without imposing uniform constraints on the sound pressure amplitude at each measurement point in the bright area. This results in uneven sound pressure distribution in the bright area and severe deterioration of spatial flatness. While the flatness control method improves the uniformity of the bright area by setting a target sound pressure of equal amplitude, its optimization direction is fundamentally in conflict with the acoustic contrast target, leading to a decrease in the ability to suppress sound pressure in the dark area and a significant increase in the workload of the loudspeaker array. Summary of the Invention
[0004] This invention provides a sound control method, device, equipment, medium, and program product based on audio detection. This invention solves the technical problem of insufficient transfer function modeling accuracy leading to the accumulation of sound field control errors in the prior art, and ensures that the listener accurately perceives the spatial direction of the sound source.
[0005] In a first aspect, embodiments of this application provide a sound control method based on audio detection, comprising: A sweep frequency excitation signal is injected into each speaker unit, a first sound pressure response signal is collected, and the time-domain average of the first sound pressure response signal collected multiple times is taken to obtain a second sound pressure response signal; Based on the second sound pressure response signal and the frequency sweep excitation signal, the optimal loudspeaker weight vector at each frequency point is solved; The optimal speaker weight vector is converted into time-domain filter coefficients for each channel and each speaker unit is driven to output the corresponding target audio signal.
[0006] Optionally, in a first implementation of the first aspect of the present invention, the step of injecting a sweep excitation signal into each speaker unit, acquiring a first sound pressure response signal, and taking the time-domain average of multiple acquired first sound pressure response signals to obtain a second sound pressure response signal includes: The frequency sweep excitation signal is injected into each speaker unit, and the first sound pressure response signal is synchronously acquired at each measurement point by the bright area microphone array and the dark area microphone array. The second sound pressure response signal is obtained by performing time-domain averaging on the first sound pressure response signal collected multiple times.
[0007] Optionally, in a second implementation of the first aspect of the present invention, the step of solving for the optimal loudspeaker weight vector at each frequency point based on the second sound pressure response signal and the frequency sweep excitation signal includes: The second sound pressure response signal and the swept frequency excitation signal are subjected to discrete Fourier transform respectively. The frequency domain result of the second sound pressure response signal is multiplied by the frequency point of the regularized inverse spectrum of the swept frequency excitation signal and then subjected to inverse discrete Fourier transform to obtain the room impulse response signal from each loudspeaker unit to each measurement point. The room impulse response signal is subjected to discrete Fourier transform, and the sound pressure transfer function matrix of the bright area and the sound pressure transfer function matrix of the dark area are assembled according to the measurement points of the bright area and the measurement points of the dark area, respectively. Based on the measurement points in the bright area, construct the particle velocity transfer function vector at the center point of the bright area, and combine the particle velocity transfer function vector at the center point of the bright area with the corresponding row in the sound pressure transfer function matrix of the bright area to form a sound intensity transfer function pair. Based on the bright area sound pressure transfer function matrix, the dark area sound pressure transfer function matrix, and the sound intensity transfer function pair, the optimal loudspeaker weight vector at each frequency point is solved.
[0008] Optionally, in a third implementation of the first aspect of the present invention, the step of solving for the optimal loudspeaker weight vector at each frequency point based on the bright area sound pressure transfer function matrix, the dark area sound pressure transfer function matrix, and the sound intensity transfer function pair includes: A pressure matching error term is constructed based on the sound pressure transfer function matrix of the bright area and the target sound pressure vector of equal amplitude. A flatness deviation term is constructed based on the mean square relative deviation of the sound pressure power at each measurement point in the bright area relative to the average sound pressure power in the bright area. An objective function is constructed based on the pressure matching error term and the flatness deviation term. The lower bound of the component of the sound intensity at the center point of the bright area in the direction of the target propagation, the upper bound of the sound pressure power at each measurement point in the bright area, and the upper bound of the total sound pressure power in the dark area corresponding to the sound pressure transfer function matrix of the dark area are used as constraints. A joint optimization model is constructed based on the objective function and the constraints. The joint optimization model is iteratively solved, and the optimal loudspeaker weight vector at each frequency point is obtained by terminating the solution with the relative change of the loudspeaker weight vector between two adjacent frequencies being lower than the convergence threshold.
[0009] Optionally, in a fourth implementation of the first aspect of the present invention, the iterative solution of the joint optimization model, with the termination condition being that the relative change in the loudspeaker weight vector between two adjacent iterations is less than a convergence threshold, to obtain the optimal loudspeaker weight vector at each frequency point, includes: The mean value of the sound pressure power at each measurement point in the flatness deviation term is fixed as the calculated value at the current iteration point. The sound intensity direction constraint is expanded in first order Taylor at the current iteration point. After linearization, auxiliary real variables are introduced to transform the joint optimization model into a convex problem. Using the Tikhonov regularized least squares solution as the initial loudspeaker weight vector, the convex subproblem is solved iteratively. The optimal loudspeaker weight vector at each frequency point is obtained when the relative change of the loudspeaker weight vector between two adjacent iterations is lower than the convergence threshold.
[0010] Optionally, in a fifth implementation of the first aspect of the present invention, the step of converting the optimal loudspeaker weight vector into time-domain filter coefficients for each channel and driving each loudspeaker unit to output the corresponding target audio signal includes: The optimal loudspeaker weight vector at each frequency point is conjugate symmetrically extended and then subjected to inverse discrete Fourier transform to obtain the time-domain filter coefficients of each channel. The time-domain filter coefficients of each channel are loaded into the DSP processor. The input raw audio signal is convolved with the time-domain filter coefficients for each channel to obtain the driving signal for each channel and drive each speaker unit to output the corresponding target audio signal.
[0011] Secondly, embodiments of this application provide an audio control device based on audio detection, comprising: The acquisition module is used to inject a sweep frequency excitation signal into each speaker unit, acquire the first sound pressure response signal, and take the time-domain average of the first sound pressure response signal acquired multiple times to obtain the second sound pressure response signal; The solution module is used to solve for the optimal loudspeaker weight vector at each frequency point based on the second sound pressure response signal and the frequency sweep excitation signal. The output module is used to convert the optimal speaker weight vector into time-domain filter coefficients for each channel and drive each speaker unit to output the corresponding target audio signal.
[0012] Thirdly, embodiments of this application provide a computer device, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, it implements the steps of the above-described audio control method based on audio detection.
[0013] Fourthly, embodiments of this application provide a readable storage medium storing a computer program that, when executed by a processor, implements the steps of the audio control method based on audio detection described above.
[0014] Fifthly, embodiments of this application provide a computer program product, which includes a computer program that, when executed by a processor, implements the steps of the audio control method based on audio detection described above.
[0015] In one of the solutions provided by the aforementioned audio control methods, devices, equipment, media, and program products based on audio detection, sound pressure response signals are synchronously acquired at various measurement points in both bright and dark areas using a microphone array. Frequency domain deconvolution modeling is then performed in conjunction with a frequency sweep excitation signal. This accurately constructs the sound pressure transfer function matrix for the bright area, the sound pressure transfer function matrix for the dark area, and the sound intensity transfer function pair at the center point of the bright area within the frequency domain. This solves the technical problem of insufficient transfer function modeling accuracy leading to the accumulation of sound field control errors in existing technologies. This invention incorporates the pressure matching error term and the flatness deviation term into a unified joint objective function. Simultaneously, it introduces three types of hard constraints: a lower bound constraint on the sound intensity direction, an upper bound constraint on the sound pressure amplitude in the bright area, and a sound pressure suppression constraint in the dark area. Within a single optimization framework, it simultaneously considers the accuracy of sound pressure reconstruction in the bright area, the spatial consistency of sound pressure in the bright area, and the sound pressure suppression effect in the dark area. The introduction of the sound intensity direction constraint ensures that the direction of the synthesized sound energy propagation in the bright area is precisely aligned with the target location, effectively eliminating the virtual sound image offset problem caused by neglecting the directionality of sound intensity in existing technologies, and ensuring the listener's accurate perception of the spatial direction of the sound source. This invention transforms a non-convex optimization problem into a sequential convex subproblem through a continuous convex approximation strategy and solves iteratively using the interior point method. While ensuring the convergence of the solution, the optimal loudspeaker weight vector is directly converted into the time-domain filter coefficients of each channel through inverse transformation and loaded into the DSP processor for real-time driving. Attached Figure Description
[0016] To more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings used in the description of the embodiments of the present invention will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0017] Figure 1 This is a schematic diagram of a sound control system based on audio detection in one embodiment of the present invention; Figure 2 This is a schematic flowchart of an audio control method based on audio detection in one embodiment of the present invention; Figure 3 yes Figure 2 A schematic diagram of the implementation process of step S10; Figure 4 yes Figure 2 A schematic diagram of the implementation process of step S20; Figure 5 yes Figure 4 A schematic diagram of the implementation process of step S24; Figure 6 yes Figure 2 A schematic diagram of the implementation process of step S30; Figure 7 This is a schematic diagram of a sound control device based on audio detection in one embodiment of the present invention; Figure 8 This is a schematic diagram of the structure of a computer device according to an embodiment of the present invention. Detailed Implementation
[0018] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0019] It should be understood that, when used in this specification and the appended claims, the term "comprising" indicates the presence of the described features, integrals, steps, operations, elements, and / or components, but does not exclude the presence or addition of one or more other features, integrals, steps, operations, elements, components, and / or collections thereof. It should also be understood that, as used in this specification and the appended claims, the term "and / or" refers to any combination of one or more of the associated listed items and all possible combinations, and includes such combinations.
[0020] Furthermore, in the description of this invention and the appended claims, the terms "first," "second," "third," etc., are used only to distinguish descriptions and should not be construed as indicating or implying relative importance.
[0021] References to "one embodiment" or "some embodiments" as described in this specification mean that one or more embodiments of the invention include a specific feature, structure, or characteristic described in connection with that embodiment. Therefore, the phrases "in one embodiment," "in some embodiments," "in other embodiments," "in still other embodiments," etc., appearing in different parts of this specification do not necessarily refer to the same embodiment, but rather mean "one or more, but not all, embodiments," unless otherwise specifically emphasized. The terms "comprising," "including," "having," and variations thereof mean "including but not limited to," unless otherwise specifically emphasized.
[0022] It should be understood that the sequence number of each step in the following embodiments does not imply the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present invention.
[0023] To illustrate the technical solution of the present invention, specific embodiments are described below.
[0024] To address the problems mentioned above in the background art, embodiments of this application propose a sound control method, apparatus, device, medium, and program product based on audio detection. The sound control method based on audio detection provided by the embodiments of this invention can be applied to, for example... Figure 1 The audio detection-based audio control system shown includes a client and a server.
[0025] In one embodiment, such as Figure 2 As shown, an audio control method based on audio detection is provided, which is then applied to... Figure 1 The following steps are used as an example to illustrate the audio control system based on audio detection: S10: Inject a sweep frequency excitation signal into each speaker unit, collect the first sound pressure response signal, and take the time-domain average of the collected first sound pressure response signals to obtain the second sound pressure response signal; S20: Based on the second sound pressure response signal and the frequency sweep excitation signal, solve for the optimal loudspeaker weight vector at each frequency point; S30: Convert the optimal speaker weight vector into time-domain filter coefficients for each channel and drive each speaker unit to output the corresponding target audio signal.
[0026] In this embodiment, the control unit sequentially selects each speaker unit according to its channel number and injects a sweep excitation signal into the selected speaker unit. The sweep excitation signal can be a logarithmic sinusoidal sweep signal with a start frequency of 20Hz, an end frequency of 20000Hz, and a duration of 10s. The sampling rate is set to 48000Hz, which can cover the conventional audible frequency band and provide sufficient frequency sampling basis for frequency domain transfer function estimation. During the excitation of each speaker unit, the bright area microphone array and the dark area microphone array synchronously collect the sound pressure response of each measurement point. The sound pressure response between the same speaker unit and the same measurement point can be repeatedly collected 3 times, and the repeated collection results are averaged point by point in the time domain to obtain the second sound pressure response signal after the random background noise is weakened.
[0027] In this embodiment, the second sound pressure response signal and the frequency sweep excitation signal are fed into the frequency domain analysis process. The room impulse response from each speaker unit to each measurement point is extracted through discrete Fourier transform, regularized inverse spectrum multiplication, and inverse discrete Fourier transform. Then, based on the spatial assignment of the bright and dark area measurement points, the bright area sound pressure transfer function matrix and the dark area sound pressure transfer function matrix are assembled respectively. Finally, the sound intensity transfer function pair is formed by combining the sound pressure difference relationship near the center measurement point of the bright area. After completing the above modeling, a solution process for the speaker complex weight vector is established for each frequency point to make the synthesized sound pressure in the bright area as close as possible to the equal-amplitude target sound pressure, while constraining the flatness of the bright area sound pressure distribution, the sound intensity component in the target propagation direction, the upper limit of the bright area sound pressure, and the total sound pressure power in the dark area. The solution process can use a continuous convex approximation strategy for iteration, for example, setting the convergence threshold to a relative change of no more than 10 between two adjacent speaker weight vectors. -5 The maximum number of iterations is set to 50.
[0028] In this embodiment, after obtaining the optimal speaker weight vector at each frequency point, the frequency domain weight sequence is conjugate symmetrically extended and inverse discrete Fourier transform is performed to generate time-domain filter coefficients corresponding to each speaker channel. The filter coefficients can be loaded into the multi-channel filtering module of the DSP processor. During the playback stage, the input raw audio signal is convolved with the corresponding time-domain filter coefficients for each channel to obtain the driving signal for each channel. Then, the corresponding speaker unit is driven by the power amplifier to output the target audio signal, thereby forming a sound field with target directionality and good flatness in the bright area, while maintaining sound pressure suppression effect in the dark area.
[0029] In one embodiment, such as Figure 3 As shown, step S10 specifically includes the following steps: S11: Inject the frequency sweep excitation signal into each speaker unit, and simultaneously collect the first sound pressure response signal at each measurement point by the bright area microphone array and the dark area microphone array; S12: Perform time-domain averaging on the first sound pressure response signal collected multiple times to obtain the second sound pressure response signal.
[0030] In this embodiment, the control unit establishes the excitation sequence according to the channel number of the speaker array, selects only one speaker unit in any acquisition cycle, and injects a logarithmic sinusoidal frequency sweep excitation signal into the selected speaker unit. The frequency sweep excitation signal can be expressed as: ; in, ,in, For a moment The corresponding sweep excitation signal amplitude, Sampling time, The sweep frequency is set to start at 20Hz. The sweep termination frequency is set to 20000Hz. The sweep duration is set to 10 seconds. This sweep range covers the conventional audible frequency band, and the 10-second duration provides sufficient excitation period for the low-frequency band. The sampling rate can be set to 48000Hz. In the... During the output of the frequency sweep excitation signal by each speaker unit, the bright area microphone array and the dark area microphone array maintain the same sampling clock, synchronously recording the first sound pressure response signal at the bright area measurement point and the dark area measurement point, respectively. The acquisition channel maintains consistency in trigger time, sampling rate, buffer length, and timestamp, so that the response signals at different measurement points can correspond to the same excitation time axis; after completing the first... After the data acquisition for each speaker unit is completed, the control unit shuts down the current speaker unit and switches to the next speaker unit until the response relationship between all speaker units and all measurement points has been acquired.
[0031] In this embodiment, to reduce the impact of random background noise on the transfer function estimation, the first sound pressure response signal between the same speaker unit and the same measurement point is repeatedly acquired three times. Before performing time-domain averaging, start trigger point correction and sampling length truncation are performed to avoid time shifts between different acquisition rounds. The time-domain averaging calculation can be performed as follows: ; in, For the first The microphone measurement points are for the first... The second sound pressure response signal generated by each speaker unit For the first The first sound pressure response signal obtained from the second acquisition. Number the microphone measurement points. Number the speaker unit. Number the number of times the data was collected repeatedly.
[0032] In one embodiment, such as Figure 4 As shown, step S20 specifically includes the following steps: S21: Perform discrete Fourier transform on the second sound pressure response signal and the sweep frequency excitation signal respectively, and multiply the frequency domain result of the second sound pressure response signal with the regularized inverse spectrum of the sweep frequency excitation signal frequency by frequency point, and then perform inverse discrete Fourier transform to obtain the room impulse response signal from each loudspeaker unit to each measurement point. S22: Perform Discrete Fourier Transform on the room impulse response signal and assemble the sound pressure transfer function matrix of the bright area and the sound pressure transfer function matrix of the dark area according to the measurement points of the bright area and the measurement points of the dark area respectively; S23: Construct the particle velocity transfer function vector of the center point of the bright area based on the measurement points of the bright area, and combine the particle velocity transfer function vector of the center point of the bright area with the corresponding row in the sound pressure transfer function matrix of the bright area to form a sound intensity transfer function pair; S24: Based on the sound pressure transfer function matrix of the bright area, the sound pressure transfer function matrix of the dark area, and the sound intensity transfer function pair, solve for the optimal loudspeaker weight vector at each frequency point.
[0033] In this embodiment, for each combination of loudspeaker units and measurement points, frequency domain deconvolution is first performed based on the second sound pressure response signal and the sweep excitation signal of the full sampling length to obtain the room impulse response signal from the loudspeaker unit to the measurement point. After obtaining the room impulse response signal, the effective response segment containing the direct sound and the main early reflections is extracted, and the effective response segment is zero-padded to 4096 points for discrete Fourier transform. The sampling rate is maintained at 48000Hz, thus forming a frequency sampling interval of approximately 11.72Hz in the frequency domain weighting solution stage. To avoid the low energy of the sweep excitation signal at individual frequency points, which would lead to an excessively small denominator for the deconvolution integral, the system constructs a regularized inverse spectrum: ; in, For the first Regularized inverse spectrum at discrete angular frequency points The frequency domain result of the frequency sweep excitation signal, for The conjugate quantity, For the first Discrete angular frequency points, Let be the regularization coefficient and take The regularization coefficient, based on the maximum power spectrum of the swept excitation signal, can suppress numerical instability while having minimal impact on the amplitude estimation of the transfer function. The frequency domain result of the second sound pressure response signal is multiplied frequency-by-frequency by the regularized inverse spectrum, and an inverse discrete Fourier transform is performed to obtain the room impulse response signal: ,in, For the first The number of speaker units up to the number Room impulse response signal at each measurement point, This is the frequency domain result of the second sound pressure response signal. The first 512 points of the room impulse response signal are taken as the effective response segment. The time window corresponding to these 512 points at a sampling rate of 48000Hz is approximately 10.67ms. This time window is suitable for near-field desktop or cabin personal audio scenarios where direct sound and major early reflections dominate. In scenarios with larger spaces, obvious low-frequency room modes, or strong influence from distant boundary reflections, the length of the effective response segment can be adjusted according to the response energy attenuation threshold and the upper limit of real-time processing delay.
[0034] The truncated room impulse response signal is zero-padding to 4096 points and then subjected to discrete Fourier transform to obtain the frequency domain transfer function values from each loudspeaker unit to each measurement point. The measurement points belonging to the bright area are arranged in rows to form the bright area sound pressure transfer function matrix, and the measurement points belonging to the dark area are arranged in rows to form the dark area sound pressure transfer function matrix, so that the bright area sound pressure reconstruction and the dark area sound pressure suppression can be expressed separately at the same frequency point.
[0035] Select the center measurement point of the bright area, and construct the particle velocity transfer function vector based on the sound pressure difference relationship between adjacent microphone measurement points in the bright area: ; in, For the first The speaker unit in the first The transfer function of particle velocity in the x-direction formed at a pair of adjacent measurement points For the first The number of speaker units up to the number The sound pressure transfer function at each measurement point in the bright area air density and take , The distance between the measurement points of microphones in adjacent bright areas. The unit is the imaginary unit. The resulting vector of particle velocity transfer function at the center point of the bright region, together with the corresponding row in the sound pressure transfer function matrix of the bright region, constitutes a sound intensity transfer function pair.
[0036] After completing the above modeling, the target sound pressure reconstruction relationship is described by the sound pressure transfer function matrix of the bright area, the leakage sound pressure suppression relationship is described by the sound pressure transfer function matrix of the dark area, and the sound intensity transfer function pairs are used to describe the sound energy propagation direction relationship of the center of the bright area. At each frequency point, the complex weight vector corresponding to the number of loudspeaker units is solved.
[0037] In one embodiment, such as Figure 5 As shown, step S24 specifically includes the following steps: S241: Construct a pressure matching error term based on the sound pressure transfer function matrix of the bright area and the target sound pressure vector of equal amplitude; construct a flatness deviation term based on the mean square relative deviation of the sound pressure power at each measurement point in the bright area relative to the average sound pressure power in the bright area; and construct an objective function based on the pressure matching error term and the flatness deviation term. S242: The lower bound of the component of the sound intensity at the center point of the bright area in the direction of the target propagation, the upper bound of the sound pressure power at each measurement point in the bright area, and the upper bound of the total sound pressure power in the dark area corresponding to the sound pressure transfer function matrix of the dark area are used as constraints, and a joint optimization model is constructed based on the objective function and the constraints. S243: Iteratively solve the joint optimization model and terminate the solution with the relative change of the loudspeaker weight vector between two adjacent loudspeaker weight vectors being lower than the convergence threshold to obtain the optimal loudspeaker weight vector at each frequency point.
[0038] In this embodiment, at each discrete frequency point, the loudspeaker complex weight vector is used as the object to be solved. First, a pressure matching error term is constructed using the bright area sound pressure transfer function matrix and the equal-amplitude target sound pressure vector to make the synthesized sound pressure at each measurement point in the bright area as close as possible to the target sound pressure distribution. The pressure matching error term is written as: ; in, This is the normalized pressure matching error term. Let be the complex weight vector of the loudspeaker. Let be the sound pressure vector of the target with constant amplitude. A flatness deviation term is constructed based on the deviation of the sound pressure power at each measurement point in the bright area from the average sound pressure power of the bright area. This ensures that the sound pressure power distribution among different measurement points in the bright area remains consistent. The flatness deviation term is written as: ; in, This is the normalized flatness deviation term. The number of measurement points in the bright area. The sound pressure transfer function matrix in the bright region is the first... row transpose vector Let be the average sound pressure power in the bright region. The two dimensionless error terms are combined into a joint objective function: ; in, For the joint objective function, For pressure matching weighting coefficients, The flatness weighting coefficients are both initially set to 0.5, and their sum is kept to 1. This setting ensures that the target sound pressure reconstruction in the bright area and the spatial flatness of the bright area have the same optimization status in the initial solution stage. After completing the objective function construction, the sound intensity direction, the upper limit of the sound pressure in the bright area, and the suppression of the sound pressure in the dark area are included in the constraints. Among them, the sound intensity component of the center point of the bright area in the target propagation direction is not lower than the lower bound of the sound intensity, and the lower bound of the sound intensity is determined according to... set up, A value of 0.8 indicates that the sound intensity component in the target propagation direction is not less than 80% of the theoretical sound intensity of a plane wave of constant amplitude. The upper limit of the sound pressure power at each measurement point in the bright area is determined jointly based on the rated output capability of the loudspeaker unit, the linear operating range of the power amplifier, and the target listening sound pressure level. The upper limit of the total sound pressure power in the dark area is determined based on the target acoustic contrast, the target sound pressure power in the bright area, and the number of measurement points in the dark area. This ensures that the optimal weight vector satisfies the sound pressure reconstruction requirements in the bright area while providing a quantifiable constraint on the leakage sound pressure in the dark area.
[0039] After forming a joint optimization model based on the objective function and the aforementioned constraints, the joint optimization model is iterated by solving each frequency point independently. During the iteration, the flatness correlation and sound intensity direction correlation are updated according to the current weight vector, and the relative change of the loudspeaker weight vector between two adjacent iterations is compared with the convergence threshold, which is set to 10. -5 The maximum number of iterations is set to 50 to balance numerical stability and computational cost in the solution process. When the iteration meets the termination condition, the optimal loudspeaker weight vector at the corresponding frequency point is output.
[0040] In one embodiment, step S243 specifically includes the following steps: The mean value of the sound pressure power at each measurement point in the flatness deviation term is fixed as the calculated value at the current iteration point. The sound intensity direction constraint is expanded by first-order Taylor at the current iteration point. After linearization, auxiliary real variables are introduced to transform the joint optimization model into a convex subproblem. Using the Tikhonov regularized least squares solution as the initial loudspeaker weight vector, the convex subproblem is solved iteratively. The optimal loudspeaker weight vector at each frequency point is obtained when the relative change of the loudspeaker weight vector between two adjacent iterations is lower than the convergence threshold.
[0041] In this embodiment, during the iterative solution phase, the non-convex part of the joint optimization model is transformed into a sequential convex subproblem that is easier to solve numerically. The initial loudspeaker weight vector is formed using the Tikhonov regularized least squares solution. ; in, This is the initial speaker weight vector. This is the conjugate transpose of the sound pressure transfer function matrix in the bright region. Let be the regularization coefficient and take , For matrix The largest eigenvalue, This is the identity matrix. The regularization value is based on the largest eigenvalue, which can suppress drastic weight fluctuations caused by ill-conditioned matrix structures, while keeping the regularization bias within a small range. (Continue to the next step...) After the next iteration, the sound pressure power at each measurement point in the bright area is calculated based on the loudspeaker weight vector from the previous round, and the average sound pressure power in the bright area is fixed as a constant in the current iteration, specifically: , ; in, For the first The weights obtained from the previous round in the next iteration are... Sound pressure power at each measurement point in the bright area This represents the average sound pressure power in the corresponding bright area. (By fixing...) The denominator of the flatness deviation term, which originally varied with the weights, is now a constant. The flatness deviation term can be expressed as follows within the current iteration: ; The system introduces auxiliary real variables for each measurement point in the bright area, establishing a convex constraint relationship between the sound pressure power term and the auxiliary real variables, allowing the flatness deviation term to participate in the solution using a second-order cone programming approach. For the sound intensity direction constraint, the system performs a first-order Taylor expansion on the sound intensity components at the center point of the bright area at the current iteration point: ; in, Let be the complex gradient of the sound intensity component with respect to the weight vector, and calculate it using the following formula: ; Therefore, the original intensity direction constraint containing bilinear terms is transformed into a linear inequality constraint regarding the loudspeaker weight vector. After linearizing the flatness deviation term and the intensity direction constraint, the pressure matching error term, flatness deviation term, upper limit of sound pressure power in the bright area, upper limit of sound pressure power in the dark area, and the linearized intensity direction constraint are combined to form the current wheel convex problem, which is then solved using the interior-point method. The iteration complexity of a single interior-point method can be reduced by... Estimation. After each round of solving the convex subproblem, a new loudspeaker weight vector is obtained, and the relative change is used to determine whether the outer iteration has terminated: The relative convergence threshold is taken as The maximum number of outer iterations is 50. When the relative change is lower than the threshold, it is determined that the weight vector of the current frequency point has stabilized, and the optimal speaker weight vector of the current frequency point is output. If the relative convergence condition is not met even after reaching the maximum number of iterations, the last round weight vector that satisfies the constraints and has a better objective function value is output.
[0042] In one embodiment, such as Figure 6 As shown, step S30 specifically includes the following steps: S31: After performing conjugate symmetric extension on the optimal loudspeaker weight vector at each frequency point, apply inverse discrete Fourier transform to obtain the time-domain filter coefficients of each channel. S32: Load the time-domain filter coefficients of each channel into the DSP processor, perform convolution operation on the input raw audio signal with the time-domain filter coefficients for each channel, obtain the driving signal of each channel, and drive each speaker unit to output the corresponding target audio signal.
[0043] In this embodiment, the frequency domain weight sequence of each channel at all positive frequency points is extracted according to the speaker channel number, and the frequency domain weight sequence is conjugate symmetric extended so that the filter coefficients after the inverse discrete Fourier transform are a real number sequence. The extension relationship can be expressed as: ; in, For the first The speaker channel in the first Optimal frequency domain weights at each frequency point Let 4096 be the number of points for the Discrete Fourier Transform. For the first The frequency domain consists of discrete angular frequency points. Through the above extension process, the amplitude and phase control results at the positive frequency points can be mapped to the negative frequency points, and the frequency domain sequence satisfies the symmetry conditions required for the inverse transformation of the real signal. A 4096-point inverse discrete Fourier transform is applied to the extended complete frequency domain weight sequence to obtain the time domain weight sequence corresponding to each speaker channel. The center 512 points are truncated from the time domain weight sequence as an effective filter segment, and a Hanning window is applied to reduce the sidelobe leakage caused by the truncation, thereby forming a 512-point time domain filter coefficient vector for each speaker channel. When the 512-point length matches the 48000Hz sampling rate, the filter group delay is approximately 5.33ms under center alignment or approximately linear phase implementation. The delay level is suitable for the real-time playback scenario of a personal audio system. At the same time, this length can accommodate the early acoustic response characteristics preserved in the aforementioned room impulse response modeling, so that the bright area sound pressure matching, flatness control, sound intensity direction constraint, and dark area sound pressure suppression results obtained from frequency domain optimization can be implemented in the form of time domain filtering.
[0044] The filter coefficients of all speaker channels are written to the multi-channel coefficient register or on-chip memory of the DSP processor. During the writing process, the channel numbers are kept consistent with the physical connection relationships of the speakers to avoid misalignment between the channel order during frequency domain calculation and the actual driving order. Simultaneously, amplitude normalization and fixed-point quantization are performed on the filter coefficients to prevent overflow during multiplication and addition operations within the processor. A buffer switching method is used during coefficient updates to ensure a smooth transition between old and new coefficients in the playback chain. After the input raw audio signal enters the processor, single-channel audio samples are copied to the filtering branches of each channel, and each channel undergoes a convolution operation according to the corresponding time-domain filter coefficients. The convolution relationship can be expressed as: In the formula, For the first The speaker channel in the first The drive signal output at each sampling time. For the first The first speaker channel Each time-domain filter coefficient For the original audio signal in delay The input sampled value after each sampling point This represents the total number of speaker channels. Through the convolution operation described above, the original audio signal is assigned different amplitude and phase adjustment values in each channel. The drive signals for each channel are then converted from digital to analog and amplified before being sent to the corresponding speaker units. The speaker array is then superimposed in space to form the target sound field. Since the time-domain filter coefficients of each channel are derived from the optimal speaker weight vector at each frequency point, the output sound field can inherit the target direction sound intensity, bright area sound pressure consistency, and dark area sound pressure suppression characteristics obtained in the frequency domain optimization stage. This ensures that the target audio signal within the listening area maintains a relatively stable sound image direction and sound pressure distribution, while reducing the perceptible leakage sound pressure in non-target areas.
[0045] It should be understood that the sequence number of each step in the above embodiments does not imply the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present invention.
[0046] In one embodiment, an audio control device based on audio detection is provided, which corresponds one-to-one with the audio control method based on audio detection described in the above embodiments. For example... Figure 7 As shown, the audio control device based on audio detection includes: The acquisition module 701 is used to inject a sweep frequency excitation signal into each speaker unit, acquire the first sound pressure response signal, and take the time-domain average of the acquired first sound pressure response signals to obtain the second sound pressure response signal. The solver module 702 is used to solve the optimal loudspeaker weight vector at each frequency point based on the second sound pressure response signal and the frequency sweep excitation signal. The output module 703 is used to convert the optimal speaker weight vector into the time-domain filter coefficients of each channel and drive each speaker unit to output the corresponding target audio signal.
[0047] It should be noted that the information interaction and execution process between the above-mentioned devices / units are based on the same concept as the method embodiments of this application. For details on their specific functions and technical effects, please refer to the method embodiments section, and they will not be repeated here.
[0048] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the above-described division of functional units and modules is merely an example. In practical applications, the above functions can be assigned to different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above. The functional units and modules in the embodiments can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit. Furthermore, the specific names of the functional units and modules are only for easy differentiation and are not intended to limit the scope of protection of this application. The specific working process of the units and modules in the above system can be referred to the corresponding process in the foregoing method embodiments, and will not be repeated here.
[0049] This application also provides a computer device, such as... Figure 8 As shown, the computer device includes: at least one processor, a memory, and a computer program stored in the memory and executable on the at least one processor. When the processor executes the computer program, it implements the steps in any of the above method embodiments, or when the processor executes the computer program, it implements the functions of each module / unit in the above device embodiments.
[0050] For example, the computer program may be divided into one or more modules / units, which are stored in the memory and executed by the processor to complete this application. The one or more modules / units may be a series of computer program instruction segments capable of performing a specific function, which describe the execution process of the computer program in the computer device.
[0051] Those skilled in the art will understand that Figure 8 The computer device described is merely an example and does not constitute a limitation on the computer device. It may include more or fewer components than shown, or combine certain components, or different components. For example, the computer device may also include input / output devices, network access devices, buses, etc.
[0052] The aforementioned processor can be a Central Processing Unit (CPU), or other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), or field-programmable gate arrays (FPGAs). Programmable Gate Array (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. A general-purpose processor can be a microprocessor or any conventional processor.
[0053] The memory can be an internal storage unit of the computer device, such as a hard drive or RAM. The memory can also be an external storage device of the computer device, such as a plug-in hard drive, Smart Media Card (SMC), Secure Digital (SD) card, or Flash Card. Furthermore, the memory can include both internal and external storage units of the computer device.
[0054] This application also provides a readable storage medium storing a computer program, which, when executed by a processor, can implement the steps in the above-described method embodiments.
[0055] This application provides a computer program product that, when run on an electronic device, enables the electronic device to implement the steps described in the various method embodiments above.
[0056] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, all or part of the processes in the methods of the above embodiments of this application can be implemented by a computer program instructing related hardware. The computer program can be stored in a computer-readable storage medium, and when executed by a processor, it can implement the steps of the various method embodiments described above. The computer program includes computer program code, which can be in the form of source code, object code, executable files, or certain intermediate forms. The computer-readable medium can include at least: any entity or device capable of carrying computer program code to a photographing device / terminal device, a recording medium, a computer memory, a read-only memory (ROM), a random access memory (RAM), an electrical carrier signal, a telecommunication signal, and a software distribution medium. Examples include USB flash drives, portable hard drives, magnetic disks, or optical disks. In some jurisdictions, according to legislation and patent practice, computer-readable media cannot be electrical carrier signals or telecommunication signals.
[0057] In the above embodiments, the descriptions of each embodiment have different focuses. For parts that are not described in detail or recorded in a certain embodiment, please refer to the relevant descriptions of other embodiments.
[0058] Those skilled in the art will recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.
[0059] In the embodiments provided in this application, it should be understood that the disclosed apparatus / devices and methods can be implemented in other ways. For example, the apparatus / device embodiments described above are merely illustrative. For instance, the division of modules or units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between apparatuses or units may be electrical, mechanical, or other forms.
[0060] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.
[0061] The above-described embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit them. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of this application, and should all be included within the protection scope of this application.
Claims
1. An audio detection-based sound control method, characterized by, include: A sweep frequency excitation signal is injected into each speaker unit, a first sound pressure response signal is collected, and the time-domain average of the first sound pressure response signal collected multiple times is taken to obtain a second sound pressure response signal; Based on the second sound pressure response signal and the frequency sweep excitation signal, the optimal loudspeaker weight vector at each frequency point is solved; The optimal speaker weight vector is converted into time-domain filter coefficients for each channel and each speaker unit is driven to output the corresponding target audio signal.
2. The audio detection-based sound control method of claim 1, wherein, The process of injecting a sweep excitation signal into each speaker unit, acquiring a first sound pressure response signal, and taking the time-domain average of multiple acquired first sound pressure response signals to obtain a second sound pressure response signal includes: The frequency sweep excitation signal is injected into each speaker unit, and the first sound pressure response signal is synchronously acquired at each measurement point by the bright area microphone array and the dark area microphone array. The second sound pressure response signal is obtained by performing time-domain averaging on the first sound pressure response signal collected multiple times.
3. The audio detection-based sound control method of claim 2, wherein, The step of solving for the optimal loudspeaker weight vector at each frequency point based on the second sound pressure response signal and the frequency sweep excitation signal includes: The second sound pressure response signal and the swept frequency excitation signal are subjected to discrete Fourier transform respectively. The frequency domain result of the second sound pressure response signal is multiplied by the frequency point of the regularized inverse spectrum of the swept frequency excitation signal and then subjected to inverse discrete Fourier transform to obtain the room impulse response signal from each loudspeaker unit to each measurement point. The room impulse response signal is subjected to discrete Fourier transform, and the sound pressure transfer function matrix of the bright area and the sound pressure transfer function matrix of the dark area are assembled according to the measurement points of the bright area and the measurement points of the dark area, respectively. Based on the measurement points in the bright area, construct the particle velocity transfer function vector at the center point of the bright area, and combine the particle velocity transfer function vector at the center point of the bright area with the corresponding row in the sound pressure transfer function matrix of the bright area to form a sound intensity transfer function pair. Based on the bright area sound pressure transfer function matrix, the dark area sound pressure transfer function matrix, and the sound intensity transfer function pair, the optimal loudspeaker weight vector at each frequency point is solved.
4. The audio control method based on audio detection as described in claim 3, characterized in that, The step of solving for the optimal loudspeaker weight vector at each frequency point based on the bright area sound pressure transfer function matrix, the dark area sound pressure transfer function matrix, and the sound intensity transfer function pair includes: A pressure matching error term is constructed based on the sound pressure transfer function matrix of the bright area and the target sound pressure vector of equal amplitude. A flatness deviation term is constructed based on the mean square relative deviation of the sound pressure power at each measurement point in the bright area relative to the average sound pressure power in the bright area. An objective function is constructed based on the pressure matching error term and the flatness deviation term. The lower bound of the component of the sound intensity at the center point of the bright area in the direction of the target propagation, the upper bound of the sound pressure power at each measurement point in the bright area, and the upper bound of the total sound pressure power in the dark area corresponding to the sound pressure transfer function matrix of the dark area are used as constraints. A joint optimization model is constructed based on the objective function and the constraints. The joint optimization model is iteratively solved, and the optimal loudspeaker weight vector at each frequency point is obtained by terminating the solution with the relative change of the loudspeaker weight vector between two adjacent frequencies being lower than the convergence threshold.
5. The audio control method based on audio detection as described in claim 4, characterized in that, The iterative solution of the joint optimization model, terminating when the relative change in the loudspeaker weight vector between two adjacent iterations is less than a convergence threshold, yields the optimal loudspeaker weight vector for each frequency point, including: The mean value of the sound pressure power at each measurement point in the flatness deviation term is fixed as the calculated value at the current iteration point. The sound intensity direction constraint is expanded in first order Taylor at the current iteration point. After linearization, auxiliary real variables are introduced to transform the joint optimization model into a convex subproblem. Using the Tikhonov regularized least squares solution as the initial loudspeaker weight vector, the convex subproblem is solved iteratively. The optimal loudspeaker weight vector at each frequency point is obtained when the relative change of the loudspeaker weight vector between two adjacent iterations is lower than the convergence threshold.
6. The audio control method based on audio detection as described in claim 1, characterized in that, The step of converting the optimal speaker weight vector into time-domain filter coefficients for each channel and driving each speaker unit to output the corresponding target audio signal includes: The optimal loudspeaker weight vector at each frequency point is conjugate symmetrically extended and then subjected to inverse discrete Fourier transform to obtain the time-domain filter coefficients of each channel. The time-domain filter coefficients of each channel are loaded into the DSP processor. The input raw audio signal is convolved with the time-domain filter coefficients for each channel to obtain the driving signal for each channel and drive each speaker unit to output the corresponding target audio signal.
7. A sound control device based on audio detection, characterized in that, The steps for implementing the audio detection-based audio control method as described in any one of claims 1 to 6 include: The acquisition module is used to inject a sweep frequency excitation signal into each speaker unit, acquire the first sound pressure response signal, and take the time-domain average of the first sound pressure response signal acquired multiple times to obtain the second sound pressure response signal; The solution module is used to solve for the optimal loudspeaker weight vector at each frequency point based on the second sound pressure response signal and the frequency sweep excitation signal. The output module is used to convert the optimal speaker weight vector into time-domain filter coefficients for each channel and drive each speaker unit to output the corresponding target audio signal.
8. A computer device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the steps of the audio control method based on audio detection as described in any one of claims 1 to 6.
9. A readable storage medium storing a computer program, characterized in that, When the computer program is executed by the processor, it implements the steps of the audio detection-based sound control method as described in any one of claims 1 to 6.
10. A computer program product, characterized in that, The computer program product includes a computer program that, when executed by a processor, implements the steps of the audio detection-based sound control method as described in any one of claims 1 to 6.