Method for providing a spatialized sound field

By using sparse loudspeaker arrays and digital signal processing, and by utilizing beamforming and HRTF-optimized signal processing, the problem of inaccurate virtual sound source localization in sparse transducer arrays was solved, achieving stable spatialized sound reproduction and sound field control.

CN115715470BActive Publication Date: 2025-11-18COMHEAR INC +1
View PDF 81 Cites 0 Cited by

Patent Information

Application Number
CN202080097794.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2019-12-30
Filing Date
2020-12-30
Publication Date
2025-11-18
Estimated Expiration
2040-12-30

AI Technical Summary

Technical Problem

Existing technologies struggle to provide accurate virtual sound source localization and spatialized sound reproduction for multiple listeners in unstable listening environments, especially in sparse transducer arrays. Furthermore, existing methods often require high-order cancellation or rely on specific listening positions and headphone usage, making them unsuitable for varying acoustic environments.

Method used

By employing a sparse speaker array combined with digital signal processing (DSP), and optimizing signal processing through beamforming and head-related transfer function (HRTF), spatial sound reproduction is achieved, adapting to changes in the head and environment of different listeners, reducing crosstalk, and enhancing sound field control.

Benefits of technology

It enables stable virtual sound source localization and spatial sound reproduction for multiple listeners in varying acoustic environments, reducing crosstalk and enhancing sound field control and listening experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115715470B_ABST
    Figure CN115715470B_ABST
Patent Text Reader

Abstract

A signal processing system and method for delivering spatialized sound from a sparse loudspeaker array to a user's ear by optimizing the sound waveforms. The system can provide a listening area within a room or space to provide spatialized sound to create a 3D audio effect. In a binaural mode, a binary loudspeaker array provides a target beam for a user's ear.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to digital signal processing for controlling loudspeakers, and more specifically, to a signal processing method for controlling a sparse loudspeaker array to deliver spatialized sound. Background Technology

[0002] For various purposes, every reference, patent, patent application or other specifically identified information is explicitly incorporated into this document in its entirety by reference.

[0003] Spatialized sound is useful for a range of applications, including virtual reality, augmented reality, and modified reality. Such systems typically consist of audio and video devices that provide three-dimensional perception of virtual audio and visual objects. One challenge in creating such systems is updating the audio signal processing scheme for an unstable listener so that the listener perceives the expected sound image, especially when using sparse transducer arrays.

[0004] Sound reproduction systems that attempt to give listeners a sense of space try to make them feel that the sound is coming from a location where a real sound source may not exist. For example, when a listener is seated in the "optimal position" in front of a good two-channel stereo system, it is possible to present a virtual sound field between the two speakers. If two identical signals are passed to the two speakers facing the listener, the listener should feel that the sound is coming from a location directly in front of him or her. If the input to one of the speakers is increased, the virtual sound source will be biased towards that speaker. This principle is called amplitude stereo, and it has been the most commonly used technique for mixing two-channel material since the two-channel stereo format was first introduced.

[0005] However, amplitude stereo itself cannot create an accurate virtual image beyond the angle spanned by the two speakers. In fact, even between the two speakers, amplitude stereo only works properly when the angle spanned by the speakers is 60 degrees or less.

[0006] Virtual source imaging systems work by optimizing the sound waves (amplitude, phase, delay) at the listener's ears. A real sound source generates a certain interaural time difference and level difference at the listener's ears, which the auditory system uses to locate the sound source. For example, a sound source on the listener's left is louder in the left ear and arrives earlier than in the right. Virtual source imaging systems are designed to accurately reproduce these cues. In practice, loudspeakers are used to reproduce a desired signal in the area surrounding the listener's ears. The loudspeaker input is determined by the characteristics of the desired signal, and the desired signal must be determined by the characteristics of the sound emitted by the virtual sound source. Therefore, a typical approach to sound localization is to determine the head-related transfer function (HRTF) representing the listener's binaural perception, as well as the influence of the listener's head, and then invert the HRTF and the sound processing and transfer chain back to the head to produce the optimized "desired signal." By limiting binaural perception to spatialized sound, acoustic emissions can be optimized to produce sound. For example, the HRTF then simulates the auricle of the ear. Barreto, Armando, and Navarun Gupta. “Auricular dynamic modeling for audio spatialization”, WSEAS Journal of Acoustics and Music 1, No. 1 (2004): 77-82.

[0007] Typically, a single transducer delivers sound optimally to only a single head, and optimizing for multiple listeners requires very high-order cancellation, ensuring that the sound intended for one listener is effectively canceled out at another. Accurate multi-user spatialization is difficult outside of an anechoic chamber unless headphones are used.

[0008] Binaural technology is commonly used for the reproduction of virtual sound images. Binaural technology is based on the principle that if the sound reproduction system can generate the same sound pressure at the listener's eardrum as the sound pressure produced by the real sound source, then the listener will not be able to distinguish the difference between the virtual image and the real sound source.

[0009] For example, typical discrete surround sound systems assume specific speaker settings to generate an optimal listening point where the auditory imaging is stable and robust. However, not all areas are suited to the appropriate specifications of such systems, further minimizing the already small optimal point. To achieve binaural technology on the speakers, it is necessary to eliminate crosstalk, which prevents the signal intended for one ear from being heard by the other. However, such crosstalk cancellation, typically achieved by time-invariant filters, only applies to specific listening positions, and the sound field can only be controlled at the optimal location.

[0010] A digital sound projector is a transducer or speaker array that is controlled to emit audio input signals in a controlled manner within a space in front of the array. Typically, sound is emitted in beams, directed in any direction within the half-space in front of the array. By using carefully chosen reflection paths derived from the room's characteristics, the listener will perceive the sound beam emitted by the array as if it originated from the location of its last reflection. If the last reflection occurred in a rear corner, the listener will perceive the sound as if it originated from a sound source behind him or her. However, human perception also involves echo processing, meaning that secondary and higher reflections should physically correspond to the listener's accustomed environment; otherwise, the listener may perceive distortion.

[0011] Therefore, if a person is looking for the sensation of sound coming from the left front of the listener in a rectangular room, the listener will expect a slightly delayed echo from the back, as well as further second-order reflections from another wall, each acoustically colored by the characteristics of the reflecting surfaces.

[0012] One application of digital sound projectors is to replace conventional discrete surround sound systems, which typically employ several separate speakers placed at different locations around the listener's position. Digital sound projectors create true surround sound at the listener's location by generating beams for each channel of the surround sound audio signal and directing the sound beams in the appropriate direction, without requiring additional speakers or wiring. Such a system is described in U.S. Patent Publication No. 2009 / 0161880 to Hooley et al., the disclosure of which is incorporated herein by reference.

[0013] Crosstalk cancellation is, in a sense, the ultimate problem of sound reproduction because an efficient crosstalk canceller provides complete control over the sound field at multiple “target” locations. The goal of a crosstalk canceller is to reproduce the desired signal at a single target location while perfectly eliminating sound at all other target locations. The basic principle of crosstalk cancellation using only two speakers and two target locations has been known for over 30 years. Atal and Schroeder’s US3,236,949 (1966) used physical reasoning to determine how a crosstalk canceller works, comprising only two speakers symmetrically placed in front of a single listener. To reproduce the short pulse only in the left ear, the left speaker first emits a positive pulse. This pulse must be canceled out in the right ear by a slightly weaker negative pulse emitted by the right speaker. Then, this negative pulse must be canceled out in the left ear by another, even weaker positive pulse emitted by the left speaker, and so on. Atal and Schroeder’s model assumes free-field conditions. The effects of the listener’s torso, head, and outer ear on the incoming sound waves are neglected.

[0014] In order to control the delivery of binaural signals or “target” signals, it is necessary to know how the listener’s torso, head, and auricles (outer ear) modify the incoming sound waves according to the location of the sound source. This information can be obtained by measuring a “dummy head” or a human subject. The result of such measurements is called the “head correlation transfer function,” or HRTF.

[0015] Significant differences in HRTF exist between listeners, especially at high frequencies. This large statistical variation in HRTF between listeners is one of the main problems with virtual source imaging on headphones. Headphones can control the reproduced sound very well. There is no "crosstalk" (sound doesn't circle around the head to the opposite ear), and the acoustic environment doesn't modify the reproduced sound (room reflections don't interfere with direct sound). However, unfortunately, when headphones are used for reproduction, the virtual image often feels too close to the head, sometimes even inside the head. This phenomenon is particularly difficult to avoid when attempting to place the virtual image directly in front of the listener. It seems necessary to compensate not only for the listener's own HRTF but also for the response of the headphones used for reproduction. Furthermore, the entire sound field moves with the listener's head (unless head tracking and sound field resynthesis are used, which requires significant additional processing power). On the other hand, spatialized loudspeaker reproduction using linear transducer arrays provides natural listening conditions, but requires compensation for crosstalk and also needs to account for reflections from the acoustic environment.

[0016] Comhear MyBeam TM A linear array employs digital signal processing (DSP) on identical, equidistant, independently powered, and perfectly phase-aligned loudspeaker elements to produce constructive and destructive interference. See US 9,578,440. The loudspeakers are intended to be placed in a linear array parallel to the listener's interaural axis, in front of the listener.

[0017] Beamforming, or spatial filtering, is a signal processing technique used in sensor arrays for directional signal transmission or reception. This is achieved by combining elements in an antenna array such that signals at specific angles experience constructive interference, while other signals experience destructive interference. To achieve spatial selectivity, beamforming can be used at both the transmitting and receiving ends. The improvement compared to omnidirectional reception / transmission is known as the array's directivity. Adaptive beamforming is used to detect and estimate the signal of interest at the sensor array output by means of optimal (e.g., least-squares) spatial filtering and interference suppression.

[0018] The Mybeam speaker is active—it contains its own amplifier and I / O, can be configured to include ambient monitoring for automatic volume adjustment, and can adjust its beamforming focus based on the listener's distance. It operates in several different modes, including binaural (transaural), single beamforming optimized for voice and privacy, near-field coverage, far-field coverage, and multiple listeners. In binaural mode with near-field or far-field coverage, Mybeam renders exceptionally clear standard PCM stereo music or video signals (compressed or uncompressed sources), a very wide and detailed soundstage, excellent dynamic range, and a strong sense of surround sound (the speaker's pictorial musicality is partly due to the speaker array's accurate phase calibration). The speaker operates at sampling rates up to 96kHz and 24-bit precision, reproducing high-resolution and high-definition audio with exceptional fidelity. High-resolution 3D audio imaging is readily perceived when reproducing PCM stereo signals with binaural processing. Height information and the 180-degree frontal view are rendered well, and rear imaging is achieved for some sources. The reference form factor includes 12-speaker, 10-speaker, and 8-speaker versions, with widths ranging from approximately 8 to 22 inches.

[0019] US5,862,227 discloses a spatialized sound reproduction system. This system employs z-domain filters and optimizes the coefficients of filters H1(z) and H2(z) to minimize a cost function given by [value missing]. Where E[·] is the expectation operator, and em(n) represents the error between the expected signal and the reproduced signal near the position of the magnetic head. The cost function can also have a term that is the sum of the squared magnitudes of the filter coefficients used in the penalty filters H1(z) and H2(z) to improve the conditions of the inversion problem.

[0020] Another spatialized sound reproduction system is disclosed in US6,307,941. Exemplary embodiments may use any combination of (i) FIR and / or IIR filters (digital or analog) and (ii) spatially shifted signals (e.g., coefficients), which are generated using any of the following methods: raw impulse response acquisition; balanced model order reduction; Hankel norm modeling; least squares modeling; modified or unmodified Prony method; minimum phase reconstruction; iterative pre-filtering; or critical band smoothing.

[0021] US9,215,544 relates to sound spatialization for multi-channel encoding of binaural reproduction on two loudspeakers. A summation process from multiple channels is used to define the signals from the left and right loudspeakers.

[0022] US7,164,768 provides a directional channel audio signal processor.

[0023] US8,050,433 provides an apparatus and method for eliminating crosstalk between a two-channel loudspeaker and the listener's two ears in a stereo generation system.

[0024] US9,197,977 and 9,154,896 relate to a method and apparatus for processing audio signals to create “4D” spatialized sound, which uses two or more loudspeakers and features multiple reflection modeling.

[0025] ISO / IEC FCD 23003-2:200x, Spatial Audio Object Coding (SAOC), Coding of Moving Images and Audio, ISO / IEC JTC1 / SC29 / WG11N10843, July 2009, London, UK, discusses stereo downmixing from audio streams in MPEG audio format. The encoding conversion is performed in two steps: In the first step, based on information from the rendering matrix, the object parameters (OLD, NRG, IOC, DMG, DCLD) from the SAOC bitstream are encoded into the spatial parameters (CLD, ICC, CPC, ADG) of the MPEG surround bitstream. In the second step, the object downmixing is modified based on parameters derived from the object parameters and the rendering matrix to form a new downmixed signal.

[0026] The signal and parameter calculations are performed per processing band m and parameter time slot l. The input signal to the code converter is a stereo downmixer, represented as...

[0027] The data available in the code converter are the covariance matrix E and the rendering matrix M. ren The mixing matrix D. The covariance matrix E is an approximation of the original signal matrix multiplied by its complex conjugate transpose, SS. * ≈E, where S=s n,k The elements E of the matrix are obtained from the objects OLD and IOC. in and Rendering matrix M ren The size 6×N is obtained through matrix multiplication Y = y n,k =M ren S determines the target rendering of the audio object S. The size of the downmixing weight matrix D is 2×N, and the downmixing signal is determined in the form of a two-row matrix by matrix multiplication X = DS.

[0028] The element d of the matrix ij (i = 1, 2; j = 0...N-1) From the dequantized DCLD and DMG parameters DMG was obtained from j =D DMG (j,l) and DCLD j =D DCLD(j,l).

[0029] The code converter is based on the rendering matrix M ren The described target rendering determines the parameters of the MPEG surround decoder. The six-channel target covariance is denoted by F and is given by... The code conversion process can be conceptually divided into two parts. In one part, three-channel rendering is performed on the left, right, and center channels. During this stage, parameters for undermixing modifications and prediction parameters for the TTT box used in the MPS decoder are obtained. In the other part, CLD and ICC parameters (OTT parameters, left-front-left-surround, right-front-right-surround) for rendering between the front and surround channels are determined. Spatial parameters are determined to control the rendering of the left and right channels, which consist of the front and surround signals. These parameters describe the parameters used for MPS decoding. TTT The prediction matrix of the TTT box (used for the CPC parameters of the MPS decoder) and the downmixer matrix G. TTT The prediction matrix for target rendering is obtained from the modified undermixing. A3 is a rendering matrix of decreasing size (3×N), correspondingly describing the rendering of the left, right, and center channels. It is A3 = D 36 M ren Use by The limited 6 to 3 part undermixing matrix D 36 Obtained.

[0030] Adjusting the partial downmixing weights w p p = 1, 2, 3 are the values ​​that make w p (y 2p-1 +y 2p The energy of y is equal to the energy required to reach the limiting factor. 2p-1 || 2 +‖y 2p || 2 sum.

[0031] w3 = 0.5, where f i,j Let F represent the elements. This is used to estimate the desired prediction matrix C. TTT Given the preprocessing matrix G, we constrain the prediction matrix C3 to a size of 3×2 that results in target rendering C3X ≈ A3S. This type of matrix is ​​obtained by considering the normal equation C3(DED). * )≈A3ED * And exported.

[0032] Given a target covariance model, the solution to the normal equation produces the best possible waveform match for the target output. G and C TTT Now we solve the system of equations C. TTT G = C3 is obtained. To avoid the calculation term J = (DED)* ) -1 Numerical issues arose, so J was modified. First, the eigenvalues ​​λ of J were calculated. 1,2 Solve for det(J-λ) 1,2 I) = 0. Sort the eigenvalues ​​in descending order (λ1 ≥ λ2), and calculate the eigenvector corresponding to the larger eigenvalue according to the equation above. It must lie in the positive x-plane (the first element must be positive). The second eigenvector is obtained from the first eigenvector by rotating it by -90 degrees:

[0033] Calculate the weighted matrix W = (D·diag(C3)) based on the mixing matrix D and the prediction matrix C3. Because C TTT It is a function of the MPEG surround sound prediction parameters c1 and c2 (as defined in ISO / IEC 23003-1:2007), so C TTT G = C3 can be rewritten as follows to find the stationary point or point of the function. Use Γ=(D) TTT C3)W(D TTT C3) * and b = GWC3v, where and v = (1 1-1). If Γ cannot provide a unique solution (det(Γ) < 10) -3 If the distance is too large, then select the point closest to the point through which the TTT (Time To Trip) occurs. As a first step, select γ = [γ...]. i,1 γ i,2 The element contains the row i with the highest energy Γ, therefore γ i,1 2 +γ i,2 2 ≥γ j,1 2 +γ j,2 2 j = 1, 2. Then determine a solution such that in

[0034] If obtained and The solution exceeds the limit. The permissible range of the prediction coefficient (as defined in ISO / IEC 23003-1:2007) is then... The calculation is as follows. First, limit the point set, x p for:

[0035]

[0036] and distance function,

[0037] Then, the prediction parameters are defined according to the following formula: The prediction parameters are constrained by the following conditions: Where λ, γ1, and γ2 are defined as follows:

[0038]

[0039] For the MPS decoder, CPC uses D CPC_1 =c1(l,m) and D CPC_2 =c2(l,m) is provided. The rendering parameters between the front and surround channels can be directly estimated from the target covariance matrix F.

[0040] Where (a,b) = (1,2) and (3,4).

[0041] For each OTT box h, the MPS parameters are in the form and Provided by China.

[0042] The stereo downmixer X was processed into a modified downmixer signal. Where G = D TTT C3 = D TTT M ren ED * J. Final stereo output from the SAOC code converter It is generated by mixing X with decorrelated signal components, according to: Among them, the relevant signal X d This is as described in this article and based on the following mixing matrix G. Mod It is calculated using P2.

[0043] First, the rendering upmixing error matrix is ​​limited to... Where A diff =D TTT A3-GD and in addition, the predicted signal The covariance matrix is ​​constrained to be

[0044] Gain vector g vec It can then be calculated as follows:

[0045]

[0046] And the mixing matrix G Mod will be given as

[0047] Similarly, the mixing matrix P2 is given as:

[0048] To export v R and Wd The characteristic equation that needs to be solved for R is: det(R-λ) 1,2 Given that I) = 0, the eigenvalues ​​λ1 and λ2 are given. The corresponding eigenvector v of R is... R1 and v R2 It can be calculated by solving a system of equations: (R-λ) 1,2 I)v R1,R2 =0. Sort the eigenvalues ​​in descending order (λ1≥λ2), and calculate the eigenvector corresponding to the larger eigenvalue according to the equation above. It must lie in the positive x-plane (the first element must be positive). The second eigenvector is obtained from the first eigenvector by rotating it by -90 degrees: Merge P1 = (1 1)G, R d Based on: The calculation gives And finally, it's a mixing matrix.

[0049] Go to relevant signal X d Created by the decorrelation described in ISO / IEC 23003-1:2007. Therefore, decorrFunc() represents the decorrelation process:

[0050] The SAOC code converter allows the mixing matrices P1, P2, and prediction matrix C3 to be calculated using an alternative scheme for the higher frequency range. This alternative scheme is particularly useful for undermixed signals where the upper frequency range is encoded by a non-waveform-preserving coding algorithm, such as SBR in efficient AAC. For the upper limit parameter band defined by bsTttBandsLow≤pb<numBands, P1, P2, and C3, the following alternative scheme should be used for calculation:

[0051] Accordingly, the target vector for mixing energy is defined under the specified energy level:

[0052] and help matrix

[0053] Then calculate the gain vector.

[0054] This ultimately yields a new prediction matrix.

[0055] For the SAOC system's decoder mode, the output signal of the downmixing preprocessing unit (represented in the mixing QMF domain) is fed into the corresponding synthesis filter bank, as described in ISO / IEC 23003-1:2007, to produce the final output PCM signal. Downmixing preprocessing includes mono, stereo, and subsequent binaural processing (if required).

[0056] Output signal The mono downmix signal X and the decorrelated mono downmix signal X d Calculated as Remove the relevant mono downmix signal X d Calculated as X d =decorrFunc(X). In the case of binaural output, the overmixing parameters G and P2, and rendering information are derived from the SAOC data. The Head-Related Transfer Function (HRTF) parameters are applied to the downmixed signal X (and X). d This generates binaural output. Target B-ear rendering matrix A l,m The size 2×N is determined by the elements Composition. Each element Both are derived from HRTF parameters and have elements Rendering matrix Exported. Target binaural rendering matrix A l,m This represents the relationship between all audio input objects y and the desired binaural output.

[0057]

[0058] The HRTF parameters for each processing band m are provided by and The spatial location of the HRTF parameters is given by index i. These parameters are described in ISO / IEC 23003-1:2007.

[0059] Upmixing parameter G l,m and Calculated as and

[0060] Gain of left and right output channels and They are respectively and Having elements The expected covariance matrix F l,m The size 2×2 is given as F l,m =A l,m E l,m (A l,m ) * Scalar v l,m Calculated as v l,m =D l E l,m (D l ) * +ε. Has elements The undermixing matrix Dl The size 1×N can be found to be

[0061] Having elements matrix E l,m From the following relationship Export. Inter-channel phase difference. Give as Interchannel coherence Calculated as Rotation angle α l,m and β l,m Give as

[0062]

[0063] In stereo output, the "x-1-b" processing mode can be applied without using HRTF information. This can be achieved by exporting all elements of the rendering matrix A. To complete, to produce: In mono output mode, the "x-1-2" processing mode can be applied to the following items:

[0064] In the stereo-to-bin "x-2-b" processing mode, the upmixing parameter G l,m and Calculated as

[0065] The corresponding gain of the left and right output channels and They are respectively

[0066] Having elements The expected covariance matrix F l,m,x The size 2×2 is given as F l,m,x =A l,m E l,m,x (A l ,m ) * Elements with binaural signals The covariance matrix C l,m The size 2×2 is estimated to be in

[0067]

[0068] The corresponding scalar v l,m,x and v l,m Calculated as v l,m,x =D l,x E l,m (D l,x ) *+ε,v l,m =(D l,1 +D l,2 E l,m (D l,1 +D l,2 ) * +ε.

[0069] Having elements The undermixing matrix D l,x The size 1×N can be found to be

[0070] Having elements Stereo submix matrix D l The size 2×N can be found to be

[0071] Having elements matrix E l,m,x Derived from the following relationship

[0072] Having elements matrix E l,m Give as Inter-channel phase difference Give as and The calculation method is as follows Rotation angle α l,m and β l,m Give as

[0073] In the case of stereo output, stereo preprocessing is applied directly as described above. In the case of mono output, the MPEG SAOC system applies stereo preprocessing using a single active rendering matrix entry.

[0074] The audio signal is defined for each time slot n and each mixing sub-band k. The corresponding SAOC parameters are defined for each parameter time slot l and processing band m. Table A.31 of ISO / IEC 23003-1:2007 specifies the subsequent mapping between the mixing domain and the parameter domain. Therefore, all calculations are performed relative to a specific time / band index, and a corresponding dimension is implicit for each introduced variable. The mixing process on OTN / TTN is determined by a predictive mode or M... Energy The energy pattern is represented by matrix M. In the first case, M is the product of two matrices utilizing the downmixing information and the CPC for each EAO channel. It is defined in the "parameter domain" by... It means that among them It is the inverse of the extended undermixed matrix. And C implies CPC. Extended undermixing matrix coefficient m j and n j The downmixing value of each EAOj in the right and left downmix channels is represented as m. j =d 1,EAO(j) n j =d 2,EAO(j) In stereo, the extended downmixing matrix yes

[0075]

[0076] For mono, it becomes

[0077]

[0078] For stereo downmixing, each EAOj holds two CPCcs. j,0 and c j,1 Output matrix C

[0079]

[0080] CPC is derived from the transmitted SAOC parameters, namely OLD, IOC, DMG, and DCLD. For a specific EAO channel, j = 0...N EAO -1CPC can be estimated using the following formula.

[0081]

[0082] In the following discussion of energy value P Lo P Ro P LoRo P LoCo,j and P RoCo,j As described in [the text].

[0083]

[0084] Parameter OLD L 、OLD R and IOC LR Corresponding to regular objects, and can be exported using undermix information:

[0085]

[0086] CPC is subject to subsequent constraint functions:

[0087]

[0088] Using weighting factors

[0089] The constrained CPC becomes

[0090] The output of TTN elements generates

[0091]

[0092] Where X represents the input signal of the SAOC decoder / code converter.

[0093] In stereo mode, the extended downmixing matrix A matrix is

[0094]

[0095] And for mono, it becomes

[0096] For mono downmixing, only one coefficient c is used. j Generate to predict an EAOj

[0097]

[0098] Based on the relationships provided above, all matrix elements c can be obtained from the SAOC parameters. j For mono downmixing, the output signal Y of the OTN element generates... In stereo mode, matrix M Energy Obtain from the corresponding OLD according to the following equation.

[0099]

[0100] The output of TTN elements generates

[0101]

[0102] The modification of the equation for a mono signal leads to

[0103]

[0104] The output of TTN elements generates The corresponding OTN matrix M in stereo mode Energy Can be exported as

[0105]

[0106] Therefore, the output signal Y of the OTN element generates Y=M Energy d0.

[0107] For the mono case, the OTN matrix M Energy Simplified to

[0108]

[0109] Julius O. Smith III, Physical Audio Signal Processing for Virtual Instruments and Audio Effects, Center for Computational Music and Acoustics Research (CCRMA), Department of Music, Stanford University, Stanford, California 94305, USA, December 2008 (Beta), considers the requirements of acoustically simulating concert halls or other listening spaces. It assumes we only need the response of one or more discrete listening points (“ears”) in the space, due to the acoustic energy of one or more discrete point sound sources. A single delay line in series with an attenuation scaling or low-pass filter can be used to simulate the direct signal propagating from the sound source to the listener's ear. Each ray of sound reaching the listening point via one or more reflections can be simulated using a delay line and some scaling factor (or filter). Two rays create a feedforward comb filter. More generally, a tapped delay line FIR filter can simulate many reflections. Each tap produces an echo with appropriate delay and gain, and each tap can be filtered independently to simulate air absorption and lossy reflections. In principle, tapped delay lines can accurately simulate any reverberant environment because reverberation is actually composed of many sound propagation paths from each sound source to each listening point. Compared to other techniques, tapped delay lines are computationally expensive and only handle a single “point-to-point” transfer function—from a point source to an ear—and are dependent on the physical environment. Typically, the filters should also include filtering through the auricle of the ear so that each echo can be perceived as coming from the correct angle of arrival in 3D space; in other words, at least some reverberant reflections should be spatialized so that they appear to come from their natural orientation in 3D space. Similarly, the filters will change if there are any changes in the listening space, including the location of the source or the listener. The basic architecture provides a set of signals, s1(n), s2(n), s3(n), ..., fed to a set of filters (h... 11 h 12 h 13 ), (h 21 h 22 h 23 ), ... and then add them together to form composite signals y1(n) and y2(n), representing the signals from both ears. Each filter h ij Both can be implemented as tapped delay line FIR filters. In the frequency domain, it is convenient to represent the input-output relationship using a transfer function matrix.

[0110]

[0111] By h ij (n) represents the impulse response of the filter from source j to ear i, and the two output signals are calculated by six convolutions:

[0112] Where M ij Indicates the FIR filter h ij The order of the filter. Due to the many filter coefficients h ij The fact that (n) is zero (at least for small n) makes implementing them as tapped delay lines more efficient, and makes the internals sparse. For greater accuracy, each tap can contain a low-pass filter that simulates air absorption and / or spherical diffusion losses. For large n, the impulse response is not sparse, and very expensive FIR filters must be used, or cheaper IIR filters must be used to approximate the tail of the impulse response.

[0113] For music, a typical reverberation time is about one second. Let's assume we choose exactly one second as the reverberation time. At an audio sampling rate of 50kHz, each filter requires 50,000 multiplications and additions per sample, or 2.5 billion multiplications and additions per second. Processing three sources and two listening points (ears), we reach 30 billion operations per second for the reverberator. While using FFT convolutions instead of direct convolutions can improve these figures (at the cost of introducing throughput latency, which can be problematic for real-time systems), an accurate implementation of all relevant point-to-point transfer functions in the reverberation space remains computationally very expensive.

[0114] Although tapped delay line FIR filters can provide an accurate model of any point-to-point transfer function in a reverberant environment, they are rarely used in practice for this purpose due to their extremely high computational cost. While there are dedicated commercial products that achieve reverberation via direct convolution of the input signal with the impulse response, the vast majority of artificial reverberation systems use other methods to synthesize post-reverberation more economically.

[0115] One drawback of point-to-point transfer function models is that some or all of the filters must be changed when anything moves. Conversely, if the computational model is of the entire acoustic space, then the sound source and listener can be moved as needed without affecting the room simulation below. Furthermore, we can use a “virtual simulation head” as the listener, equipped with an auricular filter, so that all 3D directions of reverberation can be captured in both extracted ear signals. Therefore, there are compelling reasons to consider a complete 3D model of the desired acoustic listening space. Let’s briefly estimate the computational requirements for a “powerful” acoustic simulation of a room. It is generally accepted that an audio signal requires a bandwidth of 20kHz. Since the speed of sound is approximately one foot per millisecond, a 20kHz sine wave has a wavelength of approximately 1 / 20 of a foot, or half an inch. Because, according to basic sampling theory, we must sample at a rate more than twice the highest frequency in the signal, we need a “grid point” spacing of no more than a quarter inch in our simulation. At this grid density, simulating a typical 12'×12'×8' home room would require over 100 million grid points. Using finite difference or waveguide mesh techniques, averaging the mesh points can be achieved as multiplication-free computations; however, because it involves waves moving back and forth in six spatial directions, each sample requires approximately 10 additions. Therefore, running such a room simulator at an audio sampling rate of 50 kHz would require 50 billion additions per second, which is comparable to simulating three sources and two ears.

[0116] Due to perceptual limitations, the impulse response of a reverberation chamber can be divided into two parts. The first part, called early reflections, consists of the relatively sparse first echoes in the impulse response. The remainder, called late reverberation, has very dense echoes and is best characterized statistically in some way. Similarly, the frequency response of a reverberation chamber can be divided into two parts. The low-frequency range consists of a relatively sparse distribution of resonant modes, while at higher frequencies, these modes are so dense that they are best characterized statistically as random frequency responses with certain (regular) statistical properties. Early reflections are a specific target of spatialization filters, ensuring that echoes originate from the correct direction in 3D space. Early reflections are well known to have a strong influence on spatial perception, i.e., the listener's perception of the shape of the listening space.

[0117] All poles of the lossless prototype reverberator lie on the unit circle in the z-plane, and its reverberation time is infinite. To set the reverberation time to the desired value, we need to slightly shift the poles within the unit circle. Furthermore, we want the high-frequency poles to be more damped than the low-frequency poles. This type of transformation can be achieved by replacing the z-plane. -1 ←G(z)z -1We obtain this, where G(z) represents the filtering of each sample in the propagation medium (a low-pass filter with a gain not exceeding 1 at all frequencies). Therefore, to set the reverberation time in a feedback delay network (FDN), we need to find G(z) that moves the poles to the desired locations and then design the low-pass filter. Place them at the output (or input) of each delay line. All pole radii in the reverb should vary smoothly with frequency.

[0118] Let t 60 (ω) represents the desired reverberation time at the radian frequency ω, and H is set to... i (z) represents the transfer function of a low-pass filter placed in series with delay line i. The problem we now consider is how to design these filters to produce the desired reverberation time. We will use H... i (z) The ideal amplitude response is specified based on the expected reverberation time at each frequency, and then a low-order approximation of this ideal specification is obtained using conventional filter design methods. Since replacement introduces losses z... -1 ←G(z)z -1 We need to find out its effect on the radius of the poles of the lossless prototype. Let Let z represent the i-th pole. (Recall that all poles of the lossless prototype lie on the unit circle). If the per-sample loss filter G(z) is zero-phase, then replace z. -1 ←G(z)z -1 This will only affect the radius of the poles, and not their angles. If the magnitude response of G(z) is close to 1 along the unit circle, then we obtain the i-th pole from... Move to Approximate value of, where

[0119] In other words, when z -1 Replace with G(z)z -1 When, where G(z) is zero phase, and |G(e^(z)| ... jω | Approaching (but less than) 1, frequency ω i A pole on the unit circle of radius is moved approximately along a radial line in the complex plane to a radius of . The point. The pole we expect is at a certain frequency ω. i The radius on the t is what we expect. 60 (ω i ): Therefore, the ideal single-sample filter G(z) satisfies

[0120] Therefore, with length M i A low-pass filter with series delay lines should be approximately This means Take 20log 10 Both sides gave

[0121] Now that we have specified the ideal delay line filter H i (e jωT Therefore, any number of filter design methods can be used to find a low-order H that provides a good approximation. i (z). The example includes the Matlab functions invfreqz and stmcb. Since the change in reverberation time is typically very smooth with respect to ω, the filter H... i (z) can be of a very low order.

[0122] Early reflections should be spatialized by including a head-related transfer function (HRTF) at each tap of the early reflection delay line. Later reverberation may also require some form of spatialization. The true... Diffuse Field It consists of the sum of plane waves propagating in all directions in 3D space. Spatialization can also be applied to later reflections, although the implementation differs because these are handled statistically.

[0123] See also US10,499,153; 9,361,896; 9,173,032; 9,042,565; 8,880,413; 7,792,674; 7,532,734; 7,379,961; 7,167,566; 6,961,439; 6,694,033; 6,668,061; 6,442,277; 6,185,152; 6,009,396; 5,943,427; 5,987,142; 5,841,879; 5,661,812; 5,465,302; 5,459,790; 5,272,757; 20010031051; 20 020150254; 20020196947; 20030059070; 20040141622; 20040223620; 20050114121; 20050135643; 20050271212; 20060045275; 20060056639; 20070109977; 20070286427; 20070294061; 20080004866; 20080025534; 20080137870; 20080144794; 20080304670; 20080306720; 20090046864; 20 090060236; 20090067636; 20090116652; 20090232317; 20090292544; 20100183159; 20100198601; 20100241439; 20100296678; 20100305952; 20110009771; 20110268281; 20110299707; 20120093348; 20120121113; 20120162362; 20120213375; 20120314878; 20130046790; 20130163766; 20 140016793;20140064526;20150036827;20150131824;20160014540;20 160050508;20170070835;20170215018;20170318407;20180091921;20 180217804; 20180288554; 20180288554; 20190045317; 20190116448; 20 190132674;20190166426;20190268711;20190289417;20190320282;WO 00 / 19415; WO 99 / 49574; and WO 97 / 30566.

[0124] Naef, Martin, Oliver Staadt, and Markus Gross. “Spatialized Audio Rendering for Immersive Virtual Environments.” Proceedings of the ACM Symposium on Virtual Reality Software and Technology, pp. 65–72. ACM, 2002, disclosed feedback from graphics processing units for spatialized audio signal processing. Lauterbach, Christian, Anish Chandak, and Dinesh Manocha. “Interactive Sound Rendering in Complex and Dynamic Scenes Using Frustum Tracking.” IEEE Transactions on Visualization and Computer Graphics 13, No. 6 (2007): 1672–1679 also uses graphic style analysis for audio processing. Murphy, David, and Flaithrí Neff. “Spatial Sound in Computer Games and Virtual Reality.” Game Sound Technology and Player Interaction: Concepts and Developments, pp. 287–312. IGI Global, 2011, discusses spatialized audio in computer games and VR environments. Begault, Durand R., and Leonard J. Trejo. “3D Sound in Virtual Reality and Multimedia.” (2000), NASA / TM-2000-209606 discusses various implementations of spatial audio systems. See also Begault, Durand, Elizabeth M. Wenzel, Martine Godfroy, Joel D. Miller, and Mark R. Anderson. “Applying Spatial Audio to Human-Machine Interfaces: 25 Years of NASA Experience.” Audio Engineering Society Conference: 40th International Conference: Spatial Audio: Sound in Perceiving Space. Audio Engineering Society, 2010.

[0125] Herder and Jens, “Optimizing the Spatialization Resource Management of Sound by Clustering”, Journal of 3D Imaging, 3D Forum Association, Vol. 13, No. 3, pp. 59-65, 1999, which discusses algorithms for simplifying spatial audio processing.

[0126] Verron, Charles, Mitsuko Aramaki, Richard Kronland-Martinet, and Grégory Palone, “3-D Immersive Synthesizers for Ambient Sounds,” IEEE Transactions on Audio, Speech and Language Processing 18, Vol. 6 (2009): 1550-1561, discusses spatialized sound synthesis.

[0127] Malham, David G, and Anthony Myatt. “3-D Sound Spatialization Using Two-Channel Stereo Technology,” Computer Music Journal 19, No. 4 (1995): 58-70 discusses the use of surround sound technology (the use of 3D sound fields). See also Hollerweger and Florian, “Spatialization of Surround Sound in Multi-User Virtual Environments,” PhD dissertation, Institute for Electronic Music and Acoustics (IEM), Center for Electronic Arts Technology (CREATE), 2006.

[0128] McGee, Ryan, and Matthew Wright, “Sound Element Spatializer,” ICMC. 2011z.; and McGee, Ryan, “Sound Element Spatializer” (MSThesis, U.California Santa Barbara 2010), introduced the Sound Element Spatializer (SES), a novel system for rendering and controlling spatial audio. SES offers a variety of 3D sound rendering techniques and allows for arbitrary speaker configurations utilizing any number of moving sound sources.

[0129] Transear audio processing is discussed below:

[0130] Baskind, Alexis, Thibaut Carpentier, Markus Noisternig, Olivier Warusfel, and Jean-Marc Lyzwa, “Binaural and Transaural Spatialization Techniques in Multichannel 5.1 Production” (Application of Binaural and Transaural Playback Techniques in 5.1 Music Production), 27th Tonmeist Ertigung-VDT International Conference, November 2012

[0131] Bosun, Xie, Liu Lulu, and Chengyun Zhang. Transaural reproduction of spatial surround sound using four actual loudspeakers. In Proceedings of the International Conference on Noise and Noise, Vol. 259, No. 9, pp. 61-69. Institute of Noise Control Engineering, 2019.

[0132] Casey, Michael A., William G. Gardner, and Sumit Basu. “Visually controlled beamforming and cross-auditory rendering for artificial life interactive video environments (alive).” Audio Engineering Society Congress 99. Audio Engineering Society, 1995.

[0133] Cooper, Duane H., and Jerald L. Bauck. Prospects for transear recording. Journal of the Audio Engineering Society, 37, Vol. 1 / 2 (1989): 3-19.

[0134] Fazi, Filippo Maria, and Eric Hamdan. “Stage Compression in Transear Audio.” Audio Engineering Society Convention No. 144. Audio Engineering Society, 2018.

[0135] Gardner and William Grant. Transaural 3D Audio. Perceptual Computing Division, MIT Media Lab, 1995.

[0136] Glasal, Ralph, Two-Channel Stereo: Replacing Stereo with Concert Hall Realism, 2nd Edition (2015).

[0137] Greff "The Use of Parametric Arrays in Transear Applications." Proceedings of the 20th International Congress of Acoustics, pp. 1-5. 2010.

[0138] Guastavino, Catherine, Véronique Larcher, Guillaume Catusseau, and Patrick Boussard. “Spatial Audio Quality Assessment: Comparing Transaural, Surround, and Stereo Sound.” Georgia Institute of Technology, 2007.

[0139] Guldenschuh, Markus, and Alois Sontacchi. "Applications of Transear Focusing Sound Reproduction." Presented at the 6th Euroconductivity Conference (INO) in 2009.

[0140] Guldenschuh, Markus, and Alois Sontacchi. "Transauronic Stereo in Beamforming Methods." In Proc. DAFx, Vol. 9, pp. 1-6. 2009.

[0141] Guldenschuh, Markus, Chris Shaw, and Alois Sontacchi. “Evaluation of transauricular beamformers.” 27th General Assembly of the International Council for Aeronautical Sciences (ICAS2010). Nizza, Frankreich, pp. 2010-10.

[0142] Guldenschuh and Markus. “Transauricular Beamforming.” PhD dissertation, Master’s thesis, Graz University of Technology, Graz, Austria, 2009.

[0143] Hartmann, William M., Brad Rakerd, Zane D. Crawford, and Peter Xinya Zhang. "Transear experiments and a modified duplex theory for low-frequency pitch localization." Proceedings of the Acoustical Society of America, 139, Vol. 2 (2016): 968-985.

[0144] Ito, Yu, and Yoichi Haneda. “Study on beamforming transear systems using circular loudspeaker arrays.” Proc. 23rd Int. Cong. Acoustics (2019).

[0145] Johannes, Reuben, and Woon-Seng Gan. “3D Sound with Transear Audio Beam Projection.” 10th Western Pacific Conference on Acoustics, Beijing, China, Paper, Vol. 244, No. 8, pp. 21-23. 2009.

[0146] Jost, Adrian, and Jean-Marc Jot. “Transaural 3D Audio with User-Controlled Calibration.” Proceedings of the COST-G6 Digital Audio Effects Conference, DAFX2000, Verona, Italy, 2000.

[0147] Kaiser, Fabio. "Transaural Audio - Reproducing Binaural Signals via Speakers." PhD dissertation, Graz University of Music and Art / College of Music and Art / IRCAM, March 2011.

[0148] LIU, Lulu, and Bosun XIE. "Limitations of static transaural reproduction by two front speakers." (2019)

[0150] Méaux, Eric, and Sylvain Marchand. “Synthetic Transear Audio Rendering (STAR): A Perceptual Approach to Sound Spatialization.” 2019.

[0151] Samejima, Toshiya, YoSasaki, Izumi Taniguchi, and Hiroyuki Kitajima. “A robust transear sound reproduction system based on feedback control.” Acoustics Science and Technology, Vol. 31, No. 4 (2010): 251-259.

[0152] Simon Galvez, Marcos F., and Filippo Maria Fazi. “Speaker arrays for transear reproduction” (2015).

[0153] Simon Gálvez, Marcos Felipe, Miguel Blanco Galindo, and Filippo Maria Fazi. “Study on reflection and reverberation effects in low-channel-number transauricular systems.” In Proceedings of the International Conference on Noise and Noise, Vol. 259, No. 3, pp. 6111-6122. Institute of Noise Control Engineering, 2019.

[0154] Villegas, Julián, and Takaya Ninagawa. “A transear filter based on pure data with range control”. (2016)

[0155] en.wikipedia.org / wiki / Perceptual-based_3D_sound_localization

[0156] Duraiswami, Grant, Mesgarani, Shamma, Enhanced intelligibility in multilingual environments. Proceedings of the International Conference on Auditory Displays (ICAD'03), 2003.

[0157] Shohei Nagai, Shunichi Kasahara, JunRe kimot, Proceedings of the 6th International Conference on Augmented Humans, “Directed Communication Using Spatial Sound in Human-Telepresentation”, Singapore 2015, ACM New York, USA, ISBN: 978-1-4503-3349-8.

[0158] Siu-Lan Tan, Annabel J. Cohen, Scott D. Lipscomb, Roger A. Kendall, "Multimedia Music Psychology", Oxford University Press, 2013. Summary of the Invention

[0159] In one aspect of the invention, a system and method for three-dimensional (3-D) audio technology are provided to create complex immersive auditory scenes that immerse the listener using a sparse linear (or curved) array of acoustic transducers. A sparse array is an array with discontinuous intervals relative to an idealized channel model (e.g., four or fewer sound emitters), where the sound emitted from the transducers is internally modeled in a higher dimension and then reduced or superimposed. In some cases, the number of sound emitters is four or more, derived from a larger number of channels in the channel model, such as more than eight.

[0160] The three-dimensional sound field is modeled based on mathematical and physical constraints. The system and method provide multiple loudspeakers, i.e., free-field sound transmission transducers that emit sound into the space containing the ears of the target listener. These systems are controlled in real time by a sophisticated multi-channel algorithm.

[0161] The system may assume a fixed relationship between the sparse speaker array and the listener's ears, or it may employ a feedback system to track the movement and position of the listener's ears or head.

[0162] The algorithm employed provides highly localized audio through a loudspeaker array, thereby offering surround sound imaging and sound field control. Typically, the loudspeakers in a sparse array seek to operate in a wide-angle scattering mode of emission, rather than a more traditional “beaming mode,” in which each transducer emits a narrow-angle sound field toward the listener. That is, the transducer emission mode is wide enough to avoid spatial pauses in the sound.

[0163] In some cases, the system supports multiple listeners in the environment, although in such cases either an enhanced stereo operating mode or head tracking is employed. For example, when two listeners are in the environment, nominally identical signals are provided to each listener's left and right ears, regardless of their orientation in the room. In a significant implementation, this requires multiple transducers to cooperate in canceling left-ear emissions at each listener's right ear and right-ear emissions at each listener's left ear. However, a trial-and-error approach can be used to reduce the requirement for at least one pair of transducers per listener.

[0164] Typically, spatial audio is normalized not only for transear audio amplitude control but also for group delay to ensure that the correct sound is perceived at the correct time in each ear. Therefore, in some cases, the signal may represent a trade-off between fine amplitude and delay control.

[0165] Therefore, the source content can be virtually rotated to various angles, allowing different dynamically changing sound fields to be generated for different listeners based on their positions.

[0166] A signal processing method is provided for delivering spatialized sound in various ways using deconvolution filters to deliver discrete left / right ear audio signals from a speaker array. The method can be used to provide private listening areas in public spaces, address multiple listeners with discrete sound sources, provide spatialization of source material for a single listener (virtual surround sound), and enhance the intelligibility of conversations in noisy environments using spatial cues, to name just a few applications.

[0167] In some cases, microphones or microphone arrays can be used to provide feedback on sound conditions at voxels in a space, such as at or near the listener's ears. While initially appearing equivalent to headphones, where one could simply use a single transducer for each ear, this technology does not force the listener to wear headphones, and the result is more natural. Furthermore, microphones (one or more) can be used to initially understand room conditions, which are then no longer needed, or selectively used only for a portion of the environment. Finally, microphones can be used to provide interactive voice communication.

[0168] In binaural mode, the speaker array generates two emitted signals, typically aimed at the primary listener's ears, one discrete beam for each ear. The shapes of these beams are designed using convolution or inverse filtering methods so that the beam from one ear contributes almost no energy to the listener's other ear. This provides a convincing virtual surround sound via binaural signals. In this mode, binaural sources can be accurately rendered without headphones. A virtual surround sound experience is delivered without the need for physical discrete surround speakers. Note that in a real environment, echoes from walls and surfaces colorize and delay sound, and natural sound emissions provide these environment-related cues. While the human ear has some ability to distinguish between sounds coming from the front and back due to the shape of the ear and head, the key characteristic of most source material is time and sound coloration. Therefore, the activity of the environment can be simulated by delay filters in processing, emitting delayed sound from the same array that has a beam pattern substantially the same as the primary sound signal.

[0169] In one aspect, a method for generating binaural sound from a speaker array is provided, wherein multiple audio signals are received from multiple sources, and each audio signal is filtered using a head-related transfer function (HRTF) based on the listener's position and orientation relative to the transmitter array. The filtered audio signals are combined to form a binaural signal. In a sparse transducer array, it may be desirable to provide cross-signaling between the respective binaural channels, although cross-signaling may not be necessary if the array's directivity is sufficient to provide physical isolation of the listener's ears and the listener's position relative to the array is well defined and constrained. Typically, the audio signals are processed to provide crosstalk cancellation.

[0170] When the source signal is pre-recorded music or other processed audio, the initial processing may optionally remove processing effects that attempt to isolate the original object and its corresponding sound emissions, making the spatialization accurate for the sound field. In some cases, the spatial location inferred from the source is artificial; that is, the object's location is defined as part of the production process and does not represent its actual location. In such cases, spatialization may extend back to the original source and seek to (re)optimize the process, since the original product may not have been optimized for reproduction by the spatialization system.

[0171] In a sparse linear loudspeaker array, filtered / processed signals for multiple virtual channels are processed separately, and then combined (e.g., summed) for each corresponding virtual loudspeaker into a single loudspeaker signal. The loudspeaker signal is then fed to the corresponding loudspeaker in the loudspeaker array and transmitted to the listener through the corresponding loudspeaker.

[0172] The summation process can correct the timing alignment of the corresponding signals. That is, the original complete array signal has a time delay for the corresponding signal in each ear. When summed without compensation, in order to produce a composite signal, the signal will contain multiple incremental time-delay representations arriving at the ears at different times, representing the same point in time. Therefore, spatial compression leads to temporal expansion. However, since the time delays are programmed according to an algorithm, timing alignment can be restored through algorithmic compression.

[0173] The result is that spatialized sound has accurate timing, phase alignment, and spatialized sound complexity to reach each ear.

[0174] In another aspect, a method is provided for receiving at least one audio signal, filtering each audio signal through a set of spatialized filters (each input audio signal is filtered through different sets of spatialized filters, which may be interactive or ultimately combined), wherein a separate spatialized filter path segment is provided for each speaker in a speaker array, such that each input audio signal is filtered through different spatialized filter segments, the filtered audio signals of each corresponding speaker are summed to a speaker signal, each speaker signal is transmitted to a corresponding speaker in the speaker array, and the signal is delivered to one or more areas of space (typically occupied by one or more listeners accordingly).

[0175] In this way, the complexity of the acoustic signal processing path is simplified to a set of parallel stages using combiners to represent array positions. An alternative approach for providing spatialized audio from two loudspeakers offers an object-based processing algorithm whose beam follows the audio path between corresponding sources, leaving the scattering object and reaching the listener's ear. The latter approach offers more arbitrary algorithmic complexity and lower consistency across processing paths.

[0176] In some cases, filters can be implemented as recurrent neural networks or deep neural networks, which typically simulate the same spatialized process but without explicit discrete mathematical functions, and seek the best overall effect rather than optimizing each effect serially or in parallel. The network can be a holistic network that receives sound input and produces sound output, or a channelized system where each channel can represent space, band, delay, source object, etc., processed using different networks, and the network outputs are combined. Furthermore, neural networks or other statistically optimized networks can provide coefficients for general signal processing chains, such as digital filters, which can have finite impulse response (FIR) and / or infinite impulse response (IIR) characteristics, leakage paths to other channels, dedicated time and delay equalizers (where direct implementation via FIR or IIR filters is undesirable or inconvenient).

[0177] More typically, audio data is processed using discrete digital signal processing algorithms based on physical (or virtual) parameters. In some cases, the algorithm can be adaptive, based on automatic or manual feedback. For example, the microphone might detect distortions due to resonance or other effects that are not inherently compensated for in the basic algorithm. Similarly, a general HRTF can be used, which adjusts based on the actual parameters of the listener's head.

[0178] In another aspect, a loudspeaker array system for generating localized sound includes: an input receiving multiple audio signals from at least one source; a computer having a processor and memory that determines whether the multiple audio signals should be processed by an audio signal processing system; a loudspeaker array including multiple loudspeakers; wherein the audio signal processing system includes: at least one head correlation transfer function (HRTF) that senses or estimates the spatial relationship between a listener and the loudspeaker array; and a combiner configured to combine multiple processing channels to form a loudspeaker drive signal. The audio signal processing system implements a spatialization filter; wherein the loudspeaker array delivers corresponding loudspeaker signals (or beamforming loudspeaker signals) to one or more listeners via multiple loudspeakers.

[0179] Beamforming means that the transducer's emission is not omnidirectional or cardioid, but has an emission axis, with a spacing between the left and right ears greater than 3dB, preferably greater than 6dB, more preferably greater than 10dB, and even higher spacing can be achieved by utilizing active cancellation between transducers.

[0180] Multiple audio signals can be processed by a digital signal processing system that incorporates binauralization before being delivered to one or more listeners via multiple speakers.

[0181] A listener head tracking unit may be provided, which adjusts the binaural processing system and the acoustic processing system based on changes in the position of one or more listeners.

[0182] The binaural processing system may also include a binaural processor, which calculates the left HRTF and right HRTF in real time, or synthesizes the HRTF.

[0183] The method of this invention employs an algorithm that allows beam delivery by using deconvolution or inverse filters and physical or virtual beamforming, the beams being configured to generate binaural sound—sound for each ear—without headphones. In this way, a virtual surround sound experience can be delivered to the system's listener. The system avoids the use of classic two-channel "crosstalk cancellation" to provide superior speaker-based binaural sound imaging.

[0184] Binaural 3D sound reproduction is a sound pre-production achieved through headphones. On the other hand, transaural 3D sound reproduction is a sound pre-production achieved through speakers. See, Kaiser, Fabio. “Transaural Audio – Reproduction of Binaural Signals through Speakers.” PhD dissertation, Graz University of Music and Arts / Academy of Music and Arts / IRCAM, March 2011. Transaural audio is a three-dimensional sound spatialization technique capable of reproducing binaural signals through speakers. It is based on eliminating the sound path between the speaker and the listener's ears.

[0185] Psychoacoustic research has shown that well-recorded stereo signals and binaural recordings contain cues that help create robust, detailed 3D auditory images. One implementation of 3D spatialized audio, called "MyBeam" (Comhear, San Diego, California), preserves key psychoacoustic cues by focusing the left and right channel signals onto the appropriate ears, while avoiding crosstalk through precise beamforming directionality.

[0186] In summary, these cues are known as the Head Related Transfer Function (HRTF). In short, the HRTF component cues are the interaural time difference (ITD, the time difference of sound arrival between two locations), the interaural intensity difference (IID, the difference in sound intensity between two locations, sometimes called ILD), and the interaural phase difference (IPD, the phase difference of the wave arriving at each ear, depending on the frequency of the sound wave and the ITD). Once the listener's brain analyzes the IPD, ITD, and ILD, it can determine the location of the sound source with relatively high accuracy.

[0187] This invention provides a method for optimizing beamforming and controlling small linear loudspeaker arrays to generate spatial, localized, and binaural or transaural virtual surround or 3D sound. The signal processing method allows the small loudspeaker array to deliver sound in various ways using highly optimized inverse filters, delivering a narrow sound beam to the listener while producing negligible artifacts. Unlike earlier compact beamforming audio techniques, this method does not rely on ultrasound or high-power amplification. The technique can be implemented using low-power techniques, generating 98 dB SPL at one meter while utilizing approximately 20 watts of peak power. In loudspeaker applications, the primary use case allows sound from small (10-inch–20-inch) linear loudspeaker arrays to be focused with a narrow beam:

[0188] • Guide the voice in a highly understandable way where it is needed and effective;

[0189] Limit sound in unwanted or potentially disruptive locations.

[0190] • Provides high-definition, controllable audio imaging based on non-headphones, where stereo or binaural signals are directed to the listener's ears to produce a vivid 3D listening experience.

[0191] In the case of microphone applications, the basic use case allows sound from a microphone array (ranging from a few small diaphragms to dozens of 1D, 2D, or 3D arrangements) to be captured in narrow beams. These beams can be dynamically manipulated and can cover many speakers and sound sources within their coverage pattern, amplifying desired sources and providing cancellation or suppression of unwanted sources.

[0192] In multipoint teleconference or video conferencing applications, the technology allows for distinct spatialization and positioning of each participant in the meeting, providing a significant improvement over existing technologies where the sound from each speaker spatially overlaps. This overlap makes it difficult to distinguish between different participants without requiring each participant to identify themselves each time they speak, which detracts from the natural feel of face-to-face conversation. Furthermore, the invention can be extended to use video analytics or motion sensors to provide real-time beam control and tracking of the listener's position, thus continuously optimizing binaural or spatialized audio delivery as the listener moves within a room or in front of the speaker array.

[0193] The system is likely smaller and more portable than most (if not all) similar speaker systems. Therefore, it can be used not only in fixed structural installations, such as in rooms or virtual reality caves, but also in private vehicles such as cars, public transportation such as buses, trains, and airplanes, and open areas such as office cubicles and open-plan classrooms.

[0194] The technology is relative to MyBeam TM This is an improvement because it offers similar applications and advantages while requiring fewer speakers and amplifiers. For example, the method virtualizes a 12-channel beamforming array as two channels. Typically, the algorithm downmixes each pair of six channels (designed to drive a set of six equally spaced speakers) into a single speaker signal for a speaker mounted in the middle of those six speakers. Typically, the virtual line array is 12 speakers, with two real speakers located between elements 3-4 and 9-10.

[0195] The real speakers are mounted directly at the center of each group of 6 virtual speakers. If (s) is the center-to-center distance between the speakers, then the distance from the center of the array to the center of each real speaker is: A = 3*s.

[0196] The left speaker is offset by -A from the center, and the right speaker is offset by A.

[0197] The main algorithm is to simply downmix the six virtual channels, applying limiters and / or compressors to prevent saturation or clipping. For example, the left channel is:

[0198] L 输出 = Limitation (L1+L2+L3+L4+L5+L6).

[0199] However, due to variations in the location of the audio source, the delay between speakers needs to be considered, as described below. In some cases, the phase of some drivers can be altered to limit peaking while avoiding clipping or limiting distortion.

[0200] Because six speakers are combined into one at different locations, the variation in propagation distance, i.e., the delay for the listener, can be significant, especially at higher frequencies. The delay can be calculated based on the variation in travel distance between the virtual and physical speakers.

[0201] In this discussion, we will only focus on the left side of the array. The right side is similar, but reversed.

[0202] To calculate the distance from the listener to each virtual speaker, assume the speakers n are numbered from 1 to 6, where 1 is the speaker closest to the center and 6 is the leftmost. The distance from the center of the array to the speaker is: d = ((n-1) + 0.5) * s.

[0203] Using the Pythagorean theorem, the distance from the speaker to the listener can be calculated as follows:

[0204]

[0205] The distance from the actual speaker to the listener is

[0206] The sample delay of each speaker can be calculated using the distance difference between two listeners. This can be converted into samples (assuming a sound speed of 343 m / s and a sampling rate of 48 kHz).

[0207]

[0208] This can result in significant delays between listener distances. For example, if the speaker-to-speaker distance is 38mm and the listener distance from the array is 500mm, the delay from the virtual leftmost speaker (n=6) to the actual speakers would be:

[0209]

[0210] Although the delay may seem small, the amount of delay is significant, especially at higher frequencies, where there may only be 3 or 4 samples in the entire cycle.

[0211] Table 1

[0212] speaker Delay relative to a real speaker 1 -2 2 -1 3 -1 4 1 5 2 6 4

[0213] Therefore, when combining the signals used for the virtual speaker into the physical speaker signal, time offset is preferably compensated based on the displacement of the virtual speaker relative to the physical speaker. This can be done at different points in the signal processing chain.

[0214] Therefore, this technology provides downmixing of spatialized audio virtual channels to maintain the delay coding of the virtual channels while minimizing the number of physical drivers and amplifiers required.

[0215] With similar sound output, the power of each speaker will naturally increase with downmixing, leading to peak power handling limitations. Assuming the amplitude, phase, and delay of each virtual channel are important information, the ability to control peaking is limited. However, given that clipping or limiting is particularly incongruous, controlling other variables helps achieve high power ratings. Control can be facilitated by manipulating the delay; for example, in a speaker system with a lower 30Hz range, a 125ms delay can be applied to allow for the calculation of all significant echo and peak-clipping mitigation strategies. Such delay can be reduced when video content is also being rendered. However, delay is not required.

[0216] In some cases, the listener is not centered relative to the physical speaker transducer, or multiple listeners are scattered throughout the environment. Furthermore, the peak power of the physical transducer caused by the proposed downmixing may exceed limitations. In such and other cases, the downmixing algorithm can be adaptive or flexible, and can provide different mappings from the virtual transducer to the physical speaker transducer.

[0217] For example, the distribution of virtual transducers to physical speaker transducers in a virtual array may be unbalanced due to listener position or peak levels. For instance, in an array of 12 virtual transducers, 7 virtual transducers might be used for downmixing the left physical transducer, and 5 virtual transducers for the right physical transducer. This has the effect of shifting the sound axis and also moves additional effects of adaptively allocated transducers to another channel. If a transducer is out of phase relative to other transducers, peaks will be canceled out, while if they are in phase, constructive interference will occur.

[0218] Reassignment can be a virtual transducer at the boundary between groups, or it can be a discontinuous virtual transducer. Similarly, adaptive assignment can be more than one virtual transducer.

[0219] Furthermore, the number of physical transducers can be an even or odd number greater than 2, and is typically less than the number of virtual transducers. In the case of three physical transducers typically located on the nominal left, middle, and right, the allocation between virtual and physical transducers can be adaptive in terms of group size, group transitions, group continuity, and the possibility of group overlap (i.e., portions of the same virtual transducer signal are represented in multiple physical channels) based on the listener's (or multiple listener's) location, spatialization effects, peak amplitude attenuation issues, and listener preferences.

[0220] The system can employ various techniques to achieve optimal HRTF. In the simplest case, the optimal prototype HRTF is used regardless of the listener and environment. In other cases, the characteristics of one or more listeners are determined by login, direct input, camera, biometric measurements, or other means, as well as a customized or selected HRTF chosen or calculated for one or more specific listeners. This is typically achieved during the filtering process, independent of the downmixing process, but in some cases, customization can be implemented as a post-processing or partial post-processing of spatial filtering. That is, in addition to downmixing, processes following the main spatial filtering and virtual transducer signal creation can be implemented to adapt or modify the signal according to one or more listeners, the environment, or other factors, separate from downmixing and timing adjustments.

[0221] As mentioned above, limiting the peak amplitude is potentially important because a set of dummy transducer signals (e.g., six), time-aligned and summed, could result in a peak amplitude potentially six times higher than the peak of any single dummy transducer signal. One approach to this problem is to simply limit the combined signal or use a compander (nonlinear amplitude filter). However, these introduce distortion and interfere with spatialization effects. Other options involve some phase shift of the dummy sensor signals, but this can also lead to auditory artifacts and requires a delay. Another option provided is to assign dummy transducers (especially those near the transition between groups) to the downmixing group based on phase and amplitude. While this can also be achieved with a delay, it is also possible to shift the group assignment almost instantaneously, which could result in position artifacts rather than harmonic distortion artifacts. These techniques can also be combined to minimize perceived distortion by spreading the effect across various peak reduction options.

[0222] Therefore, the aim is to provide a method for generating transaural spatialized sound, comprising: receiving audio signals representing spatial audio objects; filtering each audio signal using a spatialization filter to generate a virtual audio transducer signal array representing a virtual audio transducer array of spatialized audio; separating the virtual audio transducer signal array into subsets, each subset comprising multiple virtual audio transducer signals, each subset being used to drive physical audio transducers located within a physical location range of the respective subset; time-shifting the respective virtual audio transducer signals of the respective subset based on the time difference of arrival of sound from the nominal position of the respective virtual audio transducer and the physical position of the physical audio transducer relative to the target ear of a listener; and combining the time-shifted respective virtual speaker signals of the respective subsets into physical audio transducer drive signals.

[0223] Another objective is to provide a system for generating transaural spatialized sound, comprising: an input configured to receive an audio signal representing a spatial audio object; a spatialized audio data filter configured to process each audio signal to generate a virtual audio transducer signal array representing a virtual audio transducer array of spatialized audio, the virtual audio transducer signal array being divided into subsets, each subset including multiple virtual audio transducer signals, each subset being used to drive physical audio transducers located within a physical location range of the respective subset; a time delay processor configured to time-shift the respective virtual audio transducer signals of the respective subset based on the time difference of arrival of sound from the nominal position of the respective virtual audio transducer and the physical position of the corresponding physical audio transducer relative to the target ear of a listener; and a combiner configured to combine the time-shifted respective virtual speaker signals of the respective subsets into physical audio transducer drive signals.

[0224] Another object is to provide a system for generating spatialized sound, comprising: an input configured to receive an audio signal representing a spatial audio object; at least one automation processor configured to: process each audio signal through a spatialization filter to generate a virtual audio transducer signal array representing a virtual audio transducer array of spatialized audio, the virtual audio transducer signal array being divided into subsets, each subset including a plurality of virtual audio transducer signals, each subset being used to drive physical audio transducers located within a physical location range of the respective subset; time-shifting the respective virtual audio transducer signals of the respective subset based on the time difference of arrival of sound from the nominal position of the respective virtual audio transducer and the physical position of the corresponding physical audio transducer relative to the target ear of a listener; and combining the time-shifted respective virtual speaker signals of the respective subset into a physical audio transducer drive signal; and at least one output port configured to present the physical audio transducer drive signal of the respective subset.

[0225] The method may further include reducing the peak amplitude of the corresponding virtual audio transducer signal of the combined time offset to reduce saturation distortion of the physical audio transducer.

[0226] Filtering may include processing at least two audio channels using a digital signal processor. Filtering may also include using a graphics processing unit configured to act as an audio signal processor to process at least two audio channels.

[0227] The virtual audio transducer signal array can be a linear array of 12 virtual audio transducers. The virtual audio transducer array can be a linear array with at least three times the number of virtual audio transducer signals as driving signals for the physical audio transducers. The virtual audio transducer array can be a linear array with at least six times the number of virtual audio transducer signals as driving signals for the physical audio transducers.

[0228] Each subset can be a non-overlapping adjacent group of virtual audio transducer signals. Each subset can be at least six non-overlapping adjacent groups of virtual audio transducer signals. Each subset can have a virtual audio transducer whose location overlaps with the represented location range of another subset of virtual audio transducer signals. The overlap can be a single virtual audio transducer signal.

[0229] The virtual audio transducer signal array can be a linear array with 12 virtual audio transducer signals, divided into two non-overlapping groups. Each group has 6 adjacent virtual audio transducer signals, which are combined accordingly to form 2 physical audio transducer drive signals. The corresponding physical audio transducer in each group can be located between the 3rd and 4th virtual audio transducers in the adjacent groups of 6 virtual audio transducer signals.

[0230] Physical audio transducers can have non-directional emission modes. Virtual audio transducer arrays can be modeled for directivity. Virtual audio transducer arrays can be phased arrays of audio transducers.

[0231] Filtering can include crosstalk cancellation. A reentrant data filter can be used for filtering.

[0232] The method may further include receiving a signal representing the position of the listener's ear. The method may also include tracking the listener's movement and adjusting the filter based on the tracked movement.

[0233] The method may also include adaptively allocating virtual audio transducer signals to appropriate subsets.

[0234] The method may further include adaptively determining the listener's head-related transfer function and filtering based on the adaptively determined head-related transfer function.

[0235] The method may further include sensing features of the listener's head and adjusting the head-related transfer function based on those features.

[0236] Filtering can include time-domain filtering or frequency-domain filtering.

[0237] The physical audio transducer drive signal can be delayed by at least 25ms relative to the received audio signal representing the spatial audio object.

[0238] The system may also include a peak amplitude reduction filter, limiter, or compander configured to reduce saturation distortion of the physical audio transducer of the combined time-off virtual audio transducer signal.

[0239] The system may also include a phase rotator configured to rotate the relative phase of at least one virtual audio transducer signal.

[0240] A spatialized audio data filter may include a digital signal processor configured to process at least two audio channels. The spatialized audio data filter may also include a graphics processing unit configured to process at least two audio channels.

[0241] Spatialized audio data filters can be configured to perform crosstalk cancellation. Spatialized audio data filters may include reentrant data filters.

[0242] The system may also include an input port configured to receive a signal indicating the listener's ear position.

[0243] The system may also include an input configured to receive a signal that tracks the movement of a listener, wherein the spatialized audio data filter is adaptively dependent on the tracked movement.

[0244] The virtual audio transducer signal can be adaptively assigned to the appropriate subset.

[0245] Spatialized audio data filters can rely on a head-related transfer function that is adaptively determined by the listener.

[0246] The system may also include an input port configured to receive signals including sensed features of the listener’s head, wherein a head-related transfer function is adjusted based on the features.

[0247] Spatialized audio data filters may include time-domain filters and / or frequency-domain filters. Attached Figure Description

[0248] Figure 1A This is a diagram illustrating the operation of the Wave Field Synthesis (WFS) mode for private listening.

[0249] Figure 1B This diagram illustrates the use of WFS mode in multi-user, multi-location audio applications.

[0250] Figure 2 This is a block diagram showing the WFS signal processing chain.

[0251] Figure 3 This is a diagram illustrating an exemplary arrangement of control points for WFS mode operation.

[0252] Figure 4 This is a diagram of a first embodiment of a signal processing scheme for WFS mode operation.

[0253] Figure 5 This is a diagram of a second embodiment of the signal processing scheme used for WFS mode operation.

[0254] Figures 6A to 6E It is a set of polar coordinate plots, which respectively show the measured performance of a prototype loudspeaker array with the beam turned 0 degrees at frequencies of 10000Hz, 5000Hz, 2500Hz, 1000Hz and 600Hz.

[0255] Figure 7A This is a diagram illustrating the basic principle of binaural mode operation.

[0256] Figure 7B This is a diagram illustrating binaural mode operation for spatialized sound presentation.

[0257] Figure 8 This is a block diagram illustrating an exemplary binaural mode processing chain.

[0258] Figure 9 This is a diagram of a first embodiment of a signal processing scheme for binaural mode.

[0259] Figure 10 This is a diagram illustrating an exemplary arrangement of control points for binaural mode operation.

[0260] Figure 11 This is a block diagram of a second embodiment of the signal processing chain for binaural mode.

[0261] Figure 12A and Figure 12B Accordingly, analog frequency-domain and time-domain representations of the predicted performance of an exemplary loudspeaker array in binaural modes, measured in the left and right ears, are shown.

[0262] Figure 13 The relationship between the virtual speaker array and the physical speakers is shown. Detailed Implementation

[0263] In binaural mode, the speaker array provides two sound outputs, directed towards the primary listener's ear. The inverse filter design approach is derived from mathematical simulations, where a near-real-world model of the speaker array is created, and virtual microphones are placed throughout the target sound field. An objective function is created or requested across these virtual microphones. The inverse problem is solved using regularization, creating a stable and realizable inverse filter for each speaker element in the array. For each array element, the source signal is convolved with these inverse filters.

[0264] In the second beamforming or wavefield synthesis (WFS) mode, the transform processor array represents sound signals from multiple discrete sources to separate physical locations within the same approximate region. The masking signal can also be dynamically adjusted in amplitude and time to provide optimal masking and a lack of intelligibility of the signal of interest to the listener.

[0265] The WFS mode also uses an inverse filter. Instead of pointing two beams at the listener's ears, this mode points or directs multiple beams at different locations around the array.

[0266] The technology relates to a digital signal processing (DSP) strategy that allows both binaural rendering and WFS / sound beamforming to be used individually or simultaneously. As described above, virtual spatialization is then combined for a small number of physical transducers, such as two or four.

[0267] For binaural and WFS modes, the signal to be reproduced is filtered using a set of digital filters. These filters can be generated by numerically solving an inverse electroacoustic problem. The specific parameters of the particular inverse problem to be solved are described below. However, typically, digital filter design is based on a least-squares minimization principle, i.e., a cost function of the type J = E + βV.

[0268] The cost function is the sum of two terms: performance error E, which measures the reproduction of the desired signal at the target point; and effort cost βV, a quantity proportional to the total power input to all speakers. The positive real number β is a regularization parameter that determines the weights assigned to the effort term. Note that, according to this implementation, the cost function can be applied after the summation, and optionally after the execution of the limiter / peak reduction function.

[0269] By varying β from zero to infinity, the solution gradually shifts from minimizing only the performance error to minimizing only the effort cost. In practice, this regularization is achieved by limiting the loudspeaker's power output to the ill-conditioned frequencies of the inversion problem. This is done without affecting the system's performance at the well-conditioned frequencies of the inversion problem. This prevents spikes from appearing in the spectrum of the reproduced sound. If needed, frequency-dependent regularization parameters can be used to selectively attenuate peaks.

[0270] Wavefield synthesis / beamforming modes

[0271] The WFS (Wireless Sound Field) audio signal is generated for a linear array of virtual loudspeakers, defining several separate sound beams. In WFS mode operation, narrow beams can be used to direct different source content from the loudspeaker array at different angles to minimize leakage to adjacent areas during listening. For example... Figure 1A As shown, private listening is made possible using adjacent beams of music and / or noise delivered by speaker array 72. The direct sound beam 74 is received by the target listener 76, while a masking noise beam 78 (which can be music, white noise, or some other signal different from the main sound beam 74) is directed around the target listener to prevent unintentional eavesdropping by others in the surrounding area. The masking signal can also be dynamically adjusted in amplitude and time to provide optimal masking and intelligibility of the signal of interest to the listener, as illustrated in the diagram containing the DRCEDSP block later.

[0272] When virtual speaker signals are combined, a significant portion of the spatial sound cancellation capability is lost; however, for direct (i.e., non-reflective) sound paths, it is at least theoretically possible to optimize the sound at each ear of the listener.

[0273] In WFS mode, the array provides signals from multiple discrete sources. For example, three people can listen to three different sources around the array with minimal interference between their signals. Figure 1B An exemplary configuration of WFS mode for multi-user / multi-location applications is shown. With only two speaker transducers, complete control over each listener is impossible, although acceptable (relative to stereo audio improvement) is available through optimization. As shown, array 72 defines discrete sound beams 73, 75, and 77 for each of listeners 76a and 76b, each beam containing different sound content. Although both listeners are shown receiving the same content (each of the three beams), different content may be delivered to one or the other at different times. When the array signals are summed, some directionality is lost and, in some cases, inverted. For example, in the case of summing the signals of a set of 12 speaker arrays into a 4-speaker signal, the direction cancellation signal may not be canceled in most locations. However, preferably, appropriate cancellation is preferably available for the listener in the optimal location.

[0274] WFS mode signals are generated through a DSP chain, such as Figure 2As shown in the diagram, discrete source signals 801, 802, and 803 are each convolved with an inverse filter of each of the speaker array signals. An inverse filter is a mechanism that allows optimization of a localized audio beam for a specific location based on specifications in a mathematical model used to generate the filter. This can be calculated in real-time to provide instantaneously optimized beam control capabilities, allowing users to track the array with audio. In the example shown, the speaker array 812 has twelve elements, thus twelve filters 804 for each source. The resulting filtered signals corresponding to the same nth speaker signal are added at a combiner 806, whose result is fed to a multi-channel sound card 808 having a DAC corresponding to each of the twelve speakers in the array. The twelve signals are then split into channels, either two or four, and the members of each subset are time-adjusted for the positional difference between the physical location of the corresponding array signal and the corresponding physical transducer, summed, and then subjected to a limiting algorithm. The limiting signals are then amplified using a Class D amplifier 810 and delivered to one or more listeners via two or four speaker arrays 812.

[0275] Figure 3 This illustrates how to generate a spatialized filter. First, assume a given relative arrangement of N array elements. Define a set of M virtual control points 92, where each control point corresponds to a virtual microphone. The control points are arranged on a semicircle surrounding the N loudspeaker arrays 98, centered on the center of the loudspeaker arrays. The radius of the arc 96 can be scaled with the size of the array. The control points 92 (virtual microphones) are uniformly arranged on the arc, with a constant angular distance between adjacent points.

[0276] Calculate the M×N matrix H(f), which represents the electroacoustic transfer function between each loudspeaker and each control point in the array, as a function of frequency f, where H p , l corresponds to the transfer function between the 1st loudspeaker (in N loudspeakers) and the p-th control point 92. These transfer functions can be measured or analyzed from the acoustic radiation model of the loudspeakers. An example of the model is given by an acoustic monopole, given by the following equation:

[0277] Where c is the speed of sound propagation, f is the frequency, and r p,l It is the distance between the first speaker and the p-th control point.

[0278] Instead of correcting the time delay after the array signal has been fully finite, the correct speaker position can be used when generating the signal to avoid refinite the signal.

[0279] As is known in the art, a more advanced analytical radiation model for each loudspeaker can be obtained through multipole expansion. (See, for example, V. Rokhlin, “Diagonal Form of Translation Operators for Three-Dimensional Helmholtz Equations,” Applied and Computational Harmonic Analysis, 1:82-93, 1993).

[0280] Let M elements define the vector p(f), representing the target sound field at the location identified by control point 92, as a function of frequency f. The target field has several options. One possibility is to assign a value of 1 to control point(s) that identify the direction(s) of the desired sound beam(s), and assign a value of zero to all other control points.

[0281] Digital filter coefficients, defined in the frequency (f) domain or the digital sampling (z) domain, are N elements of the vector a(f) or a(z) output by the filter computation algorithm. Filters may have different topologies, such as FIR, IIR, or other types. For each frequency f or sample parameter z, the coefficients are calculated by minimizing, for example, the following cost function: J(f) = ‖H(f)a(f)-p(f)‖. 2 +β‖a(f)‖ 2 The linear optimization problem is used to compute vector a. The symbol ‖...‖ denotes the L... 2 The norm is used, and β is a regularization parameter whose value can be limited by the designer. Standard optimization algorithms can be used to numerically solve the above problems.

[0282] Now for reference Figure 4 The system input is any set of audio signals (from A to Z), referred to as sound source 102. The system output is a set of audio signals (from 1 to N) driving N units of speaker array 108. These N signals are called "speaker signals".

[0283] For each sound source 102, the input signal is filtered through a set of N digital filters 104, with one digital filter 104 for each loudspeaker in the array. These digital filters 104 are referred to as “spatialization filters”, which are generated by the algorithm disclosed above and vary as a function of the position of one or more listeners and / or the expected direction of the sound beam to be generated.

[0284] Digital filters can be implemented as finite impulse response (FIR) filters; however, other filter topologies can achieve higher efficiency and better response modeling, such as feedback or reentrant infinite impulse response (IIR) filters. Filters can be implemented in traditional DSP architectures, or in graphics processing units (GPUs, developer.nvidia.com / vrworks-audio-s dk-depth) or audio processing units (APUs, www.NVIDIA.com / en-us / drivers / APU / ). Advantageously, acoustic processing algorithms are presented as ray tracing, transparency, and scattering models.

[0285] For each sound source 102, the audio signal filtered by the nth digital filter 104 (i.e., corresponding to the nth speaker) is added at the combiner 106 to the audio signal corresponding to a different sound source 102 but to the same nth speaker. The summed signal is then output to the speaker array 108.

[0286] Figure 5 It shows Figure 4 An alternative embodiment of the binaural signal processing chain includes the use of optional components, including a psychoacoustic bandwidth extension processor (PBEP) and a dynamic range compressor and extender (DRCE), which provide more sophisticated dynamic range and masking control, environment-specific filtering algorithm customization, room equalization, and distance-based attenuation control.

[0287] PBEP 112 allows the listener to perceive sound information contained in the lower part of the audio spectrum by generating higher-frequency sound material (using higher-frequency sounds to provide the perception of lower frequencies). Since PBE processing is non-linear, it is important that it occurs before the spatialization filter 104. Inserting a non-linear PBEP block 112 after the spatial filter would severely degrade the generation of sound beams.

[0288] It is important to emphasize that the PBEP 112 is used to compensate for the poor directivity of the speaker array at lower frequencies (psychoacoustic), rather than to compensate for the poor bass response of the individual speakers themselves, as is commonly done in existing technology applications.

[0289] The DRCE114 in the DSP chain provides loudness matching for the source signals, ensuring sufficient relative masking of the output signals of array 108. In binaural rendering mode, the DRCE used is a 2-channel block that performs the same loudness correction on both incoming channels.

[0290] Similar to the PBEP block 112, the DRCE 114 is important because its processing is non-linear, and therefore it appears before the spatialization filter 104. Inserting the non-linear DRCE block 114 after the spatial filter 104 would severely degrade the generation of the sound beam. However, without this DSP block, the psychoacoustic performance of the DSP chain and array would also be reduced.

[0291] Another optional component is a listener tracking device (LTD) 116, which allows the device to receive information about the location of one or more listeners and dynamically adjust the spatialization filter in real time. LTD 116 can be a video tracking system that detects head movement of the listener, or it can be another type of motion sensing system known in the art. LTD 116 generates a listener tracking signal, which is input into a filter calculation algorithm 118. Adaptation can be achieved by recalculating the digital filter in real time or by loading different filter banks from a pre-calculated database. Alternative user positioning includes radar (e.g., heartbeat) or lidar tracking, RFID / NFC tracking, breathing sounds, etc.

[0292] Figures 6A to 6E It is a polar coordinate energy radiation map of the radiation pattern of a prototype array, which is driven by a DSP scheme operating in WFS mode at five different frequencies of 10,000 Hz, 5,000 Hz, 2,500 Hz, 1,000 Hz and 600 Hz, and measured with a microphone array with beams turned at 0 degrees.

[0293] Binaural mode: The DSP used for binaural mode includes the convolution of the audio signal to be reproduced with a set of digital filters representing the head-related transfer function (HRTF).

[0294] Figure 7A The basic method used in binaural mode operation is illustrated, wherein an array of speaker positions 10 is configured to generate specially shaped audio beams 12 and 14, which can be delivered to the listener's ears 16L and 16R, respectively. Using this mode, the beams themselves eliminate crosstalk. However, this is not feasible after summarizing and demonstrating with a smaller number of speakers.

[0295] Figure 7B The illustration depicts a hypothetical video conference call involving multiple locations and parties. When a party in New York is speaking, the sound appears to be delivered from a direction coordinated with the video image of a speaker in video display 18. When a participant in Los Angeles speaks, the sound can be delivered in coordination with the position of the speaker's image in the video display. Real-time binaural encoding can also be used to deliver convincing spatial audio headsets, avoiding the noticeable sound misalignment often found in existing headset setups.

[0296] like Figure 8 The binaural signal processing chain shown in the figure consists of multiple discrete sources. In the example shown, there are three sources: sources 201, 202, and 203. These are then convolved with binaural head-related transfer function (HRTF) encoded filters 211, 212, and 213, which correspond to the desired virtual transmission angle from the nominal speaker position to the listener. Each sound source has two HRTF filters, one for the left ear and one for the right ear. The resulting HRTF-filtered signals for the left ear are all summed together to generate the input signal corresponding to the sound to be heard by the listener's left ear. Similarly, the HRTF-filtered signals for the listener's right ear are summed together. The resulting left and right ear signals are then convolved accordingly with inverse filter banks 221 and 222, where each virtual speaker element in the virtual speaker array has one filter. Then, through further spatiotemporal transformation, combination, and limiting / peak reduction, the virtual loudspeaker signals are combined into real loudspeaker signals, and the resulting combined signals are sent to the corresponding loudspeaker elements via a multi-channel sound card 230 and a Class D amplifier 240 (one for each physical loudspeaker) for audio transmission to the listener via the loudspeaker array 250. In binaural mode, the invention generates sound signals that feed the virtual linear array. The virtual linear array signals are combined into loudspeaker drive signals. The loudspeakers provide two sound beams to the primary listener's ears—one for the left ear and one for the right ear.

[0297] Figure 9 A binaural mode signal processing scheme with binaural modes from sound source A to Z is shown. (See reference...) Figure 8 The system input is a set of sound source signals 32 (A to Z), and the system output is correspondingly a set of speaker signals 38 (1 to N). For each sound source 32, the input signal is filtered by two digital filters 34 (HRTF-L and HRTF-R) representing left and right head correlation transfer functions, which are calculated for the angle intended to be rendered to the listener for a given sound source 32. For example, a speaker's voice can be rendered as a plane wave arriving at 30 degrees to the listener's right. The HRTF filters 34 can be obtained from a database or calculated in real time using a binaural processor. After HRTF filtering, the processed signals corresponding to different sound sources but corresponding to the same ear (left or right ear) are combined together at a combiner 35. This generates two signals, hereinafter referred to accordingly as "Total Binaural Signal - Left" or "TBS-L" and "Total Binaural Signal - Right" or "TBS-R".

[0298] Each of the two total binaural signals, TBS-L and TBS-R, is filtered through a set of N digital filters 36, each for one loudspeaker, calculated using the algorithm disclosed below. These filters are referred to as “spatialization filters.” For clarity, it is emphasized that the spatialization filter bank used for the right total binaural signal is different from the spatialization filter bank used for the left total binaural signal.

[0299] The filtered signals corresponding to the same nth virtual speaker but for two different ears (left and right) are added together at combiner 37. These are the virtual speaker signals, which are fed to the combiner system, which in turn feeds to the physical speaker array 38.

[0300] The computation algorithm for the spatialized filter 36 used in the binaural mode is similar to that used in the WFS mode described above. The main difference from the WFS case is that only two control points are used in the binaural mode. These control points correspond to the positions of the listener's ears, and are as follows: Figure 10 The arrangement shown indicates that the distance between the two points 42 representing the listener's ears is in the range of 0.1m and 0.3m, while the distance between each control point and the center 46 of the speaker array 48 can be scaled depending on the size of the array used, but is typically in the range of 0.1m and 3m.

[0301] A 2×N matrix H(f) is computed as a function of frequency f using elements of the electroacoustic transfer function between each loudspeaker and each control point. As mentioned above, these transfer functions can be either measured or calculated analytically. A 2-element vector p is defined. This vector can be [1,0] or [0,1], depending on whether the spatialized filter is computed for the left or right ear accordingly. The filter coefficients for a given frequency f are obtained by minimizing the following cost function J(f) = ||H(f)a(f)-p(f)|| 2 +β||a(f)|| 2 Calculate the N elements of the vector a(f). If multiple solutions are possible, choose the L elements corresponding to a(f). 2 The solution for the minimum norm.

[0302] Figure 11 It shows Figure 9An optional embodiment of the binaural signal processing chain includes optional components, including a psychoacoustic bandwidth extension processor (PBEP) and a dynamic range compressor and extender (DRCE). PBEP 52 allows the listener to perceive sound information contained in the lower part of the audio spectrum by generating higher-frequency sound material (using higher-frequency sounds to provide the perception of lower frequencies). Because PBEP processing is non-linear, it is important that it occurs before the spatialization filter 36. If a non-linear PBEP block 52 is inserted after the spatial filter, its effect will severely reduce the generation of sound beams.

[0303] It is important to emphasize that the PBEP 52 is used to compensate for the poor directivity of the speaker array at lower frequencies (psychoacoustic), rather than to compensate for the poor bass response of individual speakers themselves.

[0304] The DRCE 54 in the DSP chain provides loudness matching for the source signals, ensuring sufficient relative masking of the output signals of array 38. In binaural rendering mode, the DRCE used is a 2-channel block that performs the same loudness correction on both incoming channels.

[0305] Similar to PBEP block 52, the DRCE 54 process is non-linear, making its placement before spatialization filter 36 crucial. Inserting a non-linear DRCE block 54 after spatial filter 36 would severely degrade sound beamforming. However, without this DSP block, the psychoacoustic performance of the DSP chain and array would also be reduced.

[0306] Another optional component is a listener tracking device (LTD) 56, which allows the device to receive information about the location of one or more listeners and dynamically adjust the spatialization filter in real time. LTD 56 can be a video tracking system that detects head movement of the listener, or it can be another type of motion sensing system known in the art. LTD 56 generates a listener tracking signal, which is input into a filter calculation algorithm 58. Adaptation can be achieved by recalculating the digital filter in real time or by loading different filter banks from a pre-calculated database.

[0307] Figure 12A and Figure 12B The simulated performance of the algorithm for binaural mode is shown. Figure 12A The simulated frequency domain signals at the target locations of the left and right ears are shown, while Figure 12B The time-domain signal is shown. Both figures clearly illustrate the capability to target one ear, in this case, the left ear, with the desired signal while minimizing the signal detected in the listener's right ear.

[0308] WFS and binaural mode processing can be combined into a single device to produce overall sound field control. Such methods combine the advantages of directing selected sound beams towards a target listener, such as for privacy or enhanced intelligibility, with individual control over the mixing of sound delivered to the listener's ears to produce surround sound. The device can use either binaural or WFS modes alternatively or in combination to process audio. Although not specifically shown herein, the use of both WFS and binaural modes will be determined by… Figure 5 and Figure 11 The block diagram shows that their corresponding outputs are combined by combiners 37 and 106 in the signal summation step. The use of both WFS and binaural modes can also be achieved through combination. Figure 2 and Figure 8 The block diagram illustrates that their corresponding outputs are summed together at the last summation block before the multi-channel sound card 230.

[0309] Example

[0310] A 12-channel spatialized virtual audio array is implemented according to US9,578,440. This virtual array provides signals to drive a linear or curved, equally spaced array, such as 12 speakers, located in front of the listener. The virtual array is divided into two or four sections. In the case of two sections, for example, six signals for the "left" section are directed to the left physical speaker, and for example, six signals for the "right" section are directed to the right physical speaker. The virtual signals are summed using at least two intermediate processing steps.

[0311] The first intermediate processing step compensates for the time difference between the nominal location of the virtual loudspeakers and the physical location of the loudspeaker transducers. For example, the virtual loudspeaker closest to the listener is assigned a reference delay, while more distant virtual loudspeakers are assigned an increased delay. Typically, the position of the virtual array causes the time difference between adjacent virtual loudspeakers to vary incrementally, although this allows for more rigorous analysis. At a 48kHz sampling rate, the difference between the nearest and farthest virtual loudspeakers could be, for example, four cycles.

[0312] The second intermediate processing step limits the signal peaks to avoid overdriving the physical speakers or causing significant distortion. This limitation can be frequency-selective, so only one band is affected by this process. This step should be performed after delay compensation. For example, a compander can be used. Alternatively, assuming only rare peaks, a simple limitation can be used. In other cases, more complex peak reduction techniques can be used, such as phase shifting of one or more channels, typically based on a predicted peak value of the signal that is slightly delayed from its real-time presentation. Note that this phase shift alters the time delay of the first intermediate processing step; however, a trade-off is necessary when the physical limitations of the system are met. For a virtual line array of 12 speakers and 2 physical speakers, the physical speaker positions are between elements 3-4 and 9-10. If (s) is the center-to-center distance between speakers, then the distance from the center of the array to the center of each real speaker is: A = 3s. The left speaker is offset from the center by -A, and the right speaker is offset by A.

[0313] The second intermediate processing step primarily utilizes limiters and / or compressors or other processing to provide downmixing for six virtual channels with peak reduction, applying this processing to prevent saturation or clipping. For example, the left channel is: L 输出 =Restriction (L1+L2+L3+L4+L5+L6),

[0314] And the correct channel is R. 输出 = Limit (R1+R2+R3+R4+R5+R6).

[0315] Before downmixing, it's necessary to consider the delay difference between the virtual speakers and the listener's ear compared to the physical speaker transducers and the listener's ear. This delay is particularly significant at higher frequencies due to the increased ratio of the length of the virtual speaker array to the sound wavelength. To calculate the distance from the listener to each virtual speaker, assume the speakers n are numbered 1 to 6, where 1 is the closest to the center and 6 is the farthest. The distance from the array center to the speaker is: d = ((n-1) + 0.5) * s. Using the Pythagorean theorem, the distance from the speaker to the listener can be calculated as:

[0316] The distance from the actual speaker to the listener is

[0317] The sample delay of each speaker can be calculated using the distance difference between two listeners. This can be converted into samples (assuming a sound speed of 343 m / s and a sampling rate of 48 kHz).

[0318]

[0319] This can result in significant delays between listener distances. For example, if the distance between virtual array speakers is 38mm and the listener is 500mm from the array, the delay from the leftmost virtual speaker (n=6) to the actual speaker is: One sample.

[0320] At higher audio frequencies, i.e., 12kHz, a complete wave period is 4 samples, with a difference equal to a 360° phase shift. See Table 1.

[0321] Therefore, when combining the signals used for the virtual speaker into the physical speaker signal, it is preferable to compensate for the time offset based on the displacement of the virtual speaker relative to the physical speaker. The time offset can also be performed within the spatialization algorithm, rather than as post-processing.

[0322] This invention can be implemented in software, hardware, or a combination of both. It can also be embodied as computer-readable code on a computer-readable medium. The computer-readable medium can be any data storage device capable of storing data that can subsequently be read by a computing device. Examples of computer-readable media include read-only memory, random access memory, CD-ROM, magnetic tape, optical data storage devices, and carrier waves. The computer-readable medium can also be distributed across a network-connected computer system, enabling the computer-readable code to be stored and executed in a distributed manner.

[0323] Many features and advantages of the invention are apparent from the written description, and therefore the appended claims are intended to cover all such features and advantages. Furthermore, since many modifications and variations will readily occur to those skilled in the art, it is not intended to limit the invention to the exact configurations and operations shown and described. Therefore, all suitable modifications and equivalents may be considered to fall within the scope of the invention.

Claims

1. A method for generating transauricular spatialized sound, comprising: Receive audio signals representing spatial audio objects; Each audio signal is filtered by a spatialization filter to generate a virtual audio transducer signal array for representing a virtual audio transducer array. The virtual audio transducer signal array is divided into subsets, each subset including multiple virtual audio transducer signals, and each subset is used to drive physical audio transducers located within the physical location range of the corresponding subset. Based on the time difference of arrival of sound from the nominal position of the corresponding virtual audio transducer and the physical position of the corresponding physical audio transducer relative to the listener's target ear, the corresponding virtual audio transducer signal of the corresponding subset is time-shifted. as well as The corresponding virtual speaker signals with time offsets of the corresponding subsets are combined into physical audio transducer drive signals.

2. The method of claim 1, further comprising reducing the peak amplitude of the corresponding virtual audio transducer signal of the combined time offset to reduce saturation distortion of the physical audio transducer.

3. The method of claim 1, wherein the filtering comprises processing at least two audio channels using a digital signal processor.

4. The method of claim 1, wherein the filtering includes processing at least two audio channels using a graphics processing unit configured to act as an audio signal processor.

5. The method according to claim 1, wherein the virtual audio transducer signal array is a linear array of 12 virtual audio transducers.

6. The method of claim 1, wherein the virtual audio transducer array is a linear array having at least three times the number of virtual audio transducer signals as physical audio transducer drive signals.

7. The method of claim 1, wherein the virtual audio transducer array is a linear array having at least six times the number of virtual audio transducer signals as physical audio transducer drive signals.

8. The method of claim 1, wherein each subset is a non-overlapping adjacent group of virtual audio transducer signals.

9. The method of claim 6, wherein each subset is a non-overlapping adjacent group of at least 6 virtual audio transducer signals.

10. The method of claim 1, wherein each subset has a virtual audio transducer, the location of which overlaps with the representation location range of another subset of the virtual audio transducer signal.

11. The method of claim 10, wherein the overlap is a virtual audio transducer signal.

12. The method of claim 1, wherein the virtual audio transducer signal array is a linear array of 12 virtual audio transducer signals, divided into two non-overlapping groups, each group having 6 adjacent virtual audio transducer signals, which are respectively combined to form 2 physical audio transducer drive signals.

13. The method of claim 12, wherein the corresponding physical audio transducer of each group is located between the third and fourth virtual audio transducers of adjacent groups of the six virtual audio transducer signals.

14. The method of claim 1, wherein the physical audio transducer has a non-directional emission mode.

15. The method of claim 14, wherein the virtual audio transducer array is modeled for directivity.

16. The method of claim 14, wherein the virtual audio transducer array is a phased array of audio transducers.

17. The method of claim 1, wherein the filtering includes crosstalk cancellation.

18. The method of claim 1, wherein the filtering is performed using a reentrant data filter.

19. The method of claim 1, further comprising receiving a signal indicating the position of the listener's ear.

20. The method of claim 1, further comprising tracking the movement of the listener and adjusting the filtering based on the tracked movement.

21. The method of claim 1, further comprising adaptively allocating virtual audio transducer signals to a corresponding subset.

22. The method of claim 1, further comprising adaptively determining the listener's head-related transfer function and filtering based on the adaptively determined head-related transfer function.

23. The method of claim 22, further comprising sensing features of the listener's head and adjusting the head-related transfer function based on the features.

24. The method of claim 1, wherein the filtering includes time-domain filtering.

25. The method of claim 1, wherein the filtering includes frequency domain filtering.

26. The method of claim 1, wherein the physical audio transducer drive signal is delayed by at least 25 milliseconds relative to the received audio signal representing the spatial audio object.

27. The method according to claim 1, further comprising: Adaptively determine the listener's head-related transfer function; Filtering is performed based on an adaptively determined head-related transfer function; Sensing features of the listener's head; as well as The head-related transfer function is adjusted based on the described features.

28. A system for generating transear spatialized sound, comprising: The input terminal is configured to receive an audio signal representing a spatial audio object; A spatialized audio data filter is configured to process each audio signal to generate a virtual audio transducer signal array representing a virtual audio transducer array of spatialized audio. The virtual audio transducer signal array is divided into subsets, each subset comprising multiple virtual audio transducer signals, each subset being used to drive physical audio transducers located within a physical location range of the corresponding subset. A time delay processor is configured to time-shift the corresponding virtual audio transducer signals of a corresponding subset based on the time difference of arrival of sound from the nominal position of the corresponding virtual audio transducer and the physical position of the corresponding physical audio transducer relative to the target ear of the listener. as well as A combiner configured to combine the time-offset virtual speaker signals of the respective subsets into physical audio transducer drive signals.

29. The system of claim 28, further comprising at least one of the following: A peak amplitude reduction filter is configured to reduce the saturation distortion of the physical audio transducer in the combined time offset of the corresponding virtual audio transducer signal; A limiter configured to reduce the saturation distortion of the physical audio transducer in the corresponding virtual audio transducer signal of the combined time offset; A compander configured to reduce the saturation distortion of the physical audio transducer in relation to the corresponding virtual audio transducer signal of the combined time offset; and A phase rotator configured to rotate the relative phase of at least one virtual audio transducer signal.

30. The system of claim 28, further comprising a peak amplitude reduction filter configured to reduce saturation distortion of the physical audio transducer of the combined time-off corresponding virtual audio transducer signal.

31. The system of claim 28, further comprising a limiter configured to reduce saturation distortion of the physical audio transducer of the combined time-offset corresponding virtual audio transducer signals.

32. The system of claim 28, further comprising a compander configured to reduce saturation distortion of the physical audio transducer of the combined time-shifted corresponding virtual audio transducer signals.

33. The system of claim 28, further comprising a phase rotator configured to rotate the relative phase of at least one virtual audio transducer signal.

34. The system of claim 28, wherein the spatialized audio data filter includes a digital signal processor configured to process at least two audio channels.

35. The system of claim 28, wherein the spatialized audio data filter includes a graphics processing unit configured to process at least two audio channels.

36. The system of claim 28, wherein the virtual audio transducer signal array is a linear array of 12 virtual audio transducer signals.

37. The system of claim 28, wherein the virtual audio transducer signal array is a linear array having at least three times the number of virtual audio transducer signals as physical audio transducer signals.

38. The system of claim 28, wherein the virtual audio transducer signal array is a linear array having at least six times the number of virtual audio transducer signals as physical audio transducer drive signals.

39. The system of claim 28, wherein each subset is a non-overlapping adjacent group of virtual audio transducer signals.

40. The system of claim 39, wherein each subset is a non-overlapping adjacent group of at least 6 virtual audio transducer signals.

41. The system of claim 28, wherein each subset has a virtual audio transducer signal having a representation position that overlaps with the position range of another subset of the virtual audio transducer signal.

42. The system of claim 41, wherein the overlap is a virtual audio transducer signal.

43. The system of claim 28, wherein the virtual audio transducer signal array is a linear array of 12 virtual audio transducer signals, divided into two non-overlapping groups, each group having 6 adjacent virtual audio transducer signals, which are combined to form 2 corresponding physical audio transducer drive signals.

44. The system of claim 43, wherein the corresponding physical audio transducer of each group is located between the third and fourth virtual audio transducers of adjacent groups of the six virtual audio transducer signals.

45. The system of claim 28, wherein the physical audio transducer has a non-directional emission mode.

46. ​​The system of claim 45, wherein the spatialized audio data filter is configured to model the virtual audio transducer array for directionality.

47. The system of claim 46, wherein the virtual audio transducer array is a phased array of audio transducers.

48. The system of claim 28, wherein the spatialized audio data filter is configured to perform crosstalk cancellation.

49. The system of claim 28, wherein the spatialized audio data filter includes a reentrant data filter.

50. The system of claim 28, further comprising an input port configured to receive a signal indicating the ear position of the listener.

51. The system of claim 28, further comprising an input configured to receive a signal that tracks the movement of the listener, wherein the spatialized audio data filter adaptively depends on the tracked movement.

52. The system of claim 28, wherein the virtual audio transducer signal is adaptively assigned to a corresponding subset.

53. The system of claim 28, wherein the spatialized audio data filter depends on a head-related transfer function adaptively determined by the listener.

54. The system of claim 53, further comprising an input port configured to receive a signal including sensed features of the listener’s head, wherein the head-related transfer function is adjusted according to the features.

55. The system of claim 28, wherein the spatialized audio data filter comprises a time-domain filter.

56. The system of claim 28, wherein the spatialized audio data filter comprises a frequency domain filter.

57. The system of claim 28, wherein the physical audio transducer drive signal is delayed by at least 25 milliseconds relative to the received audio signal representing the spatial audio object.

58. A system for generating spatialized sound, comprising: The input terminal is configured to receive an audio signal representing a spatial audio object; At least one automation processor, said at least one automation processor being configured to: Each audio signal is processed by a spatialization filter to generate a virtual audio transducer signal array representing a virtual audio transducer array. The virtual audio transducer signal array is divided into subsets, each subset including multiple virtual audio transducer signals, and each subset is used to drive physical audio transducers located within the physical location range of the corresponding subset. Based on the time difference of arrival of sound from the nominal position of the corresponding virtual audio transducer and the physical position of the corresponding physical audio transducer relative to the listener's target ear, the corresponding virtual audio transducer signal of the corresponding subset is time-shifted. as well as The corresponding virtual speaker signals with time offsets of the corresponding subsets are combined into physical audio transducer drive signals; as well as At least one output port, the at least one output port being configured to present a corresponding subset of the physical audio transducer drive signals.

59. The system of claim 58, further comprising a peak amplitude reduction filter configured to reduce saturation distortion of the physical audio transducer of the combined time-off corresponding virtual audio transducer signal.

60. The system of claim 58, further comprising a limiter configured to reduce saturation distortion of the physical audio transducer of the combined time-offset corresponding virtual audio transducer signals.

61. The system of claim 58, further comprising a compander configured to reduce saturation distortion of the physical audio transducer of the combined time-off corresponding virtual audio transducer signals.

62. The system of claim 58, further comprising a phase rotator configured to rotate the relative phase of at least one virtual audio transducer signal.

63. The system of claim 58, wherein the spatialization filter includes a digital signal processor configured to process at least two audio channels.

64. The system of claim 58, wherein the spatialization filter includes a graphics processing unit configured to process at least two audio channels.

65. The system of claim 58, wherein the virtual audio transducer signal array is a linear array of 12 virtual audio transducer signals.

66. The system of claim 58, wherein the virtual audio transducer signal array is a linear array having at least three times the number of virtual audio transducer signals as physical audio transducer drive signals.

67. The system of claim 58, wherein the virtual audio transducer signal array is a linear array having at least six times the number of virtual audio transducers as physical audio transducer drive signals.

68. The system of claim 58, wherein each subset is a non-overlapping adjacent group of virtual audio transducer signals.

69. The system of claim 68, wherein each subset is a non-overlapping adjacent group of at least six virtual audio transducer signals.

70. The system of claim 58, wherein each subset has a virtual audio transducer signal having a representation position that overlaps with the position range of another subset of the virtual audio transducer signal.

71. The system of claim 70, wherein the overlap is a virtual audio transducer signal.

72. The system of claim 58, wherein the virtual audio transducer signal array is a linear array of 12 virtual audio transducer signals, divided into two non-overlapping groups, each group having 6 adjacent virtual audio transducer signals, which are combined to form 2 corresponding physical audio transducer signals.

73. The system of claim 72, wherein the corresponding physical audio transducer of each group is located between the third and fourth virtual audio transducers of adjacent groups of the six virtual audio transducer signals.

74. The system of claim 58, wherein the physical audio transducer has a non-directional emission mode.

75. The system of claim 74, wherein the spatialization filter is configured to model the virtual audio transducer signal array for directionality.

76. The system of claim 75, wherein the virtual audio transducer signal array is a phased array of audio transducers.

77. The system of claim 58, wherein the spatialization filter is configured to perform crosstalk cancellation.

78. The system of claim 58, wherein the spatialization filter comprises a reentrant data filter.

79. The system of claim 58, further comprising an input port configured to receive a signal indicating the ear position of the listener.

80. The system of claim 58, further comprising an input configured to receive a signal that tracks the movement of the listener, wherein the spatialization filter adaptively depends on the tracked movement.

81. The system of claim 58, wherein the virtual audio transducer signal array is adaptively assigned to a corresponding subset.

82. The system of claim 58, wherein the spatialization filter depends on a head-related transfer function adaptively determined by the listener.

83. The system of claim 82, further comprising an input port configured to receive a signal including sensed features of the listener’s head, wherein the head-related transfer function is adjusted according to the features.

84. The system of claim 58, wherein the spatialization filter comprises a time-domain filter.

85. The system of claim 58, wherein the spatialization filter comprises a frequency domain filter.

86. The system of claim 58, wherein the physical audio transducer drive signal is delayed by at least 25 milliseconds relative to the received audio signal representing the spatial audio object.

Citation Information

Patent Citations

  • Enhanced virtual stereo reproduction for unmatched transaural loudspeaker systems

    US10499153B1

  • Stereo to enhanced spatialisation in stereo sound HI-FI decoding process method and apparatus

    US20010031051A1

  • Audio user interface with selective audio field expansion

    US20020150254A1

  • System and method for localization of sounds in three-dimensional space

    US20020196947A1

  • Method and apparatus for producing spatialized audio signals

    US20030059070A1