Method and system for optimizing the behavior of an audio playback system
By optimizing FIR filters and addressing linear and nonlinear components in loudspeaker systems, the method enhances the playback experience of audio systems, ensuring accurate sound reproduction and improved perceptual rendering.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- EX MACHINA SOUNDWORKS LLC
- Filing Date
- 2024-03-26
- Publication Date
- 2026-04-10
AI Technical Summary
Conventional audio playback systems fail to optimize the perceptual rendering of audio output by headphones or other playback systems, leading to a sub-optimal listening experience for music or audio related to visual media.
A method for optimizing finite impulse response (FIR) filters for loudspeaker systems by separately addressing linear and nonlinear components, using digital signal processing to enhance the performance of loudspeakers and headphones, and applying head-related transfer functions to achieve accurate sound imaging.
The method provides an optimized playback experience that meets or exceeds performance expectations by enhancing the linear and nonlinear behaviors of loudspeaker systems and accurately reproducing sound as if it were coming from specific locations, improving the perceptual rendering of audio.
Smart Images

Figure 2026511007000001_ABST
Abstract
Description
Technical Field
[0001] Cross - Reference to Related Applications
[0001] This application claims priority to U.S. Provisional Patent Application No. 63 / 454,841, filed on March 27, 2023, entitled "Methods and Systems for Optimizing Behavior of Audio Playback Systems", which is incorporated herein by reference.
Background Art
[0002] Background
[0002] This disclosure relates to methods for optimizing the behavior of audio playback systems. More particularly, the methods and systems described herein relate to functions for independently optimizing linear behavior from non - linear behavior in a loudspeaker system. The methods and systems described herein may further relate to functions for optimizing and applying filters to provide a perceptual rendering of audio optimized for playback by headphones.
[0003]
[0003] In the design process of conventional loudspeaker systems, the linear and non - linear characteristics of the system are considered together, and in many cases, improving one aspect compensates for a performance degradation in another aspect.
[0004]
[0004] Conventional audio playback systems do not optimize the perceptual rendering of audio output by headphones or other playback systems, causing a sub - optimal listening experience for music or audio related to visual media such as movies or television content. Therefore, there is a need for a technical solution to reproduce the experience of sound coming from any collection of sound sources at any location, including the playback of audio from stereo speakers.
Summary of the Invention
Means for Solving the Problems
[0005] Brief Summary
[0005] In one embodiment, a method for perceptually rendering audio for playback by headphones includes optimizing a first finite impulse response (FIR) filter associated with a first channel of audio input of audio associated with a sound source for application to a first transducer of output headphones. The method further includes optimizing a second FIR filter associated with a second channel of audio input of audio associated with a sound source for application to a second transducer of output headphones, the second FIR filter being modified to include at least one coefficient defined based on the relationship between the first transducer and the second transducer.
[0006]
[0006] In another embodiment, a method for perceptually rendering audio for playback by a playback speaker system includes optimizing a first finite impulse response (FIR) filter associated with a first channel of audio input of audio associated with a sound source for application to a first speaker in the playback speaker system for output from a first transducer of a first speaker. The method includes optimizing a second finite impulse response (FIR) filter associated with a second channel of audio input of audio associated with a sound source for application to a second speaker in the playback speaker for output from at least one transducer of a second speaker, further including modifying the second FIR filter to include at least one coefficient defined based on the relationship between the first transducer and the second transducer.
[0007]
[0007] In yet another embodiment, a method for independently optimizing the linear component from the nonlinear component in a loudspeaker design includes receiving identification information of at least one design specification relating to the nonlinear component of the loudspeaker system. The method includes defining at least one characteristic of at least one hardware component of the loudspeaker system that satisfies the received identification information. The method includes optimizing at least one linear component of the loudspeaker system, the optimization further includes optimizing the finite impulse response (FIR) filter of the loudspeaker system, and in combination with the application of the optimized FIR filter, the execution of the loudspeaker system including the selected nonlinear component satisfies a threshold level of performance for the loudspeaker system.
[0008] Brief explanation of the drawing
[0008] The above and other purposes, aspects, features, and advantages of this disclosure will be more clearly and better understood by referring to the following description as interpreted in conjunction with the attached drawings. [Brief explanation of the drawing]
[0009] [Figure 1A]
[0009] This is a block diagram showing one embodiment of a system for independently optimizing linear components from nonlinear components in loudspeaker design. [Figure 1B]
[0010] This block diagram shows one embodiment of a system for independently optimizing linear components from nonlinear components in loudspeaker design. [Figure 2]
[0011] This flowchart illustrates one embodiment of a method for independently optimizing the linear component from the nonlinear component in loudspeaker design. [Figure 3]
[0012] This flowchart illustrates one embodiment of a method for perceptually rendering audio for playback through headphones. [Figure 4]
[0013] This flowchart illustrates one embodiment of a method for perceptually rendering audio for playback by a playback speaker system. [Modes for carrying out the invention]
[0010] Detailed explanation
[0014] Referring next to Figure 1, a block diagram shows one embodiment of a system for independently optimizing the linear behavior of a loudspeaker system from its nonlinear behavior. In outline, the system 100 includes a computer 106a that communicates with a user computer 102 and optionally with a computer 106b. Computer 106a can run an optimization engine 103. The user computer 102 can run a client interface 105. Computer 106a can transmit data specifying the optimal design of a loudspeaker system, including one or more loudspeakers, to an optional computer 106b associated with a loudspeaker manufacturer. By generating design specifications for use in manufacturing optimized loudspeakers, the methods and systems described herein provide a technical solution to the problem of optimizing the linear behavior of a loudspeaker system separately from its nonlinear behavior in order to create a loudspeaker system that incorporates one or more design requirements without compromising performance in terms of linear or nonlinear behavior. System 100 may include functions for performing experiential steps (e.g., determining one or more measurements of the behavior of a loudspeaker system), functions for performing optimization steps, and functions for directing the deployment of the optimized system.
[0011]
[0015] The optimization engine 103 can be provided as a software component. The optimization engine 103 may also be provided as a hardware component. The computing device 106a can execute the optimization engine 103.
[0012]
[0016] For the sake of simplification, Figure 1 shows the optimization engine 103 and the client interface 105 as separate modules, but it should be understood that this does not limit the architecture to a specific implementation. For example, these components may be contained within a single circuit or software function, or they may be distributed across multiple computing devices.
[0013]
[0017] When used herein, linear behavior includes components of behavior that can be extracted from or contribute to the measurement of the impulse response (e.g., measured amplitude response and timing behavior), and such components may be invariant with respect to power. Linear behavior can be controlled by software. Nonlinear characteristics may include strain and dispersion without limitation. With the methods and systems described herein, designers of loudspeaker systems can consider only the power-dependent components of the design (including strain and dispersion without limitation), and these power-dependent components can be controlled by hardware design.
[0014]
[0018] Using digital signal processing to correct the linear behavior of a loudspeaker system allows for the optimization of nonlinear characteristics in the physical design of the loudspeaker and its components without sacrificing the linear performance of the finished system. Linearization can be implemented as a finite impulse response (FIR) filter on a digital signal processor within the loudspeaker system. Optimized coefficients for the FIR filter can be derived, as illustrated in Figure 2 below.
[0015]
[0019] Next, referring to Figure 2, the flowchart outlines one embodiment of Method 200 for independently optimizing linear behavior from nonlinear behavior in loudspeaker design. Method 200 includes receiving identification information of at least one design specification relating to the nonlinear behavior of a loudspeaker system (202). Method 200 includes defining at least one characteristic of at least one hardware component of the loudspeaker system that satisfies the received identification information (204). Method 200 includes optimizing at least one linear behavior of the loudspeaker system, the optimization further includes optimizing the FIR filter of the loudspeaker system, and in combination with the application of the optimized FIR filter, the execution of the loudspeaker system including the specified nonlinear behavior satisfies a threshold level of loudspeaker system performance (206). Furthermore, in combination with the application of the optimized FIR filter, the execution of a loudspeaker system designed to include at least one hardware component having at least one defined characteristic satisfies a threshold level of loudspeaker system performance.
[0016]
[0020] Next, referring in more detail to Figure 2 and in relation to Figure 1, Method 200 includes receiving identification information for at least one design specification relating to the nonlinear behavior of the loudspeaker system (202). The optimization engine 103 can receive identification information for at least one design specification from the user computing device 102. Alternatively, the optimization engine 103 can provide a user interface displayed on the computing device 106a to directly accept one or more design specifications.
[0017]
[0021] The design specifications may include specifications for the loudspeaker performance threshold levels for each of the one or more components within the loudspeaker. The computer 106a can provide an enumeration of one or more components in the loudspeaker system, and for each component in the enumeration, the computer 106a can provide one or more attributes of the loudspeaker system that the user can associate with and specify performance threshold levels. Using one or more received design specifications, the computer 106a can automatically (e.g., without human intervention) assign performance threshold levels to other attributes not addressed by the received design specifications.
[0018]
[0022] Below are examples of design specifications that the optimization engine 103 may receive for nonlinear components, which may result in the specifications generated and displayed in Figure 1B above. As an example, the design specifications may indicate that the acoustic center of the loudspeaker should be located at the midpoint of the y-axis of the narrowest possible front baffle (considering manufacturing tolerances and a maximum corner radius of 40 mm). This transducer may be either a full-range or two-way coaxial design. If this transducer cannot generate a sufficient sound pressure level with 2% THD or less at all frequencies required for the intended end application of the loudspeaker, it may be assisted by an additional low-frequency transducer. In any case, the full-range or coaxial transducer must meet the minimum SPL requirements for its application at 1 / 4 of its maximum output when crossovered with the low-frequency transducer at frequencies below 300 Hz.
[0019]
[0023] As another example, design specifications may indicate that full-range or coaxial transducers should be self-sealed within a sealed package, or physically isolated from the back pressure waves of low-frequency transducers via internal chambering or a separate sealed enclosure mounted inside the front baffle. In addition, transducers should not share an enclosure with any other electronic components. In the design of a powered loud speaker, all electronic components should be housed in a separate, internally sealed compartment, and all electrical connections between the acoustic enclosure and the electronic enclosure should be airtight using seals or gaskets.
[0020]
[0024] For example, the design specifications may indicate that the crossover should use a filter slope that is neither shallower than fourth order nor steeper than eighth order. All crossovers should employ high-precision digital filters, and in the case of a two-way coaxial design, the woofer / midrange and tweeter should be powered independently from their respective amplifier channels. In designs employing low-frequency transducers, they can be connected in series and / or parallel to a single amplifier channel, provided the net impedance load is at least 4 ohms.
[0021]
[0025] As an example, the design specifications may indicate that SPL and low-frequency extension requirements stipulate that the full-range or coaxial transducer should be supplemented by two or four additional dedicated low-frequency transducers. In designs intended to operate within a 2π space, such as a Sofit-mounted loudspeaker, these transducers should be radially equidistant from the full-range or coaxial transducer on the front baffle so that the design maintains at least theta angle axial symmetry, and should be inset-mounted or rear-mounted with a constant radius relative to the front baffle to the minimum depth required so that the apex of the driver surround is coplanar with or behind the front surface of the front baffle. In designs intended to operate within a 4π space, such as freestanding bookshelf loudspeakers, tower loudspeakers, or monitor loudspeakers, two low-frequency transducers should be symmetrically mounted on the side baffle along the same φ plane as the coaxial or full-range transducer, or four low-frequency transducers should be mounted on the side baffle in φ-angle symmetrical pairs above and below the full-range or coaxial transducer, such that the acoustic center of each transducer is equidistant from the acoustic center of the full-range or coaxial transducer.
[0022]
[0026] For example, the design specifications may indicate that acoustic methods for recovering backwave energy, such as Helmholtz resonators, transmission lines, or passive radiators, should not be employed that cause an acoustic group delay or phase angle deviation exceeding 30° at any frequency with respect to the forward pressure wave of the affected transducer.
[0023]
[0027] As an example, the design specifications may indicate that transducer design and selection should focus only on distortion, SPL potential, deflection angle / off-axis output mean, and eigenmode behavior. On-axis amplitude and group delay behavior should be ignored. Full-range or coaxial transducers should be optimized to a conical constant directivity index (at least ±3 dB for on-axis amplitude behavior from 300 Hz to 10 kHz) not narrower than 60° × 60°, using additional acoustic lenses or waveguides as needed.
[0024]
[0028] As another example, the design specifications may indicate that the linear behavior of the loudspeaker should be optimized by using digital signal processing with an FIR filter that simultaneously controls the amplitude and phase response of the speaker measured before the addition of the FIR. Measurements shall be performed in an acoustically controlled measurement environment. The final measurement dataset to be corrected may be generated as the average of a series of measurements taken at multiple points between 0° and 30° from the central axis, or it may be taken at a single point that best fits the average of the aforementioned output averages from 0° to 30°. The distance between the microphone and the acoustic center should remain constant across multiple measurements and should be selected such that the output from the speaker transducer is evenly integrated into a coherent wavefront.
[0025]
[0029] The optimization engine can identify points in the acoustic output of a loudspeaker system (or a component within a loudspeaker system) where the output of at least one transducer in the loudspeaker system satisfies a threshold level for wavefront integration.
[0026]
[0030] By taking the output of the driver and cabinet selection and using that output to optimize one or more linear components (such as the coefficients of filters applied to the audio stream by the digital signal processing component without limitation), Method 200 can provide a customized loudspeaker system with optimal performance.
[0027]
[0031] Method 200 includes defining at least one characteristic of at least one hardware component of a loudspeaker system that satisfies the received identification information (204).
[0028]
[0032] Optimizing the linear behavior of a loudspeaker system may involve characterizing the linear behavior using the speaker's impulse response, which can capture both amplitude and timing behavior. The speaker's impulse response represents the speaker's output when a unit impulse is input. The impulse response can be measured by inputting a series of repeating sine sweeps, each covering the relevant frequency range of the corrected loudspeaker system. The speaker system's acoustic output is captured by a microphone. From this, the transfer function (and therefore the impulse response) can be extracted by calculating the difference in amplitude and phase between the input signal and the captured acoustic output.
[0029]
[0033] The measured amplitude and phase of the acoustic output of a loudspeaker system can be determined by both the linear behavior of the system and the position where the output is captured. The distance between the capture point and the loudspeaker can be selected so that the outputs of the individual transducers in the loudspeaker system are sufficiently integrated into a coherent wavefront. When selected in this way, the system can be determined to remain near the minimum distance at which this integration is achieved, because increasing the distance from this minimum can cause the measurement to be affected by the acoustic characteristics of the room in which the measurement is taken. The influence of the room can also be reduced by using a semi-anechoic or fully anechoic chamber.
[0030]
[0034] The system can also consider the angle from the speaker's central axis, i.e., the angle from the axis passing through the speaker's acoustic center and perpendicular to the speaker's front baffle. The angle that the line from the microphone to the acoustic center makes with this central axis can also have a significant impact on the measured response, because the energy decreases at higher frequencies as the angle from the central axis increases. A loudspeaker can be characterized by taking multiple impulse response measurements with a microphone at multiple angles from the central axis ranging from 0° (on the central axis) to 30° (in the horizontal, vertical, or a combination thereof) while maintaining the same distance to the acoustic center, and averaging the measured impulse responses into a single time-domain representation of linear behavior. A loudspeaker can also be characterized by measuring the impulse response at a single angle selected to closely track the average of the speaker's output within the range of 0° to 30°.
[0031]
[0035] Next, referring to Figure 1B, a block diagram shows one embodiment of the loudspeaker specifications in a loudspeaker system generated by the optimization engine 103. In the design example shown in Figure 1B, a full-range transducer covering a passband of 300 Hz to 25 kHz is crossed over to four small low-frequency transducers using a digital 8th-order Linkwitz-Riley filter. The full-range is acoustically isolated from the backwaves of the low-frequency transducers using an extruded square tube sealed at both ends with foam gaskets. Electrical connections to the gasket-sealed amplifier are independently provided for the full-range, top, and bottom pairs of low-frequency transducers. The low-frequency drivers are installed in symmetrical pairs above and below the full-range and consist of 16Ω voice coils, and these drivers are connected in parallel to a single amplifier channel so that the net impedance load is 4Ω. The electronics compartment houses an integrated stereo amplifier module that also supplies auxiliary power to the floating-point digital signal processor and its data converter, regulator, and other peripheral circuits. The full-range driver is rear-mounted inside the front baffle, which is machined into the shape of a flattened spherical conical waveguide for the full-range driver, helping to meet the Paradigm's coverage angle requirements. The driver is designed or selected solely for its nonlinear performance characteristics; for example, the full-range driver would not achieve a satisfactory on-axis amplitude response with conventional designs without proprietary corrective digital signal processing. The onboard floating-point DSP hosts a proprietary set of FIR coefficients for controlling the linear component of the loudspeaker's behavior. These coefficients compensate for the linearity of the group delay sum caused by both the physical alignment of the low-frequency transducer with respect to the loudspeaker's acoustic center and the impedance curve of the driver motor structure, as well as the resulting integral amplitude sum of the hemispherical wavefront.
[0032]
[0036] Referring again to Figure 2, Method 200 comprises optimizing at least one linear behavior of the loudspeaker system, the optimization further comprises optimizing a finite impulse response (FIR) filter of the loudspeaker system, and in combination with the application of the optimized FIR filter, the execution of the loudspeaker system including the specified nonlinear behavior satisfies a threshold level of performance for the loudspeaker system (206).
[0033]
[0037] Optimizing the FIR filter of a loudspeaker system may involve identifying at least one coefficient for use in the mathematical representation of the FIR filter, which includes multiple coefficients, and modifying the FIR filter to include at least one identified coefficient. In one embodiment, using the impulse response measured as described above, the system can represent the impulse response as a vector Xm and identify a filter F that corrects the speaker behavior represented by Xm. If the speaker behavior with the filter is defined as Yt, this is given by the convolution relation: Yt = Xm * F. Since Xm is a known quantity and Yt can be defined to represent the desired behavior of the entire system, this equation can be solved to generate the filter F.
[0034]
[0038] The optimized FIR filter can be stored in the firmware of the loudspeaker within the loudspeaker system. The digital signal processor (DSP) chip in the loudspeaker system can access the FIR filter and apply it to the audio stream. The FIR filter can be applied to the audio stream for playback by the loudspeaker system. The FIR filter can be applied to the audio stream in real time, that is, immediately before or during playback of the audio stream.
[0035]
[0039] Therefore, FIR filters can be customized for one or more loudspeakers in a speaker system exhibiting one or more nonlinear behaviors. Since hosting a DSP within a loudspeaker system is expensive, most loudspeakers are analog, and if a conventional loudspeaker includes a DSP, it does not have sufficient resources to customize the FIR filter applied to the audio stream by the DSP, nor does it have the resources to perform such customization in real time during or before playback. Thus, in contrast to conventional systems that typically do not even support the use of a DSP, the methods and systems described herein are linked to the customization of FIR filters accessed by the DSP, resulting in a loudspeaker design enhanced by such customization. When a DSP applies an optimized FIR filter, the behavior of the entire loudspeaker system, including linear and nonlinear behaviors, can provide an optimized playback experience that meets or exceeds expectations for one or more design specifications. Accordingly, in some embodiments, the optimization engine 103 can generate a design for a loudspeaker system that satisfies one or more nonlinear behavior specifications and includes an optimized FIR filter accessible by a DSP in the loudspeaker system, the execution of which enables optimized performance within the constraints specified by the received design specifications.
[0036]
[0040] The methods and systems described herein may further relate to the ability to apply filters to provide a perceptual rendering of audio optimized for playback through headphones. When listening to music, we often perceive the singer's voice as coming from the center, even though the actual sound is coming from speakers or headphones. This "phantom center" is an example of "imaging," the illusion that sound is coming from a sound source other than the actual physical transducer. Imaging is an important aspect of the listening experience of reproduced sound, whether it is music or sound associated with visual media such as movies or television content. The perceptual experience of imaging depends heavily on how the sound is reproduced. In particular, imaging from speakers typically feels as if the sound is coming from in front of the listener, while imaging from headphones often feels as if the sound is coming from inside the listener's head. In one embodiment, the methods and systems described herein provide a technique for rendering the perceptual experience of listening in front of speakers to headphones by modifying the audio stream with digital signal processing. This technique can be extended to reproduce the experience of sound from any set of sound sources at any location, but for simplicity of explanation, we will start with a single sound source at a single location.
[0037]
[0041] Referring next to Figure 3, the method 300 for perceptually rendering audio for playback by headphones includes optimizing a first finite impulse response (FIR) filter associated with a first channel of audio input of audio associated with a sound source for application to a first transducer of output headphones (302). The method 300 also includes optimizing a second FIR filter associated with a second channel of audio input of audio associated with a sound source for application to a second transducer of output headphones, further including modifying the second FIR filter to include at least one coefficient defined based on the relationship between the first transducer and the second transducer (304).
[0038]
[0042] Next, referring in detail to Figure 3 and in relation to Figures 1A, 1B, and 2, a method 300 for perceptually rendering audio for playback by headphones includes optimizing a first finite impulse response (FIR) filter associated with a first channel of audio input of audio associated with a sound source for application to a first transducer of headphones for output (302).
[0039]
[0043] Method 300 includes optimizing a second FIR filter associated with a second channel of audio input of audio associated with a sound source for application to a second transducer of output headphones, further comprising modifying the second FIR filter to include at least one coefficient defined based on the relationship between the first transducer and the second transducer (304).
[0040]
[0044] Optimizing the first and second FIR filters may involve identifying at least one coefficient for use in the mathematical representation of the optimized FIR filter, which includes multiple coefficients. Optimizing the first and second FIR filters may also involve modifying the optimized FIR filters to include at least one identified coefficient. The optimization may be performed as described above in relation to Figures 1A, 1B, and 2.
[0041]
[0045] The first and second audio input channels can be generated for each sound source.
[0042]
[0046] A sound source can be associated with an object at any position in space. A sound source can be associated with a defined audio channel within a loudspeaker system.
[0043]
[0047] When sound is played on speakers, the transducers are in front of the listener, and the sound from each speaker interacts with the listener's head and reaches both ears. When the same sound is played on headphones, the transducers are to the sides of the ears, and the sound from both sides of the headphones interacts less with the listener's head and reaches only one ear. This difference is an observation that underlies common methods for reproducing the speaker experience on headphones. By measuring the acoustic effect that the listener's head has on sound, it is possible to incorporate this effect into the audio stream before it is played back through headphones, theoretically delivering the same sound to the listener's ears as if it were coming from speakers. This measured effect is known as the head-relationship transfer function (HRTF) and can also be expressed as the head-relationship impulse response (HRIR).
[0044]
[0048] When a listener's HRTF is accurately measured, the auditory illusion created by applying the measured HRTF to an audio stream will be highly accurate for that listener. However, there are individual differences in the size and shape of people's heads and ears, which significantly affect the HRTF. Therefore, an HRTF that reproduces an accurate experience for one listener may not be as effective for another. In the common deployment of HRTF-based spatial audio, a generic HRTF is available. However, if the listener's own HRTF does not closely match the generic HRTF, the imaging will not be rendered accurately. Some platforms offer the option to personalize the HRTF profile, which typically requires imaging or mapping the listener's ears and / or head and then calculating the HRTF profile from those images. The methods and systems described herein do not rely on accurate matching of the HRTF profile and do not require imaging or mapping of the listener's physiological functions.
[0045]
[0049] While many conventional HRTF-based methods for spatial audio directly convolve generic or personalized HRTFs to generate imaging effects, the methods and systems described herein can instead rely on mathematical operations within a plurality of transfer functions that describe the transfer function of the head relative to the sound source as a specific location. Treating these functions as a set of functions under function synthesis allows the optimization engine 103 to represent physical operations on the HRTF and manipulate the transfer functions to optimize the final playback. Some of the transfer functions represent sound propagating through the air over a specific distance, while others in the group may relate to the left or right ear at a specific angle. As will be understood by those skilled in the art, due to the symmetry of the head, the HRTF is generally symmetrical in the left-right direction, so that the transfer function measured at the right ear from a sound source at a specific distance horizontally and vertically from the central axis is equal to the measurement at the left ear at the same angle vertically from the central axis and at the same distance from the sound source, but the horizontal angle is reversed (e.g., 90 degrees to the right is -90 degrees to the left). Accordingly, the optimization engine 103 can specify a function that represents the difference in sound in one ear when the sound source moves from beside the ear to a specified position, and use the difference in response between the left and right ears to determine how to render the output stream while providing a perceptual experience of the sound coming from the sound source and moving from the left ear to the right ear. In other words, to achieve perceptual rendering, the optimization engine 103 can apply a function that represents the relationship between the elements of the HRTF group (e.g., distance from the sound source, left ear, and right ear, as discussed above), use the associative properties of function synthesis to determine the relationship between these elements that represent the sound transmitted as input to the headphones providing input to the left and right ears, and specify a function to apply to the audio for each ear. Thus, the method and implementation of the system described herein can provide a more robust rendering than the rendering provided when using conventional HRTFs, which are attenuated by variations in the listener's head position.
[0046]
[0050] In some embodiments, the transfer functions in the transfer function group can be transformed into impulse responses by using the Fourier transform. Since the functions for the left and right ears are known impulse responses, and given the symmetrical properties of the transfer functions for the ears discussed above, the optimization engine 103 can specify a function that associates a specific location (e.g., sound source) at a particular time point with respect to one ear rather than both ears, then represent that function as an FIR filter, identify the coefficients as described above with respect to Figure 2, and then do the same to identify the coefficients of the FIR filter for the other ear. Since multiple FIR filters can be applied to an audio stream, the DSP can apply multiple FIR filters with optimized coefficients when playing back the audio stream. As an example and without limitation, the above example illustrates identifying the coefficients to use in an FIR filter applied to the sound played for the right ear (e.g., by the right transducer of the headphones) and identifying the coefficients to use in an FIR filter applied to the sound played for the left ear (e.g., by the left transducer of the headphones) when the sound source is at a specific distance from the headphones during a particular time series. Furthermore, since multiple FIR filters may be applied, for each additional audio channel, the optimization engine 103 may generate an FIR filter and associated optimization coefficients (for example, for each of the two transducers in the headphones), and then instruct the DSP applying the FIR filters to apply each of the generated FIR filters to the audio stream.
[0047]
[0051] Accordingly, the optimization engine 103 can optimize at least one coefficient of a first FIR filter associated with at least one channel of the audio input to the headphones for the output of the first transducer of the headphones. The optimization engine 103 can specify a relationship between the reproduction of sound at a specific distance (and / or angle) from the sound source by the first transducer and the reproduction of sound at a distance from the sound source by the second transducer, such that the first and second transducers have a symmetrical relationship with respect to each other. The optimization engine 103 can then optimize at least one coefficient of a second filter associated with at least one channel of the audio input to the headphones for the output of the second transducer of the headphones. In embodiments where the audio input includes multiple channels, the optimization engine 103 can optimize a pair of FIR filters (one for the first transducer and one for the second transducer) for each channel of the audio input. The optimization may be performed as described above with respect to Figure 2.
[0048]
[0052] An FIR filter with optimized coefficients can be applied to a previously recorded audio stream without requiring re-recording in order to benefit from the optimization.
[0049]
[0053] An FIR filter with optimized coefficients can be integrated into the headphone hardware so that the filter is applied to the audio as the headphone transducer reproduces sound for the wearer.
[0050]
[0054] The optimization engine 103 can scalably generate one or more FIR filters having optimized coefficients.
[0051]
[0055] An FIR filter with optimized coefficients can be integrated into a streaming audio platform, for example, as a plug-in to a distribution platform that streams audio, or as a plug-in to a playback application that receives an audio stream. The application of an FIR filter with optimized coefficients can be integrated into a processing step in a method performed when preparing to stream audio to a receiver. As an example, a hardware accelerator can perform the functions of the optimization engine 103. As an example and without limitation, the optimization engine 103 can be provided as a standalone software program or as a plug-in to existing software used in the production stage of a film or music, in which case one use case includes enabling engineers and / or producers to listen to the sound as an end user might hear it and make production decisions accordingly, such engineers and / or producers can listen to the sound from a location further away from the production stage, while requiring less bandwidth than conventional systems typically require.
[0052]
[0056] The methods and systems described herein can be implemented to optimize the FIR filter applied during the audio playback process using headphones, as described above. The methods and systems described herein can also be implemented to optimize the FIR filter applied during the audio playback process using a playback speaker system including multiple speakers.
[0053]
[0057] Therefore, and with reference to Figure 4, a method 400 for perceptually rendering audio for playback by a playback speaker system may include optimizing a first finite impulse response (FIR) filter associated with a first channel of audio input of audio associated with a sound source for application to a first speaker in the playback speaker system for output from a first transducer of a first speaker (402). The method 400 may also include optimizing a second finite impulse response (FIR) filter associated with a second channel of audio input of audio associated with a sound source for application to a second speaker in the playback speaker for output from at least one transducer of a second speaker, further including modifying the second FIR filter to include at least one coefficient defined based on the relationship between the first transducer and the second transducer (404). The optimization of the first and second FIR filters may be carried out as described above with respect to Figures 1 to 3.
[0054]
[0058] In general, in the context of audio processing, including perceptual rendering and HRTFs, functions can be represented as filters, and function synthesis can be represented as convolutions. The precise behavior of a filter can be fully described in either form: in the time domain as an impulse response, i.e., as the coefficients of an FIR filter, or in the frequency domain as a combination of the filter's amplitude and phase response. In the context of perceptual rendering, or audio processing in general, it is sometimes desirable to derive a filter with a specific phase response along with a controlled and / or mitigated amplitude response, or vice versa. This is particularly important in perceptual rendering, where phase behavior is necessary for imaging and spatialization effects, while excessive amplitude deviation can impair the listener's experience. Therefore, the methods and systems described herein include a function for separating the phase and amplitude behavior of a function represented by an FIR filter, and can yield a filter that matches or approximates the phase response of the original filter while controlling and / or mitigating the amplitude response. This technique can also be used to yield a filter that matches or approximates the amplitude response of the original filter while controlling and / or mitigating the phase response.
[0055]
[0059] Filters and functions are equivalently represented in the frequency domain or the time domain, and operations in one domain correspond to operations in the other domain. Time domain representation A t Filter A has a corresponding frequency domain representation. For simplicity, the frequency domain representation of A is given by the ordered pair of its amplitude response and phase response (μ A , φ A ) we consider. A filter with a frequency domain representation (0,0) does not affect either the amplitude or the phase, i.e., it becomes a Dirac impulse.
[0056]
[0060] Convolution in the time domain is equivalent to addition in the frequency domain. That is, for filters B and C = A*B, the following equation holds: μ C =μ A +μ B φ C =φ A+φ B (1) Another operation that the system can perform, particularly on a FIR filter, is time reversal. Reversal in the time domain corresponds to reversal of the phase response in the frequency domain, while the amplitude response remains unchanged. Therefore
Number
Number
Number
Number
[0058]
[0062] When we take the time inversion of B, the corresponding frequency domain representation (μ B ,-φ B ) = -μ A ,φ A ) has
number
number
[0059]
[0063] Dual-amplitude filter M 2 A Similarly, dual-phase filter P 2 A and / or R 2 A This can be applied when they represent stackable effects. If the underlying functions H1 and H2 represent signals propagated through space over a certain distance d, the resulting filter A represents the effect of that distance d on the signal. Thus, a dual-phase filter P 2 A and / or R 2 A P 2 A In this case, the minimum amplitude behavior, or R 2 AIn this case, a phase effect of distance 2d can be represented with defined and controlled amplitude behavior. These filters can also be applied multiple times to further double the represented distance, i.e., P 2 A and / or R 2 A Applying this filter n times represents the phase effect of sound propagating through the air over a distance of n*2d. Such a filter can be applied when perceptually rendering a controllable distance parameter.
[0060]
[0064] In some embodiments, the system performs a convolution route to generate a filter that reflects only one instance of the original phase or amplitude response. Since convolution in the time domain corresponds to addition in the frequency domain, convolving a filter with itself generates a filter whose phase and amplitude responses are doubled. The system can isolate the desired filter by convolving a filter with itself. Generalizing to filters F and G such that G*G=F given F, the system can apply numerical estimation methods, including gradient descent-based methods, genetic / evolutionary algorithms, and general Monte Carlo methods without restriction, to solve the approximation of G, where F is the frequency response (0, 2φ). A ) has P 2 A In this case, the resulting estimated value of G is the approximate frequency response (0, φ A ) has P A This means that F is M 2 A In that case, the system has an approximate frequency response (μ A M having ,0) A An estimate of R can be generated. Since the rectangular function is negative outside the desired band and 0 inside the desired band, rect / 2 = rect, 2 A From the frequency response (rect, φ A ) filter R A The same process can be used to derive the result.
[0061]
[0065] The above phase separation technique can be used for any target response. One technique for generating such a target response in an impulse is to use zero-phase filtering. Dirac Impulse I d and a set of IIR filters F1, F2, ..., F n Starting from there, the system uses each filter F i Impulse I d Apply it twice, once in the forward time direction and once in the reverse time direction, F i By embedding twice the amplitude response into the impulse, the phase response is canceled out, thereby resulting in a filtered impulse I a is the F of the filter i It has twice the amplitude response and zero phase. The same zero-phase filtering process can be applied to the representation of H1 in the inverse derivation method to separate the phase and introduce the intended amplitude response. In both the filter inversion and inverse derivation methods, the zero-phase filtering process is applied to P before taking the convolution route. 2 A It can be applied as a filter instead. Filter F i The choice allows for any amplitude response of the resulting filter. Note that the amplitude is doubled by the zero-phase technique defining the target impulse and then halved again by the subsequent convolution route, so the resulting output filter will have nearly the same amplitude response as collectively introduced by the filter. This is a more generalized and flexible form of the process. The sinc target is filters F1, ..., F n This can be considered a special case of any amplitude target, where the set of filters collectively represents an ideal "brickwall" filter at the cutoff frequency. The target response can also be modified with a series of all-pass filters to introduce the desired phase behavior into the resulting FIR. Combined with the application of a zero-phase filter, the system can achieve arbitrarily defined phase and amplitude behavior.
[0062]
[0066] As those skilled in the art will understand, perceptual rendering typically depends on the accuracy of empirical data. By deriving measurements as described above, the system can use the above functions to (i) identify amplitude and / or phase components, and (ii) modify one or more FIR filters to identify, address, and remove identified components that impair the audio quality level. Accordingly, the methods described herein may include methods for rendering audio for playback by an output device, the methods including optimizing a first FIR filter associated with a first channel of audio input of audio associated with a sound source for application to a first transducer of the output device, optimizing a second FIR filter associated with a second channel of audio input of audio associated with a sound source for application to a second transducer of the output device, further including modifying the second FIR filter to include at least one coefficient defined based on the relationship between the first transducer and the second transducer, and modifying the first FIR filter, the modifications including modifying the amplitude component of the first FIR filter by applying the inverse of the first FIR filter to the first FIR filter. Modifying the amplitude component may include removing the amplitude component. Modifying the amplitude component may include separating the amplitude component.
[0063]
[0067] The methods described herein may further include methods for rendering audio for playback by an output device, the methods including optimizing a first FIR filter associated with a first channel of audio input of audio associated with a sound source for application to a first transducer of the output device, optimizing a second FIR filter associated with a second channel of audio input of audio associated with a sound source for application to a second transducer of the output device, further including modifying the second FIR filter to include at least one coefficient defined based on the relationship between the first transducer and the second transducer, and modifying the second FIR filter, the modifications including modifying the amplitude component of the second FIR filter by applying the inverse of the second FIR filter to the second FIR filter. Modifying the amplitude component may include removing the amplitude component. Modifying the amplitude component may include separating the amplitude component.
[0064]
[0068] Similarly, a method for rendering audio for playback by an output device may include optimizing a first FIR filter associated with a first channel of audio input of audio associated with a sound source for application to a first transducer of the output device; optimizing a second FIR filter associated with a second channel of audio input of audio associated with a sound source for application to a second transducer of the output device, further comprising modifying the second FIR filter to include at least one coefficient defined based on the relationship between the first and second transducers; and modifying the second FIR filter, which includes applying a derivative of the second FIR filter to the second FIR filter to modify the phase component of the second FIR filter. A method for perceptually rendering audio for playback by an output device may include optimizing a first FIR filter associated with a first channel of audio input of audio associated with a sound source for application to a first transducer of the output device, optimizing a second FIR filter associated with a second channel of audio input of audio associated with a sound source for application to a second transducer of the output device, further including modifying the second FIR filter to include at least one coefficient defined based on the relationship between the first transducer and the second transducer, and modifying the first FIR filter, which includes applying a derivative of the first FIR filter to the first FIR filter to modify the phase component of the first FIR filter. Modifying the phase component may include removing the amplitude component. Modifying the phase component may include separating the amplitude component.
[0065]
[0069] The methods described herein may further include methods for rendering audio for playback by an output device, regardless of whether the rendering is perceptual rendering or rendering of another kind. Thus, the methods for rendering may include optimizing a first FIR filter associated with a first channel of audio input of audio associated with a sound source for application to a first transducer of an output device, optimizing a second FIR filter associated with a second channel of audio input of audio associated with a sound source for application to a second transducer of an output device, further including modifying the second FIR filter to include at least one coefficient defined based on the relationship between the first and second transducers, and modifying the second FIR filter, which includes modifying the amplitude components of the second FIR filter by applying the inverse of the second FIR filter to the second FIR filter. The methods described herein may include methods for rendering audio for playback by an output device, the method including optimizing a first FIR filter associated with a first channel of audio input of audio associated with a sound source for application to a first transducer of the output device, optimizing a second FIR filter associated with a second channel of audio input of audio associated with a sound source for application to a second transducer of the output device, further including modifying the second FIR filter to include at least one coefficient defined based on the relationship between the first transducer and the second transducer, and modifying the first FIR filter, the modification including applying the inverse of the first FIR filter to the first FIR filter to modify the amplitude components of the first FIR filter.Similarly, a method for rendering audio for playback by an output device (which does not need to be perceptual rendering) may include optimizing a first finite impulse response (FIR) filter associated with a first channel of audio input of audio associated with a sound source for application to a first transducer of the output device, optimizing a second FIR filter associated with a second channel of audio input of audio associated with a sound source for application to a second transducer of the output device, further comprising modifying the second FIR filter to include at least one coefficient defined based on the relationship between the first and second transducers, and modifying the second FIR filter, which includes applying a derivative of the second FIR filter to the second FIR filter to modify the phase component of the second FIR filter. A method for rendering audio for playback by an output device (regardless of whether the rendering is perceptual rendering) may include optimizing a first finite impulse response (FIR) filter associated with a first channel of audio input of audio associated with a sound source for application to a first transducer of the output device, optimizing a second FIR filter associated with a second channel of audio input of audio associated with a sound source for application to a second transducer of the output device, further including modifying the second FIR filter to include at least one coefficient defined based on the relationship between the first and second transducers, and modifying the first FIR filter, which includes modifying the phase component of the first FIR filter by applying a derivative of the first FIR filter to the first FIR filter. Modifying the phase component may include removing the amplitude component. Modifying the phase component may include separating the amplitude component.
[0066]
[0070] As described herein, filters and functions can be derived based on the intended relationships. The system can select one of several methods to identify one or more solutions or approximate solutions in a time-efficient manner. Unless otherwise specified, approximate solutions or representations are acceptable for any part of the described process.
[0067]
[0071] Accordingly, the methods and systems described herein provide a function to improve the perceptual experience of listening to audio by reproducing sound from one or more sound sources differently by applying different filters having coefficients optimized for playback from different transducers. The methods and systems described herein can provide a function to improve the perceptual experience of listening to audio in games and virtual reality and / or augmented reality applications.
[0068]
[0072] In some embodiments, the system 100 includes a non-temporary computer-readable medium containing computer program instructions stored in tangible form on the non-temporary computer-readable medium, the instructions being executable by at least one processor to perform each of the steps of the method described above.
[0069]
[0073] It should be understood that the systems described above may provide any or more of these components, and these components may be provided on an independent machine or, in some embodiments, on multiple machines in a distributed system. The phrases “in one embodiment” and “in another embodiment” generally mean that the specific features, structures, steps, or characteristics that follow the phrase are included in at least one embodiment of the Disclosure, and may be included in multiple embodiments of the Disclosure. Such phrases may, but not necessarily, refer to the same embodiment. However, the scope of protection is defined by the appended claims, and the embodiments referred to herein are examples.
[0070]
[0074] As used in various embodiments of this disclosure, the terms “A or B,” “at least one of A and B,” “at least one of A or B,” or “one or more of A and B” include any and all combinations of words listed with them. For example, “A or B,” “at least one of A and B,” or “at least one of A or B” may mean (1) including at least one A, (2) including at least one B, (3) including A or B, or (4) including both at least one A and at least one B.
[0071]
[0075] Any step or action disclosed herein as being performed or executable by a computer or other machine may be performed automatically by a computer or other machine, whether or not it is expressly disclosed herein as such. Steps or actions performed automatically may be performed solely by a computer or other machine without human intervention. Steps or actions performed automatically may operate based solely on input received from a computer or other machine, and not from a human being, for example. Steps or actions performed automatically may be initiated by signals received from a computer or other machine, and not from a human being, for example. Steps or actions performed automatically may provide output to a computer or other machine, and not from a human being, for example.
[0072]
[0076] In this specification, terms such as "optimize" and "optimal" may be used, but in practice, embodiments of the present invention may include methods for producing an output that is not optimal or is not known to be optimal, but is still useful. For example, embodiments of the present invention can produce an output that approximates the optimal solution within a certain range of error. Consequently, terms such as "optimize" and "optimal" in this specification should be understood to refer not only to processes that produce the optimal output, but also to processes that produce an output that approximates the optimal solution within a certain range of error.
[0073]
[0077] The systems and methods described above can be implemented as methods, devices, or products using programming and / or engineering techniques to generate software, firmware, hardware, or any combination thereof. The techniques described above can be implemented by one or more computer programs running on a programmable computer including a processor, a storage medium readable by the processor (including, for example, volatile and non-volatile memory and / or memory elements), at least one input device, and at least one output device. To perform the described functions and generate outputs, program code can be applied to inputs received using the input device. Outputs can be provided to one or more output devices.
[0074]
[0078] Each computer program included in the attached claims may be implemented in any programming language, such as assembly language, machine language, a high-level procedural programming language, or an object-oriented programming language. The programming language may be, for example, LISP, PROLOG, PERL, C, C++, C#, JAVA, Python, Rust, Go, or any compiled or interpreted programming language.
[0075]
[0079] Each such computer program may be implemented by a computer program product that is tangibly embodied in a machine-readable storage device for execution by a computer processor. The steps of the method may be executed by a computer processor that runs a program tangibly embodied on a computer-readable medium in order to perform the functions of the method and system described herein by operating on inputs and producing outputs. Suitable processors include, for example, both general-purpose microprocessors and dedicated microprocessors. Generally, processors receive instructions and data from read-only memory and / or random-access memory. Storage devices suitable for tangibly embodiing computer program instructions include, for example, all forms of computer-readable devices, firmware, programmable logic, and hardware (e.g., integrated circuit chips, electronic devices, computer-readable non-volatile storage devices, non-volatile memory such as semiconductor memory devices including EPROMs, EEPROMs, and flash memory devices, magnetic disks such as built-in hard disks and removable disks, magneto-optical disks, and CD-ROMs). Any of the above can be supplemented by or incorporated into specially designed ASICs (Application-Specific Integrated Circuits) or FPGAs (Rewritable Gate Arrays). Computers can also generally receive programs and data from storage media such as internal disks (not shown) or removable disks. These elements are found in conventional desktop or workstation computers, as well as in other computers suitable for running computer programs that implement the methods described herein, and such computers can be used with any digital print engine or marking engine, display monitor, or other raster output device capable of producing color or grayscale pixels on paper, film, display screen, or other output media. Computers can also receive programs and data (including, for example, instructions for storage on non-temporary computer-readable media) from a second computer that provides access to the program via network transmission lines, wireless transmission media, spatially propagating signals, radio waves, infrared signals, etc.
[0076]
[0080] While certain embodiments of methods and systems for independently optimizing linear components from nonlinear components in a loudspeaker system, and for perceptually rendering audio for playback, will be apparent to those skilled in the art that other embodiments incorporating the concepts of the present disclosure may be used. Therefore, the present disclosure should not be limited to certain embodiments, but rather should be limited only by the spirit and scope of the appended claims.
Claims
1. A method for perceptually rendering audio for playback through headphones, To optimize a first finite impulse response (FIR) filter associated with a first channel of audio input of audio associated with a sound source, for application to a first transducer of output headphones, and Optimizing a second FIR filter associated with a second channel of the audio input of the audio associated with the sound source for application to a second transducer of the output headphones, further comprising modifying the second FIR filter to include at least one coefficient defined based on the relationship between the first transducer and the second transducer. Methods that include...
2. The first FIR filter can be optimized as follows: Identifying at least one coefficient for use in the mathematical representation of the first FIR filter which includes multiple coefficients, and Modifying the first FIR filter to include at least one identified coefficient. The method according to claim 1, further comprising:
3. The method according to claim 1, wherein the first channel and the second channel of the audio input are generated for each sound source.
4. The method according to claim 1, wherein the sound source is related to an object located at any position in space.
5. The method according to claim 1, wherein the sound source is associated with a defined audio channel in a loudspeaker system.
6. The method according to claim 1, further comprising modifying the first FIR filter, which includes applying the inverse of the first FIR filter to the first FIR filter to modify the amplitude component of the first FIR filter.
7. The method according to claim 1, further comprising modifying the second FIR filter, which includes applying the inverse of the second FIR filter to the second FIR filter to modify the amplitude component of the second FIR filter.
8. The method according to claim 1, further comprising modifying the first FIR filter, which includes applying a derivative of the first FIR filter to the first FIR filter to modify the phase component of the first FIR filter.
9. The method according to claim 1, further comprising modifying the second FIR filter, which includes applying a derivative of the second FIR filter to the second FIR filter to modify the phase component of the second FIR filter.
10. A method for perceptually rendering audio for playback by a playback speaker system, To optimize a first finite impulse response (FIR) filter associated with a first channel of audio input of audio associated with a sound source, for application to the first speaker in a playback speaker system for the output of the first transducer of the first speaker, and Optimizing a second finite impulse response (FIR) filter associated with a second channel of the audio input of the audio associated with the sound source, for application to the second speaker in the playback speaker for the output of at least one transducer of the second speaker, further comprising modifying the second FIR filter to include at least one coefficient defined based on the relationship between the first transducer and the second transducer. Methods that include...
11. Optimizing a third finite impulse response (FIR) filter for a third speaker in the playback speaker for output from at least one transducer of the third speaker, the optimization relating to a third channel of the audio input of the audio relating to the sound source, further comprising modifying the third FIR filter to include at least one coefficient defined based on the relationship between the first transducer and the second transducer. The method according to claim 10, further comprising:
12. The method according to claim 10, wherein the first channel and the second channel of the audio input are generated for each sound source.
13. The method according to claim 10, wherein the sound source is related to an object located at any position in space.
14. The method according to claim 10, wherein the sound source is associated with a defined audio channel in a loudspeaker system.
15. The method according to claim 10, further comprising modifying the first FIR filter, which includes modifying the amplitude component of the first FIR filter by applying the inverse of the first FIR filter to the first FIR filter.
16. The method according to claim 10, further comprising modifying the second FIR filter, which includes modifying the amplitude component of the second FIR filter by applying the inverse of the second FIR filter to the second FIR filter.
17. The method according to claim 10, further comprising modifying the first FIR filter, which includes applying a derivative of the first FIR filter to the first FIR filter to modify the phase component of the first FIR filter.
18. The method according to claim 10, further comprising modifying the second FIR filter, which includes applying a derivative of the second FIR filter to the second FIR filter to modify the phase component of the second FIR filter.