Hybrid Near / Far-Field Speaker Virtualization
The hybrid near-field/far-field speaker virtualization method addresses the limitations of home theater systems by generating synchronized near-field and far-field signals for enhanced spatial audio reproduction, overcoming the constraints of limited speaker configurations.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2025-12-03
- Publication Date
- 2026-03-10
AI Technical Summary
Home theater systems struggle to accurately reproduce three-dimensional sound due to limited speaker configurations, which hinder the creation of a profound sense of nearness or distance from the listener, and traditional speaker virtualization algorithms fall short in reproducing certain 3D sound elements using stereo or 5.1 surround systems.
A hybrid near-field/far-field speaker virtualization method that generates near-field and far-field signals based on source audio, using weighted linear combinations and speaker characteristics, and synchronizes these signals for playback through near-field and far-field speakers to enhance spatial audio reproduction.
Enhances the listening experience by adding height, depth, and spatial information, ensuring synchronized and artifact-free playback of audio content using a combination of near-field and far-field speakers.
Smart Images

Figure 2026041871000001_ABST
Abstract
Description
[Technical Field]
[0001] CROSS-REFERENCE TO RELATED APPLICATIONS This application claims priority to U.S. Provisional Application No. 62 / 903,975, filed September 23, 2019, U.S. Provisional Application No. 62 / 904,027, filed September 23, 2019, and U.S. Provisional Application No. 63 / 077,517, filed September 11, 2020, each of which is incorporated by reference herein in its entirety.
[0002] Technical Field The present disclosure relates generally to audio signal processing. [Background technology]
[0003] A typical cinema soundtrack contains many different sound elements corresponding to on-screen, off-screen, unseen implied elements and images, dialogue, noises, and sound effects, which emanate from different on-screen elements and are combined with background music and environmental effects to create an overall audience experience. The artistic intent of the creators and producers represents a desire to have these sounds reproduced in a manner that corresponds as closely as possible to what is shown on the screen in terms of source location, intensity, movement, and other similar parameters.
[0004] Traditional channel-based audio systems send audio content in the form of speaker feeds to individual speakers in a playback environment, such as a stereo or 5.1 system. To further improve the listener experience, some home theater systems use object-based audio to provide a three-dimensional (3D) spatial presentation of sound using audio objects. An audio object is an audio signal with an associated parametric source description of apparent source location (e.g., 3D coordinates), apparent source width, and other parameters.
[0005] Home theater systems have fewer speakers than movie theaters and are therefore less able to reproduce 3D sound according to the creator's artistic intent. In fact, a drawback of all listening environments is the peripheral nature of the listening environment, which limits their ability to create a profound sense of nearness or distance from the listener. Speaker virtualization algorithms are often used in home theater systems to reproduce sound at various locations in the playback environment where no physical speakers exist. However, some 3D sound cannot be reproduced using stereo speakers alone, or even a 5.1 surround system. These are the most common speaker layouts found in home theater systems. Summary of the Invention [Means for solving the problem]
[0006] An embodiment for hybrid near-field / far-field speaker virtualization is disclosed. In one embodiment, a method includes receiving, using a media source device, a source signal including at least one of channel-based audio or audio objects; generating, using the media source device, one or more near-field gains and one or more far-field gains based on the source signal and a mixing mode; generating, using the media source device, a far-field signal based, at least in part, on the source signal and the one or more far-field gains; rendering, using a speaker virtualizer, the far-field signal into an audio playback environment for reproduction of far-field acoustic audio through far-field speakers; generating, using the media source device, a near-field signal based on the source signal and the one or more near-field gains; transmitting the near-field signal to a near-field reproduction device or an intermediate device coupled to the near-field reproduction device before providing the far-field signal to the far-field speaker; and providing the far-field signal to the far-field speaker.
[0007] In one embodiment, the method further includes: filtering the source signal into a low-frequency signal and a high-frequency signal; generating two sets of near-field gains including a near-field low-frequency gain and a near-field high-frequency gain; generating two sets of far-field gains including a far-field low-frequency gain and a far-field high-frequency gain; generating the near-field signal based on a weighted linear combination of the low-frequency signal and the high-frequency signal, wherein the low-frequency signal is weighted by the near-field low-frequency gain and the high-frequency signal is weighted by the near-field high-frequency gain; and generating the far-field signal based on a weighted linear combination of the low-frequency signal and the high-frequency signal, wherein the low-frequency signal is weighted by the far-field low-frequency gain and the high-frequency signal is weighted by the far-field high-frequency gain.
[0008] In one embodiment, the mixed mode is based, at least in part, on the layout of the far-field speakers in the audio reproduction environment and one or more characteristics of the far-field speakers or the near-field speakers coupled to the near-field reproduction device.
[0009] In one embodiment, the mixing mode is surround sound rendering, and the method further includes: setting the one or more near field gains and the one or more far field gains to include all surround channel-based audio or surround audio objects in the near field signal and all front channel-based audio or front audio objects in the far field signal.
[0010] In one embodiment, the method further includes: determining, based on the near-field and far-field speaker characteristics, that the far-field speaker is more capable of reproducing low frequencies than the near-field speaker; and: setting the one or more near-field gains and the one or more far-field gains to include all low-frequency channel-based audio or low-frequency audio objects in the far-field signal.
[0011] In one embodiment, the method further includes determining that the source signal includes a distance effect; and setting the one or more near-field gains and the one or more far-field gains to be a function of a normalized distance between the far-field speaker and a specified position in the audio reproduction environment.
[0012] In one embodiment, the method further includes: determining that the source signal includes channel-based audio or audio objects for enhancing a particular type of audio content in the source signal; and setting the one or more near-field gains and the one or more far-field gains to include the channel-based audio or audio objects for enhancing the particular type of audio content in the near-field signal.
[0013] In one embodiment, the particular type of audio content is dialogue content.
[0014] In one embodiment, the source signal is received along with metadata including the one or more near field gains and the one or more far field gains.
[0015] In one embodiment, the metadata includes data indicating that the source signal can be used for hybrid speaker virtualization using the far-field speakers and the near-field speakers.
[0016] In an embodiment, the near field signal, or the rendered near field signal, and the rendered far field signal include an inaudible marker signal to assist in synchronous overlay of the near field acoustic audio with the far field acoustic audio.
[0017] In one embodiment, the method further comprises: acquiring head pose information of a user in the audio playback environment; and rendering the near-field signal using the head pose information.
[0018] In one embodiment, equalization is applied to the rendered near-field signal to compensate for the frequency response of the near-field speakers.
[0019] In one embodiment, the near field signal or the rendered near field signal is provided to the near field reproduction device over a wireless channel.
[0020] In one embodiment, the step of providing the near field signal or the rendered near field signal to the near field reproduction device further comprises: using the media source device to transmit the near field signal or the rendered near field signal to an intermediate device coupled to the near field reproduction device.
[0021] In one embodiment, equalization is applied to the rendered far-field signal to compensate for the frequency response of the near-field speakers.
[0022] In one embodiment, to assist in synchronous overlay of the near-field audio with the far-field audio, a timestamp associated with the near-field signal or a rendered near-field signal is provided by the media source device to the near-field playback device or an intermediate device.
[0023] In one embodiment, generating the far-field signal and the near-field signal based, at least in part, on the source signal and the one or more far-field gains further includes: storing the source signal in a buffer of the media source device; retrieving a first set of frames of the source signal stored in a first location in the buffer, the first location corresponding to a first time; generating, using the media source device, the far-field signal based, at least in part, on the first set of frames and the one or more far-field gains; retrieving a second set of frames of the source signal stored in a second location in the buffer, the second location corresponding to a second time earlier than the first location; and generating, using the media source device, the near-field signal based, at least in part, on the second set of frames and the one or more near-field gains.
[0024] In one embodiment, a method includes the steps of: receiving, in an audio reproduction environment, near-field signals transmitted by a media source device, the near-field signals including a weighted linear combination of low-frequency and high-frequency channel-based audio or audio objects for projection through near-field speakers located in the audio reproduction environment and proximate to or inserted into a user's ears; converting, using one or more processors, the near-field signals into digital near-field data; buffering, using the one or more processors, the digital near-field data; capturing, using one or more microphones, far-field acoustic audio projected by far-field speakers; converting the far-field audio into digital far-field data using one or more processors; buffering the digital far-field data using the one or more processors; determining a time offset using the one or more processors and buffer contents; adding a set of local time offsets to the time offsets using the one or more processors to generate a total time offset; and using the one or more processors to start playing the near-field data through the near-field speakers using the total time offset, thereby causing near-field sound data projected by the near-field speakers to be synchronously overlaid with the far-field sound audio.
[0025] In one embodiment, a method includes: receiving, using a media source device, a source signal including at least one of channel-based audio or audio objects; generating, using the media source device, a far-field signal based, at least in part, on the source signal; rendering, using the media source device, the far-field signal for playback through far-field speakers into an audio playback environment; generating, using the media source device, one or more near-field signals based, at least in part, on the source signal; transmitting the near-field signals to a near-field playback device or an intermediate device coupled to the near-field speakers before providing the far-field signals to the far-field speakers; and providing the rendered far-field signals to the far-field speakers for projection into the audio playback environment.
[0026] In one embodiment, the near field signal comprises enhanced dialogue.
[0027] In one embodiment, there are at least two near-field signals sent to the near-field reproduction device or the intermediate device, a first near-field signal being rendered into near-field acoustic audio for playback through a near-field speaker of the near-field device, and a second near-field signal being used to assist in synchronizing the far-field acoustic audio with the first near-field signal.
[0028] In one embodiment, there are at least two near-field signals sent to the near-field reproduction device, a first near-field signal containing dialogue content in a first language and a second near-field signal containing dialogue content in a second language different from the first language.
[0029] In some embodiments, the near-field signal and the rendered far-field signal include inaudible marker signals to assist in synchronous overlay of the near-field acoustic audio with the far-field acoustic audio.
[0030] In one embodiment, the method further includes the steps of: receiving near-field signals transmitted by a media source device in an audio playback environment using a wireless receiver; converting the near-field signals into digital near-field data using one or more processors; buffering the digital near-field data using the one or more processors; capturing far-field acoustic audio projected by far-field speakers using one or more microphones; converting the far-field acoustic audio into digital far-field data using the one or more processors; buffering the digital far-field data using the one or more processors; determining a time offset using the one or more processors and the buffer contents; adding a set of local time offsets to the time offsets using the one or more processors to generate a total time offset; and using the one or more processors to start playing the near-field data through the near-field speakers using the total time offset, whereby the near-field acoustic data projected by the near-field speakers is synchronously overlaid with the far-field acoustic audio.
[0031] In one embodiment, the method further includes: capturing a target sound from the audio playback environment using one or more microphones of the near-field playback device; converting the captured target sound into digital data using the one or more processors; generating an anti-sound by inverting the digital data using a filter that approximates an electro-acoustic transfer function using the one or more processors; and canceling the target sound using the anti-sound using the one or more processors.
[0032] In one embodiment, the far-field acoustic audio includes a first dialogue in a first language that is a target voice, and the canceled first dialogue is replaced with a second dialogue in a second language that is different from the first language, and the second language dialogue is included in the secondary near-field signal.
[0033] In one embodiment, the far-field acoustic audio includes a first commentary that is a target audio, and the canceled first commentary is replaced with a second commentary that is different from the first commentary, and the second commentary is included in a secondary near-field signal.
[0034] In one embodiment, the far-field sound audio is the target sound cancelled by the anti-sound to mute the far-field sound audio.
[0035] In one embodiment, differences between the cinema rendering and the near-field playback device rendering of one or more audio objects are included in the near-field signal and used to render the near-field sound audio, whereby the one or more audio objects included in the cinema rendering but not the near-field playback device rendering are excluded from the near-field sound audio rendering.
[0036] In one embodiment, weighting is applied as a function of object-to-listener distance in the audio reproduction environment, so that one or more specific sounds intended to be heard closely by the listener are conveyed only in the near-field signal, which is used to cancel the same specific sound or sounds in the far-field acoustic audio.
[0037] In one embodiment, the near-field signals are modified by the listener's head-related transfer functions (HRTFs) to provide enhanced spatiality.
[0038] In one embodiment, an apparatus comprises: one or more processors; and a memory storing instructions that, when executed by the one or more processors, cause the one or more processors to perform any of the aforementioned methods.
[0039] In one embodiment, a non-transitory computer-readable storage medium having stored thereon instructions that, when executed by one or more processors, cause the one or more processors to perform any of the aforementioned methods.
[0040] Certain embodiments disclosed herein provide one or more of the following advantages: An audio playback system that includes near-field and far-field speaker virtualization enhances a user's listening experience by adding height, depth, or other spatial information that is missing, incomplete, or imperceptible when audio is rendered for playback using only far-field speakers. [Brief explanation of the drawings]
[0041] In the accompanying drawings referenced below, various embodiments are illustrated as block diagrams, flowcharts, and other figures. Each block in a flowchart or block may represent a module, program, or portion of code, including one or more executable instructions for performing specified logical functions. These blocks are shown in a particular sequence for performing method steps, but they need not necessarily be performed in the exact sequence illustrated. For example, they may be performed in reverse sequence or simultaneously, depending on the nature of the respective operations. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations thereof, may be implemented by a dedicated software-based or hardware-based system for performing the specified functions / operations, or by a combination of dedicated hardware and computer instructions.
[0042] [Figure 1] 1 illustrates an audio playback environment including hybrid near-field / far-field speaker virtualization for audio enhancement, according to one embodiment.
[0043] [Figure 2] FIG. 1 is a flow diagram of a processing pipeline for hybrid near-field / far-field speaker virtualization for audio enhancement, according to one embodiment.
[0044] [Figure 3] 1 illustrates a timeline for wireless transmission of near field signals, including early transmission of near field signals, according to an embodiment.
[0045] [Figure 4A] FIG. 2 is a block diagram of a processing pipeline for determining a total time offset for synchronizing playback of near-field sound audio with far-field sound audio, according to an embodiment.
[0046] [Figure 4B] FIG. 2 is a block diagram of a processing pipeline for synchronizing playback of near-field sound audio with far-field sound audio, according to an embodiment.
[0047] [Figure 5] FIG. 1 is a flow diagram of a process for hybrid near / far-field speaker virtualization for audio enhancement, according to one embodiment.
[0048] [Figure 6] FIG. 1 is a flow diagram of a process for synchronizing playback of near-field sound audio with far-field sound audio, according to an embodiment.
[0049] [Figure 7] FIG. 10 is a flow diagram of an alternative process for synchronizing playback of near-field sound audio with far-field sound audio, according to an embodiment.
[0050] [Figure 8] FIG. 10 is a flow diagram of another alternative process for synchronizing playback of near-field sound audio with far-field sound audio, according to an embodiment.
[0051] [Figure 9] FIG. 7 is a block diagram of a media source device architecture for implementing the features and processes described with reference to FIGS. 1-6, according to one embodiment.
[0052] [Figure 10] FIG. 7 is a block diagram of a near field reproduction device architecture for implementing the features and processes described with reference to FIGS. 1-6, according to one embodiment.
[0053] The same reference symbols used in the various drawings indicate like elements. DETAILED DESCRIPTION OF THE INVENTION
[0054] Nomenclature and definitions The following description is directed to certain implementations for purposes of describing some innovative aspects of the present disclosure and examples of contexts in which these innovative aspects may be implemented. However, the teachings herein may be applied in a variety of different ways. Moreover, the described embodiments may be implemented in a variety of hardware, software, firmware, etc. For example, aspects of the present application may be embodied, at least in part, in an apparatus, a system including two or more devices, a method, a computer program product, etc.
[0055] Accordingly, aspects of the present application may take the form of hardware, software (including firmware, resident software, microcode, etc.), and / or a combination of software and hardware. The disclosed embodiments may be referred to herein as "circuits," "modules," or "engines." Some aspects of the present application may take the form of a computer program product embodied in one or more non-transitory medium(s) having computer-readable program code embodied therein. Such non-transitory medium(s) may include, for example, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. Thus, the teachings of the present disclosure are not intended to be limited to the implementations shown in the drawings and / or described herein, but have broad applicability.
[0056] As used herein, the following terms have the following associated meanings: The term "channel" refers to an audio signal plus metadata where the location is coded as a channel identifier (eg, left front or right top surround).
[0057] The term "channel-based audio" refers to audio formatted for playback through a predefined set of speaker zones (e.g., 5.1, 7.1, 9.1, etc.) with associated nominal positions.
[0058] The term "audio object" or "object-based audio" means one or more audio signals with parametric source descriptions, such as apparent source position (e.g., 3D coordinates), apparent source width, etc.
[0059] The term "audio playback environment" means any open, partially enclosed, or fully enclosed area, such as a room, that can be used for the playback of audio content, alone or together with video or other content, and may be embodied in a home, cinema, theater, auditorium, studio, game console, etc.
[0060] The term "rendering" refers to mapping audio object position data to specific channels.
[0061] The term "binaural" rendering means that left and right (L / R) binaural signals are sent to the L / R ears. Binaural rendering can use generic or personalized head-related transfer functions (HRTFs), and aspects of HRTFs, such as interaural level and time differences, to improve the sense of spatialization.
[0062] The term "media source device" refers to any device that plays media content (e.g., audio, video) contained in a bitstream or stored on a medium (e.g., Ultra-HD or Blu-ray, DVD), including, but not limited to, television systems, set-top boxes, digital media receivers, surround sound systems, portable computers, tablet computers, etc.
[0063] The term "far-field speaker" refers to any loudspeaker that is wired or wirelessly connected to a media source device, that is located at a fixed physical location in the audio playback environment, and that is not located near or inserted into the listener's ears, including, but not limited to, stereo speakers, surround speakers, low frequency enhancement (LFE) devices, sound bars, etc.
[0064] The term "near-field speaker" refers to any loudspeaker that is embedded in or coupled to a near-field reproduction device and is positioned near or inserted into the listener's ears.
[0065] The term "near-field reproduction device" refers to any device that includes or is coupled to a near-field speaker, including, but not limited to, headphones, earbuds, headsets, earphones, smart glasses, game controllers / devices, augmented reality (AR), virtual reality (VR) headsets, hearing aids, bone conduction devices, or any other means of providing sound in close proximity to a user's ears. The near-field reproduction device may be two devices, for example, a pair of truly wireless earbuds. Alternatively, the near-field reproduction device may be a single device for use with two ears, such as a pair of headphones with two ear cups. The near-field reproduction device may also be designed for use with only one ear.
[0066] In some embodiments, the near-field reproduction device includes at least one microphone for capturing sounds near the user, which may include far-field acoustic audio. There may be one microphone for each ear. The microphone may be located at a central point, such as on a headphone band on the head, or at a central point where wires from each ear converge. There may also be multiple microphones, for example, one inside or near each ear.
[0067] In some embodiments, the near-field reproduction device may include conventional elements for performing signal processing on microphone and other audio data, including an analog-to-digital converter (ADC), a central processing unit (CPU), a digital signal processor (DSP), and memory. The near-field reproduction device may also include conventional elements for audio reproduction, such as a digital-to-analog converter (DAC) and an amplifier.
[0068] In one embodiment, the near-field reproduction device includes at least one near-field speaker, ideally one near-field speaker proximate each ear, which may include a balanced armature, a traditional dynamic driver, or a bone conduction transducer.
[0069] In one embodiment, the near-field playback device includes a link to a media source system device or an intermediate device (e.g., a personal mobile device) for receipt of the near-field signal. The link may be a radio frequency (RF) link such as Wi-Fi, Bluetooth, or Bluetooth Low Energy (BLE), or the link may be a wire.
[0070] In one embodiment, near-field signals are transmitted over the link in many well-known formats, such as analog signals or digitally encoded signals, which may be encoded using codecs such as Opus, AAC, or G.772 to reduce the required data bandwidth.
[0071] In one embodiment, the near-field reproduction device may make microphone measurements of ambient audio, including far-field sound audio (defined below), while also receiving near-field signals over the link. Using signal processing (described below), the near-field reproduction device can determine a time offset between the far-field sound audio and the near-field sound audio (defined below). The time offset is then used to reproduce the near-field sound audio from the near-field speakers, synchronously overlapping with the far-field sound audio projected by the far-field speakers into the audio reproduction environment.
[0072] The term "intermediate device" refers to a device coupled between a media source device and a near-field reproduction device and configured to process and / or render audio signals received from the media source device and transmit the processed / rendered audio signals to the near-field reproduction device via a wired or wireless connection.
[0073] In one embodiment, the intermediate device is a personal mobile device such as a smartphone, which typically includes a larger battery and more computing power than can fit within the near-field reproduction device. Thus, the personal device can be conveniently used in conjunction with the near-field reproduction device to reduce the power required by the near-field reproduction device and thereby extend its battery life. To this end, some of the components within the near-field reproduction device may be preferentially located within the personal mobile device.
[0074] For example, if the link between the near-field reproduction device and the personal mobile device is a wire, the ear device may not require an ADC, CPU or DSP, DAC, or amplifier because the microphone and speaker signals are measured, processed, or generated entirely within the personal mobile device and transmitted along the wire. In this case, the near-field reproduction device may be similar to headphones with a microphone. If the simple headphones do not have a microphone, it may be possible to measure far-field audio using a microphone on the personal mobile device. However, this is not ideal because users often carry their mobile devices in their pockets or bags, which muffles the far-field audio.
[0075] If the communication link between the near-field reproduction device and the personal mobile device is wireless, the near-field reproduction device may include components for signal measurement, processing, and generation. Depending on the relative power efficiency of computation versus communication over the link, it may be more power-efficient to keep all signal processing within the ear device or to continually offload measurements to the personal mobile device for processing. While the overall system has the computational power to perform signal processing, this power may be distributed among the components.
[0076] In one embodiment, the personal mobile device can receive a near-field signal from an entertainment device via a relatively energy-intensive RF protocol and retransmit it to the near-field reproduction device via a relatively low-energy protocol. Some examples of high-energy protocols include cellular radio and WiFi. Some examples of relatively low-energy protocols include Bluetooth and Bluetooth Low Energy (BLE). If the near-field reproduction device is a wired headphone, the personal mobile device can receive a secondary stream from the entertainment device via the RF protocol and transmit it via a wire to the near-field reproduction device.
[0077] In some embodiments, the personal mobile device may provide a screen or control for a graphical user interface (GUI).
[0078] In one embodiment, the personal mobile device may be a charging carry case for the near field reproduction device.
[0079] The term "source signal" includes a bitstream of audio content or audio and other content (e.g., audio and video), where the audio content may include frames of audio samples and associated metadata, and each audio sample is associated with a channel (e.g., left, right, center, surround) or audio object. The audio content may include, for example, music, dialogue, and sound effects.
[0080] "Far-field acoustic audio" means audio projected from far-field loudspeakers into an audio reproduction environment.
[0081] The term "near-field acoustic audio" refers to audio that is projected from near-field speakers into the user's ears (e.g., earbuds) or in close proximity to the user's ears (e.g., headphones).
[0082] Overview The following detailed description is directed to hybrid near-field / far-field speaker virtualization for audio enhancement. In one embodiment, a media source device located in an audio playback environment receives a time-domain source signal containing channel-based audio, object-based audio, or a combination of channel-based audio and object-based audio. A crossover filter within the media source device filters the source signal into low-frequency and high-frequency time-domain signals. Near-field and far-field signals are generated that are weighted linear combinations of the low-frequency and high-frequency time-domain signals, and the contributions of the low-frequency and high-frequency time-domain signals to the near-field and far-field signals are determined by sets of near-field and far-field gains, respectively. In one embodiment, the gains are generated by a blending algorithm that takes into account the far-field speaker layout and the characteristics of the far-field and near-field speakers.
[0083] The near-field and far-field signals are routed to near-field and far-field audio processing pipelines, respectively, where the signals are rendered into near-field and far-field signals that optionally undergo post-processing treatment such as equalization or compression. In one embodiment, low-frequency content (e.g., <40 Hz) is cross-filtered and sent directly to the LFE unit, bypassing the near-field and far-field signal processing pipelines.
[0084] After any post-processing has been applied, the rendered far-field signal is fed to a far-field speaker feed, resulting in far-field acoustic audio being projected into the audio reproduction environment. Prior to projection of the far-field acoustic audio, and after any post-processing has been applied, the rendered near-field signal is fed to a wireless transmitter for wireless transmission to a near-field reproduction device for playback through near-field speakers. The near-field speakers project near-field acoustic audio that is overlaid on and synchronized with the far-field acoustic audio.
[0085] In some embodiments, the rendered near-field signal is received by the intermediate device over a first wireless communication link (e.g., a WiFi or Bluetooth communication link) and further processed before being transmitted to the near-field reproduction device over a second wireless communication channel (e.g., a Bluetooth channel). In some embodiments, the near-field signal is rendered by the near-field reproduction device or the intermediate device rather than by the media source device.
[0086] In one embodiment, the total time offset used to synchronize the far-field audio and the near-field audio is calculated at the near-field reproduction device or the intermediate device. For example, multiple samples of the far-field audio may be captured by one or more microphones at the near-field reproduction device or the intermediate device and stored in a first buffer at the near-field reproduction device or the intermediate device. Similarly, multiple samples of the rendered (or unrendered) near-field signal received over the wireless link may be stored in a second buffer at the near-field reproduction device or the intermediate device. The first and second buffer contents are then correlated to determine the time offset between the two signals.
[0087] In one embodiment, a local time offset is calculated that accounts for local signal processing at the near-field reproduction device and / or intermediate device, and the time required to transmit audio from the intermediate device to the near-field reproduction device over a wireless communication channel. The local time offset is added to the time offset resulting from the correlation to determine a total time offset. The total time offset is then used to synchronize near-field acoustic audio with far-field acoustic audio for substantially artifact-free enhanced audio reproduction.
[0088] Example of an illustrative audio playback environment FIG. 1 illustrates an audio reproduction environment 100 including hybrid near-field / far-field speaker virtualization for audio enhancement, according to one embodiment. The audio reproduction environment 100 includes a media source device 101, a far-field speaker 102, an LFE device 108, an intermediate device 110, and a near-field reproduction device 105. One or more microphones 107 are attached to or embedded in the near-field reproduction device 105 and / or the intermediate device 110. A wireless transceiver 106 is shown attached to or embedded in the near-field reproduction device 105, while wireless transceivers 103 and 109 are shown attached to or embedded in the far-field speaker 102 (or alternatively, the media source device 101) and the LFE device 108, respectively. A wireless transceiver (not shown) is embedded in the intermediate device 110.
[0089] It should be understood that audio reproduction environment 100 is merely one example environment for hybrid far-field speaker virtualization, and that other audio reproduction environments are applicable to the disclosed embodiments, including, but not limited to, more or fewer speakers, different types of speakers or speaker arrays, more or fewer microphones, and more or fewer (or different) near-field reproduction devices or intermediate devices. For example, audio reproduction environment 100 can be a gaming environment with multiple players, each with their own near-field reproduction device.
[0090] In FIG. 1 , users 104 are viewing media content (e.g., a movie) played through media source device 101 (e.g., a television) and far-field speakers 102 (e.g., a sound bar), respectively. The media content is contained in frames of a source signal that includes a combination of channels and audio objects. In one embodiment, the source signal can be provided over a wide area network (e.g., the Internet) coupled to a digital media receiver (not shown) through a WiFi connection. The digital media receiver (DMR) is coupled to media source device 101 using, for example, an HDMI port and / or an optical link. In another embodiment, the source signal can be received into a television set-top box and into media source device 101 through a coaxial cable. In yet another embodiment, the source signal is extracted from a broadcast signal received through an antenna or satellite dish. In other embodiments, a media player provides the source signal, and the source signal is retrieved from a storage medium (e.g., an Ultra-HD, Blu-ray, or DVD disc) and provided to media source device 101.
[0091] During playback of the source signal, the far-field speaker 102 projects far-field acoustic audio into the audio reproduction environment 100. Additionally, low-frequency content (e.g., sub-bass frequency content) in the source signal is provided to the LFE unit 108, which in this example is "paired" with the far-field speaker 102, for example, using a Bluetooth pairing protocol. The wireless transmitter 103 transmits a radio frequency (RF) signal having the low-frequency content (e.g., sub-bass frequency content) into the audio reproduction environment 100, where it is received by a wireless receiver 109 attached to or embedded in the LFE unit 108 and projected by the LFE unit 108 into the audio reproduction environment 100.
[0092] For certain media content, the described exemplary audio playback environment 100 may not be adept at handling certain types of audio content. For example, certain sound effects may be encoded in an allocentric or egocentric frame of reference as ceiling objects located above the user 104. Far-field speakers 102, such as the sound bar shown in FIG. 1, may not be able to render these ceiling objects as intended by the content creator. For such content, a near-field playback device 105 can be used to play binaurally rendered near-field signals according to the content creator's intent. For example, for better results, the sound effect of a helicopter flying overhead may be rendered for playback on the stereo near-field speakers of the near-field playback device 105 rather than the far-field speakers 102.
[0093] There are several problems with the audio reproduction environment 100. As will be explained below with reference to Figure 3, the combination of acoustic propagation time, radio transmission time, and signal processing time can result in the far-field and near-field acoustic audio being out of sync. A solution to this problem is explained with reference to Figures 4A and 4B.
[0094] Another problem associated with audio reproduction environment 100 is ear occlusion by near-field speakers due to their structure (e.g., closed-back headphones) or frequency response (e.g., poor low-frequency response). Occlusion can be mitigated by using low-occlusion earbuds or other open-back headphones. The frequency response of near-field speakers can be compensated for using equalization (EQ). For example, an average or calibrated EQ profile (e.g., an EQ profile that is the inverse or mirror image of the near-field speaker's natural frequency response profile) can be applied to the rendered near-field speaker input signal before sending the signal to the near-field speaker feed.
[0095] In a single-user embodiment, near-field reproduction device 105 communicates with media source device 101 via wireless transceivers 103, 106 to provide data indicative of near-field speaker characteristics, such as the frequency response of the near-field speakers, and / or audio occlusion data, which is used by an equalizer in media source device 101 to adjust the EQ of the rendered far-field signal. For example, if the audio occlusion data indicates that the near-field speakers attenuate audio data in certain frequency bands (e.g., high frequency bands) by 3 dB, then these frequency bands can be boosted by approximately 3 dB in the rendered far-field signal.
[0096] In one embodiment, at least a portion of the rendered near-field speaker input signals are equalized based at least in part on an average target equalization based on many instances of the same near-field speaker type to compensate for non-flatness of the near-field speakers. For example, the rendered near-field signals for a set of headphones may be attenuated by 3 dB for a frequency band in light of the average target equalization because the average target equalization would result in boosting the rendered far-field signal for that frequency band by 3 dB more than necessary for the audio occlusion caused by the set of headphones.
[0097] In embodiments where latency is a factor, ambient sounds in the listening environment are captured using one or more microphones in an intermediate device or headphones and compensated for in the headphones using the inverse of the occlusion. The end result of the above processing is that the near-field speakers project near-field acoustic audio that is synchronously overlaid with the far-field acoustic audio projected by the far-field speakers 102. Thus, for certain audio content, the near-field speakers can be used to enhance the listening experience of the user 104 by adding height, depth, or other spatial information that is missing, incomplete, or imperceptible when such audio content is rendered for playback using only the far-field speakers 102.
[0098] Exemplary Signal Processing Pipeline 2 is a flow diagram of a processing pipeline 200 for hybrid near-field / far-field virtualization to enhance audio, according to one embodiment. A source signal s(t) is input to a crossover filter 201 and a gain generator 210. The source signal may include channel-based audio, object-based audio, or both channel-based and object-based audio. The output of the crossover filter 201 (e.g., a high-pass filter) is a low-frequency signal lf(t) and a high-frequency signal hf(t). The crossover filter 201 may be configured to filter out any desired crossover frequency f c For example, f c may be 100 Hz, giving a low frequency signal lf(t) containing frequencies below 100 Hz and a high frequency signal hf(t) containing frequencies above 100 Hz.
[0099] In one embodiment, gain generator 210 generates two far-field gains Gf(t), Gf'(t) and two near-field gains Gn(t), Gn'(t). The gains Gf(t) and Gn(t) are applied to the high-frequency signal hf(t), and the gains Gf(t) and Gn'(t) are applied to the low-frequency signal l f(t) in far-field mixing module 202 and near-field mixing module 207, respectively. Note that the superscript "'" indicates low frequency.
[0100] In some embodiments, the gains may be determined according to the amplitude panning method described, for example, in section 2, pages 3-4, of "Audio Reproduction Environment 100," in Audio Reproduction Environment 100. In some embodiments, other methods may be used to pan audio objects in the far field, such as methods involving the synthesis of corresponding acoustic plane waves or spherical waves, as described in "Audio Reproduction Environment 100." In some implementations, at least some of the gains may be frequency dependent. Both the near-field gain and the far-field gain may be related to object or channel position and the far-field speaker layout in the audio reproduction environment 100. [Non-Patent Document 1] V. Pulkki, Compensating Displacement of Amplitude-Panned Virtual Sources, Audio Engineering Society (AES) International Conference on Virtual, Synthetic and Entertainment Audio [Non-patent document 2] D. de Vries, Wave Field Synthesis, AES Monograph 1999
[0101] In one embodiment, rather than splitting the source signal s(t) into near-field and far-field signals, the source signal s(t) includes two channels (L / R stereo channels) that are pre-rendered for playback on a near-field playback device using the methods described above. These "ear" tracks can also be created using a manual process. For example, in a movie theater embodiment, objects can be marked as "ear" or "near" during the content creation process. Because of the way movie theater audio is packaged, these tracks are pre-rendered and provided as part of a digital cinema package (DCP). Other parts of the DCP can include channel-based audio and full Dolby Atmos® channels. In a home entertainment embodiment, content can be provided in two separate pre-rendered "ear" tracks. The "ear" tracks can be temporally offset relative to the other audio and video tracks when stored. In this way, two reads of media data from storage are not required to send the audio to the near-field playback device early.
[0102] Example Mixed Mode In general, Gf(t) = Gf'(t) and Gn(t) = Gn'(t). However, if the far-field speakers 206-1 to 206-n have a higher ability to reproduce low frequencies, all audio content can be routed to the far-field speaker virtualizer 203 by setting Gn'(t) = 0 and Gf'(t) = 1.
[0103] For traditional surround rendering using channel-based audio, if only front speakers (e.g., L / R stereo speakers and an LFE device) are present, the mixing function can route all surround channels to the near-field speaker virtualizer 208 by applying Gn(t)=1.0 and Gf(t)=0.0, and route all front speaker channels (e.g., L / R speaker channels) to the far-field speaker virtualizer 203 by applying Gn(t)=0.0 and Gf(t)=1.0.
[0104] To render the distance effect, both the far-field speaker virtualizer 203 and the near-field speaker virtualizer 208 are mixed as a function of the (normalized) distance r to the center of the audio reproduction environment 100 (e.g., the center of the room or the preferred listening position of the user 104) such that Gn(t)=1.0-r and Gf(t)=sqrt(1.0-Gn(t)*Gn(t)), for r between 0.0 (100% near field) and 1.0 (100% far field).
[0105] In one embodiment, a percentage of the audio content can be played through far-field and near-field speakers to provide an enhancement layer (e.g., a dialogue enhancement layer), where the audio object or center channel is rendered with Gf(t)=1.0 and Gn(t)>0.0.
[0106] In one embodiment, the output of the far-field mixing module 202 is a far-field signal f(t), which is a weighted linear combination of high- and low-frequency signals hf(t), lft(t), where the weights are the far-field gains Gf(t), Gf′(t): f(t)=Gf'(t)*lf(t)+Gf(t)*hf(t) [1]
[0107] The far-field signal f(t) is input to a far-field speaker virtualizer 203, which generates a rendered far-field signal F(t). The rendered far-field signal F(t) can be generated using any desired speaker virtualization algorithm utilizing any number of physical speakers, including but not limited to vector-based amplitude panning (VBAP) and multiple-direction amplitude panning (MDAP).
[0108] The rendered far-field signal F(t) is input to an optional far-field post-processor 204 to apply any desired post-processing (e.g., equalization, compression) to the rendered far-field signal F(t). The rendered, optionally post-processed, far-field signal F(t) is then input to an audio subsystem 205, which is coupled to far-field speakers 206-1 through 206-n. The audio subsystem 205 includes various electronics (e.g., amplifiers, filters) for generating electrical signals to drive far-field speakers 206-1 through 206-n. In response to the electrical signals, far-field speakers 206-1 through 206-n project far-field acoustic audio into audio reproduction environment 100. In one embodiment, the far-field processing pipeline described above is implemented, in whole or in part, in software running on a central processing unit and / or digital signal processor.
[0109] Referring now to the near-field processing pipeline of FIG. 2, the output of the near-field mixing module 207 is the near-field signal n(t), which is a weighted linear combination of the high- and low-frequency signals hf(t), lf(t), where the weights are the near-field gains Gn(t), Gn′(t): n(t)=Gn'(t)*lf(t)+Gn(t)*hf(t) [2]
[0110] In one embodiment, the near-field signal n(t) is input directly to the wireless transceiver 103, which encodes and transmits the near-field signal n(t) over a wireless communication channel to the near-field reproduction device 105 or intermediate device 110. The near-field signal is delivered to the near-field reproduction device and becomes near-field acoustic audio that is reproduced through near-field speakers located proximate the user's ears.
[0111] In some embodiments, the near-field signal is an augmentation of some or all of the far-field acoustic audio. For example, the near-field signal may contain only dialogue, whereby the effect of listening to the far-field acoustic audio and the near-field acoustic audio together is enhanced, resulting in more audible dialogue. Alternatively, the near-field signal may provide a mix of dialogue and background (e.g., music, effects, etc.), whereby the net effect is a personalized, more immersive experience.
[0112] In some implementations, near-field signals include sounds intended to be perceived as close to the listener, such as user proximity sounds in spatial sound systems. In such systems, audio objects, such as the sound of an airplane flying overhead through a scene, are rendered to a set of speakers within the audio playback environment based on audio object coordinates that may change over time, so that the audio object source appears to move within the audio playback environment. However, because sound system speakers are typically located at the periphery of a room or theater, they have limited ability to create a sense of depth of proximity or distance from the listener. This is typically solved by panning the audio to and through speakers near the user's ears.
[0113] In some embodiments, near-field signals may include sounds that are intended to be perceived close to the listener for artistic reasons, such as movie sounds occurring on or around a particular character in the movie, where heartbeats, breathing, rustling clothes, footsteps, whispers, and other sounds that are close to the character may be heard close to the listener, eliciting an emotional connection, empathy, or personal identification with that character.
[0114] In some embodiments, the near-field signals may include sounds intended to be played near the listener to increase the size of the optimal listening position in a room equipped with a spatial audio system. The near-field signals may be synchronized with the far-field acoustic audio so that audio objects panned to or through the user's location are compensated for acoustic travel time from the far-field speakers.
[0115] In some embodiments, the near-field signal includes sounds used to correct for imperfections in the room acoustics. For example, the near-field signal may be an exact copy of the rendered far-field signal. The far-field audio is sampled with a microphone in the near-field reproduction device and compared to the near-field signal in the near-field reproduction device or an intermediate device. If the far-field audio is found to be defective in some way, for example by lacking certain frequency components due to the user's position in the room, those frequency components may be enhanced before playback in the near-field speakers.
[0116] Aspects of the near-field signal may be customizable by the user to suit their preferences. Some options for customization may include choosing between types of near-field signals, adjusting loudness equalization in two or more frequency bands, or spatializing the near-field signal. Types of near-field signals may include dialogue only, dialogue, music, a combination of effects, or an alternative language track.
[0117] Near-field signals can be generated in a variety of ways. One method is intentional authoring, where one or more possible near-field signals for a particular piece of entertainment content can be authored as part of the media creation process. For example, a clean (i.e., isolated, devoid of other sounds) dialogue track can be created. Alternatively, spatial audio objects can be intentionally panned through coordinates that cause them to be rendered to near-field speakers proximal to the user. Alternatively, an artistic choice can be made to position certain sounds, such as those occurring on or around an empathetic protagonist, closer to the user.
[0118] Another method for near-field signal generation is to do so automatically or algorithmically during media content generation. For example, the center channel in a 5.1 or similar audio mix often contains dialogue, and the L and R channels typically contain the main parts of all other sounds, so L+C+R can be used as the near-field signal. Similarly, if the goal of the near-field signal is to provide enhanced dialogue, deep learning or other methods known in the art can be used to extract clean dialogue.
[0119] Near-field signals can also be created automatically or algorithmically during media playback. Many of the entertainment devices mentioned above can use internal computational resources, such as a central processing unit (CUP) or digital signal processor (DSP), to extract dialogue or combine channels for use as near-field signals. The far-field acoustic audio signals and near-field signals may contain inserted signals or data for the purpose of improving time offset calculations; for example, marker signals may be simple ultrasonic tones or may be modulated to convey information or improve detectability, as described in more detail below.
[0120] In an alternative embodiment, the near-field signal n(t) is input to the near-field speaker virtualizer 208, which generates a rendered near-field signal N(t). The rendered near-field signal N(t) can be generated, for example, using a binaural (stereophonic) rendering algorithm that uses head-related transfer functions (HRTFs). In one embodiment, the near-field speaker virtualizer 208 receives the near-field signal n(t) and the head pose of the user 104, and generates and outputs the rendered near-field signal N(t). The head pose of the user 104 can be determined based on real-time input from the far-field speakers 206-1 through 206-n or a head tracking device (e.g., camera, Bluetooth tracker) that outputs the orientation and possibly the head position of the user 104 relative to the audio reproduction environment 100.
[0121] In one embodiment, the rendered near-field signal N(t) is input to an optional near-field post-processor 209 to apply any desired post-processing (e.g., equalization) to the rendered near-field signal N(t). For example, equalization can be applied to compensate for imperfections in the frequency response of the near-field speakers. The rendered or optionally post-processed near-field signal N(t) is then input to a wireless transceiver 103, which encodes and transmits the rendered near-field signal N(t) to the near-field reproduction device 105 or the intermediate device 110 over a wireless communication channel.
[0122] As will be explained in more detail below, the near-field signal n(t), or rendered near-field signal N(t), is transmitted earlier than the projection of the far-field audio to allow for synchronous overlay of the far-field audio and the near-field audio. In the following, embodiments are described in which the near-field signal n(t) is transmitted to a near-field reproduction device or intermediate device 110.
[0123] In one embodiment, the wireless transceiver 103 is a Bluetooth or WiFi transceiver, or uses a custom wireless technology / protocol. In one embodiment, the near field processing pipeline described above with reference to Figure 2 may be implemented completely or partially in software running on a central processing unit and / or digital signal processor.
[0124] In one embodiment, the near-field reproduction device 105 and / or the intermediate device 110, rather than the media source device 101, includes a near-field speaker virtualizer 208 and a near-field post-processor 209. In this embodiment, the gains Gn(t), Gf(t) and the near-field signal n(t) are transmitted by the wireless transceiver 103 to the near-field reproduction device 105 or the intermediate device 110. The intermediate device 110 then renders the near-field signal n(t) into a rendered near-field signal N(t) and transmits the rendered signal to the near-field reproduction device 105 (e.g., headphones, earbuds, or a headset). The near-field reproduction device 105 projects near-field acoustic audio near or into the ears of the user 104 through near-field speakers embedded in or coupled to the near-field reproduction device 105.
[0125] In one embodiment, the gains Gn(t) and Gf(t) are pre-computed at a headend or other network-based content service provider or distributor and sent as metadata in one or more layers (e.g., the transport layer) of the bitstream to media source device 101, where the source signal and gains are demultiplexed, decoded, and applied to the audio content of the source signal. This allows audio content creators to create different versions of their audio content for use with hybrid near-field / far-field speaker virtualization on different speaker layouts in different audio playback environments. Additionally, the metadata can include one or more flags (e.g., one or more bits) that indicate to a decoder that the bitstream includes far-field and near-field gains, making it suitable for use with hybrid near-field / far-field speaker virtualization.
[0126] In some embodiments, one or both of the near-field and far-field signals may be generated on a network computer and delivered to a media source device, with the far-field signals optionally being further processed before being projected from the far-field speakers, and the near-field signals optionally being further processed before being sent to a near-field reproduction device or intermediate device as described above.
[0127] Early transmission of near-field signals FIG. 3 shows an example timeline for wireless transmission of near-field signal n(t), illustrating the benefits of early transmission, according to one embodiment. The timeline illustrates propagation time of far-field acoustic audio versus near-field wireless transmission latency and signal processing time. The far-field acoustic audio begins propagating away from the far-field speakers 206-1 through 206-n at t=0 and arrives at the user 104's location at t=10 ms (assuming a distance of approximately 3 meters from the far-field speakers 206-1 through 206-n). The timeline shown in FIG. 3 is a nonlinear scale by a factor of 10, where negative numbers indicate times earlier than t=0 (e.g., −0.01 is 10 ms before t=0). To enable synchronization, the wireless transmission of near-field signal n(t) should be received and decoded, and all synchronization signal processing and rendering completed, before or just at the same time the far-field acoustic audio arrives at the microphone 107 of the near-field reproduction device 105 or intermediate device 110.
[0128] Referring to Figure 3, timeline (a) shows how a custom wireless protocol (not commonly used in consumer electronics) can provide short transmission latency, allowing the rendered near-field signal to be available in time. Timeline (b) shows that universal protocols (e.g., WiFi, Bluetooth) do not deliver the near-field signal in time. Timeline (c) shows how wireless transmission can begin arbitrarily earlier than t = 0 seconds, compensating for any transmission latency and accounting for any signal processing time, allowing for synchronization of far-field and near-field audio.
[0129] The transmission, decoding, and signal processing time required to deliver and synchronize near-field signals can be significant. Wireless transmission methods commonly used in consumer electronics, such as Wi-Fi and Bluetooth, have latencies ranging from tens to hundreds of milliseconds. Furthermore, wireless transmissions often encode audio using digital codecs that compress the digital information to minimize the required bandwidth. Once received, some signal processing time is required to decode the encoded signal and recover the audio signal. Signal processing for synchronization, described in detail below, can require millions of computational operations. Depending on the speed of the processor used, decoding and signal processing can also require long periods of time, especially in battery-powered endpoint devices that may have limited computational power.
[0130] Sound travels one meter in just under three milliseconds. A user in a home living room or movie theater may be anywhere from one meter to several tens of meters from the far-field speaker, so the expected sound travel time ranges from approximately 3 ms to 100 ms. If the near-field signal n(t), and the subsequent processing of that signal, requires a time longer than the travel time of the far-field acoustic audio, the near-field signal n(t) will arrive too late, and synchronization of the near-field acoustic audio with the far-field acoustic audio will be impossible.
[0131] In situations where users are much farther away from the far-field speakers, such as in a large concert venue, it may be possible for the near-field signal n(t) to reach those users in a time sufficient to allow synchronization. Furthermore, if the wireless protocol is less ubiquitous or perhaps a custom-built technology, the wireless transmission latency can be shorter than the travel time of the far-field acoustic audio. However, using a wireless protocol not already built into most consumers' personal mobile devices requires secondary equipment for wireless reception.
[0132] A better solution is to deliver the near-field signal n(t) using a common wireless protocol, but sufficiently earlier than the far-field acoustic audio is expected to arrive at the near-field reproduction device 105. For example, if transmission through a Wi-Fi router results in a worst-case latency of 250 ms, decoding and synchronization require 20 ms, and the expected acoustic travel time is 10 ms, then transmission of the near-field signal n(t) to the near-field reproduction device 105 (or intermediate device 110) should be more than 260 ms before the rendered far-field signal F(t) is provided to the speaker feeds of the far-field speakers 206-1 through 206-n. Such early transmission of the near-field signal n(t) provides sufficient time for synchronization in the near-field reproduction device 105 (or intermediate device 110). In practice, advance times of 300 ms to 1000 ms are useful.
[0133] Note that early transmission of the near-field signal n(t) is not possible at live events, where stage sounds (vocals, instruments, etc.) propagate immediately outside and then almost simultaneously through amplifiers and speakers, and any electronic recording and radio transmission can only begin after the moment of sound generation. However, at a "live" event, some or all sounds can be transmitted wirelessly immediately and then delayed before being played back through speakers, so that the radio transmission has time to be received and used. This can be particularly useful for stage sounds that do not propagate acoustically immediately, such as electronic instruments, or when speaker volume is sufficiently loud to mask any stage sounds. Early transmission of live events to users not present at the live event is also possible. For example, viewers of a football game on their home entertainment system may receive the entertainment content at home only after several seconds of delay due to network censorship delays, signal processing delays, broadcast and transmission equipment delays, etc. Typically, such delays accumulate and easily amount to at least several seconds.
[0134] There are several ways to transmit the near-field speaker signal n(t) early. In one embodiment, the media source device 101, which receives or plays media and delivers far-field acoustic audio, has a buffer containing the source signal. This buffer is read twice: once from a first location in the buffer to deliver the far-field speaker input signal F(t) and possibly associated video, and a second time from a second location in the buffer, a desired advance time later, to deliver the near-field signal n(t) to the near-field playback device 105 or intermediate device 110. The order of reading these two buffers can be switched; only their relative positions within the buffer are important. In one embodiment, there can be multiple buffers, such as one buffer for the rendered far-field signal F(t) and one buffer for the near-field signal n(t).
[0135] In another embodiment, media source device 101 is configured to ingest a source signal including audio and video content. The ingested source signal is buffered to allow for a specified delay. The near-field signal n(t) is sent to near-field reproduction device 105, where it is projected through near-field speakers as near-field acoustic audio. After the specified delay, the audio and video are read from the buffer, and the audio is processed as described above to generate far-field acoustic audio.
[0136] Discovery methods In one embodiment, the near-field reproduction device 105 (with optional intermediate device 110) includes hardware or software to understand when the near-field signal n(t) is available. This can be as simple as listening for multicast packets on a Wi-Fi network. This can also be accomplished using various methods of zero-configuration networking protocols, such as Apple Bonjour®.
[0137] Time stamp transmission for synchronization There are well-known methods by which wired or wireless networked devices can share information to synchronize their clocks. Two examples are the Network Time Protocol (NTP) and the IEEE 1588 Precision Time Protocol (PTP). If the media source device 101 and the near-field playback device 105 (or intermediate device 110) synchronize their clocks using such methods, time-stamped audio packets can be played back synchronously by each device at the agreed-upon time.
[0138] In a more detailed example, a DMR (e.g., an Apple® TV DMR) and an intermediate device (e.g., a smartphone) have synchronized clocks using NTP. Frames of near-field signal n(t) are transmitted using WiFi from the DMR to the intermediate device 500 ms before the same frames are played through media source device 101 (e.g., a television) via a High-Definition Multimedia Interface (HDMI®) and / or optical link. Each frame of near-field signal n(t) includes a timestamp that indicates to intermediate device 110 the exact time at which the frame should be played into the user's ear. Intermediate device 110 adjusts the time required to transmit near-field signal n(t) from intermediate device 110 to near-field playback device 105 to play the audio frame at the indicated time.
[0139] The use of timestamps does not guarantee that near-field acoustic audio will be played in sync with far-field acoustic audio. This is because timestamps do not automatically account for several sources of time error, such as processing time at the media source device 101 for playing the far-field acoustic audio, wireless signal transmission latency from the intermediate device 110 to the near-field reproduction device 105, and acoustic transmission time of the far-field acoustic audio from the far-field speakers 206-1 through 206-n to the position of the user 104 within the audio generation environment 100. Nevertheless, using timestamps reduces the range of possible delay times that need to be explored, thereby reducing computation time and power consumption. Timestamps can also provide a second-best delay time for synchronization if acoustic synchronization fails. In combination with more precise time offset determination, described below, timestamps can provide a close estimate, a known-good fallback when acoustic synchronization fails, and reduced complexity and power consumption.
[0140] Determining the Time Offset To avoid a negative listening experience, the near-field audio is synchronously reproduced together with the far-field audio by the near-field reproduction device 105. Small time differences between the near-field audio and the far-field audio, on the order of a few milliseconds, can cause noticeable and unpleasant spectral coloration. As the time difference approaches 10-30 ms and gets closer, the spectral coloration spreads to lower frequencies and then becomes comb-filtered. The user 104 then hears two copies of the audio content. With a small delay, this can sound like a close echo, and with a large delay, it can sound like a distant echo. With even larger time delays, listening to the copies of the audio content creates a cognitive load that is very unpleasant.
[0141] To avoid these negative effects, the near-field sound audio is synchronously overlaid with the far-field sound audio by the near-field reproduction unit 105. In one embodiment, a total time offset between the far-field sound audio and the near-field sound audio is determined to indicate which segments of the near-field sound audio should be sent to the near-field speakers to achieve the synchronous overlay. The total time offset determination is achieved using one or more of the methods described with reference to FIG. 4A.
[0142] Exemplary Methods for Determining Time Offsets FIG. 4A is a block diagram of a processing pipeline 400a for determining a total time offset for synchronizing playback of near-field acoustic audio with far-field acoustic audio, according to one embodiment. In the near-field reproduction device 105 (or intermediate device 110), one or more microphones 107 capture samples of the far-field acoustic audio projected by the far-field speakers 206-1 through 206-n. The samples are captured and processed by an analog front end (AFE) and digital signal processor (DSP) 401a to generate digital far-field data that is stored in a far-field data buffer 403b. In one embodiment, the AFE may include a preamplifier and an analog-to-digital converter (ADC). Prior to receiving the far-field acoustic audio (see FIG. 3), the near-field signal n(t) is received by the wireless transceiver 106 and processed using the AFE / DSP 401b. AFE / DSP 401b includes, for example, circuitry for demodulating / decoding near-field signal n(t), which is converted into digital near-field data that is stored in near-field data buffer 403b.
[0143] The far-field data and near-field data stored in buffers 403a and 403b, respectively, are then compared using a correlation technique. In one embodiment, buffers 403a and 403b each store one second of data. The time offset between the contents of buffers 403a and 403b is determined by correlator 404, which correlates the far-field data stored in buffer 403a against the near-field data stored in buffer 403b. Correlation can be achieved by correlator 404 using brute force in the time domain, or it can be performed in the frequency domain after transforming the buffered data into the frequency domain using, for example, a fast Fourier transform (FFT). In one embodiment, correlator 404 can implement the well-known generalized cross correlation with phase transform (GCC-PHAT) algorithm in the time or frequency domain.
[0144] In one embodiment, the near-field signal n(t) and the rendered far-field signal F(t) include an inaudible, high-frequency marker signal. Such a marker signal may be a simple ultrasonic tone or may be modulated to convey information or improve detectability. For example, the marker signal may be above 18.5 kHz, which is in the frequency range that most people cannot hear but that most audio devices pass. Because such a marker signal is common to both the far-field acoustic audio and the near-field signal, it can be used to improve the time offset calculation between the far-field acoustic audio and the near-field signal. In one embodiment, the marker signal is extracted by AFE / DSP 401a and AFE / DSP 401b using marker signal extractors 402a and 402b, respectively, so that the marker signal is not played back from the near-field speakers. In an embodiment, marker signal extractors 402a and 402b are low-pass filters that filter out the high-frequency, inaudible time marker signal provided to correlator 404.
[0145] The output of the correlator 404 is a time offset and a confidence indicator. The time offset is the time between the arrival of the far-field acoustic audio at the microphone 107 of the near-field reproduction device 105 or intermediate device 110 and the arrival of the near-field signal n(t) at the near-field reproduction device 105. The time offset indicates which portion of the buffer 403b to play through the near-field speaker of the near-field reproduction device 105 and is approximately sufficient for a perfect synchronous overlay of the near-field acoustic audio onto the far-field acoustic audio.
[0146] The total time offset can be determined by adding an additional fixed local time offset 405 to the time offset output by the correlator 404. The local time offset includes the additional time required to send the near-field signal n(t) from the intermediate device 110 to the near-field regenerator 105, including but not limited to packet transmission time, propagation delay, and processing delay. This local offset time can be accurately measured by the intermediate device 110.
[0147] In one embodiment, the total time offset determination described above is continuous, rather than occurring only once during a startup or setup step. For example, the total time offset can be calculated once per second, or several times per second. This duty cycle allows synchronization to adapt to the changing position of the user 104 within the audio reproduction environment 100. While the total time offset calculation shown in FIG. 4A occurs in the near-field reproduction device 105 or intermediate device 110, in principle, the total time offset calculation could occur in the media source device 101 in certain applications, such as applications with a single near-field reproduction device 105.
[0148] In one embodiment, the correlator 404 also outputs a confidence indicator to know when to trust that synchronization has been achieved. One suitable confidence indicator is the known Pearson correlation coefficient between buffers 403a, 404b shifted by a time offset value, which outputs an indicator of linear correlation, where "1" is an overall positive linear correlation, "0" is no linear correlation, and "-1" is an overall negative linear correlation.
[0149] 4B is a block diagram of a processing pipeline 400b for synchronizing near-field audio playback with far-field audio playback, according to one embodiment. In one embodiment, a synchronizer 406 receives as input the digital near-field data from buffer 403b and the total time offset and confidence indicators output from processing pipeline 403a, and applies the total time offset to the rendered near-field signal to synchronize the near-field audio playback with the far-field audio playback. In one embodiment, the total time offset is used only if its corresponding confidence indicator indicates a positive linear correlation between the contents of buffers 403a, 403b (i.e., exceeds a positive threshold). If the confidence indicator does not indicate a linear correlation (i.e., is below a positive threshold), the synchronizer 406 does not apply the total time offset to the rendered near-field signal N(t). Alternatively, a predetermined total time offset can be used.
[0150] In one embodiment, synchronizer 406 performs a calculation or operation that provides a pointer into near-field data buffer 403b that corresponds to the exact sample in the rendered near-field signal from which to begin playback. Playing back the rendered near-field signal may mean retrieving a frame from buffer 403b starting at the pointer position. The pointer position may represent a single audio sample. The frame boundaries of the audio data retrieved from buffer 403b may or may not be aligned with the frame boundaries used when placing or storing data in buffer 403b, so audio can be played back from any time instant.
[0151] In some operating scenarios, the synchronization algorithms described herein may cause some samples in the buffer to be played more than once or skipped. This can occur when a listener moves closer to or farther away from a far-field speaker. In such cases, blending operations can be performed to make audio artifacts (e.g., repetitions or skips) inaudible or less noticeable.
[0152] The near-field signal n(t) and the far-field audio generated from the rendered far-field signal F(t) have a temporal correspondence, such that each contains or provides audio that is intended to be heard simultaneously when synchronized with the other. For example, the far-field audio may be the full audio of a war movie, including dialogue partially obscured by loud acoustic noise. The near-field signal n(t) or the user-proximal audio generated therefrom may contain the same dialogue, but "clean" or unobscured by noise. The temporal correspondence in this example is a multitude of precisely simultaneous dialogues. Time intervals, such as the exact time between two utterances or other audio events, can have the same length in each signal.
[0153] Secondary near-field signals In some embodiments, the near-field signals may include audio signals intended for playback in the ears and secondary near-field signals for additional purposes. One use of the secondary near-field signals is to provide additional information to improve synchronization. For example, if the ear channels of the near-field signals are sparse, there are not many signals common to both the near-field signals and the far-field audio. In that case, synchronization is difficult or rare. In that case, the secondary near-field signals provide additional signals common to the far-field audio, and synchronization operates on the secondary near-field signals to synchronously overlap the far-field audio with the near-field audio.
[0154] In another embodiment, the secondary near-field signals include alternative content intended for in-ear playback. This content may not be common to the far-field acoustic audio. For example, the far-field acoustic audio may include at least English dialogue for a movie, and the secondary near-field signals may include dialogue in an alternative language. Synchronization operates on the far-field acoustic audio and the near-field signals, but the secondary near-field signals are played in-ear. In some implementations, the alternative content may include auditory descriptions of scenes and actions for visually impaired users.
[0155] Synchronized Stream Cancellation Early delivery and synchronization present unique opportunities for active noise cancellation (ANC). Traditional ANC in ear devices relies on microphones to measure the target sound to be canceled. Latency and time response issues always exist. After the sound is measured, it arrives at the eardrum very quickly, during which time the anti-sound must be calculated and generated. This is often impossible, especially at high frequencies. However, if the target sound is part of the near-field signal or a secondary near-field signal and is also part of the far-field audio, the target sound can be actively canceled, i.e., removed from the far-field audio, without some of the drawbacks of typical ANC. Examples of such target sounds include dialogue, sounds intended to be shared throughout a theater with multiple seating positions, and loud, dynamic, non-dialogue sounds (e.g., music, explosions) that cause masking for people with hearing impairments.
[0156] ANC microphones typically face outward for feedforward cancellation and / or are located inside the earcup or ear canal for feedback cancellation. In both feedforward and feedback cancellation, the sound to be canceled is measured by the microphone. An analog-to-digital converter (ADC) converts the microphone signal into digital data. An algorithm then inverts that sound using a filter that approximates the associated electroacoustic transfer function to generate an anti-sound that can destructively interfere with ambient sounds. The filter may be adaptive to perform well under changing conditions. The anti-sound is then converted back to an analog signal by a digital-to-analog converter (DAC). An amplifier reproduces the anti-sound in the ear using a transducer, typically a dynamic driver or balanced armature.
[0157] Every component in this system requires time to operate. Each stage, including the microphone, ADC, filter, DAC, and speaker amplifier, can require tens of microseconds or more to operate. The overall latency can be on the order of 100 microseconds or more. This latency severely impairs active noise cancellation by reducing the available phase margin at higher frequencies. For example, a 100 microsecond delay is 10% of one period of a 1 kHz sound wave.
[0158] If the near-field signal or secondary near-field signal components are the sound to be canceled, the early delivery of these signals constitutes prior knowledge of the sound to be canceled. The output of the noise cancellation filters can be pre-calculated and all other system component delays can be compensated for, so the operating delays of these filters and system components are not significant. This is a different situation from general noise cancellation, where there is no prior knowledge of the sound to be canceled.
[0159] In one embodiment, synchronized stream cancellation is used to remove dialogue from far-field acoustic audio, which can then be replaced with dialogue in an alternative language. Active voice cancellation targets the original dialogue transmitted to the ear device in a near-field signal to remove the original dialogue from the far-field acoustic audio. A dialogue track in an alternative language transmitted via a secondary near-field signal can be played instead.
[0160] In one embodiment, synchronized stream cancellation is used to select among possible commentary in sports content. Far-field audio includes, for example, "home" commentary for a football game. Individual viewers of this game can choose to hear commentary for the "away" team instead. The "home" commentary in the far-field audio is delivered via a near-field signal to a near-field playback device and is subject to audio cancellation. A secondary near-field signal delivers the "away" commentary to individual viewers.
[0161] In one embodiment, synchronized stream cancellation is used to essentially mute the entire far-field acoustics audio. For example, a viewer is watching entertainment media and the far-field acoustics audio is being played in a room. The near-field signal contains a copy of the far-field acoustics audio and is subject to voice cancellation. This mode may be useful when a viewer wants to hear people nearby.
[0162] In one embodiment, synchronized stream cancellation is used to correct spatial audio in a spatial audio entertainment system. For example, in a movie theater with a surround sound system, some users may have near-field playback devices as disclosed herein, while others may not. Users without near-field playback devices can be given the full, typical movie theater experience. Thus, the rendered far-field signal contains the complete spatial audio object sound. The near-field signal includes a user-near channel in which the spatial audio object is panned through the user's near-field playback device. The rendering of the same spatial audio object in a movie-only system and the near-field signal may be substantially different, so that users with near-field playback devices have their spatial audio experience diminished by extraneous room sounds. In one embodiment, the difference between the movie theater far-field signal rendering of an audio object and the near-field device rendering of the same audio object can be included in a secondary near-field signal and subjected to sound cancellation in the near-field playback device or an intermediate device.
[0163] In some implementations, weighting is applied as a function of the distance of objects to the listener in the audio playback environment, so that audio objects intended to be heard close to the listener are conveyed only in the near-field signal, and the secondary near-field signal cancels out sound from common audio objects shared by the entire audience in a theater, for example. This allows for the location of sounds very close to the listener (or even inside their head) in a way that is not possible with shared audio signals.
[0164] In another embodiment, synchronized stream cancellation uses a combination of near-field signals and secondary near-field signals to compensate for non-ideal seating positions in a theater with surround sound (or other 3D sound technology), such as near any boundary of the sound signal space, i.e., near one side of the room, in a back corner, etc. In this way, the listener can receive a perceptual rendering that is much closer to the mixing engineer's intent.
[0165] In one embodiment, synchronized stream cancellation uses an algorithm, such as a least mean squares (LMS) adaptive filter algorithm, to construct a filter that matches the captured microphone signal containing the far-field audio with the near-field signal. That filter can then be inverted and applied to the near-field signal to generate an anti-tone. The anti-tone is then played at the correct instant to cancel the portion of the far-field audio that is common to the near-field signal.
[0166] In an alternative embodiment, the algorithms and filters are designed to target all sounds that are not common to the far-field acoustic audio and the near-field signal. In this embodiment, the filters target all sounds that are not in the near-field signal, canceling all sounds except those in the near-field signal, and the user hears only those sounds in the near-field signal. For example, if the near-field signal is a copy of the far-field signal, extraneous room sounds, such as conversation or kitchen sounds, can be canceled in the near-field reproduction device or intermediate device. In some embodiments, far-field acoustic audio is captured by one or more microphones in the near-field device or intermediate device and partially rendered in the near-field reproduction device to compensate for any occlusion of the ear canal by the near-field speaker. If it is desired to improve the user's experience of ambient sound, it may not be desirable to block all ambient sound in the audio reproduction environment. For example, some earbuds partially occlude most people's ears. The occlusion attenuates and potentially colors the user's perception of ambient sound in an undesirable way. To correct for this, in one embodiment the effect of the occlusion is measured and the missing portion of the ambient sound is added back into the near-field signal before being rendered for playback through a near-field playback device.
[0167] 5 is a flow diagram of a process 500 for hybrid near-field / far-field speaker virtualization for audio enhancement, according to one embodiment. Process 500 can be implemented, for example, by the media source device architecture described with reference to FIG.
[0168] Process 500 begins by obtaining (501) a source signal. The source signal may include channel-based audio, object-based audio, or a combination of channel-based audio and object-based audio. The source signal may be provided by a media source device such as a television system, a set-top box, or a DMR. The source signal may also be a bitstream received from a network or a storage device (e.g., an Ultra-HD, Blu-ray, or DVD disc).
[0169] Process 500 continues by generating far-field and near-field gains based on the source signal, the far-field speaker layout, and the far-field and near-field speaker characteristics (502). For example, if an audio object in the audio content of the source signal is located above the user's head and the media source device is a soundbar, gains are calculated such that the entire audio object is included in the rendered near-field speaker input signals, thereby allowing it to be binaurally rendered by a near-field playback device or intermediate device.
[0170] Process 500 continues by generating 503 far-field and near-field signals using the gains. For example, the far-field and near-field signals can be weighted linear combinations of the low-frequency and high-frequency signals output by the cross filters, where the weights are the low-frequency and high-frequency gains.
[0171] Process 500 continues by rendering the far-field signals and, optionally, post-processing the rendered far-field signals (505). For example, the far-field signals (e.g., VBAP) can be rendered using any known algorithm, and the near-field signals can be binaurally rendered using HRTFs. In one embodiment, the near-field signals are rendered / post-processed at the media source device before being sent to the near-field playback device.
[0172] The process 500 continues by early transmitting the near-field signal to a near-field reproduction device or intermediate device (506) and transmitting the rendered far-field signal to a far-field speaker feed (507). For example, the near-field signal is transmitted to the near-field reproduction device or intermediate device to provide sufficient time to calculate a total time offset for synchronization with the far-field acoustic audio, as described with reference to FIGS. 3 and 4A and 4B.
[0173] 6 is a flow diagram of a process for synchronizing playback of near-field sound audio with far-field sound audio, according to one embodiment. Process 600 can be implemented, for example, by the near-field playback device architecture described with reference to FIG.
[0174] Process 600 begins by receiving 601 a previously transmitted near-field signal. For example, as described with reference to Figures 1 and 2, the near-field signal containing first channel-based audio and / or audio objects can be received over a wired or wireless channel.
[0175] Process 600 continues by receiving (602) far-field acoustic audio. For example, a rendered far-field signal including second channel-based audio and / or audio objects is captured by one or more microphones. Process 600 continues by converting (603) the microphone output to digital far-field data and the near-field signal to digital near-field data, as described with reference to FIG. 4A, and storing (604) the digital far-field data and the digital near-field data in a buffer.
[0176] The process 600 continues by determining (605) a total time offset and an optional reliability indicator using the buffer contents and adding a local time offset, as described with reference to FIG. 4A.
[0177] The process 600 continues by beginning playback 606 of the near-field data through the near-field speakers using a total time offset such that the near-field sound data projected by the near-field speakers is synchronously overlaid with the far-field sound. In one embodiment, the synchronization is applied based on a confidence metric indicative of correlation.
[0178] 7 is a flow diagram of an alternative process 700 for synchronizing playback of near-field sound audio with far-field sound audio, according to one embodiment. Process 700 can be implemented, for example, by the media source device architecture described with reference to FIG.
[0179] The process 700 begins by receiving (701) a source signal containing at least one of channel-based audio or audio objects using a media source device, as described with reference to FIG.
[0180] Process 700 continues by generating a far-field signal based at least in part on the source signal using a media source device, as described with reference to FIG.
[0181] Process 700 continues by rendering (703) a far-field signal into the audio playback environment using the media source device for playback of far-field acoustic audio through far-field speakers, as described with reference to FIG. 2.
[0182] Process 700 continues by generating (704) one or more near-field signals based at least in part on the source signal using a media source device, as described with reference to FIG.
[0183] The process 700 continues by transmitting (705) the near-field signal to a near-field reproduction device or intermediate device coupled to a near-field speaker, as described with reference to FIG. 2, before providing the far-field signal to the far-field speaker.
[0184] Process 700 continues by providing the rendered far-field signals to far-field speakers (706) for projection into the audio reproduction environment, as described with reference to FIG.
[0185] 8 is a flow diagram of another alternative process 800 for synchronizing playback of near-field sound audio with far-field sound audio, according to an embodiment. Process 800 can be implemented, for example, by the near-field playback device architecture described with reference to FIG.
[0186] Process 800 can begin by receiving (801) a near-field signal transmitted by a media source device in an audio playback environment using a wireless receiver, as described with reference to FIG. 4A.
[0187] Process 800 continues by converting 802 the near-field signals into digital near-field data using one or more processors, as described with reference to FIG. 4A.
[0188] Process 800 continues by buffering (803) the digital near field data using the one or more processors as described with reference to FIG. 4A.
[0189] Process 800 continues by capturing 804 far-field acoustic audio projected by the far-field speakers using one or more microphones, as described with reference to FIG. 4A.
[0190] Process 800 continues by converting 805 the far-field acoustic audio into digital far-field data using the one or more processors as described with reference to FIG. 4A.
[0191] Process 800 continues by buffering 806 the digital far-field data using the one or more processors as described with reference to FIG. 4A.
[0192] Process 800 continues by determining (807) a time offset using the one or more processors and buffer contents as described with reference to FIG. 4A.
[0193] Process 800 continues by adding a set of local time offsets to the time offset using the one or more processors to generate a total time offset (808), as described with reference to FIG. 4A.
[0194] Process 800 continues by using the one or more processors to begin playing the near-field data through the near-field speakers using a full time offset, as illustrated in Figure 4B, so that the near-field acoustic data projected by the near-field speakers is synchronously overlaid with the far-field acoustic audio (809).
[0195] FIG. 9 is a block diagram of a media source device architecture 900 for implementing the features and processes described with reference to FIGS. 1-8 , according to one embodiment. The architecture 900 includes a wireless interface 901, an input user interface 902, a wired interface 903, an I / O port 904, a speaker array 905, an audio subsystem 906, a power interface 907, LED indicators 908, a logic and control unit 909, a memory 910, and an audio processor 912. Each of these components is coupled to one or more buses 913. The memory 910 further includes a buffer 914 for use as described with reference to FIG. 2 . The architecture 900 can be implemented in a television system, a set-top box, a DMR, a personal computer, a surround sound system, etc.
[0196] The wireless interface 901 includes a wireless transceiver chip or chipset and one or more antennas for receiving wireless communications from wireless routers (e.g., WiFi routers), remote controls, wireless near-field playback devices, wireless intermediate devices, and any other devices that wish to communicate with the media source device.
[0197] The input user interface 902 includes input mechanisms, such as mechanical buttons, switches, and / or touch interfaces, that allow a user to control and manage the media source device.
[0198] The wired interface 903 contains circuitry for handling communications from various I / O ports 904 (e.g., Bluetooth, WiFi, HDMI, optical), and the audio subsystem 906 contains audio amplifiers and other circuitry necessary to drive the speaker array 905.
[0199] The speaker array 905 can include any number, size, and type of speakers, whether located together in a single housing or in separate housings.
[0200] The power interface 907 includes a power manager and circuitry for regulating power from an AC outlet or a USB port or any other power source.
[0201] LED indicators 908 provide visual feedback to the user for various operations of the device.
[0202] Logic and control devices 909 include a central processing unit, microcontroller device, or any other circuitry for controlling the various functions of the media source device.
[0203] The memory 910 can be any type of memory, such as RAM, ROM, and flash memory.
[0204] The audio processor 912 may be a DSP that implements a codec and prepares audio content for output through the speaker array 905 .
[0205] FIG. 10 is a block diagram of a near-field reproduction device architecture 1000 for implementing the features and processes described with reference to FIGS. 1-8 , according to one embodiment. The architecture 1000 includes a wireless interface 1001, a user interface 1002, a haptic interface 1003, an audio subsystem 1004, a speaker 1005, a microphone 1006, an energy storage / battery charger 1007, an input power interface / protection circuit 1008, a sensor 1009, a memory 1010, and an audio processor 1011. Each of these components is coupled to one or more buses 1013. The memory 1010 further includes a buffer 1012. The architecture 1000 can be implemented in headphones, earbuds, earphones, headsets, gaming hardware, smart glasses, headgear, AR / VR goggles, smart speakers, chair speakers, various automotive interior trim pieces, and the like.
[0206] The wireless interface 1001 includes a wireless transceiver chip and one or more antennas for receiving / transmitting wireless communications to / from the media source device and / or intermediate device and any other devices that wish to communicate with the near field playback device.
[0207] The input user interface 1002 includes input mechanisms such as mechanical buttons, switches, and / or touch interfaces that allow a user to control and manage the endpoint device.
[0208] The haptic interface 1003 includes a haptic engine for providing force feedback to the user, and the audio subsystem 1004 includes an audio amplifier and any other circuitry required to drive the speaker 1005.
[0209] Speakers 1004 may include stereo speakers such as those found in headphones, earbuds, and the like.
[0210] The audio subsystem 1004 includes circuitry (e.g., preamplifiers, ADCs, filters) for processing signals from one or more microphones 1006 .
[0211] Input power interface / protection circuitry 1008 includes circuitry for regulating power from an energy storage unit 1007 (eg, a rechargeable battery), a USB port, a charging mat, a charging dock, or any other power source.
[0212] The sensors 1009 may include motion sensors (eg, accelerometers, gyros) and biosensors (eg, fingerprint detectors).
[0213] The memory 1010 can be any type of memory, such as RAM, ROM, and / or flash memory.
[0214] Buffer 1012 (e.g., buffers 403a, 403b of FIG. 4A) may be generated from a portion of memory 1010 and used to store audio data for determining the total time offset, as described above with reference to FIG. 4A.
[0215] While this document contains many specific implementation details, these should not be construed as limitations on the scope of what may be claimed, but rather as descriptions of features that may be specific to particular embodiments. Certain features that are described herein in the context of separate embodiments may also be implemented in combination in a single embodiment. Conversely, various features that are described in the context of a single embodiment may also be implemented in multiple embodiments separately or in any suitable subcombination. Furthermore, while features may be described above as acting in certain combinations and may even initially be claimed as such, one or more features from a claimed combination may, in some cases, be carved out of that combination, and the claimed combination may be directed to a subcombination or variation of the subcombination. The logic flow depicted in the figures does not require the particular order or sequential sequence shown to achieve desirable results. Additionally, other steps may be provided or steps may be removed from the described flow, and other components may be added to or removed from the described system. Accordingly, other implementations are within the scope of the following claims.
[0216] Several aspects will be described. [Aspect 1] receiving, using a media source device, a source signal including at least one of channel-based audio or audio objects; generating, using the media source device, one or more near-field gains and one or more far-field gains based on the source signal and a mixed mode; generating, using the media source device, a far-field signal based at least in part on the source signal and the one or more far-field gains; Rendering the far-field signals for playback into a far-field acoustic audio playback environment through far-field speakers using a speaker virtualizer; generating, using the media source device, a near-field signal based on the source signal and the one or more near-field gains; transmitting the near-field signal to a near-field reproduction device or an intermediate device coupled to the near-field reproduction device before providing the far-field signal to the far-field speaker; providing the far-field signal to the far-field speaker. method. [Aspect 2] filtering the source signal into a low frequency signal and a high frequency signal; generating two sets of near-field gains including a near-field low-frequency gain and a near-field high-frequency gain; generating two sets of far-field gains including a far-field low-frequency gain and a far-field high-frequency gain; generating the near-field signal based on a weighted linear combination of the low-frequency signal and the high-frequency signal, wherein the low-frequency signal is weighted by the near-field low-frequency gain and the high-frequency signal is weighted by the near-field high-frequency gain; generating the far-field signal based on a weighted linear combination of the low-frequency signal and the high-frequency signal, wherein the low-frequency signal is weighted by the far-field low-frequency gain and the high-frequency signal is weighted by the far-field high-frequency gain. 2. The method of embodiment 1. Aspect 3 3. The method of claim 1 or 2, wherein the mixed mode is based, at least in part, on the layout of the far-field speakers in the audio reproduction environment and one or more characteristics of the far-field speakers or near-field speakers coupled to the near-field reproduction device. Aspect 4 The blending mode is surround sound rendering, and the method further comprises: setting the one or more near field gains and the one or more far field gains to include all surround channel-based audio or surround audio objects in the near field signal and all front channel-based audio or front audio objects in the far field signal; The method of embodiment 3. Aspect 5 determining, based on the near-field and far-field speaker characteristics, that the far-field speaker is more capable of reproducing low frequencies than the near-field speaker; setting the one or more near-field gains and the one or more far-field gains to include all of the low-frequency channel-based audio or low-frequency audio objects in the far-field signal. 5. The method of embodiment 3 or 4. Aspect 6 determining that the source signal includes a distance effect; setting the one or more near-field gains and the one or more far-field gains to be a function of a normalized distance between the far-field speaker and a designated position in the audio reproduction environment. 6. The method of any one of embodiments 3 to 5. Aspect 7 determining that the source signal includes channel-based audio or audio objects for enhancing a particular type of audio content in the source signal; and setting the one or more near-field gains and the one or more far-field gains to include in the near-field signal the channel-based audio or audio objects for enhancing the particular type of audio content. 7. The method of any one of embodiments 3 to 6. Aspect 8 8. The method of claim 7, wherein the particular type of audio content is dialogue content. Aspect 9 Aspect 9. The method of any one of aspects 1-8, wherein the source signal is received along with metadata including the one or more near-field gains and the one or more far-field gains. Aspect 10 10. The method of claim 9, wherein the metadata includes data indicating that the source signal can be used for hybrid speaker virtualization using the far-field speakers and the near-field speakers. Aspect 11 A method according to any one of aspects 1 to 10, wherein the near field signal, or the rendered near field signal and the rendered far field signal, includes an inaudible marker signal to assist in synchronous overlay of the near field acoustic audio with the far field acoustic audio. Aspect 12 acquiring head pose information of a user in the audio playback environment; and rendering the near field signal using the head pose information. 12. The method of any one of embodiments 1 to 11. Aspect 13 13. The method of any one of aspects 1-12, wherein equalization is applied to the rendered near-field signal to compensate for a frequency response of the near-field speaker. Aspect 14 14. The method of any one of aspects 1-13, wherein the near field signal or the rendered near field signal is provided to the near field reproduction device over a wireless channel. Aspect 15 The step of providing the near field signal or the rendered near field signal to the near field reproduction device may further comprise: using the media source device to transmit the near field signal or the rendered near field signal to an intermediate device coupled to the near field reproduction device. 15. The method of any one of embodiments 1 to 14. Aspect 16 16. The method of any one of aspects 1-15, wherein equalization is applied to the rendered far-field signal to compensate for a frequency response of the near-field speaker. Aspect 17 17. A method according to any one of aspects 1 to 16, wherein a timestamp associated with the near-field signal or the rendered near-field signal is provided by the media source device to the near-field playback device or an intermediate device to assist in synchronized overlay of the near-field audio with the far-field audio. Aspect 18 Generating the far field signal and the near field signal based at least in part on the source signal and the one or more far field gains comprises: storing the source signal in a buffer of the media source device; retrieving a first set of frames of the source signal stored in a first location in the buffer, the first location corresponding to a first time; generating, using the media source device, the far-field signal based at least in part on the first set of frames and the one or more far-field gains; retrieving a second set of frames of the source signal stored at a second location in the buffer, the second location corresponding to a second time earlier than the first location; and generating, using the media source device, the near field signal based at least in part on the second set of frames and the one or more near field gains. 18. The method of any one of embodiments 1 to 17. Aspect 19 receiving, in an audio reproduction environment, a near-field signal transmitted by a media source device, the near-field signal comprising a weighted linear combination of low-frequency and high-frequency channel-based audio or audio objects for projection through near-field speakers located in the audio reproduction environment proximate to or inserted into a user's ear; converting, using one or more processors, the near field signals into digital near field data; buffering the digital near field data using the one or more processors; capturing far-field acoustic audio projected by the far-field speaker using one or more microphones; converting the far-field audio into digital far-field data using the one or more processors; buffering the digital far-field data using the one or more processors; determining a time offset using the one or more processors and buffer contents; using the one or more processors to add a set of local time offsets to the time offset to generate a total time offset; using the one or more processors to begin playing the near-field data through the near-field speakers using the total time offset, thereby causing near-field acoustic data projected by the near-field speakers to be synchronously overlaid with the far-field acoustic audio. method. Aspect 20 one or more processors; and a memory storing instructions that, when executed by the one or more processors, cause the one or more processors to perform the method of any one of aspects 1 to 20. Device. Aspect 21 A non-transitory computer-readable storage medium storing instructions that, when executed by one or more processors, cause the one or more processors to perform the method of any one of aspects 1 to 20. Aspect 22 receiving, using a media source device, a source signal including at least one of channel-based audio or audio objects; generating, using the media source device, a far-field signal based at least in part on the source signal; using the media source device, rendering the far-field signal for reproduction of far-field acoustic audio into an audio reproduction environment for far-field acoustic audio through far-field speakers; generating, using the media source device, one or more near-field signals based at least in part on the source signal; transmitting the near-field signal to a near-field reproduction device or an intermediate device coupled to the near-field reproduction device before providing the far-field signal to the far-field speaker; providing the rendered far-field signal to the far-field speaker for projection into the audio reproduction environment. method. Aspect 23 23. The method of embodiment 22, wherein the near-field signal comprises enhanced dialogue. Aspect 24 24. The method of claim 22 or 23, wherein there are at least two near-field signals sent to the near-field reproduction device or the intermediate device, a first near-field signal being rendered into near-field sound audio for playback through a near-field speaker of the near-field reproduction device, and a second near-field signal being used to assist in synchronizing the far-field sound audio with the first near-field signal. Aspect 25 A method according to any one of aspects 22 to 24, wherein there are at least two near-field signals sent to the near-field reproduction device, a first near-field signal including dialogue content in a first language, and a second near-field signal including dialogue content in a second language different from the first language. Aspect 26 A method described in any one of aspects 22 to 25, wherein the near-field signal and the rendered far-field signal include an inaudible marker signal to support synchronous overlay of the near-field audio with the far-field audio. Aspect 27 receiving, using a wireless receiver, near-field signals transmitted by a media source device in an audio playback environment; converting, using one or more processors, the near field signals into digital near field data; buffering the digital near field data using the one or more processors; capturing far-field acoustic audio projected by the far-field speaker using one or more microphones; converting the far-field acoustic audio into digital far-field data using the one or more processors; buffering the digital far-field data using the one or more processors; determining a time offset using the one or more processors and buffer contents; using the one or more processors to add a set of local time offsets to the time offset to generate a total time offset; using the one or more processors to initiate playback of the near-field data through near-field speakers using the total time offset, whereby near-field acoustic data projected by the near-field speakers is overlaid in synchronization with the far-field acoustic audio. method. Aspect 28 capturing a target sound from the audio playback environment using one or more microphones of the near-field playback device; converting the captured target sound into digital data using the one or more processors; generating anti-speech using the one or more processors by inverting the digital data using a filter that approximates an electro-acoustic transfer function; and using the one or more processors to cancel out the target sound using the anti-sound. 28. The method according to embodiment 27. Aspect 29 The method of claim 28, wherein the far-field acoustic audio includes a first dialogue in a first language that is the target voice, and the canceled first dialogue is replaced with a second dialogue in a second language that is different from the first language, and the second language dialogue is included in a secondary near-field signal. Aspect 30 A method as described in aspect 28 or 29, wherein the far-field acoustic audio includes a first commentary that is the target audio, and the canceled first commentary is replaced with a second commentary that is different from the first commentary, and the second commentary is included in a secondary near-field signal. Aspect 31 Aspect 31. The method of any one of aspects 28 to 30, wherein the far-field sound audio is the target sound that is canceled by the anti-sound to mute the far-field sound audio. Aspect 32 A method as described in aspect 28, wherein differences between the cinema rendering and the near-field playback device rendering of one or more audio objects are included in the near-field signal and used to render the near-field sound audio, whereby the one or more audio objects included in the cinema rendering but not the near-field playback device rendering are excluded from the rendering of the near-field sound audio. Aspect 33 A method as described in aspect 32, wherein weighting is applied as a function of the distance from objects in the audio reproduction environment to the listener, so that one or more specific sounds intended to be heard close to the listener are transmitted only in the near-field signal, and the near-field signal is used to cancel out the same specific one or more sounds in the far-field acoustic audio. Aspect 34 34. The method of any one of aspects 27-33, wherein the near-field signals are modified by a head-related transfer function (HRTF) of a listener to provide enhanced spatiality. Aspect 35 one or more processors; and a memory storing instructions that, when executed by the one or more processors, cause the one or more processors to perform the method of any one of aspects 22 to 34. Device. Aspect 36 A non-transitory computer-readable storage medium storing instructions that, when executed by one or more processors, cause the one or more processors to perform the method of any one of aspects 22 to 34.
Claims
1. receiving, using a media source device, a source signal including at least one of channel-based audio or audio objects; generating, using the media source device, one or more near-field gains and one or more far-field gains based on the source signal and a mixed mode; setting the one or more near-field gains and the one or more far-field gains to include the channel-based audio or audio objects in the near-field signal to enhance a particular type of audio content; generating, using the media source device, a far-field signal based at least in part on the source signal and the one or more far-field gains; Rendering the far-field signals into an audio reproduction environment using a speaker virtualizer for reproduction of far-field acoustic audio through far-field speakers; generating, using the media source device, a near-field signal based on the source signal and the one or more near-field gains; determining that the source signal includes channel-based audio or audio objects for enhancing the particular type of audio content in the source signal; transmitting the near-field signal to a near-field reproduction device or an intermediate device coupled to the near-field reproduction device before providing the far-field signal to the far-field speaker; providing the far-field signal to the far-field speaker. method.
2. The method of claim 1 , wherein the particular type of audio content is dialogue content.
3. The method of claim 1 , wherein the source signal is received along with metadata indicative of the one or more near-field gains and the one or more far-field gains.
4. The method of claim 3 , wherein the metadata includes data indicating that the source signal can be used for hybrid speaker virtualization using the far-field speakers and the near-field speakers.
5. The method of claim 1 , wherein the mixing mode is surround sound rendering.
6. receiving, using a media source device, a source signal including at least one of channel-based audio or audio objects; generating, using the media source device, a far-field signal based at least in part on the source signal; using the media source device to render the far-field signals into an audio reproduction environment for reproduction of far-field acoustic audio through far-field speakers; generating, using the media source device, one or more near-field signals based at least in part on the source signal, the near-field signals comprising enhanced dialogue; transmitting the near-field signals to a near-field reproduction device or an intermediate device coupled to the near-field reproduction device before providing the far-field signals to the far-field speaker, wherein at least two near-field signals are sent to the near-field reproduction device, a first of the two near-field signals containing dialogue content in a first language and a second of the two near-field signals containing dialogue content in a second language different from the first language; providing the rendered far-field signal to the far-field speaker for projection into the audio reproduction environment. method.
7. receiving, using a wireless receiver, near-field signals transmitted by a media source device in an audio playback environment; converting the near field signals into digital near field data using one or more processors; buffering the digital near field data using the one or more processors; capturing far-field acoustic audio projected by the far-field speaker using one or more microphones; using the one or more processors, converting the far-field acoustic audio into digital far-field data, the far-field acoustic audio including a first dialogue in a first language that is a target voice, and replacing the first dialogue with a second dialogue in a second language that is different from the first language, the second language dialogue being included in a secondary near-field signal, and the second dialogue being a second commentary that is different from the first commentary; using the one or more processors, buffering the digital far field data and determining buffer contents; determining a time offset using the one or more processors and the buffer contents; using the one or more processors to add a set of local time offsets to the time offset to generate a total time offset; using the one or more processors to initiate playback of the near-field data through near-field speakers using the total time offset, whereby near-field acoustic data projected by the near-field speakers is overlaid in synchronization with the far-field acoustic audio. method.