Radio adjustment of sound source
The A/V hub adjusts playback time and synchronizes audio devices wirelessly to enhance audio quality and user experience by correcting clock drift and adapting to environmental and listener positions, improving satisfaction and revenue.
Patent Information
- Application Number
- JP2025039496
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2016-12-13
- Filing Date
- 2025-03-12
- Publication Date
- 2025-06-24
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
Achieving high audio quality in an environment is difficult due to improper placement of acoustic sources and listener positioning, which can break the 'sweet spot' and deteriorate the acoustic quality, affecting user satisfaction and experience.
An A/V hub that adjusts playback time of electronic devices using wireless communication, calculating time offsets and vector distances to synchronize audio playback, and adapts to acoustic characteristics and listener positions for optimal audio experience.
Improves acoustic quality and user experience by correcting clock drift, adapting to environmental characteristics, and ensuring synchronized audio playback across devices, thereby increasing customer loyalty and revenue for providers.
Smart Images

Figure 2025094013000001_ABST
Abstract
Description
Technical Field
[0001] The described embodiments relate to communication technologies. More specifically, the described embodiments include communication technologies for wirelessly adjusting the playback time of an electronic device that outputs sound.
Background Art
[0002] Music often has a great impact on an individual's emotions and perceptions. This is thought to be the result of the connection or relationship between the areas of the brain that decode, learn, and remember music and the areas that produce emotional reactions such as the frontal lobe and limbic system. Indeed, emotions are thought to be involved in the process of interpreting music and are at the same time extremely important for the influence of music in the brain. Considering this ability of music to "move" the listener, audio quality is often an important factor in user satisfaction when listening to audio content and more generally when viewing audio / video (A / V) content.
[0003] However, it is often difficult to achieve high audio quality in an environment. For example, an acoustic source (such as a loudspeaker) may not be properly placed in the environment. Instead of or in addition to this, the listener may not be positioned in an ideal location within the environment. Specifically, in a stereo playback system, the so-called "sweet spot", where the amplitude difference and arrival time difference are small enough so that both a clear image and positioning of the original sound source are maintained, is usually restricted to a fairly small area between the loudspeakers. When the listener is outside this area, the clear image is broken and only one or the other independent audio channel output by the loudspeaker may be audible. Furthermore, achieving high audio quality in an environment usually places strong constraints on the synchronization of the loudspeakers.
[0004] As a result, when one or more of these factors are sub - optimal, the acoustic quality in the environment may deteriorate. As a result, this can negatively affect the listener's satisfaction and overall user experience when listening to audio content and / or A / V content. Summary of the Invention Means for Solving the Problems
[0005] The first group of described embodiments includes an audio / video (A / V) hub. This A / V hub includes one or more antennas and an interface circuit that communicates with an electronic device using wireless communication during operation. During operation, the A / V hub receiver receives a frame from the electronic device via wireless communication, where a given frame includes the transmission time at which a given electronic device transmitted the given frame. Next, the A / V hub stores the reception time at which the frame was received, where the reception time is based on the A / V hub's clock. Further, the A / V hub calculates the current time offset between the clock of the electronic device and the clock of the A / V hub based on the reception time and transmission time of the frame. Next, the A / V hub transmits one or more frames including audio content and playback timing information to the electronic device, where the playback time information specifies the playback time at which the electronic device plays the audio content based on the current time offset. Furthermore, the playback time of the electronic device has a temporal relationship such that the playback of the audio content by the electronic device is adjusted.
[0006] Note that the temporal relationship can have a non - zero value, and at least a portion of the electronic devices are commanded to play audio content having a phase relative to each other by using different values of the playback time. For example, the different playback times can be based on the acoustic characterization of the environment including the electronic device and the A / V hub. Alternatively or in addition, the different playback times can be based on the desired acoustic characteristics in the environment.
[0007] In some embodiments, the electronic device is positioned at a vector distance from the A / V hub, and the interface circuit determines the magnitude of the vector distance based on the transmission time and the reception time using wireless ranging. Further, the interface circuit can determine the angle of the vector distance based on the angle of arrival of the wireless signal associated with the frame received by one or more antennas during wireless communication. Additionally, different playback times can be based on the determined vector distance.
[0008] Alternatively or in addition, different playback times are based on the estimated position of the listener relative to the electronic device. For example, the interface circuit can communicate with another electronic device and calculate the estimated position of the listener based on the communication with the other electronic device. Further, the A / V hub can include an acoustic transducer that performs sound measurements of the environment including the A / V hub, and the A / V hub can calculate the estimated position of the listener based on the sound measurements. Additionally, the interface circuit can communicate with other electronic devices in the environment and receive additional sound measurements of the environment from the other electronic devices. In these embodiments, the A / V hub calculates the estimated position of the listener based on the additional sound measurements. In some embodiments, the interface circuit performs time-of-flight measurements and calculates the estimated position of the listener based on the time-of-flight measurements.
[0009] Note that the electronic device can be positioned at a non-zero distance from the A / V hub, and the current time offset can be calculated based on the transmission time and the reception time using wireless ranging by ignoring the distance.
[0010] Additionally, the current time offset can be based on a model of clock drift in the electronic device.
[0011] Another embodiment provides a computer-readable storage medium for use with an A / V hub. The computer-readable storage medium includes program modules that, when executed by the A / V hub, cause the A / V hub to perform at least a portion of the operations described above.
[0012] Another embodiment provides a method for adjusting the playback of audio content. The method includes at least a portion of the operations performed by the A / V hub.
[0013] Another embodiment provides one or more electronic devices.
[0014] A second group of described embodiments includes an audio / video (A / V) hub. The A / V hub includes one or more antennas and an interface circuit that communicates with an electronic device using wireless communication during operation. During operation, the A / V hub receives a frame from the electronic device via wireless communication. The A / V hub then stores the reception time at which the frame was received, where the reception time is based on the A / V hub's clock. Further, the A / V hub calculates a current time offset between the clock of the electronic device and the clock of the A / V hub based on the reception time of the frame and a predicted transmission time, where the predicted transmission time is based on an adjustment of the clocks of the electronic device and the A / V hub at a previous time and a predefined transmission schedule of the frame. Next, the A / V hub transmits one or more frames including audio content and playback timing information to the electronic device, where the playback timing information specifies a playback time at which the electronic device is to play the audio content based on the current time offset. Still further, the playback time of the electronic device has a temporal relationship such that the playback of the audio content by the electronic device is adjusted.
[0015] Note that the time relationship can have non-zero values, whereby at least a part of the electronic device can be commanded to play audio content having a phase with respect to each other by using different values of the playback time. For example, the different playback times can be based on the acoustic characterization of the environment including the electronic device and the A / V hub. Alternatively or in addition, the different playback times can be based on the desired acoustic characteristics in the environment.
[0016] In some embodiments, the electronic device is positioned at a vector distance from the A / V hub, and the interface circuit determines the magnitude of the vector distance based on the transmission time and reception time of the frame using wireless ranging. Further, the interface circuit can determine the angle of the vector distance based on the angle of arrival of the wireless signal associated with the frame received by one or more antennas during wireless communication. Furthermore, the different playback times can be based on the determined vector distance.
[0017] Alternatively or in addition, the different playback times are based on the estimated position of the listener relative to the electronic device. For example, the interface circuit can communicate with another electronic device and calculate the estimated position of the listener based on the communication with the other electronic device. Further, the A / V hub performs sound measurements of the environment including the A / V hub, and the A / V hub can calculate the estimated position of the listener based on the sound measurements. Furthermore, the interface circuit can communicate with other electronic devices in the environment and receive additional sound measurements of the environment from the other electronic devices. In these embodiments, the A / V hub calculates the estimated position of the listener based on the additional sound measurements. In some embodiments, the interface circuit performs time-of-flight measurements and calculates the estimated position of the listener based on the time-of-flight measurements.
[0018] Note that the adjustment of the clock of the electronic device and the clock of the A / V hub can occur during the initialization mode of operation.
[0019] Furthermore, the current time offset can be based on a model of the clock drift of the electronic device.
[0020] Another embodiment provides a computer-readable storage medium for use with an A / V hub. The computer-readable storage medium includes program modules that, when executed by the A / V hub, cause the A / V hub to perform at least a portion of the operations described above.
[0021] Another embodiment provides a method for adjusting the playback of audio content. The method includes at least a portion of the operations performed by the A / V hub.
[0022] Another embodiment provides one or more electronic devices.
[0023] This summary is provided only for the purpose of illustrating some exemplary embodiments in order to provide a basic understanding of some aspects of the subject matter described herein. Accordingly, it is to be understood that the above features are merely illustrative and are not to be construed as limiting the scope or spirit of the subject matter described herein. Other features, aspects, and advantages of the subject matter described herein will become apparent from the following detailed description, the drawings, and the claims.
Brief Description of the Drawings
[0024]
Figure 1
Figure 2
Figure 3
Figure 4
Figure 5
Figure 6
Figure 7
Figure 8
Figure 9
Figure 10
Figure 11
Figure 12
Figure 13
Figure 14
Figure 15
Figure 16
Figure 17
Figure 18
Figure 19
Figure 20
Figure 21
Figure 22
Figure 23
[0025] Note that like reference numerals refer to corresponding elements throughout the drawings. Further, multiple instances of the same element are indicated by a common prefix separated from the example number by a dash.
[0026] In a first group of embodiments, an audio / video (A / V) hub that adjusts the playback of audio content is described. Specifically, the A / V hub can calculate the current time offset between the clock of an electronic device (such as an electronic device including a speaker) and the clock of the A / V hub based on the difference between the transmission time of a frame from the electronic device and the reception time when the frame is received. For example, the current time offset can be calculated using radio ranging by ignoring the distance between the A / V hub and the electronic device. Next, the A / V hub can transmit one or more frames including audio content and playback timing information to the electronic device, and the playback timing information can specify the playback time at which the electronic device plays the audio content based on the current time offset. Furthermore, the playback time of the electronic device can have a temporal relationship such that the playback of the audio content by the electronic device is adjusted.
[0027] By adjusting the playback of audio content by an electronic device, this adjustment technique can provide an improved audio experience in an environment including an A / V hub and an electronic device. For example, this adjustment technique can correct the clock drift between the A / V hub and the electronic device. Instead of or in addition to this, this adjustment technique can correct or adapt to the acoustic characteristics of the environment and / or can be based on the desired acoustic characteristics in the environment. In addition, this adjustment technique can correct the playback time based on the estimated position of the listener relative to the electronic device. In this way, this adjustment technique can improve the acoustic quality and, more generally, the user experience when using the A / V hub and the electronic device. As a result, this adjustment technique can increase the customer loyalty and revenue of the provider of the A / V hub and the electronic device.
[0028] In a second group of embodiments, an A / V hub is described that selectively determines one or more acoustic characteristics of an environment including audio / video (A / V). Specifically, the A / V hub can detect an electronic device (such as an electronic device including a speaker) in the environment using wireless communication. The A / V hub can then determine a change state such as when the electronic device has not been detected in the environment previously and / or a change in the position of the electronic device. In response to determining the change state, the A / V hub can transition to a characterization mode. During the characterization mode, the A / V hub can provide an instruction to the electronic device to play audio content at a specified playback time, determine one or more acoustic characteristics of the environment based on acoustic measurements in the environment, and store the one or more acoustic characteristics and / or the position of the electronic device in memory.
[0029] By selectively determining one or more acoustic characteristics, this characterization technique can facilitate the improvement of the acoustic experience in an environment including an A / V hub and electronic devices. For example, the characterization technique can identify and characterize a modified environment, which can be later used to correct the effects of changes during the playback of audio content by one or more electronic devices (including the electronic devices). In this way, the characterization technique can improve the acoustic quality and, more generally, the user experience when using an A / V hub and electronic devices. As a result, this characterization technique can increase customer loyalty and revenue for providers of A / V hubs and electronic devices.
[0030] In a third group of embodiments, an A / V hub for adjusting the playback of audio content is described. Specifically, the A / V hub can calculate a current time offset between the clock of an electronic device (such as an electronic device including a speaker) and the clock of the A / V hub based on a measured sound corresponding to one or more acoustic characterization patterns, one or more times when the electronic device output the sound, and one or more acoustic characterization patterns. Next, the A / V hub can transmit one or more frames including audio content and playback timing information to the electronic device, and the playback timing information can specify a playback time for the electronic device to play the audio content based on the current time offset. Further, the playback time of the electronic device can have a temporal relationship such that the playback of the audio content by the electronic device is adjusted.
[0031] By adjusting the playback of audio content by an electronic device, this adjustment technique can provide an improved acoustic experience in an environment including an A / V hub and the electronic device. For example, the adjustment technique can correct the clock drift between the A / V hub and the electronic device. Alternatively or in addition, the adjustment technique can correct or adapt to the acoustic characteristics of the environment and / or be based on desired acoustic characteristics in the environment. In addition, the adjustment technique can correct the playback time based on the estimated position of the listener relative to the electronic device. In this way, the adjustment technique can improve the acoustic quality and, more generally, the user experience when using the A / V hub and the electronic device. As a result, the adjustment technique can increase the customer loyalty and revenue of the provider of the A / V hub and the electronic device.
[0032] In a fourth group of embodiments, an audio / video (A / V) hub that calculates an estimated position is described. Specifically, the A / V hub can calculate the estimated position of a listener relative to an electronic device (such as an electronic device including a speaker) in an environment including the A / V hub and the electronic device based on communication with another electronic device, sound measurements in the environment, and / or time-of-flight measurements. The A / V hub can then transmit to the electronic device one or more frames including audio content and playback timing information, where the playback timing information can specify the playback time at which the electronic device can play the audio content based on the estimated position. Further, the playback time of the electronic device can have a temporal relationship such that the playback of the audio content by the electronic device is adjusted.
[0033] By calculating the listener's estimated position, this characterization technique can facilitate an improved acoustic experience in an environment including an A / V hub and electronic devices. For example, the characterization technique can track changes in the position of the listener in the environment, which can later be used to correct or adapt the playback of audio content by one or more electronic devices. In this way, the characterization technique can improve the acoustic quality and, more generally, the user experience when using the A / V hub and electronic devices. As a result, the characterization technique can increase customer loyalty and revenue for providers of A / V hubs and electronic devices.
[0034] In a fifth group of embodiments, an audio / video (A / V) hub that aggregates electronic devices is described. Specifically, the A / V hub can measure the sound corresponding to the audio content output by an electronic device (such as an electronic device including a speaker). The A / V hub can then aggregate the electronic devices into two or more subsets based on the measured sound. Further, the A / V hub can determine playback timing information that can specify, for a subset, the playback time when the electronic devices of a given subset play the audio content. Next, the A / V hub can transmit one or more frames including the audio content and the playback timing information to the electronic devices, where the playback time of the electronic devices in at least a given subset has a temporal relationship such that the playback of the audio content by the electronic devices of the given subset is adjusted.
[0035] By aggregating electronic devices, this characterization technique can facilitate an improved acoustic experience in an environment including an A / V hub and electronic devices. For example, the characterization technique can aggregate electronic devices based on different audio contents, the measured acoustic delay of sound, and / or desired acoustic characteristics in the environment. In addition, the A / V hub can determine the playback volume of a subset used when the subset generates audio content in order to reduce acoustic crosstalk between two or more subsets. In this way, the characterization technique can improve the acoustic quality and, more generally, the user experience when using the A / V hub and electronic devices. As a result, the characterization technique can increase customer loyalty and revenue for providers of A / V hubs and electronic devices.
[0036] In a sixth group of embodiments, an audio / video (A / V) hub that determines equalized audio content is described. Specifically, the A / V hub can measure the sound corresponding to the audio content output by an electronic device (such as an electronic device including a speaker). Next, the A / V hub can compare the measured sound with the desired acoustic characteristics of a first position in the environment based on the acoustic transfer function of the environment in at least one of a first position, a second position of the A / V hub, and a frequency band. For example, this comparison can include calculating the acoustic transfer function of the first position based on the acoustic transfer functions of other positions in the environment and correcting the measured sound based on the calculated acoustic transfer function of the first position. Further, the A / V hub can determine equalized audio content based on the comparison and the audio content. The A / V hub can transmit one or more frames including the equalized audio content to the electronic device to facilitate the output of additional sound corresponding to the equalized audio content by the electronic device.
[0037] By determining equalized audio content, this signal processing technology can facilitate an improved acoustic experience in an environment including an A / V hub and electronic devices. For example, the signal processing can dynamically modify the audio content based on at least one of the estimated position of the listener relative to the position of the electronic device and the acoustic transfer function of the environment in a frequency band. This can enable a desired acoustic characteristic or type of audio reproduction (such as monophonic, stereophonic, or multichannel) to be achieved at the estimated position of the environment. In this way, the signal processing technology can improve the acoustic quality and, more generally, the user experience when using the A / V hub and electronic devices. As a result, the signal processing technology can increase customer loyalty and revenue for providers of A / V hubs and electronic devices.
[0038] In a seventh group of embodiments, an audio / video (A / V) hub that adjusts the playback of audio content is described. Specifically, the A / V hub can calculate a current time offset between the clock of an electronic device (such as an electronic device including a speaker) and the clock of the A / V hub based on the difference between the reception time at which a frame is received from the electronic device and the predicted transmission time of the frame. For example, the predicted transmission time can be based on the adjustment of the clocks of the electronic device and the A / V hub at a previous time and a predefined transmission schedule of the frame. The A / V hub can then transmit one or more frames including the audio content and playback timing information to the electronic device, where the playback timing information can specify a playback time at which the electronic device can play the audio content based on the current time offset. Furthermore, the playback time of the electronic device can have a temporal relationship such that the playback of the audio content by the electronic device is adjusted.
[0039] By adjusting the playback of audio content by an electronic device, this adjustment technique can provide an improved acoustic experience in an environment including an A / V hub and the electronic device. For example, the adjustment technique can correct clock drift between the A / V hub and the electronic device. Alternatively or in addition, the adjustment technique can correct or adapt to the acoustic characteristics of the environment and / or be based on desired (or target) acoustic characteristics in the environment. Additionally, the adjustment technique can correct the playback time based on the estimated position of the listener relative to the electronic device. In this way, the adjustment technique can improve the acoustic quality and, more generally, the user experience when using the A / V hub and the electronic device. As a result, the adjustment technique can increase customer loyalty and revenue for the providers of the A / V hub and the electronic device.
[0040] In the following description, an A / V hub (which may also be referred to as an "adjustment device"), an A / V display device, a portable electronic device, one or more receiving devices, and / or one or more electronic devices (such as speakers and more generally customer electronic devices) can include a radio that transmits packets or frames according to one or more communication protocols such as the Institute of Electrical and Electronics Engineers (IEEE) 802.11 standard (which may also be referred to as "Wi-Fi (registered trademark)" from the Wi-Fi (registered trademark) Alliance in Austin, Texas), Bluetooth (registered trademark) (Bluetooth Special Interest Group in Kirkland, Washington), cellular telephone communication protocols, short-range communication standards or specifications (from the NFC Forum in Wakefield, Massachusetts), and / or another type of wireless interface. For example, cellular telephone communication protocols can include second-generation mobile telecommunication technologies, third-generation mobile telecommunication technologies (such as communication protocols compliant with the International Mobile Telecommunications-2000 specifications by the International Telecommunication Union in Geneva, Switzerland), fourth-generation mobile telecommunication technologies (International Mobile Telecommunications Next Generation specifications by the International Telecommunication Union in Geneva, Switzerland), and / or other cellular telephone communication technologies, or can be compliant therewith. In some embodiments, the communication protocol includes Long Term Evolution, i.e., LTE. However, a wide variety of communication protocols (such as Ethernet) can also be used. Additionally, communication can occur via a wide variety of frequency bands. The portable electronic device, A / V hub, A / V display device, and / or one or more electronic devices can communicate using infrared communication compliant with infrared communication standards (including unidirectional or bidirectional infrared communication).
[0041] Furthermore, the A / V content in the following discussion can include video and associated audio (such as music, sound, dialogue, etc.), video only, or audio only.
[0042] Communication between electronic devices is shown in FIG. 1 and includes a portable electronic device 110 (such as a remote control or cellular phone), one or more A / V hubs (such as A / V hub 112), one or more A / V display devices 114 (such as a television, monitor, computer, and more generally, a display associated with an electronic device), one or more receiving devices (such as receiving device 116, which can be a local wireless receiver associated with a proximity A / V display device 114-1 that can receive A / V content converted frame by frame from an A / V hub 112 for display on an A / V display device 114-1), one or more speakers 118 (and more generally, one or more electronic devices including one or more speakers) and / or one or more content sources 120 associated with one or more content providers (such as a radio receiver, video player, satellite receiver, an access point providing connection to a wired network such as the Internet, a media or content source, a consumer electronic device, an entertainment device, a set-top box, over-the-top content distributed through the Internet or a network without cable involvement, a satellite or multi-system operator, a security camera, a monitoring camera, etc.). A block diagram illustrating a system 100 is shown. Note that A / V hub 112, A / V display device 114, receiving device 116, and speaker 118 may sometimes be collectively referred to as "components" within system 100. However, A / V hub 112, A / V display device 114, receiving device 116, and / or speaker 118 may also sometimes be referred to as "electronic devices".
[0043] Specifically, the portable electronic device 110 and the A / V hub 112 can communicate with each other using wireless communication, and one or more other components within the system 100 (at least one of the A / V display devices 114, the receiving device 116, one of the speakers 118, and / or one of the content sources 120, etc.) can communicate using wired and / or wireless communication. During wireless communication, these electronic devices can communicate wirelessly, transmit advertisement frames on a wireless channel, detect each other by scanning the wireless channel, establish a connection (e.g., by sending an association request), and / or transmit and receive packets or frames (which can include additional information as a payload such as an association request, and / or information indicating communication performance, data, user interface, A / V content, etc.).
[0044] As will be further described below with respect to FIG. 23, the portable electronic device 110, the A / V hub 112, the A / V display device 114, the receiving device 116, the speaker 118, and the content source 120 can include subsystems such as a networking subsystem, a memory subsystem, and a processor subsystem. In addition, one or more of the portable electronic device 110, the A / V hub 112, the receiving device 116, and / or the speaker 118, and optionally the A / V display device 114 and / or the content source 120 can include a radio 122 of the networking subsystem. For example, a radio or receiving device can be incorporated within the A / V display device, for example, radio 122-5 is included in A / V display device 114-2. Note that radio 122 can be the same instance of a radio or can be different from each other. More generally, the portable electronic device 110, the A / V hub 112, the receiving device 116, and / or the speaker 118 (and optionally the A / V display device 114 and / or the content source 120) can include (or can be included within) any electronic device having a networking subsystem that enables the portable electronic device 110, the A / V hub 112, the receiving device 116, and / or the speaker 118 (and optionally the A / V display device 114 and / or the content source 120) to communicate wirelessly with each other. This wireless communication can include transmitting advertisements on a wireless channel for the electronic devices to make initial contact or detect each other, and then exchanging data / management frames (association requests and responses) to establish a connection, configure security options (e.g., Internet Protocol security), and transmit and receive packets or frames over the connection, etc.
[0045] As can be seen in FIG. 1, the wireless signal 124 (represented by the wavy line) is transmitted from the radio 122-1 of the portable electronic device 100. These wireless signals can be received by at least one of the A / V hub 112, the receiving device 116, and / or the speaker 118 (and optionally one or more of the A / V display device 114 and / or the content source 120). For example, the portable electronic device 110 can transmit packets. These packets can be received by the radio 122-2 of the A / V hub 112. This can enable the portable electronic device 100 to transmit information to the A / V hub 112. FIG. 1 shows the portable electronic device 110 transmitting packets, but it should be noted that the portable electronic device 110 can also receive packets from one or more other components in the A / V hub 112 and / or the system 100. More generally, wireless signals can be transmitted and / or received by one or more of the components of the system 100.
[0046] In the described embodiments, the processing of packets or frames at the portable electronic device 110, A / V hub 112, receiving device 116, and / or speaker 118 (and optionally one or more of the A / V display device 114 and / or content source 120) includes receiving a wireless signal 124 having the packet or frame, decoding / extracting the packet or frame from the received wireless signal 124 to obtain the packet or frame, and processing the packet or frame to determine information (such as information associated with a data stream) included in the packet or frame. For example, the information from the portable electronic device 110 can include user interface activity information associated with a user interface displayed on the touch sensor display (TSD) 128 of the portable electronic device 110, and can be used by a user of the portable electronic device 110 to control at least one of the at least A / V hub 112, at least one of the A / V display devices 114, at least one of the speakers 118, and / or at least one of the content sources 120. (In some embodiments, instead of or in addition to the touch sensor display 128, the portable electronic device 110 includes a user interface with physical knobs and / or buttons that can be used by a user to control at least one of the A / V hub 112, one of the A / V display devices 114, at least one of the speakers 118, and / or one of the content sources 120. Alternatively, one or more of the portable electronic device 110, A / V hub 112, one or more of the A / V display devices 114, receiving device 116, one or more of the speakers 118, and / or one or more of the content sources 120 can specify communication performance regarding communication between the portable electronic device 110 and one or more other components within the system 100.)Furthermore, the information from the A / V hub 112 can include device state information (on, off, playing, rewinding, fast-forwarding, selected channel, selected A / V content, content source, etc.) regarding the current device state of at least one of the A / V display devices 114, at least one of the speakers 108, and / or one of the content sources 120, or can include user interface information of the user interface (which can be dynamically updated based on the device state information and / or user interface activity information). Additionally, the information from at least one of the A / V hub 112 and / or the content source 120 can include audio and / or video (sometimes referred to as "audio / video" or "A / V" content) that is displayed or presented on one or more of the A / V display devices 114, and display commands that specify how the audio and / or video is to be displayed or presented.
[0047] However, as described above, the audio and / or video can be transmitted between the components of the system 100 via wired communication. Thus, as shown in FIG. 1, there can be a wired cable or link such as a High-Definition Multimedia Interface (HDMI (registered trademark)) cable 126 between the A / V hub 112 and the A / V display device 114-3. The audio and / or video can be included within the HDMI (registered trademark) content or associated with the HDMI (registered trademark) content, and in other embodiments, the audio content can be included within or associated with A / V content that conforms to another format or standard used in the disclosed communication technology embodiments. For example, the A / V content can include, or conform to, H.264, MPEG-2, QuickTime video format, MPEG-4, MP4, and / or TCP / IP. Further, the video mode of the A / V content can be 720p, 1080i, 1080p, 1440p, 2000, 2160p, 2540p, 4000p, and / or 4320p.
[0048] The A / V hub 112 can determine a display command (having a display layout) for A / V content based on the format of a display of one of the A / V display devices 114, such as the A / V display device 114-1. Alternatively, the A / V hub 112 can use a pre-determined display command, or the A / V hub 112 can modify or convert the A / V content based on the display layout so that the modified or converted A / V content has an appropriate format for display on the display. Further, the display command can specify the information to be displayed on the display of the A / V display device 114-1, including the location (such as the central window, tile window, etc.) where the A / V content is to be displayed. As a result, the information to be displayed (i.e., an instance of the display command) can be based on the format of the display, such as the display size, display resolution, display aspect ratio, display contrast ratio, display type, etc. Still further, when the A / V hub 112 receives A / V content from one of the content sources 120, the A / V content and the display command can be provided to the A / V display device 114-1 so that the A / V content is displayed on the display of the A / V display device 114-1. For example, the A / V hub 112 can collect the A / V content in a buffer until a frame is received, and the A / V hub 112 can provide the complete frame to the A / V display device 114-1. Alternatively, the A / V hub 112 can provide a packet having a portion of the frame to the A / V display device 114-1 when received. In some embodiments, the display command can be provided specifically to the A / V display device 114-1 (such as when the display command changes), periodically or cyclically (such as one for every N packets or for the packets of each frame, etc.) or for each packet.
[0049] Furthermore, the communication between one or more of the portable electronic device 110, the A / V hub 112, the A / V display device 114, the receiving device 116, one or more of the speakers 118, and / or one or more content sources 120 can be characterized by a variety of performance metrics such as received signal strength indicator (RSSI), data rate, data rate discounting wireless protocol overhead (sometimes called "throughput"), error rate (packet error rate, or retry or retransmission rate), mean squared error of the equalization signal relative to the equalization target, intersymbol interference, multipath interference, signal-to-noise ratio, width of the eye pattern, time interval (the latter half of which is sometimes called the "capacity" of the channel or link), ratio of the number of bytes successfully transmitted during the time interval (such as 1 to 10 s) to the estimated maximum number of bytes that can be transmitted, and / or ratio of the actual data rate to the estimated maximum data rate (sometimes called "utilization"). Also, the performance during communication associated with different channels can be monitored individually or together (e.g., to identify dropped packets).
[0050] Communication among one of the portable electronic device 110, the A / V hub 112, the A / V display device 114, the receiving device 116, one of the speakers 118, and / or one or more of the content sources 120 of FIG. 1 can include one or more independent simultaneous data streams in different radio channels (or different equivalent communication protocols such as different Wi-Fi communication protocols) in one or more connections or links transmitted using multiple radios. Note that each of the one or more connections or links can have a separate or different identifier (such as a different service set identifier) in the wireless network of the system 100 (which can be a proprietary network or a public network). Further, the one or more simultaneous data streams are, on a dynamic or per-packet basis, to improve or maintain performance metrics even when there are transient changes (such as interference, changes in the amount of information to be transmitted, movement of the portable electronic device 100, etc.), and for channel calibration, determination of one or more performance metrics, execution of service quality characterization without interrupting communication (execution of channel estimation, execution of link quality, execution of channel calibration, and / or execution of spectral analysis associated with at least one channel), seamless handoff between different radio channels, and coordinated communication between components, etc. Services (while conforming to the communication protocol, for example, the Wi-Fi communication protocol) can be made partially or fully redundant. These features can reduce the number of packets to be retransmitted, thus reducing communication latency and avoiding communication interruptions, and can enhance the experience of one or more users who view A / V content on one or more of the A / V display devices 114 and / or listen to the audio output by one or more of the speakers 118.
[0051] As described above, the user can control at least one of the A / V hub 112, at least one of the A / V display devices 114, at least one of the speakers 118, and / or at least one of the content sources 120 via a user interface displayed on the touch sensor display 128 of the portable electronic device 110. Specifically, at a given time, the user interface can include one or more virtual icons that enable the user to activate, stop, or change the function or ability of at least one of the A / V hub 112, at least one of the A / V display devices 114, at least one of the speakers 118, and / or at least one of the content sources 120. For example, a given virtual icon of the user interface can have an associated strike area on the surface of the touch sensor display 128. When the user touches and then releases the surface within the strike area (e.g., using one or more fingers or a stylus), the portable electronic device 110 (such as a processor executing a program module) can receive user interface activity information that instructs the activation of this command or instruction from a touch screen input / output (I / O) controller coupled to the touch sensor display 128. (Alternatively, the touch sensor display 128 can respond to pressure. In these embodiments, the user can maintain contact with the touch sensor display 128 with an average contact pressure that is typically less than a threshold value such as 10 - 20 kPa, and can activate a given virtual icon by raising the average contact pressure with the touch sensor display 128 above the threshold value.) In response, the program module can command the interface circuit of the portable electronic device 110 to wirelessly communicate the user interface activity information that instructs a command or instruction to the A / V hub 112, and the A / V hub 112 can transmit the command or instruction to components of the system 100 (such as the A / V display device 114-1).This instruction or command can result in, for example, turning on or off the A / V display device 114-1, displaying A / V content from a specific content source, and performing trick modes of operation (fast forward, rewind, fast rewind, or skip). For example, the A / V hub 112 can request A / V content from the content source 120-1 and then provide the A / V content with a display command to the A / V display device 114-1, whereby the A / V display device 114-1 displays the A / V content. Alternatively or in addition, the A / V hub 112 can provide audio content associated with video content from the content source 120-1 to one or more of the speakers 118.
[0052] As described above, it is often difficult to achieve high audio quality in an environment (such as a room, building, vehicle, etc.). Specifically, achieving high audio quality in an environment usually imposes strong constraints on the adjustment of loudspeakers such as the speaker 118. For example, this adjustment may need to be maintained with an accuracy of 1 to 5 μs (exemplary values, not limiting). In some embodiments, the adjustment includes synchronization in the time domain within time or phase accuracy and / or in the frequency domain within frequency accuracy. Without proper adjustment, the acoustic quality in the environment may deteriorate due to a corresponding impact on listener satisfaction and the overall user experience when listening to audio content and / or A / V content.
[0053] This problem can be addressed by an adjustment technique of directly or indirectly adjusting the speaker 118 by the A / V hub 112. As will be described below with respect to FIGS. 2-4, in some embodiments, the adjusted playback of audio content by the speaker 118 can be facilitated using wireless communication. Specifically, since the speed of light is approximately six orders of magnitude faster than the speed of sound, the propagation delay of wireless signals in the environment (such as a room) can be ignored with respect to the desired adjustment accuracy of the speaker 118. For example, the desired adjustment accuracy of the speaker 118 can be on the order of microseconds, and the propagation delay in a typical room (e.g., exceeding a distance of up to 10-30 m) can be made one or two orders of magnitude smaller. As a result, the speaker 118 can be adjusted using techniques such as wireless ranging or radio-based distance measurement. Specifically, during wireless ranging, the A / V hub 112 can transmit a frame or packet including the transmission time and identifier of the A / V hub 112 based on the clock of the A / V hub 112, and a given one of the speakers 118 (such as speaker 118-1) can determine the arrival or reception time of the frame or packet based on the clock of the speaker 118-1.
[0054] Alternatively, the speaker 118-1 can transmit a frame or packet (which may be referred to as an "input frame") including the transmission time and identifier of the speaker 118-1 based on the clock of the speaker 118-1, and the A / V hub 112 can determine the arrival or reception time of the frame or clock based on the clock of the A / V hub 112. More generally, the distance between the A / V hub 112 and the speaker 118-1 is determined based on the product of the time of flight (the difference between the arrival time and the transmission time) and the propagation speed. However, by ignoring the physical distance between the A / V hub 112 and the speaker 118-1, i.e., assuming immediate propagation (introducing a static offset that can be ignored by fixed devices in the same room or environment), the difference between the arrival time and the transmission time can be used to dynamically track the drift or current time offset (as well as the static offset that can be ignored) in the adjustment of the clocks of the A / V hub 112 and the speaker 118-1.
[0055] The current time offset can be determined by the A / V hub 112 or provided to the A / V hub 112 by the speaker 118-1. The A / V hub 112 can transmit one or more frames (which may also be referred to as "output frames") including audio content and playback timing information to the speaker 118-1, and the playback timing information can specify the playback time at which the speaker 118-1 plays the audio content based on the current time offset. This can be repeated for the other speakers 118. Furthermore, the playback times of the speakers 118 can have a temporal relationship such that the playback of the audio content by the speakers 118 is adjusted.
[0056] In addition to correcting the drift within the clock, this adjustment technique (as well as other embodiments of the adjustment techniques described below) can provide an improved acoustic experience in an environment including the A / V hub 112 and the speakers 118. For example, the adjustment technique can correct or adapt the pre-determined or dynamically determined acoustic characteristics of the environment based on the desired acoustic characteristics in the environment (e.g., the type of playback such as monophonic, stereophonic, and / or multi-channel, the acoustic radiation pattern such as directed or diffusive, ease of penetration, etc.) and / or based on the dynamically estimated positions of one or more listeners relative to the speakers 118 (described below with respect to FIGS. 14-16). In addition, the adjustment technique can be used with respect to the dynamic aggregation of the speakers 118 into groups (described below with respect to FIGS. 17-19) and / or the difference between the audio content based on the dynamically equalized audio content to be played and the acoustic characteristics in the environment and the desired acoustic characteristics (described below with respect to FIGS. 20-22).
[0057] Note that wireless ranging (and generally wireless communication) can be performed in or within one or more bands of frequencies such as the 2 GHz wireless band, 5 GHz wireless band, ISM band, 60 GHz wireless band, ultra-wideband, etc.
[0058] In some embodiments, one or more additional communication techniques can be used to identify and / or act on multipath radio signals during the adjustment of speaker 118. For example, A / V hub 112 and / or speaker 118 can use a directional antenna, an array of antennas with known time differences of arrival, and / or angles of arrival at two receivers with known positions (i.e., trilateration or multilateration) to determine the angle of arrival (including non-line-of-sight reception).
[0059] As described below with respect to FIGS. 5-7, another way to adjust speaker 118 can be to use a scheduled transmission time. Specifically, during the calibration mode, the clocks of A / V hub 112 and speaker 118 can be adjusted. Subsequently, in the normal operation mode, A / V hub 112 can transmit a frame or packet with the identifier of A / V hub 112 at a pre-defined transmission time based on the clock of A / V hub 112. However, due to the relative drift of the clock of A / V hub 112, these packets or frames will arrive or be received at speaker 118 at a time different from the expected pre-defined transmission time based on the clock of speaker 118. Thus, by again ignoring the propagation delay, the difference between the arrival time of a given frame at a given one of speakers 118 (such as speaker 118-1) and the pre-defined transmission time can be used to dynamically track the drift or current time offset in the adjustment of the clocks of A / V hub 112 and speaker 118-1 (as well as a static offset that may be ignored and associated with the propagation delay).
[0060] Alternatively or in addition thereto, after the calibration mode, the speaker 118 can transmit a frame or packet with the identifier of the speaker 118 at a predefined transmission time based on the clock of the speaker 118. However, due to the drift of the clock of the speaker 118, these packets or frames will arrive at or be received by the A / V hub 112 at a time different from the expected predefined transmission time based on the clock of the A / V hub 112. Therefore, by ignoring the propagation delay again, the difference between the arrival time of a given frame from a given one of the speakers 118 (such as speaker 118-1) and the predefined transmission time can dynamically track the drift or current time offset in the clock adjustment of the A / V hub 112 and the speaker 118-1 (as well as an optional static offset associated with the propagation delay).
[0061] Similarly in this case, the current time offset can be determined by the A / V hub 112 or provided to the A / V hub 112 by one or more of the speakers 118 (such as speaker 118-1). Note that in some embodiments, the current time offset is further based on a model of the clock drift of the A / V hub 112 and the speaker 118. Then, the A / V hub 112 can transmit one or more frames including audio content and playback timing information to the speaker 118-1, where the playback timing information can specify the playback time at which the speaker 118-1 plays the audio content based on the current time offset. This can be repeated for other speakers 118. Furthermore, the playback times of the speakers 118 can have a temporal relationship such that the playback of the audio content by the speakers 118 is adjusted.
[0062] Furthermore, note that one or more additional communication technologies can be used in these embodiments to identify and / or implement multipath radio signals during the adjustment of the speaker 118.
[0063] As described below with respect to FIGS. 8-10, another way to adjust the speaker 118 can be to use acoustic measurements. Specifically, during the calibration mode, the clocks of the A / V hub 112 and the speaker 118 can be adjusted. Subsequently, the A / V hub 112 can output a sound corresponding to an acoustic characterization pattern that uniquely identifies the A / V hub 112 (a sequence of pulses, different frequencies, etc.) at a predefined transmission time. This acoustic characterization pattern can be output at frequencies outside the range of human hearing (such as ultrasonic frequencies). However, due to the relative drift in the clock of the A / V hub 112, the sound corresponding to the acoustic characterization pattern is measured (i.e., arrives or is received) at the speaker 118 at a time different from the expected predefined transmission time based on the clock of the speaker 118. In these embodiments, it is necessary to correct for the different times of contributions associated with acoustic propagation based on the predefined or known positions of the A / V hub 112 and the speaker 118 and / or using wireless ranging. For example, the position can be determined using triangulation and / or trilateration in a local positioning system, a global positioning system, and / or a wireless network (such as a cellular phone network or a WLAN). Thus, after correcting for the acoustic propagation delay, the difference between the arrival time of a given frame of a given one of the speakers 118 (such as speaker 118-1) and the predefined transmission time can be used to dynamically track the drift or current time offset in the adjustment of the clocks of the A / V hub 112 and the speaker 118-1.
[0064] Alternatively or in addition to this, after the calibration mode, the speaker 118 can output a sound corresponding to an acoustical characterization pattern (such as a sequence of different pulses, different frequencies, etc.) that uniquely identifies the speaker 118 at a predefined transmission time at the predefined transmission time. However, due to the relative drift in the clock of the speaker 118, the sound corresponding to the acoustical characterization pattern is measured (i.e., arrives or is received) at the A / V hub 112 at a time different from the expected predefined transmission time based on the clock of the A / V hub 112. In these embodiments, it is necessary to correct the different times of contribution associated with the acoustic propagation delay based on the predefined or known positions of the A / V hub 112 and the speaker 118 and / or using wireless ranging. Therefore, after correcting the acoustic propagation delay, the difference between the arrival time of a given frame from a given one of the speakers 118 (such as speaker 118-1) and the predefined transmission time can dynamically track the drift or current time offset in the adjustment of the clocks of the A / V hub 112 and the speaker 118-1.
[0065] Similarly in this case, the current time offset can be determined by the A / V hub 112 or provided to the A / V hub 112 by the speaker 118-1. The A / V hub 112 can then transmit one or more frames including audio content and playback timing information to the speaker 118-1, where the playback timing information can specify the playback time at which the speaker 118-1 plays the audio content based on the current time offset. This can be repeated for other speakers 118. Further, the playback times of the speakers 118 can have a temporal relationship such that the playback of the audio content by the speakers 118 is adjusted.
[0066] Although a network environment shown in FIG. 1 is described as an example, in alternative embodiments, there may be a different number or type of electronic devices. For example, some embodiments may include more or fewer electronic devices. As another example, in another embodiment, different electronic devices may be transmitting and / or receiving packets or frames. Although the portable electronic device 110 and the A / V hub 112 are illustrated by a single instance of the radio 112, in other embodiments, the portable electronic device 110 and the A / V hub 112 (and optionally the A / V display device 114, the receiving device 116, the speaker 118, and / or the content source 120) can include multiple radios.
[0067] Embodiments of communication techniques are described herein. FIG. 2 presents a flowchart showing a method 200 for adjusting the playback of audio content that can be performed by an A / V hub (FIG. 1), such as the A / V hub 112. During operation, an A / V hub (such as a processor that executes a program module of the A / V hub, for example, a control circuit or a logic circuit) can receive frames (operation 210) or packets from one or more electronic devices via wireless communication, and a given frame or packet includes the transmission time at which a given electronic device transmitted the given frame or packet.
[0068] Next, the A / V hub can store the reception time at which a frame or packet was received (operation 212), where the reception time is based on the A / V hub's clock. For example, the reception time can be added to an instance of a packet, or a frame or packet, received from one of the electronic devices by a physical layer and / or media access control (MAC) layer within or associated with the interface circuit of the A / V hub. Note that the reception time is associated with the leading edge of the packet, such as a reception time signal associated with the leading edge, or the trailing edge of the packet, such as a reception clear signal associated with the trailing edge, or a point associated with the frame or packet. Similarly, a transmission time can be added to an instance of a frame or packet transmitted by one of the electronic devices by a physical layer and / or MAC layer within or associated with the interface circuit within the electronic device. In some embodiments, the transmission and reception times are determined and added to the frame or packet by the wireless ranging capabilities of a physical layer and / or MAC layer within or associated with the interface circuit.
[0069] Furthermore, the A / V hub can calculate a current time offset between the clock of the electronic device and the clock of the A / V hub based on the reception time and transmission time of the frame or packet (operation 214). Additionally, the current time offset can be calculated by the A / V hub based on a model of clock drift of an electronic circuit, such as an electronic circuit model of a clock circuit, and / or a look-up table of clock drift as a function of time. Note that the electronic device is positioned at a non-zero distance from the A / V hub, and the current time offset can be calculated based on the transmission time and reception time using wireless ranging by ignoring the distance.
[0070] Next, the A / V hub can send one or more frames (operation 216) or packets that include audio content and playback timing information to the electronic device. It should be noted that the playback timing information specifies the playback time at which the electronic device plays the audio content based on the current time offset. Additionally, the playback time of the electronic device can have a temporal relationship such that the playback of the audio content by the electronic device is adjusted. The temporal relationship has a non-zero value, and it should be noted that at least a portion of the electronic devices can be commanded to play the audio content with a phase relative to each other by using different values of the playback time. For example, the different playback times can be based on predetermined or dynamically determined acoustic characteristics of the environment that includes the electronic device and the A / V hub. Alternatively or in addition, the different playback times can be based on desired acoustic characteristics in the environment.
[0071] In some embodiments, the A / V hub optionally performs one or more additional operations (operation 218). For example, the electronic device can be positioned at a vector distance from the A / V hub, and the interface circuit can determine the magnitude of the vector distance based on the transmission time and the reception time using wireless ranging. Further, the interface circuit can determine the angle of the vector distance based on the angle of arrival of the wireless signal associated with the frame or packet received by one or more antennas during wireless communication. Additionally, the different playback times can be based on the determined vector distance. For example, the playback time can correspond to the determined vector distance such that the sound associated with the audio content from different electronic devices at different positions in the environment reaches the position in the environment (e.g., the position of the A / V hub, the center of the environment, the user's preferred listening position, etc.) with a desired phase relationship or to achieve the desired acoustic characteristics of this position.
[0072] Alternatively or in addition, different playback times are associated with sound from different electronic devices at different positions in the environment, based on the estimated position of the listener relative to the electronic device, such that the sound can reach the estimated position of the listener with a desired phase relationship or achieve the desired acoustic characteristics of the estimated position. Techniques that can be used to determine the position of the listener are described below with respect to FIGS. 14-16.
[0073] The wireless ranging capability in the interface circuit can include the adjusted clocks of the A / V hub and the electronic device. Note that in other embodiments, the clock is not adjusted. Thus, a wide variety of radio detection techniques can be used. In some embodiments, the wireless ranging capability includes the use of transmissions through a GHz or multi-GHz bandwidth, resulting in short-duration pulses (e.g., about 1 ns).
[0074] FIG. 3 shows the connection between the A / V hub 112 and the speaker 118-1. Specifically, the interface circuit 310 of the speaker 118-1 can transmit one or more frames or packets (such as packet 312) to the A / V hub 112. The packet 312 can include a corresponding transmission time 314 based on the interface clock 316 provided by the interface clock circuit 318 within or associated with the interface circuit 310 in the speaker 118-1 when the speaker 118-1 transmits the packet 312. When the interface circuit 320 of the A / V hub 112 receives the packet 312, the packet 312 can include a reception time 322 (or the reception time 322 can be stored in the memory 324). For each packet, the corresponding reception time can be based on the interface clock 326 provided by the interface clock circuit 328 within or associated with the interface circuit 318.
[0075] Next, the interface circuit 320 can calculate the current time offset 330 between the interface clock 316 and the interface clock 326 based on the difference between the transmission time 314 and the reception time 322. Further, the interface circuit 320 can provide the current time offset 330 to the processor 332. (Alternatively, the processor 332 can calculate the current time offset 330.)
[0076] Furthermore, the processor 332 can provide the playback timing information 334 and the audio content 336 to the interface circuit 320, and the playback timing information 334 specifies the playback time when the speaker 118-1 generates the audio content 336 based on the current time offset 330. In response, the interface circuit 330 can transmit one or more frames or packets 338 including the playback timing information 334 and the audio content 336 to the speaker 118-1. (However, in some embodiments, the playback timing information 334 and the audio content 336 are transmitted using separate or different frames or packets.)
[0077] After the interface circuit 310 receives one or more frames or packets 338, it can provide the playback timing information 334 and the audio content 336 to the processor 340. The processor 340 can execute software that performs the playback operation 342. For example, the processor 340 can store the audio content 336 in a queue in the memory. In these embodiments, the playback operation 350 includes outputting the audio content 336 from the queue, including the step of driving the electroacoustic transducer at the speaker 118-1 based on the audio content 336 such that the speaker 118-1 outputs sound at the time specified by the playback timing information 334.
[0078] In an exemplary embodiment, a communication technique is used to adjust the playback of audio content by the speaker 118. This is shown in FIG. 4, which presents a diagram showing the steps of adjusting the playback of audio content by the speaker 118. Specifically, when frames or packets (such as packet 410-1) are transmitted by the speaker 118, they can include information indicating the transmission time (such as transmission time 412-1). For example, the physical layer of the interface circuit of the speaker 118 can include the transmission time within the packet 410. Note that in FIG. 4 and other embodiments below, the information of the frame or packet can be included at any position (such as the beginning, middle, and / or end).
[0079] When the packet 410 is received by the A / V hub 112, additional information indicating the reception time (such as the reception time 414-1 of the packet 410-1) can be included in the packet 10. For example, the physical layer of the interface circuit of the A / V hub 112 can include the reception time in the packet 410. Furthermore, the transmission time and the reception time can be used to track the clock drift of the A / V hub 112 and the speaker 118.
[0080] The A / V hub 112 can calculate the current time offset between the clock of the speaker 118 and the clock of the A / V hub 112 using the transmission time and the reception time. Furthermore, the current time offset can also be calculated by the A / V hub 112 based on the model of the A / V hub 112 from the clock drift of the speaker 118. For example, a model of relative or absolute clock drift can include polynomials or cubic splines (and more generally functions) having parameters that indicate or estimate the clock drift of a given speaker as a function of time based on the historical time offset.
[0081] Subsequently, the A / V hub 112 can transmit one or more packets or frames or packets including the audio content 420 and playback timing information (such as the playback timing information 418-1 in the packet 416-1) to the speaker 118, and the playback timing information indicates the playback time for the speaker 118 device to play the audio content 420 based on the current time offset. The playback time of the speaker 118 can have a temporal relationship such that the playback of the audio content 420 by the speaker 118 is adjusted so that, for example, the associated sound or wavefront arrives at the position 422 of the environment having the desired phase relationship.
[0082] Another embodiment of the adjustment in the communication technology is shown in FIG. 5, presenting a flowchart showing a method 500 for adjusting the playback of audio content. Note that the method 500 can be executed by an A / V hub such as the A / V hub 112 (FIG. 1). During operation, the A / V hub (control circuit or control logic, for example, a processor that executes a program module in the A / V hub) can receive a frame (operation 510) or a packet from an electronic device via wireless communication.
[0083] The A / V hub can store the reception time when the frame or packet is received (operation 512), and this reception time is based on the clock of the A / V hub. For example, the reception time can be added to an instance of the frame or packet received from one of the electronic devices by the physical layer and / or MAC layer within or associated with the interface circuit of the A / V hub. Note that the reception time can be associated with the leading edge or trailing edge of the frame or packet, such as a reception time signal associated with the leading edge or a reception clear signal associated with the trailing edge.
[0084] Furthermore, the A / V hub can calculate the current time offset between the clock of the electronic device and the clock of the A / V device based on the reception time and the predicted transmission time of the frame or packet (operation 514), and the predicted transmission time is based on the adjustment of the clocks of the electronic device and the A / V hub at the previous time and a predefined transmission schedule of the frame or packet (such as every 10 or 100 ms, which is an example rather than a limitation). For example, during the initialization mode, the time offset between the clock of the electronic device and the clock of the A / V hub can be eliminated (i.e., the adjustment can be set). Note that the predefined transmission time of the transmission schedule can include or be other than the WLAN beacon transmission time. Subsequently, the clocks can have a relative drift that can be tracked based on the difference between the reception time and the predicted transmission time of the frame or packet. In some embodiments, the current time offset is calculated by the A / V hub based on a model of the clock drift of the electronic device.
[0085] Next, the A / V hub can send one or more frames (operation 516) or packets containing audio content and playback timing information to the electronic device, and the playback timing information indicates the playback time at which the electronic device plays the audio content based on the current time offset. Furthermore, the playback time of the electronic device can have a temporal relationship such that the playback of the audio content by the electronic device is adjusted. The temporal relationship can have a non-zero value, and note that this causes at least some of the electronic devices to be commanded to play audio content having a phase relative to each other by using different values of the playback time. For example, the different playback times can be based on the predefined or dynamically determined acoustic characteristics of the environment including the electronic device and the A / V hub. Alternatively or in addition, the different playback times can be based on the desired acoustic characteristics in the environment.
[0086] In some embodiments, the A / V hub optionally performs one or more additional operations (operation 518). For example, an electronic device can be positioned from the A / V hub by a vector distance, and the interface circuit can determine the magnitude of the vector distance based on the transmission time and the reception time using wireless ranging. Further, the interface circuit can determine the angle of the vector distance based on the angle of arrival of a wireless signal associated with a frame or packet received by one or more antennas during wireless communication. Additionally, different playback times can be based on the determined vector distance. For example, the playback time can be such that sound associated with audio content from different electronic devices at different positions in the environment arrives at a position in the environment (e.g., the position of the A / V hub, the center of the environment, the user's preferred listening position, etc.) having a desired phase relationship, or can achieve the desired acoustic characteristics of the position, corresponding to the determined vector distance.
[0087] Alternatively or in addition, different playback times are based on the estimated position of the listener relative to the electronic device such that sound associated with audio content from different electronic devices at different positions in the environment can reach the estimated position of the listener having a desired phase relationship or can achieve the desired acoustic characteristics of the estimated position. Techniques that can be used to determine the position of the listener are described below with respect to FIGS. 14 - 16.
[0088] FIG. 6 is a diagram showing communication between the portable electronic device 110, the A / V hub 112, and the speaker 118-1. Specifically, during the initialization mode, the interface circuit 610 of the A / V hub 112 can transmit a frame or packet 612 to the interface circuit 614 of the speaker 118-1. This packet can include information 608 that adjusts the clocks 628 and 606 provided by the interface clock circuits 616 and 618, respectively. For example, this information can eliminate the time offset between the interface clock circuits 616 and 618 and / or can set the interface clock circuits 616 and 618 to the same clock frequency.
[0089] Thereafter, the interface circuit 614 can transmit one or more frames or packets (such as packet 620) to the A / V hub 112 at a predefined transmission time 622.
[0090] When the interface circuit 610 of the A / V hub 112 receives the packet 620, it can include the reception time 624 in the packet 620 (or can store the reception time 624 in the memory 626). For each packet, the corresponding reception time can be based on the interface clock 628 provided by the interface clock circuit 616 within or associated with the interface circuit 610.
[0091] The interface circuit 610 can calculate the current time offset 630 between the interface clock 628 and the interface clock 606 based on the difference between the transmission time 622 and the reception time 624. The interface circuit 610 can provide the current time offset 630 to the processor 632. (Alternatively, the processor 632 can calculate the current time offset 630.)
[0092] Furthermore, the processor 632 can provide the playback timing information 634 and the audio content 636 to the interface circuit 610, and the playback timing information 634 specifies the playback time for the speaker 118-1 to play the audio content 636 based on the current time offset 630. In response, the interface circuit 610 can transmit one or more frames or packets 638 including the playback timing information 634 and the audio content 636 to the speaker 118-1. (However, in some embodiments, the playback timing information 634 and the audio content 636 are transmitted using different or separate frames or packets.)
[0093] After the interface circuit 614 receives one or more frames or packets 638, it can provide the playback timing information 634 and the audio content 636 to the processor 640. The processor 640 can execute software to perform the playback operation 642. For example, the processor 640 can store the audio content 636 in a queue in the memory. In these embodiments, the playback operation 650 includes outputting the audio content 636 from the queue, which includes driving the electroacoustic transducer of the speaker 118-1 based on the audio content 636 when the speaker 118-1 outputs sound at the time specified by the playback timing information 634.
[0094] In an exemplary embodiment, a communication technique is used to adjust the playback of the audio content by the speaker 118. This is shown in FIG. 7, which presents a diagram showing the steps of adjusting the playback of the audio content by the speaker 118. Specifically, the A / V hub 112 can transmit a frame or packet 710 to the speaker 118 together with information (such as the information 708 in the packet 710-1) for adjusting the clock provided by the clock circuits in the A / V hub 112 and the speaker 118.
[0095] Speaker 118 can transmit frame or packet 712 to A / V hub 112 at a pre-determined transmission time. When these frames or packets are received by A / V hub 112, information indicating the reception time can be included in packet 712 (such as reception time 714-1 of packet 712-1). Using the pre-determined transmission time and reception time, the clock drifts of A / V hub 112 and speaker 118 can be tracked.
[0096] A / V hub 112 can calculate the current time offset between the clock of speaker 118 and the clock of A / V hub 112 using the pre-determined transmission time and reception time. Furthermore, this current time offset can be calculated by A / V hub 112 based on the model of A / V hub 112 from the clock drift in speaker 118. For example, a model of relative or absolute clock drift can include a polynomial or cubic spline (and more generally a function) with parameters that indicate or estimate the clock drift of a given speaker as a function of time based on the historical time offset.
[0097] Thereafter, A / V hub 112 can transmit one or more frames or packets including audio content 720 and playback timing information (such as playback timing information 718-1 in packet 716-1) that specifies the playback time for speaker 118 to play audio content 720 based on the current time offset to speaker 118. The playback time of speaker 118 can have a temporal relationship such that, for example, the playback of audio content 720 by speaker 118 is adjusted so that the associated sound or wavefront reaches position 722 in the environment with a desired phase relationship.
[0098] Another embodiment of the adjustment in the communication technology is illustrated in FIG. 8 presenting a flowchart of a method 800 for adjusting the playback of audio content. Note that the method 800 can be executed by an A / V hub such as the A / V hub 112 (FIG. 1). During operation, the A / V hub (control circuit or control logic, e.g., a processor executing program modules in the A / V hub) can measure (operation 810) the sound output by an electronic device in an environment including the A / V hub using one or more acoustic transducers of the A / V hub, and the sound can correspond to one or more acoustic characterization patterns. For example, the measured sound can include sound pressure. Note that the acoustic characterization pattern can include pulses. Further, the sound can be in a frequency range outside of human hearing, such as ultrasonic waves.
[0099] Furthermore, a given electronic device can output sound at one or more different times other than those used by the rest of the electronic devices, whereby the sound from the given electronic device can be identified or distinguished from the sound output by the rest of the electronic devices. Alternatively or in addition, the sound output by a given electronic device can correspond to a given acoustic characterization pattern that can be different from that used by the rest of the electronic devices. Thus, the acoustic characterization pattern can uniquely identify the electronic device.
[0100] Next, the A / V hub can calculate the current time offset between the clock of the electronic device and the clock of the A / V hub based on the measured sound, the one or more times when the electronic device output the sound, and the one or more acoustic characterization patterns (operation 812). For example, the A / V hub can correct the measured sound based on the acoustic characteristics of the environment, such as the acoustic delay associated with at least one specific frequency in at least one band of frequencies (such as 100 - 20,000 Hz, which is an example rather than a limitation) of the environment or an acoustic transfer function that is predetermined (or dynamically determined), and the output time can be compared with the activated output time or a predetermined output time. This enables the A / V hub to determine the original output sound without spectral filtering or distortion associated with the environment, and allows the A / V hub to accurately determine the current time offset.
[0101] The measured sound can include information indicating the one or more times when the electronic device output the sound (for example, the pulse of the acoustic characterization pattern can indicate time), and it should be noted that the one or more times can correspond to the clock of the electronic device. Alternatively or in addition, the A / V hub can optionally provide to the electronic device the one or more times when the electronic device outputs the sound via wireless communication (operation 808), and the one or more times can correspond to the clock of the A / V hub. For example, the A / V hub can transmit one or more frames or packets to the electronic device at the one or more times. Thus, the A / V hub can activate the output of the sound or output the sound at a predetermined output time.
[0102] Next, the A / V hub can transmit, using wireless communication, one or more frames (operation 814) or packets containing audio content and playback timing information, where the playback timing information specifies the playback time for the electronic device to play the audio content based on the current time offset. The playback time of the electronic device has a temporal relationship such that the playback of the audio content by the electronic device is adjusted. Note that the temporal relationship has a non-zero value and at least a portion of the electronic devices can be commanded to play the audio content having a phase relative to each other by using different values of the playback time. For example, the different playback times can be based on pre-determined or dynamically determined acoustic characteristics of the environment including the electronic device and the A / V hub. Alternatively or in addition, the different playback times can be based on the desired acoustic characteristics of the environment and / or the estimated position of the listener relative to the electronic device.
[0103] In some embodiments, the A / V hub optionally performs one or more additional operations (operation 816). For example, the A / V hub can modify the sound measured based on the acoustic transfer function of the environment in at least one band of frequencies that includes the spectral content of the acoustic characterization pattern. Note that the acoustic transfer function can be determined and accessed in advance by the A / V hub or can be determined dynamically by the A / V hub. The time delays and dispersions associated with the propagation of sound in the environment can be much larger than the desired adjustments of the clocks of the electronic device and the A / V hub, but the rising edge of the modified direct sound can be determined with sufficient accuracy to determine the current time offset between the clock of the electronic device and the clock of the A / V hub, so this correction of the filtering associated with the environment can be required. For example, the desired adjustment accuracy of the speaker 118 can be made as small as about 1 microsecond, but the sound propagation delay in a typical room (e.g., exceeding a distance of up to 10 - 30 m) can be about five orders of magnitude larger. Nevertheless, the modified measured sound can measure the leading edge of the direct sound associated with the sound pulse output from a given electronic device with an accuracy as small as a few microseconds, facilitating the adjustment of the clocks of the electronic device and the A / V hub. In some embodiments, the A / V hub determines the temperature of the environment and the calculation of the current time offset can be corrected for changes in temperature (which can affect the speed of sound in the environment).
[0104] FIG. 9 is a diagram showing communication between a portable electronic device 110, an A / V hub 112, and a speaker 118-1. Specifically, a processor 910 of the speaker 118-1 can command 912 one or more acoustic transducers 914 of the speaker 118-1 to output sound at an output time, and the sound corresponds to an acoustic characterization pattern. For example, the output time can be defined in advance (such as based on a sequence of patterns or pulses of an acoustic characterization pattern, a predefined output schedule having a scheduled output time, or a predefined interval between output times), and thus can be known to the A / V hub 112 and the speaker 118-1. Alternatively, an interface circuit 916 of the A / V hub 112 can provide a startup frame or packet 918. After an interface circuit 920 receives the startup packet 918, a command 922 can be transferred to a processor 920 of the speaker 118-1, whereby based on the command 922, the output of sound from one or more acoustic transducers 914 is activated.
[0105] Subsequently, one or more acoustic transducers 924 of the A / V hub 112 can measure 926 the sound and provide information 928 indicating the measurement to a processor 930 of the A / V hub 112.
[0106] Next, the processor 930 can calculate a current time offset 932 between a clock from a clock circuit (such as an interface clock circuit) of the speaker 118-1 and a clock from a clock circuit (such as an interface clock circuit) of the A / V hub 112 based on the information 928, one or more times when the speaker 118-1 outputs sound, and an acoustic characterization pattern associated with the speaker 118-1. For example, the processor 930 can determine the current time offset 932 based on at least two times in the acoustic characterization pattern when one or more acoustic transducers 914 of the speaker 118-1 output sound corresponding to the acoustic characterization pattern.
[0107] Furthermore, the processor 930 can provide the playback timing information 934 and the audio content 936 to the interface circuit 916. The playback timing information 934 specifies the playback time at which the speaker 118-1 plays the audio content 936 based on the current time offset 932. Note that the processor 930 can access the audio content 936 in the memory 938. In response, the interface circuit 916 can send one or more frames or packets 940 including the playback timing information 934 and the audio content 936 to the speaker 118-1. (However, in some embodiments, the playback timing information 934 and the audio content 936 are sent using different or separate frames or packets.)
[0108] After the interface circuit 920 receives one or more frames or packets 940, it can provide the playback timing information 934 and the audio content 936 to the processor 924. The processor 924 can execute software to perform the playback operation 942. For example, the processor 924 can store the audio content 936 in a queue in the memory. In these embodiments, the playback operation 942 includes outputting the audio content 936 from the queue, which includes driving one or more of the acoustic transducers 914 based on the audio content 936 such that the speaker 118-1 outputs sound at the time specified by the playback timing information 934.
[0109] In an exemplary embodiment, a communication technology is used to adjust the playback of audio content by the speaker 118. This is illustrated in FIG. 10, which shows a diagram depicting the steps of adjusting the playback of audio content by the speaker 118. Specifically, the speaker 118 can output a sound 1010 corresponding to an acoustic characterization pattern. For example, the acoustic characterization pattern associated with the speaker 118-1 can include two or more pulses 1012, and the time interval 1014 between the pulses 1012 can correspond to the clock provided by the clock circuit of the speaker 118-1. In some embodiments, the pattern or sequence of pulses of the acoustic characterization pattern can also uniquely identify the speaker 118. Although pulses 1012 are used to show the acoustic characterization pattern in FIG. 10, in other embodiments, a variety of time, frequency, and / or modulation techniques can be used, including amplitude modulation, frequency modulation, phase modulation, and the like. The A / V hub 112 can selectively initiate the output of the sound 1010 by transmitting one or more frames or packets 1016 with information 1018 indicating the time at which the speaker 118 outputs the sound 1010 corresponding to the acoustic characterization pattern to the speaker 118.
[0110] The A / V hub 112 can measure sound 1010 output by an electronic device using one or more acoustic transducers, and this sound corresponds to one or more of the acoustic characterization patterns. After measuring the sound 1010, the A / V hub 112 can calculate the current time offset between the clock of the speaker 118 and the clock of the A / V hub 112 based on the measured sound 1010, the one or more times at which the speaker 118 output sound, and the one or more acoustic characterization patterns. In some embodiments, the current time offset can be calculated by the A / V hub 112 based on a model in the A / V hub 112 from the clock drift of the speaker 118. For example, a model of relative or absolute clock drift can include polynomials or cubic splines (and more generally functions) having parameters that specify or estimate the clock drift of a given speaker as a function of time based on historical time offsets.
[0111] Next, the A / V hub 112 can send one or more frames or packets including audio content 1022 and playback timing information (such as the playback timing information 1024-1 of packet 1020-1) to the speaker 118, and the playback timing information specifies the playback time at which the speaker 118 device plays the audio content 1022 based on the current time offset. The playback time of the speaker 118 can have a temporal relationship such that, for example, the playback of the audio content 1022 by the speaker 118 is adjusted such that the associated sound or wavefront arrives at the position 1026 of the environment having the desired phase relationship.
[0112] Communication technology can include operations used to adapt adjustments to improve a listener's acoustic experience. One method is shown in FIG. 11 presenting a flowchart of method 1100 for selectively determining one or more acoustic characteristics of an environment (such as a room). Method 1100 can be executed by an A / V hub such as A / V hub 112 (FIG. 1). During operation, the A / V hub (a control circuit or control logic, e.g., a processor executing a program module in the A / V hub) can optionally detect electronic devices in the environment using wireless communication (operation 1110). Alternatively or in addition, the A / V hub can determine a change state (operation 1112), where the change state includes that an electronic device has not been previously detected in the environment and / or a change in the position of the electronic device (including a change in position that occurs long after the electronic device was first detected in the environment).
[0113] When a change state is determined (operation 1112), the A / V hub can transition to a characterization mode (operation 1114). During the characterization mode, the A / V hub can provide an instruction to the electronic device to play audio content at a specified playback time (operation 1116), determine one or more acoustic characteristics of the environment based on acoustic measurements in the environment (operation 1118), and store characterization information in a memory (operation 1120), where the characterization information includes one or more acoustic characteristics.
[0114]
[0115] Furthermore, the A / V hub can transmit to the electronic device one or more frames (operation 1122) or packets including additional audio content and playback timing information, where the playback timing information can specify a playback time for the electronic device to play the additional audio content based on one or more acoustic characteristics.In some embodiments, the A / V hub optionally performs one or more additional operations (operation 1124). For example, the A / V hub can calculate the location of electronic devices in the environment based on, for example, wireless communication. Additionally, the characterization information can include identifiers of electronic devices that may be received by the A / V hub from the electronic devices using wireless communication.
[0116] Furthermore, the A / V hub can determine one or more acoustic characteristics based at least in part on acoustic measurements performed by other electronic devices. Thus, the A / V hub can communicate with other electronic devices in the environment using wireless communication and receive acoustic measurements from the other electronic devices. In these embodiments, the one or more acoustic characteristics can be determined based on the locations of other electronic devices in the environment. Note that the A / V hub can receive the locations of other electronic devices from the other electronic devices, access pre-determined locations of the other electronic devices stored in memory, and determine the locations of other electronic devices based on, for example, wireless communication.
[0117] In some embodiments, the A / V hub includes one or more acoustic transducers, and the A / V hub performs acoustic measurements using the one or more acoustic transducers. Thus, the one or more acoustic characteristics can be determined by the A / V hub alone or with respect to acoustic measurements performed by other electronic devices.
[0118] However, in some embodiments, instead of determining the one or more acoustic characteristics, the A / V hub receives the one or more acoustic characteristics determined from one of the other electronic devices.
[0119] While acoustic characterization can be fully automated based on the changing state, in some embodiments the user can manually initiate the characterization mode or manually approve the characterization mode when a changing state is detected. For example, the A / V hub can receive user input and transition to the characterization mode based on the user input.
[0120] FIG. 12 is a diagram showing communication between the A / V hub 112 and the speaker 118-1. Specifically, the interface circuit 1210 of the A / V hub 112 can detect the speaker 118-1 through wireless communication of a frame or packet 1212 with the interface circuit 1214 of the speaker 118-1. Note that this communication can be unidirectional or bidirectional.
[0121] The interface circuit 1210 can provide the information 1216 to the processor 1218. This information can indicate the presence of the speaker 118-1 in the environment. Alternatively or in addition to this, the information 1216 can indicate the position of the speaker 118-1.
[0122] The processor 1218 can determine whether a change state 1220 has occurred. For example, the processor 1218 can determine the presence of the speaker 118-1 in an environment that did not previously exist or that the position of the previously detected speaker 118-1 has changed.
[0123] When the change state 1220 is determined, the processor 1218 can transition to the characterization mode 1222. During the characterization mode 1222, the processor 1218 can provide the instruction 1224 to the interface circuit 1210. In response to this, the interface circuit 1210 can transmit the instruction 1224 to the interface circuit 1214 in a frame or packet 1226.
[0124] After receiving packet 1226, interface circuit 1214 can provide processor 1228 with instruction 1224 that commands one or more acoustic transducers 1230 to play audio content 1232 at a specified playback time. Note that processor 1228 can access audio content 1232 in memory 1208 or can include audio content 1232 in packet 1226. Next, one or more acoustic transducers 1234 of A / V hub 112 can perform acoustic measurement 1236 of the sound corresponding to audio content 1232 output by one or more acoustic transducers 1230. Based on acoustic measurement 1236 (and / or additional acoustic measurements received from other speakers by interface circuit 1210), processor 1218 can determine one or more acoustic characteristics 1238 of the environment stored in memory 1240.
[0125] Furthermore, processor 1218 can provide interface circuit 1210 with playback timing information 1242 and audio content 1244, where playback timing information 1242 specifies the playback time at which speaker 118-1 plays audio content 1244 based at least in part on one or more acoustic characteristics 1238. In response, interface circuit 1210 can transmit one or more frames or packets 1246 that include playback timing information 1242 and audio content 1244 to speaker 118-1. (However, in some embodiments, playback timing information 1242 and audio content 1244 are transmitted using a different or separate frame or packet.)
[0126] After interface circuit 1214 receives one or more frames or packets 1246, it can provide playback timing information 1242 and audio content 1244 to processor 1228. Processor 1228 can execute software to perform playback operation 1248. For example, processor 1228 can store audio content 1244 in a queue in memory. In these embodiments, playback operation 1248 includes outputting audio content 1244 from a queue including a stage of driving one or more of acoustic transducers 1230 based on audio content 1244 when speaker 118-1 outputs sound at a time specified by playback timing information 1242.
[0127] In an exemplary embodiment, communication technology is used to selectively determine one or more acoustic characteristics of an environment (such as a room) including A / V hub 112 when a change is detected. FIG. 13 presents a diagram showing selective acoustic characterization of an environment including speaker 118. Specifically, A / V hub 112 can detect speaker 118-1 within the environment. For example, A / V hub 112 can detect speaker 118-1 based on wireless communication of one or more frames or packets 1310 with speaker 118-1. Note that the wireless communication can be unidirectional or bidirectional.
[0128] When a change state is detected (such as when the presence of speaker 118-1 is first detected, i.e., when speaker 118-1 has not been previously detected within the environment, and / or when there is a change in the position 1312 of a previously detected speaker 118-1 within the environment), A / V hub 112 can transition to a characterization mode. For example, A / V hub 112 can transition to the characterization mode when a change in the magnitude of position 1312 at the wavelength magnitude at the upper limit of human hearing at position 1312 of speaker 118-1 is detected, such as a change of 0.085, 0.017, or 0.305 m (examples, not limitations).
[0129] During the characterization mode, the A / V hub 112 provides commands within the frame or packet 1314 to the speaker 118-1 to play the audio content at a specified playback time (i.e., output the sound 1316), determines one or more acoustic characteristics of the environment based on the acoustic measurement of the sound 1316 output by the speaker 118-1, and can store in memory one or more acoustic characteristics that can include the position 1312 of the speaker 118-1.
[0130] For example, the audio content can include a pseudo-random frequency pattern or white noise in a frequency range (between 100 and 10,000 or 20,000 Hz, or two or more sub-frequency bands within the human audible range which is given by way of example and not limitation, such as 500, 1000, and 2000 Hz, etc.), an acoustic pattern having a carrier frequency that varies as a function of time in a frequency range, an acoustic pattern having spectral components in a frequency range, and / or one or more types of music (such as symphony, classical music, chamber music, opera, rock or pop music, etc.). In some embodiments, the audio content, such as a specific temporal pattern, spectral content, and / or one or more frequency tones, uniquely identifies the speaker 118-1. Alternatively or in addition, the A / V hub 112 can receive an identifier of the speaker 118-1, such as an alphanumeric code, via wireless communication with the speaker 118-1.
[0131] However, in some embodiments, the acoustic characterization is performed without the speaker 118-1 playing the audio content. For example, the acoustic characterization can be based on the acoustic energy associated with a human voice or can be performed by measuring the ambient vibration background noise for 1-2 minutes. Thus, in some embodiments, the acoustic characterization includes passive characterization (instead of the active measurements when the audio content is being played).
[0132] Furthermore, acoustic characterization may include the acoustic spectral response of the environment in a frequency range (i.e., information indicating the amplitude response as a function of frequency), the transfer function or impulse response in a frequency range (i.e., information indicating the amplitude and phase responses as a function of frequency), the room reverberation or low-frequency room modes (which can be determined by measuring the sound in the environment having nodes and antinodes as a function of position or location in the environment and in directions differing by 90° from each other), the position 1312 of the speaker 118-1, reflections (early reflections within 50 - 60 ms from the arrival of the direct sound from the speaker 118-1, and delayed reflections or echoes occurring on a long time scale that may affect intelligibility), the acoustic delay of the direct sound, the average reverberation time in a frequency range (or the persistence of the acoustic sound in the environment in a frequency range after the audio content has been interrupted), the volume of the environment (such as the size and / or dimensions of the room that can be determined optically), the background noise in the environment, the ambient sound in the environment, the temperature of the environment, the number of people in the environment (and more generally, the absorption or acoustic loss across a frequency range in the environment), a measure of acoustic liveness, whether the environment is bright or dark and / or information indicating the type of environment (lecture hall, multi-purpose room, concert hall, room size, type of room interior, etc.). For example, the reverberation time can be defined as the time of the sound pressure associated with an impulse at a particular level such as -60 dB that decays. In some embodiments, the reverberation time is a function of frequency. Note that the frequency ranges in the above examples of acoustic characteristics may be the same as or different from each other. Thus, in some embodiments, different frequency ranges can be used for different acoustic characteristics. Additionally, note that the "acoustic transfer function" in some embodiments can include the magnitude of the acoustic transfer function (sometimes referred to as the "acoustic spectral response"), the phase of the acoustic transfer function, or both of these.
[0133] As described above, the acoustic characteristics can include the position 1312 of the speaker 118-1. The position 1312 of the speaker 118-1 (including distance and direction) can be determined by the A / V hub 112 and / or in conjunction with other electronic devices (such as the speaker 118) in the environment using techniques such as triangulation, trilateration, time of flight, radio ranging, angle of arrival, etc. Further, the position 1312 can be determined by the A / V hub 112 using wireless communication (such as communication with a wireless local area network or a cellular phone network), acoustic measurements, a local positioning system, a global positioning system, etc.
[0134] The acoustic characteristics can be determined by the A / V hub 112 based on measurements performed by the A / V hub 112. However, in some embodiments, the acoustic characteristics are determined by or in conjunction with other electronic devices in the environment. Specifically, one or more other electronic devices (such as one or more other speakers 118) can perform acoustic measurements that are wirelessly transmitted to the A / V hub 112 in a frame or packet 1318. (Thus, the acoustic transducers that perform the acoustic measurements can be included in the A / V hub 112 and / or one or more other speakers 118.) As a result, the A / V hub 112 can computationally calculate the acoustic characteristics based at least in part on the acoustic measurements performed by the A / V hub 112 and / or one or more other speakers 118. Note that this computational calculation can be based on the position 1320 of one or more other speakers 118 in the environment. These positions can be received from one or more other speakers 118 in a frame or packet 1318, calculated using one of the above techniques (such as using radio ranging), and / or accessed in memory (i.e., the position 1320 can be determined in advance).
[0135] Furthermore, acoustic characterization can occur when a change state is detected, but alternatively or in addition, the A / V hub 112 can transition to a characterization mode based on user input. For example, the user can activate a virtual command icon within the user interface on the portable electronic device 110. Thus, acoustic characterization can be started automatically, manually, and / or semi-automatically (where the user interface is used to request approval from the user prior to transitioning to the characterization mode).
[0136] After determining the acoustic characteristics, the A / V hub 112 can return to its normal operating mode. In this operating mode, the A / V hub 112 can transmit one or more frames or packets (such as packet 1322) that include additional audio content 1324 (such as music) and playback timing information 1326 to the speaker 118-1, and the playback timing information 1326 can specify the playback time for the speaker 118-1 to play the additional audio content 1324 based on one or more acoustic characteristics. Thus, using acoustic characterization, one or more changes (direct or indirect) in acoustic characteristics associated with a change in the position 1312 of the speaker 118-1 can be corrected or adapted, thereby improving the user experience.
[0137] Another way to improve the acoustic experience is to adapt adjustments based on the dynamically tracked position of one or more listeners. This is shown in FIG. 14, which presents a flowchart showing a method 1400 for calculating an estimated position. Note that the method 1400 can be executed by an A / V hub such as the A / V hub 112 (FIG. 1). During operation, the A / V hub (control circuit or control logic, e.g., a processor that executes program modules in the A / V hub) can calculate the estimated position of a listener (or an electronic device associated with the listener such as the portable electronic device 110 in FIG. 1) with respect to the electronic devices in an environment that includes the A / V hub and the electronic devices (operation 1410).
[0138] Next, the A / V hub can send one or more frames (operation 1412) or packets containing the audio content and playback timing information to the electronic device, and the playback timing information specifies the playback time for the electronic device to play the audio content based on the estimated position. Furthermore, the playback time of the electronic device has a temporal relationship such that the playback of the audio content by the electronic device is adjusted. The temporal relationship can have a non-zero value such that at least a part of the electronic device is instructed to play the audio content having a phase relative to each other by using different values of the playback time. For example, the different playback times can be based on the predetermined or dynamically determined acoustic characteristics of the environment including the electronic device and the A / V hub. Alternatively or in addition, the different playback times can be based on the desired acoustic characteristics of the environment. In addition, the playback time can be based on the current time offset between the clock of the electronic device and the clock of the A / V hub.
[0139] In some embodiments, the A / V hub optionally performs one or more additional operations (operation 1414). For example, the A / V hub can communicate with another electronic device and calculate the estimated position of the listener based on the communication with the other electronic device.
[0140] Furthermore, the A / V hub can include an acoustic transducer that performs sound measurements in the environment, and calculate the estimated position of the listener based on the sound measurements. Alternatively or in addition, the A / V hub can communicate with other electronic devices in the environment, receive additional sound measurements of the environment from the other electronic devices, and calculate the estimated position of the listener based on the additional sound measurements.
[0141] In some embodiments, the A / V hub performs a time-of-flight measurement, and the estimated position of the listener is calculated based on the time-of-flight measurement.
[0142] Furthermore, the A / V hub can calculate an additional estimated position of an additional listener relative to electronic devices in the environment, and the playback time can be based on the estimated position and the additional estimated position. For example, the playback time can be based on the average of the estimated position and the additional estimated position. Alternatively, the playback time can be based on a weighted average of the estimated position and the additional estimated position.
[0143] FIG. 15 is a diagram showing communication among a portable electronic device 110, an A / V hub 112, and speakers 118 such as speaker 118-1. Specifically, the interface circuit 1510 of the A / V hub 112 can receive one or more frames or packets 1512 from the interface circuit 1514 of the portable electronic device 110. Note that the communication between the A / V hub 112 and the portable electronic device 110 can be unidirectional or bidirectional. Based on the one or more frames or packets 1512, the interface circuit 1510 and / or the processor 1516 of the A / V hub 112 can estimate the position 1518 of the listener associated with the portable electronic device 110. For example, the interface circuit 1510 can provide information 1508 based on the packet 1512 used by the processor 1516 to estimate the position 1518.
[0144] Alternatively or in addition, one or more acoustic transducers 1520 of the A / V hub 112 and / or one or more acoustic transducers 1506 of the speakers 118 can perform a measurement 1522 of the sound associated with the listener. When the speaker 118 performs the sound measurement 1522-2, one or more interface circuits 1524 of the speaker 118 (such as speaker 118-1) can transmit one or more frames or packets 1526 to the interface circuit 1510 together with information 1528 indicating the sound measurement 1522-2 based on an instruction 1530 from the processor 1532. The interface circuit 1514 and / or the processor 1516 can estimate the position 1518 based on the measured sound 1522.
[0145] The processor 1516 can command the interface circuit 1510 to transmit one or more frames or packets 1536 to the speaker 118-1 together with the playback timing information 1538 and the audio content 1540. The playback timing information 1538 specifies the playback time for the speaker 118-1 to play the audio content 1540 based at least in part on the position 1518. (However, in some embodiments, the playback timing information 1538 and the audio content 1540 are transmitted using separate or different frames or packets.) Note that the processor 1516 can access the audio content 1540 in the memory 1534.
[0146] After receiving one or more frames or packets 1536, the interface circuit 1524 can provide the playback timing information 1538 and the audio content 1540 to the processor 1532. The processor 1532 can execute software to perform the playback operation 1542. For example, the processor 1532 can store the audio content 1540 in a queue in the memory. In these embodiments, the playback operation 1542 includes the step of outputting the audio content 1540 from the queue and driving one or more of the acoustic transducers 1506 based on the audio content 1540 when the speaker 118-1 outputs sound at the time specified by the playback timing information 1538.
[0147] In an exemplary embodiment, communication technology is used to dynamically track the position of one or more listeners in the environment. FIG. 16 presents a diagram showing the step of calculating the estimated position of one or more listeners relative to the speaker 118. Specifically, the A / V hub 112 can calculate the estimated position of one or more listeners, such as the position 1610 of the listener 1612 relative to the speaker 118 in an environment including the A / V hub 112 and the speaker 118. For example, the position 1610 can be determined roughly (e.g., the nearest room, with an accuracy of 3 to 10 m, etc.) or finely (e.g., with an accuracy of 0.1 to 3 m), which are numerical examples and not limitations.
[0148] Generally, the position 1610 can be determined by the A / V hub 112 and / or with other electronic devices (such as the speaker 118) in the environment using techniques such as triangulation, trilateration, time of flight, radio ranging, angle of arrival, etc. Further, the position 1610 can be determined by the A / V hub 112 using wireless communication (such as communication with a wireless local area network or a cellular phone network), acoustic measurements, a local positioning system, a global positioning system, etc.
[0149] For example, the position 1610 of at least one listener 1612 can be estimated by the A / V hub 112 based on wireless communication (using radio ranging, time of flight measurement, angle of arrival, RSSI, etc.) of one or more frames or packets 1614 with another electronic device such as a portable electronic device 110 that can be assumed to be in proximity to the listener 1612 or worn by the listener 1612 itself. In some embodiments, wireless communication with another electronic device (such as the MAC address of a frame or packet received from the portable electronic device 110) is used as a signature or electronic fingerprint to identify the listener 1612. Note that communication between the portable electronic device 110 and the A / V hub 112 can be one-way or two-way.
[0150] During wireless ranging, the A / V hub 112 can transmit a frame or packet including the transmission time to, for example, the portable electronic device 110. When this frame or packet is received by the portable electronic device 110, the arrival time can be determined. Based on the product of the time of flight (the difference between the arrival time and the transmission time) and the propagation speed, the distance between the A / V hub 112 and the portable electronic device 110 can be calculated. This distance can be transmitted in the next transmission of a frame or packet from the portable electronic device 110 to the A / V hub 112 together with the identifier of the portable electronic device 110. Alternatively, the portable electronic device 110 can transmit a frame or packet including the transmission time and the identifier of the portable electronic device 110, and the A / V hub 112 can determine the distance between the portable electronic device 110 and the A / V hub 112 based on the product of the time of flight (the difference between the arrival time and the transmission time) and the propagation speed.
[0151] In a variant of this method, the A / V hub 112 can transmit a frame or packet 1614 that is reflected by the portable electronic device 110, and the reflected frame or packet 1614 can be used to dynamically determine the distance between the portable electronic device 110 and the A / V hub 112.
[0152] The above example shows ranging using the adjusted clocks of the portable electronic device 110 and the A / V hub 112, but in other embodiments the clocks are not adjusted. For example, even when the transmission time is unknown or unavailable, the position of the portable electronic device 100 can be estimated based on the propagation speed and arrival time data of wireless signals at several receivers at different known positions in the environment (sometimes referred to as "differential arrival time"). For example, the receiver can be at least part of another speaker 118 at a position 1616 that can be predefined or pre-determined. More generally, distance determination based on the difference in power of the RSSI relative to the original transmission signal strength (which can include corrections for absorption, refraction, shadowing and / or reflection), determination of the angle of arrival of the receiver (including listening outside the field of view) based on differential arrival time to a directional antenna or an array of antennas having known positions in the environment, distance determination based on backscattered wireless signals, and / or determination of the angle of arrival of two receivers having known positions in the environment (i.e., trilateration or multilateration), among other various radio detection techniques can be used. Note that the wireless signal can include transmissions over a GHz or multi-GHz bandwidth to generate pulses of short duration (e.g., about 1 ns, etc.) that can determine the distance within 0.305 m (e.g., 1 ft) and is illustrative and not limiting. In some embodiments, wireless ranging is facilitated using position information of one or more positions (such as position 1616) of the electronic device in the environment determined or indicated by a local positioning system, a global positioning system, and / or a wireless network.
[0153] Alternatively or in addition thereto, the position 1610 can be estimated by the A / V hub 112 based on sound measurements in an environment such as acoustic tracking of the listener 1612, based on sounds 1618 that occur, for example, when walking around, speaking, and / or breathing. The sound measurements can be performed by the A / V hub 112 (using, for example, 2 or more acoustic transducers, such as microphones that can be arranged as a phased array). However, in some embodiments, sound measurements can be performed separately or additionally by one or more electronic devices in the environment, such as the speaker 118, and these sound measurements can be wirelessly transmitted to the A / V hub 112 in frames or packets 1618, and the position 1610 is estimated using the sound measurements. In some embodiments, the listener 1612 is identified using speech recognition technology.
[0154] In some embodiments, the position 1610 is estimated by the A / V hub 112 based on sound measurements of the environment and predetermined acoustic characteristics of the environment, such as spectral response or acoustic transfer function. For example, the position 1610 can be estimated using the variation of the excitation of a predetermined room pattern when the listener 1612 moves through the environment.
[0155] Furthermore, one or more other technologies can be used to track or estimate the position 1610 of the listener 1612. For example, the position 1610 can be estimated based on an optical image of the listener 1612 at a wavelength of one band (visible light or infrared light), time-of-flight measurement (such as laser ranging), and / or an optical beam (such as an infrared beam) grid that positions the listener 1612 in a grid based on the pattern of beam intersections (and thus roughly determines the position 1610). In some embodiments, the identity of the listener 1612 is determined in an optical image using face recognition and / or gate recognition technology.
[0156] For example, in some embodiments, the position of a listener in an environment is tracked based on wireless communication with a cellular phone carried by the listener. Based on the pattern of positions in the environment, the position of furniture in the environment and / or the shape of the environment (such as the size or dimensions of a room) can be determined. This information can be used to determine the acoustic characteristics of the environment. Further, the estimated position of the listener in the environment can be regulated using the listener's historical positions. Specifically, historical information regarding the position of the listener in the environment at different times of the day can be used to assist in estimating the current position of the listener at a particular time. Thus, more generally, the position of the listener can be estimated using a combination of optical measurements, acoustic measurements, acoustic characteristics, wireless communication, and / or machine learning.
[0157] After determining the position 1610, the A / V hub 112 can transmit at least one or more frames or packets including additional audio content 1622 (such as music) and playback timing information (such as the playback timing information 1624-1 in the packet 1620-1 to the speaker 118-1) to the speaker 118, and the playback timing information 1624-1 can specify the playback time for the speaker 118-1 to play the additional audio content 1622 based on the position 1610. Thus, the change in the position 1610 can be corrected or adapted using communication technology, thereby improving the user experience.
[0158] As described above, different playback times can be based on the desired acoustic characteristics in the environment. For example, the desired acoustic characteristics can include the type of playback such as monophonic, stereophonic, and / or multichannel sound. Monophonic sound can include one or more audio signals that do not include amplitude (or level) and arrival time / phase information that replicates or simulates a directional cue.
[0159] Furthermore, stereophonic sound can include two independent audio-signal channels, and the audio signals can have specific amplitudes and phase relationships to each other such that a clear image of the original sound source exists during playback operation. Generally, the audio signals of both channels can provide coverage over much or all of the environment. By adjusting the relative amplitudes and / or phases of the audio channels, the sweet spot can be moved to trace at least the determined position of the listener. However, it is necessary to make the amplitude difference and arrival time difference (directional cue) small enough so that both the stereo image and localization are maintained. Otherwise, the image may break down and only one or one of the audio channels may be heard.
[0160] Note that the audio channels of stereophonic sound need to have an accurate absolute phase response. This means that an audio signal having a positive pressure waveform at the input to the system needs to have the same positive pressure waveform at the output from one of the speakers 118. Thus, a drum that produces a positive pressure waveform at the microphone when struck needs to produce a positive pressure waveform in the environment. Instead, if the absolute polarity is flipped to the wrong side, the audio image may not be stable. Specifically, the listener may not be able to find or perceive a stable audio image. Instead, the audio image fluctuates and can be localized at the speaker 118.
[0161] Furthermore, multi-channel sound can include left, center, and right audio channels. For example, these channels can position or mix monophonic speech enhancement and music or sound effect cues by stereo or stereo-like imaging. Thus, the three audio channels can provide coverage over most or all of the environment while maintaining the amplitude and directional cues as in the case of monophonic or stereophonic sound.
[0162] Alternatively or in addition, the desired acoustic characteristics can include an acoustic radiation pattern. The desired acoustic radiation pattern can be a function of the reverberation time in the environment. For example, the reverberation time can vary depending on the number of people in the environment, the type and amount of furniture in the environment, whether the curtains are open or closed, whether the windows are open or closed, etc. When the reverberation time is long or increasing, the desired acoustic radiation pattern can be directed so that sound is directed or emitted towards the listener (thereby reducing reverberation). In some embodiments, the desired acoustic characteristics include word intelligibility.
[0163] The above description illustrates techniques that can be used to dynamically track the position 1610 of the listener 1612 (or the portable electronic device 110), but these techniques can be used to determine the position of an electronic device (such as the speaker 118-1) in the environment.
[0164] Another way to improve the acoustic experience is to dynamically aggregate electronic devices into groups and / or adapt adjustments based on the groups. This is shown in FIG. 17, which presents a flowchart showing a method 1700 for aggregating electronic devices. Note that the method 1700 can be performed by an A / V hub such as the A / V hub 112 (FIG. 1). During operation, the A / V hub (control circuit or control logic, e.g., a processor that executes program modules in the A / V hub) can measure the sound output by an electronic device (such as the speaker 118) in the environment using one or more acoustic transducers (operation 1710), and the sound corresponds to audio content. For example, the measured sound can include sound pressure.
[0165] Next, the A / V hub can aggregate the electronic devices into two or more subsets based on the measured sound (operation 1712). Note that the different subsets can be located in different rooms within the environment. Furthermore, at least one of the subsets can play audio content different from the rest of the subsets. Additionally, the aggregation of the electronic devices into two or more subsets can be based on different audio content, the acoustic delay of the measured sound, and / or the desired acoustic characteristics in the environment. In some embodiments, the subsets and / or the electronic devices within the geographical location or area associated with the subsets are not predefined. Instead, the A / V hub can aggregate the subsets dynamically.
[0166] Furthermore, the A / V hub can determine the playback timing information for the subsets (operation 1714), where the playback timing information specifies the playback time at which the electronic devices of a given subset play the audio content.
[0167] Next, the A / V hub can use wireless communication to send one or more frames (operation 1716) or packets containing the audio content and the playback timing information to the electronic devices, where the playback time of at least the electronic devices of a given subset has a temporal relationship such that the playback of the audio content by the electronic devices of the given subset is coordinated.
[0168] In some embodiments, the A / V hub optionally performs one or more additional operations (operation 1718). For example, the A / V hub can calculate the estimated position of at least one listener relative to the electronic devices, and the aggregation of the electronic devices into two or more subsets can be based on the estimated position of at least one listener. This can help the listener have an improved acoustic experience while reducing acoustic crosstalk from other subsets.
[0169] Furthermore, the A / V hub can modify the measured sound based on a predefined (or dynamically determined) acoustic transfer function of the environment in at least one frequency band (such as 100 - 20,000 Hz, which is an example rather than a limitation). This can enable the A / V hub to determine the original output sound associated with the environment such that the time for aggregating the subset can be appropriately determined by the A / V hub, by performing spectral filtering or without distortion.
[0170] Furthermore still, the A / V hub can determine the playback volume of the subset used when the subset plays audio content, and one or two or more frames or packets can include information indicating the playback volume. For example, the playback volume of at least one of the subsets may be different from the playback volume of the remaining subsets. Alternatively or in addition to this, the playback volume can reduce the acoustic crosstalk between two or three or more subsets such that the listener is more likely to hear the sound output by the subset closest to or in proximity to the listener.
[0171] Figure 18 is a diagram showing the communication between the portable electronic device 110, the A / V hub 112, and the speaker 118. Specifically, the processor 1810 can command one or two or more acoustic transducers 1814 within the A / V hub 112 to perform a measurement 1816 of the sound associated with the speaker 118 at 1812. Next, based on the measurement 1816, the processor 1810 can aggregate the speaker 118 into two or three or more subsets 1818.
[0172] Furthermore, the processor 1810 can determine the playback timing information 1820 of the subset 1818, and the playback timing information 1820 specifies the playback time when the speaker 118 of a given subset plays the audio content 1822. Note that the processor 1810 can access the audio content 1822 within the memory 1824.
[0173] Next, the processor 1810 can instruct the interface circuit 1826 to transmit a frame or packet 1828 that includes the playback timing information 1820 and the audio content 1822 to the speaker 118. (However, in some embodiments, the playback timing information 1820 and the audio content 1822 are transmitted using separate or different frames or packets.)
[0174] After receiving one or more frames or packets 1826, the interface circuit of the speaker 118-3 can provide the playback timing information 1820 and the audio content 1822 to the processor. This processor can execute software that performs the playback operation 1830. For example, the processor can store the audio content 1822 in a queue in the memory. In these embodiments, the playback operation 1830 includes the step of outputting the audio content 1822 from the queue and driving one or more of the acoustic transducers based on the audio content 1822 when the speaker 118-3 outputs sound at the time specified by the playback timing information 1820. Note that the playback times of at least one given subset of the speakers 118 have a temporal relationship such that the playback of the audio content 1822 by the given subset of the speakers 118 is adjusted.
[0175] In an exemplary embodiment, communication technology is used to aggregate the speakers 118 into subsets. FIG. 19 presents a diagram showing the aggregation of speakers 118 that may be in the same or different rooms within the environment. The A / V hub 112 can measure the sound 1910 output by the speakers 118. Based on these measurements, the A / V hub 112 can aggregate the speakers 118 into subsets 1912. For example, the subsets 1912 can be aggregated based on sound intensity and / or acoustic delay such that adjacent speakers are grouped together. Specifically, speakers having the highest acoustic intensity or similar acoustic delay can be aggregated together. To facilitate the aggregation, the speakers 118 can wirelessly transmit and / or acoustically output identification information or acoustic characterization patterns outside the range of human hearing. For example, the acoustic characterization pattern can include pulses. However, a variety of time, frequency, and / or modulation techniques such as amplitude modulation, frequency modulation, phase modulation, etc. can be used. Alternatively or in addition, the A / V hub 112 can command each of the speakers 118 to dither the playback time or the phase of the output sound so that the A / V hub 112 can associate a particular speaker with the sound measured by the A / V hub, one at a time. Further, the sound 1910 measured using the acoustic transfer function of the environment can be corrected so that the effects (or distortions) of reflection and filtering are removed before aggregating the speakers 118. In some embodiments, the speakers 118 are aggregated based at least in part on the position 1914 of the speakers 118 that can be determined using one or more of the above-described techniques (such as wireless ranging). In this way, the subsets 1912 can be dynamically modified when one or more listeners reposition the speakers 118 within the environment.
[0176] Next, the A / V hub 112 can transmit one or more frames or packets (such as packet 1916) including additional audio content 1918 (such as music) and playback timing information 1920 to at least one of the speakers 118 in the subset 1912 (such as subset 1912-1), and the playback timing information 1920 can specify the playback time for the speaker 118 in the subset 1912-1 to play the additional audio content 1918. Therefore, for example, communication technology can be used to dynamically select the subset 1912 based on the position of the listener and / or the desired acoustic characteristics in the environment including the A / V hub 112 and the speaker 118.
[0177] Another way to improve the acoustic experience is to dynamically equalize the audio based on acoustic monitoring in the environment. FIG. 20 presents a flowchart showing a method 2000 for determining equalized audio content that can be performed by an A / V hub such as the A / V hub 112 (FIG. 1). During operation, an A / V hub (a control circuit or control logic, such as a processor that executes a program module in the A / V hub) can measure the sound output by an electronic device (such as the speaker 118) in the environment using one or more acoustic transducers (operation 2010), and the sound corresponds to the audio content. For example, the measured sound can include sound pressure.
[0178] Next, the A / V hub can compare the measured sound at the first position in the environment with the desired acoustic characteristics based on a pre-determined or dynamically determined acoustic transfer function of the environment at at least one of the first position, the second position of the A / V hub, and a frequency band (such as 100 - 20,000 kHz which is an example and not a limitation). Note that the comparison can be performed in the time domain and / or the frequency domain. To perform the comparison, the A / V hub can calculate the acoustic characteristics (such as an acoustic transfer function or a pattern response) at the first position and / or the second position, and use the calculated acoustic characteristics to correct the measured sound for filtering or distortion in the environment. Using the acoustic transfer function as an example, this calculation involves the use of the Green's function technique and can computationally calculate the acoustic response of the environment as a function of positions with one or more points or as a distributed acoustic source at a pre-defined or known position in the environment. Note that the acoustic transfer function and correction at the first position may depend on the integrated acoustic behavior of the environment (and thus the second position and / or multiple positions of acoustic sources such as speaker 118 in the environment). Thus, the acoustic transfer function can include information indicating the position in the environment where the acoustic transfer function is determined (e.g., the second position) and / or the position of the acoustic source in the environment (such as the position of at least one of the electronic devices).
[0179] Furthermore, the A / V hub can determine equalized audio content based on comparison and audio content (operation 2014). Note that the desired acoustic characteristics can be based on the type of audio reproduction such as monophonic, stereophonic and / or multichannel. Instead of or in addition to this, the desired acoustic characteristics can include an acoustic radiation pattern. The desired acoustic radiation pattern can be a function of the reverberation time in the environment. For example, the reverberation time may vary depending on the number of people in the environment, the type and amount of furniture in the environment, whether the curtains are open or closed, whether the windows are open or closed, etc. When the reverberation time is long or increasing, the desired acoustic radiation pattern can be directed to direct or emit the sound associated with the equalized audio content towards the listener (thereby reducing reverberation). As a result, in some embodiments, equalization is a complex function that modifies the amplitude and / or phase of the audio content. Furthermore, the desired acoustic characteristics can include a step of reducing room resonance or room modes by reducing the associated low-frequency energy in the acoustic content. Note that in some embodiments, the desired acoustic characteristics include word intelligibility. Thus, the target (desired acoustic characteristics) can be used to adapt the equalization of the audio content.
[0180] The A / V hub can use wireless communication to transmit one or more frames (operation 2016) or packets including the equalized audio content to an electronic device to facilitate the output of additional sound corresponding to the equalized audio content by the electronic device.
[0181] In some embodiments, the A / V hub optionally performs one or more additional operations (operation 2018). For example, the first position can include an estimated position of the listener relative to the electronic device, and the A / V hub can calculate the estimated position of the listener. Specifically, one or more of the above-described techniques for dynamically determining the position of the listener can be used to determine the estimated position of the listener. Thus, the A / V hub can calculate the estimated position of the listener based on sound measurements. Alternatively or in addition, the A / V hub can communicate with another electronic device and calculate the estimated position of the listener based on communication with the other electronic device. In some embodiments, communication with the other electronic device includes wireless ranging, and the estimated position can be calculated based on wireless ranging and the angle of arrival of wireless signals from the other electronic device. Furthermore, the A / V hub can perform time-of-flight measurements and calculate the estimated position of the listener based on the time-of-flight measurements. In some embodiments, dynamic equalization enables the "sweet spot" in the environment to be adapted based on the position of the listener. Note that the A / V hub can determine the number of listeners and / or the position of the listeners in the environment and adapt the dynamic equalization to the sound such that the desired acoustic characteristics are present when the listener (or most of the listeners) listens to the equalized audio content.
[0182] Furthermore, the A / V hub can communicate with other electronic devices in the environment and receive additional sound measurements of the environment from the other electronic devices (separately from or together with the sound measurements). The A / V hub can perform one or more additional comparisons of the additional sound measurements at the first position in the environment with the desired acoustic characteristics based on the predetermined or dynamically determined acoustic transfer function of the environment at one or more third positions (such as the position of the speaker 118) of the other electronic device and at least one frequency band, and the equalized audio content is determined based on the one or more additional comparisons. In some embodiments, the A / V hub determines one or more third positions based on communication with other electronic devices. For example, the communication with other electronic devices can include wireless ranging, and the one or more third positions can be calculated based on wireless ranging and the angle of arrival of wireless signals from the other electronic devices. Alternatively or in addition, the A / V hub can receive information indicating the third position from the other electronic devices. Thus, the position of the other electronic device can be determined using one or more of the above-described techniques for determining the position of the electronic device in the environment.
[0183] Furthermore, the A / V hub can determine playback timing information that specifies the playback time at which the electronic device plays the equalized audio content, and one or more frames or packets can include the playback timing information. In these embodiments, the playback time of the electronic device has a temporal relationship such that the playback of the audio content by the electronic device is adjusted.
[0184] FIG. 21 is a diagram showing communication between a portable electronic device 110, an A / V hub 112, and a speaker 118. Specifically, a processor 2110 can command one or more acoustic transducers 2114 of the A / V hub 112 to measure 2116 a sound associated with the speaker 118 and corresponding to audio content 2118 at 2112. The processor 2110 can compare 2120 the measured sound 2116 with a desired acoustic characteristic 2122 of the first position of the environment based on a predetermined or dynamically determined acoustic transfer function 2124 (accessible in a memory 2128) of at least one band of the first position of the environment, the second position of the A / V hub 112, and a frequency.
[0185] Furthermore, the processor 2110 can determine an equalized audio content 2126 based on the comparison 2120 and the audio content 2118 accessible in the memory 2128. Note that the processor 2110 can know in advance the audio content 2118 output by the speaker 118.
[0186] Next, the processor 2110 can determine playback timing information 2130, which specifies a playback time for the speaker 118 to play the equalized audio content 2126.
[0187] Furthermore, the processor 2110 can command an interface circuit 2132 to transmit one or more frames or packets 2134 to the speaker 118 together with the playback timing information 2130 and the equalized audio content 2160. (However, in some embodiments, the playback timing information 2130 and the audio content 2126 are transmitted using separate or different frames or packets.)
[0188] After receiving one or more frames or packets 2134, the interface circuit of one of the speakers 118 (such as speaker 118-1) can provide the playback timing information 2130 and the equalized audio content 2126 to the processor. This processor can execute software for performing a playback operation. For example, the processor can store the equalized audio content 2126 in a queue in the memory. In these embodiments, the playback operation includes the step of outputting the equalized audio content 2126 from the queue and driving one or more of the acoustic transducers based on the equalized audio content 2126 when the speaker 118-1 outputs sound at the time specified by the playback timing information 2130. Note that the playback times of the speakers 118 have a temporal relationship so that the playback of the equalized audio content 2126 by the speakers 118 is adjusted.
[0189] In an exemplary embodiment, a communication technology is used to dynamically equalize audio content. FIG. 22 presents a diagram showing the steps of determining equalized audio content using the speaker 118. Specifically, the A / V hub 112 can measure the sound 2210 corresponding to the audio content output by the speaker 118. Alternatively or in addition, at least a part of the portable electronic device 110 and / or the speaker 118 can measure the sound 2210 and provide information for instructing the measurement to the A / V hub 112 in a frame or packet 2212.
[0190] The A / V hub 112 can compare the measured sound 2210 at position 2214 in the environment (such as the dynamic position of one or more listeners, which can be the position of the portable electronic device 110) with the desired acoustic characteristics based on the predetermined or dynamically determined acoustic transfer function (or more generally, acoustic characteristics) of the environment in at least one of the bands of position 2214, the position 2216 of the A / V hub 112, the position 2218 of the speaker 118, and / or frequency. For example, the A / V hub 112 can calculate the acoustic transfer functions of positions 2214, 2216, and / or 2218. As described above, this calculation involves the use of the Green's function technique and can computationally calculate the acoustic responses at positions 2214, 2216, and / or 2218. Alternatively or in addition, this calculation can include interpolation (minimum bandwidth interpolation) of the predetermined acoustic transfer functions at different positions in the environment, positions 2214, 2216, and / or 2218. The A / V hub 112 can correct the measured sound 2210 based on the computationally calculated and / or interpolated acoustic transfer function (and more generally, acoustic characteristics).
[0191] In this way, using communication technology, it is possible to compensate for the coarse sampling when the acoustic transfer function was originally determined.
[0192] Furthermore, the A / V hub 112 can determine equalized audio content based on the comparison and the audio content. For example, the A / V hub 112 can modify the spectral content and / or phase of the audio content as a function of frequency in a frequency range (such as 100 - 10,000 or 20,000 Hz) to achieve the desired acoustic characteristics.
[0193] The A / V hub 112 can transmit one or more frames or packets containing the equalized audio content (such as music) and the playback timing information to the speaker 118 (such as a packet 2220 having the equalized audio content 2222 and the playback timing information 2224), and the playback timing information can specify the playback time at which the speaker 118 plays the equalized audio content.
[0194] In this way, the communication technology can adapt the sound output by the speaker 118 to the change of the position 2214 of one or more listeners (average or intermediate position, position corresponding to the majority of listeners, average position of the largest subset of listeners where the desired acoustic characteristics are achieved when given voice content and acoustic transfer function or acoustic characteristics of the environment, etc.). This can enable the sweet spot of stereophonic sound to track the movement of one or more listeners and / or the change in the number of listeners in the environment (which can be determined by the A / V hub 112 using one or more of the above-described techniques). Instead of or in addition to this, the communication technology can enable the sound output by the speaker 118 to adapt to the change of voice content and / or desired acoustic characteristics. For example, depending on the type of voice content (such as the type of music), one or more listeners can desire or require a large or wide sound (by branching sound waves corresponding to an acoustically extended source, clearly) or a clearly narrow or point source. Therefore, the communication technology can equalize the voice content according to the desired psychoacoustic experience of one or more listeners. The desired acoustic characteristics or the desired psychoacoustic experience can be explicitly specified by one or more of the listeners (such as by using the user interface on the portable electronic device 110), or can be indirectly determined or implied without user action (based on the type of music stored in the listening history or the preferred acoustic preferences of one or more listeners, etc.).
[0195] In some embodiments of methods 200 (FIG. 2), 500 (FIG. 5), 800 (FIG. 8), 1100 (FIG. 11), 1400 (FIG. 14), 1700 (FIG. 17) and / or 2000 (FIG. 20), there are additional or some operations. Further, the order of the operations can be changed, and / or two or more operations can be combined into a single operation. Furthermore, one or more operations can be modified.
[0196] Here, embodiments of an electronic device will be described. FIG. 23 presents a block diagram showing an electronic device 2300, such as one of the portable electronic device 110, A / V hub 112, A / V display device 114, receiving device 116, or speaker 118 of FIG. 1. This electronic device includes a processing subsystem 2310, a memory subsystem 2312, a networking subsystem 2314, an optional feedback subsystem 2334, and an optional monitoring subsystem 2336. The processing subsystem 2310 includes one or more devices configured to perform computer operations. For example, the processing subsystem 2310 can include one or more microprocessors, application specific integrated circuits (ASICs), microcontrollers, programmable logic devices, and / or one or more digital signal processors (DSPs). One or more of these components in the processing subsystem may also be referred to as a "control circuit". In some embodiments, the processing subsystem 2310 includes a "control mechanism" or "means for processing" that performs at least a portion of the operations in the communication technology.
[0197] The memory subsystem 2312 includes one or more devices that store data and / or instructions for the processing subsystem 2310 and the networking subsystem 2314. For example, the memory subsystem 2312 can include dynamic random access memory (DRAM), static random access memory (SRAM), and / or other types of memory. In some embodiments, the instructions for the processing subsystem 2310 within the memory subsystem 2312 include one or more program modules or sets of instructions (such as program module 2322 or operating system 2324) that can be executed by the processing subsystem 2310. Note that one or more computer programs or program modules can constitute a computer program mechanism. Further, the instructions for the various modules within the memory subsystem 2312 can be implemented in a high-level procedural language, an object-oriented programming language, and / or an assembly or machine language. Still further, the programming language can be compiled or translated and can be configured or configured to be executed by, for example, the processing subsystem 2310 (used synonymously in this description).
[0198] In addition, the memory subsystem 2312 can include a mechanism for controlling access to the memory. In some embodiments, the memory subsystem 2312 includes a memory hierarchy that includes one or more caches coupled to the memory of the electronic device 2300. In some of these embodiments, one or more of the caches are located in the processing subsystem 2310.
[0199] In some embodiments, the memory subsystem 2312 is coupled to one or more high-capacity mass storage devices (not shown). For example, the memory subsystem 2312 can be coupled to a magnetic or optical drive, a solid state drive, or another type of mass storage device. In these embodiments, the memory subsystem 2312 is used by the electronic device 2300 as high-speed access storage for frequently used data, while the mass storage device is used to store less frequently used data.
[0200] The networking subsystem 2314 includes one or more devices configured to be coupled to and communicate on (e.g., perform network operations on) wired and / or wireless networks, including control logic 2316, interface circuitry 2318, and associated antennas 2320. (FIG. 23 includes antenna 2320, and in some embodiments, the electronic device 2300 includes one or more nodes such as node 2308, e.g., pads that can be coupled to antenna 2320. Thus, the electronic device 2300 can or need not include antenna 2320.) For example, the networking subsystem 2314 can include a Bluetooth networking system, a cellular networking system (e.g., 3G / 4G such as UMTS, LTE), a Universal Serial Bus (USB) networking system, a networking system based on the standards described in IEEE 802.11 (e.g., a Wi-Fi networking system), an Ethernet networking system, and / or another networking system. Note that a given combination of one of interface circuitry 2318 and at least one of antennas 2320 can form a radio. In some embodiments, the networking subsystem 2314 includes a wired interface such as an HDMI (registered trademark) interface 2330.
[0201] The networking subsystem 2314 includes processors, controllers, radios / antennas, sockets / plugs, and / or other devices that are coupled to each supported networking system, communicate over each supported networking system, and process data and events of each supported networking system. Note that the mechanisms that couple to the network of each network system, communicate over the network, and process data and events on the network are sometimes collectively referred to as the "network interface" of the network system. Further, in some embodiments, there is no "network" between electronic devices. Thus, the electronic device 2300 can use the mechanisms in the networking subsystem 2314 to perform simple wireless communication between electronic devices, e.g., to transmit advertisements or beacon frames or packets and / or to scan for advertisement frames or packets transmitted by other electronic devices as described above.
[0202] Within the electronic device 2300, the processing subsystem 2310, the memory subsystem 2312, the networking subsystem 2314, the optional feedback subsystem 2334, and the optional monitoring subsystem 2336 are coupled to each other using a bus 2328. The bus 2328 can include electrical, optical, and / or electro-optical connections that the subsystems can use to transmit commands and data to each other. Although only one bus 2328 is shown for clarity, different embodiments can include different numbers or configurations of electrical, optical, and / or electro-optical connections between the subsystems.
[0203] In some embodiments, the electronic device 2300 includes a display subsystem 2326 for displaying information (such as requests to clarify the identified environment) on a display that can include a display driver, an I / O controller, and a display. Note that a wide variety of display types can be used for the display subsystem 2326, including two-dimensional displays, three-dimensional displays (such as holographic displays or volumetric displays), head-mounted displays, retinal image projectors, head-up displays, cathode ray tubes, liquid crystal displays, projection displays, electroluminescent displays, displays based on electronic paper, thin film transistor displays, high-performance addressable displays, organic light-emitting diode displays, surface field displays, laser displays, carbon nanotube displays, quantum dot displays, interference modulation displays, multi-touch touchscreens (sometimes referred to as touch sensor type displays), and / or displays based on other types of display technology or physical phenomena.
[0204] Furthermore, the optional feedback subsystem 2334 can include one or more sensor feedback mechanisms or devices such as a vibration mechanism or a vibration actuator (e.g., an eccentric rotating mass actuator or a linear resonant actuator), a light, one or more speakers, etc., and can be used to provide feedback (such as sensor feedback) to the user of the electronic device 2300. Alternatively or in addition to this, the optional feedback subsystem 2334 can be used to provide sensor input to the user. For example, one or more speakers can output sounds such as voices. One or more speakers can include an array of transducers that can be modified to adjust the characteristics of the sound output by one or more speakers such as a phase array of acoustic transducers. This capability can enable one or more speakers to modify the sound in the environment to achieve the user's desired acoustic experience by changing the equalization or spectral content, phase, and / or direction of the propagating sound waves.
[0205] In some embodiments, the optional monitoring subsystem 2336 includes one or more acoustic transducers 2338 (such as one or more microphones, a phase array, etc.) that monitor the sound in the environment including the electronic device 2300. Acoustic monitoring can enable the electronic device 2300 to acoustically characterize the environment, acoustically characterize the sound (such as the sound corresponding to voice content) output by a speaker in the environment, determine the position of the listener, determine the position of the speaker in the environment, and / or measure the sound from one or more speakers corresponding to one or more acoustic characterization patterns (which can be used to adjust the playback of voice content). In addition, the optional monitoring subsystem 2336 can include a position transducer 2340 that can be used to determine the position of a listener or an electronic device (such as a speaker) in the environment.
[0206] The electronic device 2300 can be (or can include) any electronic device having at least one network interface. For example, the electronic device 2300 can be (or can include) a desktop computer, a laptop computer, a subnotebook / netbook, a server, a tablet computer, a smartphone, a cellular phone, a smartwatch, a consumer electronic device (such as a television, a set-top box, an audio device, a speaker, a video device, etc.), a remote control, a portable computer device, an access point, a router, a switch, a communication device, a test device, and / or another electronic device.
[0207] Certain components are used to describe the electronic device 2300, but in alternative embodiments, different components and / or subsystems can be present in the electronic device 2300. For example, the electronic device 2300 can include one or more additional processing subsystems, memory subsystems, networking subsystems, and / or display subsystems. One of the antennas 2320 is shown coupled to a given one of the interface circuits 2318, but multiple antennas can also be coupled to a given one of the interface circuits 2318. For example, a 3x3 radio case can include three antennas. Additionally, one or more of the subsystems may not be present in the electronic device 2300. Furthermore, in some embodiments, the electronic device 2300 can include one or more additional subsystems not shown in FIG. 23. Also, although separate subsystems are shown in FIG. 23, in some embodiments, a part or all of a given subsystem or component can be integrated into one or more of the other subsystems or components of the electronic device 2300. For example, in some embodiments, the program module 2322 is included in the operating system 2324.
[0208] Furthermore, the circuitry and components of the electronic device 2300 can be implemented using any combination of analog and / or digital circuitry including bipolar, PMOS, and / or NMOS gates or transistors. Additionally, the signals in these embodiments can include digital signals having substantially discrete values and / or analog signals having continuous values. In addition, the components and circuitry can be single-ended or differential, and the power supply can be unipolar or bipolar.
[0209] The integrated circuit can implement some or all of the functionality of a networking subsystem 2314 such as one or more radios. Further, the integrated circuit can include hardware and / or software mechanisms used to transmit wireless signals from the electronic device 2300 and receive signals from other electronic devices at the electronic device 2300. Although radios are more generally known in the art and not described in detail, in addition to the mechanisms described herein, more generally, the networking subsystem 2314 and / or the integrated circuit can include any number of radios.
[0210] In some embodiments, the networking subsystem 2314 and / or the integrated circuit include configuration mechanisms (one or more hardware and / or software mechanisms) that configure the radio to transmit and / or receive on a given channel (e.g., a given carrier frequency). For example, in some embodiments, these configuration mechanisms can be used to switch the radio from monitoring and / or transmitting on a given channel to monitoring and / or transmitting on a different channel. (Note that as used herein, "monitoring" includes receiving signals from other electronic devices and, if possible, performing one or more processing operations on the received signals, e.g., determining whether the received signal includes an advertisement frame or packet, calculating performance metrics, performing spectral analysis, etc.). Additionally, the networking subsystem 2314 can include at least one port (such as HDMI (registered trademark) port 2332) for receiving information in a data stream and / or providing it to at least one of the A / V display devices 114 (FIG. 1), at least one of the speakers 118 (FIG. 1), and / or at least one of the content sources 120 (FIG. 1).
[0211] A Wi-Fi compliant communication protocol is used as an illustrative example, but the described embodiments can be used with a wide variety of network interfaces. Additionally, while some of the operations of the above-described embodiments are implemented in hardware or software, more generally the operations of the above-described embodiments can be implemented with a wide variety of configurations and architectures. Thus, some or all of the operations in the above-described embodiments can be performed in hardware, in software, or in both. For example, at least some of the operations in communication technology can be performed using program module 2322, operating system 2324 (such as a driver for interface circuit 2318), and / or the firmware of interface circuit 2318. Alternatively or in addition, at least some of the operations of communication technology can be performed at the physical layer, such as the hardware of interface circuit 2318.
[0212] Furthermore, the above-described embodiments include a touch sensor display of a portable electronic device that a user touches (e.g., with a finger or digit or stylus), but in other embodiments the user interface is displayed on the display of the portable electronic device and the user interacts with the user interface without contacting or touching the surface of the display. For example, the user's interaction with the user interface can be determined using time-of-flight measurement, motion sensing (such as Doppler measurement), or another non-contact measurement that can determine the direction and / or speed of the movement of the user's finger or digit (or stylus) relative to the position of one or more virtual command icons. In these embodiments, note that the user can operate a given virtual command icon by performing a gesture (such as "tapping" a finger in the air without contacting the surface of the display). In some embodiments, the user navigates the user interface and / or operates / halts the function of one of the components of system 100 (FIG. 1) using spoken commands or instructions (i.e., via voice recognition) and / or based on where the user is looking at the display of the portable electronic device 110 of FIG. 1 or the display of the A / V display device 114 of FIG. 1 (e.g., by tracking the user's line of sight or where the user is looking).
[0213] Also, although the A / V hub 112 (FIG. 1) is shown as a separate component from the A / V display device 114 (FIG. 1), in some embodiments the A / V hub and the A / V display device are combined into a single component or a single electronic device.
[0214] The above-described embodiments illustrate a communication technology with voice and / or video content (such as HDMI (registered trademark) content), but in other embodiments, the communication technology is used in any kind of context of data or information. For example, the communication technology can be used with home automation data. In these embodiments, the A / V hub 112 (FIG. 1) can facilitate communication and control among a variety of electronic devices. Therefore, using the A / V hub 112 (FIG. 1) and the communication technology, services in the so-called Internet of Things can be facilitated or implemented.
[0215] In the above description, the inventors have referred to "some embodiments". Note that "some embodiments" represents all subsets of possible embodiments, but not always the same subset of embodiments.
[0216] The above description is intended to enable those skilled in the art to use the present disclosure and is provided in the context of a particular application and its requirements. Further, the above description of the embodiments of the present disclosure is presented for purposes of illustration and description only. These are not intended to be comprehensive or to limit the present disclosure to the disclosed forms. Accordingly, many modifications and variations will be apparent to those skilled in the art, and the general principles defined herein can be applied to other embodiments and applications without departing from the spirit and scope of the present disclosure. In addition, the discussion of the above embodiments is not intended to limit the present disclosure. Therefore, the present disclosure is not intended to be limited to the illustrated embodiments, but is to be construed in accordance with a broad range that does not conflict with the principles and features disclosed herein.
Description of Reference Numerals
[0217] 100 System 110 Portable Electronic Device 112 A / V Hub 114-1 A / V Display Device 114-2 A / V Display Device 114-3 A / V Display Device 116 Receiver Device 118-1 Speaker 118-2 Speaker 118-3 Speaker 120-1 Content Source 120-2 Content Source 122-1 Radio 122-2 Radio 122-3 Radio 122-4 Radio 122-5 Radio 122-6 Radio 122-7 Radio 124 Wireless Signal 126 HDMI (Registered Trademark) Cable 128 TSD
Claims
1. one or more nodes configured to be communicatively coupled to the one or more antennas; an interface circuit communicatively coupled to the one or more nodes; An adjustment device comprising: The adjustment device comprises: receiving an input frame associated with an electronic device from the one or more nodes, a given input frame including a transmission time at which a given electronic device transmitted the given input frame; storing a receive time that the input frame was received, the receive time being based on a clock of the coordinating device; calculating a current time offset between a clock of the electronic device and a clock of the coordinating device based on the reception time and the transmission time of the input frame; transmitting to the one or more nodes one or more output frames including audio content and playback timing information for the electronic device, the playback timing information specifying a playback time at which the electronic device should play the audio content based on the current time offset; configured to execute A coordination device, wherein the playback times of the electronic devices have a temporal relationship such that playback of the audio content by the electronic devices is coordinated.
2. 2. The coordination device of claim 1, wherein the temporal relationship has a non-zero value, such that at least some of the electronic devices can be instructed to play the audio content with phase relative to each other by using different values of the play time.
3. The adjustment device of claim 2 , wherein the different play times are based on an acoustic characterization of the environment.
4. The adjustment device of claim 2 , wherein the different playback times are based on a desired acoustic characteristic of an environment.
5. the electronic device is positioned at a vector distance from the coordinating device; the interface circuitry is configured to determine a magnitude of the vector distance based on the transmit time and the receive time using radio ranging and to determine an angle of the vector distance based on an angle of arrival of a radio signal associated with the input frame; The coordination device of claim 2 , wherein the different play times are based on the determined vector distances.
6. The adjustment device of claim 2 , wherein the different play times are based on an estimated position of a listener relative to the electronic device.
7. The interface circuit further comprises: receiving a frame associated with another electronic device from the one or more nodes; calculating the estimated position of the listener based on the received frames; 7. The adjusting device according to claim 6, characterized in that it is configured to:
8. The conditioning device further comprises an acoustic transducer configured to perform sound measurements of the environment; The adjustment device of claim 6 , wherein the adjustment device is configured to calculate the estimated position of the listener based on the sound measurements.
9. the interface circuitry is further configured to receive from the one or more nodes additional sound measurements of the environment associated with another electronic device in the environment; The adjustment device of claim 6 , wherein the adjustment device is configured to calculate the estimated position of the listener based on the additional sound measurements.
10. The interface circuit includes: Perform a time-of-flight measurement, calculating the estimated position of the listener based on the time-of-flight measurements; 7. The adjusting device according to claim 6, characterized in that it is configured to:
11. the electronic device is positioned at a non-zero distance from the adjustment device; 2. The device of claim 1, wherein the current time offset is calculated based on the transmit time and the receive time using radio ranging by ignoring the distance.
12. 2. The adjustment device of claim 1, wherein the current time offset is further based on a model of clock drift of the electronic device.
13. A non-transitory computer-readable storage medium for use with a coordinating device, the computer-readable storage medium, when executed by the coordinating device, causing the coordinating device to: receiving input frames associated with electronic devices from one or more nodes of the coordinating device communicatively coupled to one or more antennas, a given input frame including a transmission time at which a given electronic device transmitted the given input frame; storing a receive time that the input frame was received, the receive time being based on a clock of the coordinating device; calculating a current time offset between a clock of the electronic device and a clock of the coordinating device based on the receiving time and the transmitting time of the input frame; transmitting to the one or more nodes one or more output frames including audio content and playback timing information for the electronic device, the playback timing information specifying a playback time at which the electronic device should play the audio content based on the current time offset; a program module for adjusting the playback of the audio content by executing one or more operations including 11. A non-transitory computer-readable storage medium for use with a coordination device, wherein the playback times of the electronic devices have a temporal relationship such that playback of audio content by the electronic devices is coordinated.
14. the time relationship has a non-zero value, such that at least some of the electronic devices are instructed to play the audio content with phase relative to each other by using different values of the play time; 14. The computer-readable storage medium of claim 13, wherein the different play times are based on an acoustic characterization of an environment.
15. the time relationship has a non-zero value, such that at least some of the electronic devices are instructed to play the audio content with phase relative to each other by using different values of the play time; 15. The computer-readable storage medium of claim 14, wherein the different playback times are based on desired acoustic characteristics in the environment.
16. the time relationship has a non-zero value, such that at least some of the electronic devices are instructed to play the audio content with phase relative to each other by using different values of the play time; the electronic device is positioned at a vector distance from the coordinating device; The one or more operations include: determining a magnitude of the vector distance based on the time of transmission and the time of reception using radio ranging; determining an angle of the vector distance based on an angle of arrival of a wireless signal associated with the input frame; wherein the different play times are based on the determined vector distance.
15. The computer-readable storage medium of claim 14.
17. the time relationship has a non-zero value, such that at least some of the electronic devices are instructed to play the audio content with phase relative to each other by using different values of the play time; 15. The computer-readable storage medium of claim 14, wherein the different play times are based on an estimated position of a listener relative to the electronic device.
18. The one or more operations include: performing a time-of-flight measurement; calculating the estimated position of the listener based on the time-of-flight measurements; 20. The computer readable storage medium of claim 18, comprising:
19. 14. The computer-readable storage medium of claim 13, wherein the current time offset is further based on a model of clock drift of the electronic device.
20. 1. A method for adjusting playback of audio content, comprising: The adjustment device allows receiving input frames associated with electronic devices from one or more nodes of the coordinating device communicatively coupled to one or more antennas, a given input frame including a transmission time at which a given electronic device transmitted the given input frame; storing a receive time that the input frame was received, the receive time being based on a clock of the coordinating device; calculating a current time offset between a clock of the electronic device and a clock of the coordinating device based on the receiving time and the transmitting time of the input frame; transmitting to the one or more nodes one or more output frames including audio content and playback timing information for the electronic device, the playback timing information specifying a playback time at which the electronic device should play the audio content based on the current time offset; Including, 13. A method according to claim 12, wherein the playback times of the electronic devices have a temporal relationship such that the playback of the audio content by the electronic devices is coordinated.
Citation Information
Patent Citations
Acoustic system, server, speaker and sound image localization checking method in acoustic system, server instrument, speaker instrument, and sound system
JP2005175744A
Audio device and audio system
JP2011091703A
Signal correction program, signal correction device and signal correction method
JP2013030985A
Wireless voice transmission system and source equipment
JP2016208285A
Sound apparatus, television receiver, speaker device, audio signal adjustment method, program, and recording medium
WO2015194326A1