Dual speaker system

Through the design and signal processing of a dual-speaker system, audio privacy switching between public and private modes is achieved, solving the audio privacy problem of headphones in noisy environments and ensuring that users can listen to audio content privately in public places.

CN116648928BActive Publication Date: 2026-05-22APPLE INC
View PDF 3 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
APPLE INC
Filing Date
2021-09-24
Publication Date
2026-05-22

AI Technical Summary

Technical Problem

Existing headphones and earphones fail to provide audio privacy in noisy environments, making it impossible for users to listen to audio content privately in public places without being overheard by those around them.

Method used

Employing a dual-speaker system, two speaker drivers share a common rear volume within the housing and adjust the phase relationship of the driver signals according to environmental conditions to generate in-phase or out-of-phase sound waves, thereby achieving omnidirectional or directional sound output and providing audio privacy.

Benefits of technology

In public mode, the audio content is ensured to be heard by the user, while in private mode, the audio content is effectively reduced or eliminated from being heard by those around the user, thus providing audio privacy protection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116648928B_ABST
    Figure CN116648928B_ABST
Patent Text Reader

Abstract

A wearable device (3) comprises a housing (11), a first loudspeaker driver (12) and a second loudspeaker driver (13), wherein both loudspeaker drivers (12, 13) are integrated within the housing (11) and arranged to project sound into the surrounding environment. Furthermore, the first loudspeaker driver (12) is closer to a wall (17) of the housing (11) than the second loudspeaker driver (13), and the first loudspeaker driver (12) and the second loudspeaker driver (13) share a common back volume (14) within the housing (11).
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Cross-reference application

[0002] This application claims the benefit of priority to U.S. Provisional Patent Application No. 63 / 083,760, filed September 25, 2020, the entire contents of which are incorporated herein by reference. Technical Field

[0003] One aspect of this disclosure relates to a dual-speaker system that provides audio privacy. Other aspects are also described. Background Technology

[0004] A headset is an audio device that includes a pair of speakers, each placed in the user's ear when worn on or around the user's head. Similar to a headset, headphones (or in-ear headsets) are two separate audio devices, each with a speaker that inserts into the user's ear. Both headsets and headphones are typically wired to a separate playback device, such as an MP3 player, which drives each speaker of the device with an audio signal to produce sound (e.g., music). Headsets and headphones provide a convenient way for users to listen to audio content privately without having to broadcast it to others nearby. Summary of the Invention

[0005] One aspect of this disclosure is an output device (such as a wearable device, headphones, or headset) comprising a housing, a first "extra-auricular" speaker driver, and a second extra-auricular speaker driver, wherein the two speaker drivers are arranged to project sound into the surrounding environment. Both speaker drivers may be integrated within the housing (e.g., as part of the housing) such that the first speaker driver is positioned closer to a wall of the housing than the second speaker driver. For example, the first speaker driver may be coupled to one wall, while the second speaker driver is coupled to another wall, wherein when the wearable device is worn on a user's head, the wall with the first speaker driver is closer to the user's ear than the other wall with the second driver.

[0006] In one aspect, the two speaker drivers may share a common rear volume within the housing. In some aspects, the common rear volume may be a sealed volume, where air within the volume cannot escape to the surrounding environment. In one aspect, the two speaker drivers may be drivers of the same type (e.g., "full-range" drivers that reproduce as many audible frequencies as possible). In another aspect, the speaker drivers may be drivers of different types (e.g., one is a "low-frequency driver" that reproduces low-frequency sounds, and the other is a full-range driver). In some aspects, the speaker drivers may project sound in different directions. For example, the front of the first speaker driver (e.g., the front of the diaphragm) may face a first direction, while the front of the second speaker driver may face a second direction different from the first direction (e.g., the two directions are opposite directions along the same axis).

[0007] In another aspect, the output device can be designed differently. For example, the output device may include an elongated tube having a first open end and a second open end, the first open end being connected to a common rear volume within a housing, and the second open end opening to the surrounding environment. Thus, air can travel between the rear volume and the surrounding environment. In one aspect, the sound output level of the rearward-radiated sound generated at the second open end of the elongated tube by at least one of the first and second speaker drivers is at least 10 dB SPL lower than the sound output level of the forward-radiated sound generated by at least one of the first and second speaker drivers.

[0008] In another aspect, the output device's housing forms an open enclosure outside the common rear volume and surrounding the front of the second speaker driver. In one aspect, the open enclosure opens to the surrounding environment through several ports through which the second speaker driver projects forward-radiated sound into the environment. In some aspects, the output device may also include an elongated tube, as described above.

[0009] In one aspect, the front of the first speaker driver faces a first direction, and the front of the second speaker driver faces a second direction. In some aspects, the first direction and the second direction are opposite directions along the same axis. In another aspect, the first direction is along the first axis and the second direction is along the second axis, wherein the first axis and the second axis are separated by less than 180° about the other axis.

[0010] Another aspect of the invention is a method performed by an output device (e.g., a programmed processor of the output device) comprising a first (e.g., extra-ear) speaker driver and a second extra-ear speaker driver, both integrated within a housing of the output device and sharing an internal volume as a rear volume. The device receives an audio signal (e.g., the audio signal may contain audio content desired by a user, such as a musical composition). The device determines a current operating mode for the output device (e.g., a "non-private" or "private" operating mode). The device generates a first driver signal and a second driver signal based on the audio signal, wherein the current operating mode corresponds to at least a portion of the first driver signal and the second driver signal (e.g., within corresponding frequency bands) being generated as in-phase or out-of-phase. The device uses the first driver signal to drive the first extra-ear speaker driver and uses the second driver signal to drive the second speaker driver.

[0011] In one aspect, the device determines the current operating mode by determining whether a person is within a threshold distance of the output device, wherein, in response to determining that the person is within the threshold distance, a first driver signal and a second driver signal are generated to be at least partially out of phase with each other. In another aspect, in response to determining that the person is not within the threshold distance, the first driver signal and the second driver signal are generated to be in phase with each other.

[0012] In one aspect, the device uses a first driver signal and a second driver signal to drive a first external-ear speaker driver and a second external-ear speaker driver, respectively, to generate a beam pattern with a main lobe in the direction of the user of the output device. In another aspect, the generated beam pattern has at least one null point away from the user of the output device.

[0013] In one aspect, the device receives a microphone signal generated by a microphone of the output device, the microphone signal including ambient noise of the surrounding environment in which the output device is located, wherein a current operating mode is determined based on the ambient noise. In another aspect, the device determines a current operating mode for the output device by determining whether ambient noise masks audio signals across one or more frequency bands; in response to ambient noise masking a first set of frequency bands in the one or more frequency bands, a first operating mode is selected, in which portions of a first driver signal and a second driver signal are generated in phase across the first set of frequency bands; and in response to ambient noise not masking a second set of frequency bands in the one or more frequency bands, a second operating mode is selected, in which portions of the first driver signal and the second driver signal are generated out of phase across the second set of frequency bands. In some aspects, the first set of frequency bands and the second set of frequency bands are non-overlapping frequency bands, such that the output device operates simultaneously in both the first operating mode and the second operating mode.

[0014] Another aspect of this disclosure is a head-mounted output device including a first external-ear speaker driver and a second external-ear speaker driver, wherein when the head-mounted output device is worn on a user's head, the first driver is closer to the user's (or intended listener's) ear than the second driver. The device also includes a processor and a memory storing instructions that, when executed by the processor, cause the output device to: receive an audio signal including noise, and use the first and second speaker drivers to generate a directional beam pattern including: 1) a noisy main lobe away from the user, and 2) a null (or notch) towards the user, wherein the sound output level of the second speaker driver is greater than the sound output level of the first speaker driver.

[0015] In one aspect, the audio signal is a first audio signal and the directional beam pattern is a first directional beam pattern, wherein the memory further has instructions for performing the following operations: receiving a second audio signal comprising audio content desired by the user (e.g., voice, music, podcast, movie soundtrack, etc.), and using a first external-ear speaker driver and a second external-ear speaker driver to generate a second directional beam pattern comprising: 1) having the audio content desired by the user and oriented towards the user's main lobe, and 2) being away from the user's null point. In some aspects, the first and second external-ear speaker drivers project forward-radiating sound toward or in the direction of the user's ear.

[0016] The above overview does not constitute an exhaustive list of all aspects of this disclosure. It is contemplated that this disclosure encompasses all systems and methods that can be practiced by all suitable combinations of the aspects outlined above and those disclosed in the detailed embodiments below and specifically pointed out in the claims. Such combinations may have specific advantages not specifically set forth in the foregoing summary. Attached Figure Description

[0017] Multiple aspects are illustrated in the accompanying drawings by way of example rather than limitation, and similar reference numerals in the drawings indicate similar elements. It should be noted that references to "a" or "an" aspect in this disclosure do not necessarily refer to the same aspect, and each refers to at least one. Furthermore, for the sake of brevity and to reduce the total number of drawings, a single drawing may be used to illustrate features of more than one aspect, and for a particular aspect, not all elements in that drawing may be necessary.

[0018] Figure 1 An electronic device with an external speaker is shown.

[0019] Figure 2A dual-speaker system with an output device is shown according to one aspect, the output device having two speaker drivers sharing a common rear volume.

[0020] Figure 3 An output device with an exhaust port is shown according to one aspect.

[0021] Figure 4 An output device with a rear chamber is shown according to one aspect.

[0022] Figure 5 An output device having both an exhaust port and a rear chamber is shown according to one aspect.

[0023] Figure 6 A block diagram of a system operating in one or more operating modes according to one aspect is shown.

[0024] Figure 7 It is a flowchart of one aspect of the process used to determine which of the two operating modes the system will operate in.

[0025] Figure 8 A system with two or more speaker drivers is shown according to one aspect, which are used to generate a noise beam pattern to mask the audio content perceived by the intended listener.

[0026] Figure 9 The graph shows the signal strength of the audio content and noise relative to one or more zones around the output device, based on several factors.

[0027] Figure 10 The diagram shows the radiation beam pattern with a null point at the intended listener's ear, based on several aspects.

[0028] Figure 11 Another radiation beam pattern is shown, which directs sound to the ear of the intended listener according to one aspect. Detailed Implementation

[0029] Various aspects of this disclosure will now be explained with reference to the accompanying drawings. Unless the shape, relative position, and other aspects of the components described in any aspect are explicitly defined, the scope of this disclosure is not limited to the components shown, which are for illustrative purposes only. Furthermore, while numerous details have been set forth, it should be understood that some embodiments may be implemented without these details. In other instances, well-known circuits, structures, and techniques have not been shown in detail so as not to obscure the understanding of the description. Moreover, unless the meaning is explicitly contrary, all scopes shown herein are to be considered to include the endpoints of each scope.

[0030] Headset-mounted devices (such as over-ear headsets) may consist of two housings (e.g., a left housing and a right housing) designed to be placed on a user's ears. Each housing may include an "internal" speaker arranged to project sound (e.g., directly) into the user's corresponding ear canal. Once placed on the user's ears, each housing acoustically isolates the user's ears from the surrounding environment, thereby preventing (or reducing) sound leakage into (and from) the housing. During use, the sound produced by the internal speaker is audible to the user, while the seal created by the housing helps prevent eavesdropping by others nearby.

[0031] In one aspect, the head-mounted device may include an "extra-ear" speaker arranged to project sound into the environment so that the user of the device can hear it. In some aspects, unlike internal speakers that direct sound into the user's ear canal while the device's housing at least partially acoustically isolates the user's ear from the surrounding environment, an extra-ear speaker can project sound into the surrounding environment (e.g., while the user's ear may not be acoustically sealed by the head-mounted device). For example, the speaker may be arranged to project sound in any direction (e.g., away from the user and / or toward the user, such as toward the user's ear). Figure 1 An example of an electronic device 6 is shown, having an external ear speaker 5 that projects sound (e.g., music) into the surrounding environment so that a user can hear it. Because the sound is projected into the environment, nearby people may be able to eavesdrop. In some situations, a user may wish to listen privately to audio content played back by an external ear speaker, such as when participating in a private telephone conversation. In this case, the user may not want others in their surrounding environment to hear the content. One way to prevent others from hearing is to reduce the speaker's sound output. However, this may adversely affect the user experience when the user is in a noisy environment and / or may not prevent eavesdropping when others are nearby. Therefore, if a user wishes to listen to private audio content or participate in a private telephone conversation using an external ear speaker, the user may need to walk away and enter a private space away from others. However, this may be impractical if a telephone call occurs when the user cannot find a private space (e.g., when the user is on an airplane or on a bus). Therefore, there is a need for an electronic system that provides audio privacy to the user.

[0032] This disclosure describes a dual-speaker system capable of operating in one or more modes (e.g., a "non-private" (first or public) operating mode and a "private" (second) operating mode). Specifically, the system includes an output device having at least two speaker drivers (a first speaker driver and a second speaker driver) arranged to project sound into the surrounding environment, each speaker driver being part of the output device at a different location (or integrated within the housing of the output device). In one aspect, the two speakers may share a common rear volume within the housing of the output device. During operation, the output device (e.g., one or more programmable processors of the output device) receives an audio signal and determines whether the device is operating in a first operating mode or a second operating mode (or is operating in either the first or second operating mode), the audio signal potentially containing audio content desired by the user (e.g., musical works, podcasts, movie soundtracks, etc.). For example, this determination may be based on whether a person is detected within a threshold distance of the output device (e.g., by performing image recognition on image data captured by the system's camera). The system processes the audio signal to generate a first driver signal driving the first speaker driver and a second driver signal driving the second speaker driver. When in the first operating mode, the two driver signals may be in phase with each other. In this scenario, the sound waves generated by the two speaker drivers can be (e.g., at least partially) in phase with each other. In one aspect, due to constructive interference, the combination of sound waves generated by the two drivers can have a larger amplitude than the original waves. However, in the second operating mode, the two driver signals may not be (e.g., completely) in phase with each other. In this case, the sound waves generated by the two drivers can interfere destructively with each other, resulting in a reduction (or elimination) of sound experienced at one or more locations in the surrounding environment, such as by someone other than the user (e.g., someone at a certain distance from the user). Thus, as described herein, by driving the speaker drivers with signals of different phases, the user of the output device can hear the audio content they expect, while potential eavesdroppers near the user may not hear it. Therefore, the private operating mode provides audio privacy for the user. In other aspects, depending on certain environmental conditions (e.g., ambient noise levels), the dual-speaker system can operate in a first operating mode for certain frequencies and simultaneously in a second operating mode for other frequencies. Further details regarding simultaneous operation in multiple operating modes are described herein.

[0033] Figure 2 A dual-speaker system with an output device is shown according to one aspect, the output device having two speaker drivers sharing a common rear volume. Specifically, the figure shows a system (or dual-speaker system) 1 including a source device 2 and an output device 3.

[0034] In one aspect, source device 2 can be a multimedia device, such as a smartphone. In another aspect, source device can be any electronic device (e.g., including memory and / or one or more processors) that can be configured to perform audio signal processing operations and / or networking operations. Examples of such devices may include desktop computers, smart speakers, electronic servers, etc. In one aspect, source device can be any wireless electronic device, such as a tablet computer, smartphone, laptop computer, etc. In another aspect, source device can be a wearable device (e.g., a smartwatch, etc.) and / or a head-mounted device (e.g., smart glasses).

[0035] Output device 3 is shown positioned close to (or adjacent to) the user's ear (e.g., within a threshold distance from the user's ear). In one aspect, the output device can be a wearable electronic device (e.g., part of a wearable electronic device) (e.g., a device designed to be worn by or worn on the user during operation of the device). For example, the output device can be a head-mounted device (HWD). For example, the output device can be a headset, such as an over-ear or full-ear headset. In the case of an over-ear headset, the output device can be a portion of the headset housing arranged to cover the user's ear, as described herein. Specifically, the output device can be the left headset housing. In one aspect, the headset can include another output device as part of the right headset housing. Thus, in one aspect, a user can have more than one output device, each performing audio signal processing operations to provide audio privacy (e.g., operating in one or more operating modes), as described herein. As another example, the output device can be an in-ear headset (headphone or earbud). In another aspect, the output device can be any HWD (or part of any HWD), such as smart glasses. For example, the output device can be part of a component (e.g., the frame) of smart glasses. In another aspect, the output device can be an HWD that (at least partially) does not cover the user's ear (or ear canal), thus exposing the user's ear to the surrounding environment. In some aspects, the output device can be other types of wearable devices.

[0036] In another aspect, output device 3 can be any electronic device configured to output sound, perform networking operations, and / or perform audio signal processing operations, as described herein. For example, output device can be (e.g., a standalone) megaphone, smart speaker, part of a home entertainment system, or part of a vehicle audio system. In some aspects, output device can be part of another electronic device, such as a laptop device, desktop device, or multimedia device such as source device 2 (as described herein).

[0037] Output device 3 includes a housing 11, a first speaker driver 12, and a second speaker driver 13. In one aspect, the output device may include more (or fewer) speaker drivers. In another aspect, the two speaker drivers may be integrated with the housing (or a portion thereof) of the output device at different locations around the output device. As shown, the two speaker drivers are located opposite each other. In one aspect, the first driver may be positioned closer to a wall of the housing than the second driver. Specifically, the first speaker driver is positioned (or coupled to) a first wall 17 (e.g., rear side) of the housing 11 of the output device, while the second speaker driver is positioned on a second wall 18 (e.g., front side) of the housing opposite wall 17. Thus, the second driver is further away from the first wall than the first driver. In another aspect, the first driver may be positioned closer to a wall than the second driver, wherein neither driver is coupled to (or positioned on) that particular wall. For example, the two speaker drivers may be coupled to another wall of the housing (not shown) that is coupled to the first wall 17. In this case, the first speaker driver may be separated from the first wall by a first (e.g., horizontal) distance, while the second speaker driver may be separated from the first wall by a second distance greater than the first distance. In some respects, the speaker drivers can be positioned differently, such as both speaker drivers being positioned on the same wall. In this case, the first speaker driver can be positioned closer to the first wall 17.

[0038] In some aspects, speaker drivers 12 and 13 may share a common rear volume 14 within the housing. Specifically, the rear volume may be an internal volume of the housing containing a volume of air and opening to the back of the diaphragm of each speaker driver. For example, the rear portion of each speaker driver located behind the diaphragm (or cone) of the driver (e.g., which may include a voice coil, magnet, black plate) may be exposed to (or within) the common rear volume 14. In this figure, the rear volume 14 is sealed within the housing of the output device, meaning that the air contained within this volume is confined within the housing. Thus, in one aspect, the rear volume 14 is an open space within the output device 3 containing a volume of air and enclosed (or sealed) within the housing of the output device. In some aspects, the rear volume may not be confined within the housing (e.g., as...). Figure 3 (as shown and described).

[0039] As described herein, the speaker driver is positioned on one or more walls of the housing 11 of the output device 3. In one aspect, the speaker driver may be arranged such that it is fixed to (or attached to) a corresponding wall of the housing. For example, speaker driver 12 may be coupled to wall 17 through an opening in the wall, such that the rear portion of the driver is exposed to the rear volume 14, while the front of the driver is exposed to the surrounding environment. In another aspect, one or more speaker drivers may be integrated into the housing, such that the driver is coupled to an interior portion of the wall. In this case, the speaker driver may be wholly (or substantially) contained within the rear volume 14.

[0040] As shown, both speaker drivers 12 and 13 are extra-ear speaker drivers arranged to project sound into the surrounding environment. In one aspect, the speaker drivers are arranged to project sound in different directions. For example, the first speaker driver 12 is arranged to project sound in one (first) direction, while the second speaker driver 13 is arranged to project sound in another (second) direction. For example, the front of the first speaker driver faces the first direction, and the front of the second speaker driver faces the second direction. In one aspect, the front of the speaker driver may be the front side of the diaphragm of the speaker driver, wherein the front side projects (or at least one) direction away from the driver that projects the forward-radiating sound generated by the speaker driver. As shown, the two speaker drivers point in opposite directions along the same (e.g., central longitudinal) axis (not shown), which extends through each driver. Thus, the first speaker driver 12 is shown projecting sound toward the user's ear, while the second speaker driver 13 is shown projecting sound away from the ear. In one aspect, the output device may be positioned differently around the user's head (and / or body). In another aspect, one speaker in the loudspeaker may be positioned with its center offset from the central longitudinal axis of the other speaker. For example, a first speaker driver 12 may be oriented along a first axis and a second speaker driver may be oriented along a second axis, wherein the two axes may be separated by less than 180° about another axis through which both the first and second axes intersect.

[0041] In one aspect, the two speaker drivers are positioned differently relative to the user (e.g., integrated within the housing of the output device). Specifically, when the output device is being worn by the user, one speaker driver may be closer to a part of the user than the other. For example, as shown, the first speaker driver 12 is closer to the user's ear than the second speaker driver 13. Further details regarding the placement of the speaker drivers are described herein.

[0042] During operation (of output device 3), the two speaker drivers generate sound waves that radiate outward (or forward). As shown, the two speaker drivers generate forward-radiating sound 15 (shown as an extended black solid curve) projected into the surrounding environment (e.g., in the direction facing the front of each respective speaker driver) and rearward-radiating sound 16 (shown as an extended black dashed curve) projected into the rear volume 14. As described herein, the sound (and more specifically, the spectral content) generated by each speaker driver can vary based on the operating mode of the output device currently in operation. Further details regarding operating modes are described herein.

[0043] Each of speaker drivers 12 and 13 can be an electrically driven driver specifically designed for sound output in a particular frequency band, such as a subwoofer, tweeter, or midrange driver. In one aspect, either driver can be a “full-range” (or “full-band”) electrically driven driver that reproduces as much of the audible frequency range as possible. In another aspect, each speaker driver can be of the same type (e.g., both drivers are full-range drivers). In another aspect, the two drivers can be different (e.g., the first driver 12 is a woofer, and the second driver 13 is a tweeter). In yet another aspect, the two speakers can produce different audio frequency ranges, with at least a portion of the two frequency ranges overlapping. For example, the first driver 12 can be a woofer, and the second driver 13 can be a full-range driver. Thus, at least a portion of the spectral content produced by the two drivers can have overlapping frequency bands, while other portions of the spectral content produced by the drivers may not overlap.

[0044] In one aspect, the output device 3 (and / or the source device 2) may include more (or fewer) components as described herein. For example, the output device may include one or more microphones. Specifically, the device may include an “external” microphone arranged to capture ambient sound and / or may include an “internal” microphone arranged to capture sound inside the output device (e.g., the housing 11 of the output device). For example, the output device may include a microphone arranged to capture rearward-radiated sound 16 within the rear volume 14. In another aspect, the output device may include one or more display screens arranged to present image data (e.g., still images and / or video). In some aspects, the output device may include more (or fewer) speaker drivers.

[0045] As shown in the figure, source device 2 is communicatively coupled to output device 3 via wireless connection 4. For example, the source device can be configured to establish a wireless connection with the output device via any wireless communication protocol (e.g., the BLUETOOTH protocol). During the established connection, the source device can exchange (e.g., transmit and receive) data packets (e.g., Internet Protocol (IP) packets) with the output device, which may include audio digital data. Alternatively, the source device can be coupled to the output device via a wired connection. In some aspects, the source device can be part of (or integrated into) the output device. For example, as described herein, at least some components of the source device (e.g., at least one processor, memory, etc.) can be part of the output device. Therefore, at least some (or all) of the operations of operating in several operating modes (and / or switching between several operating modes) can be performed by the source device, the output device, or a combination thereof (e.g., at least one processor of the source device, the output device, or a combination thereof).

[0046] As described herein, output device 3 is configured to output one or more audio signals via at least one of a first speaker driver 12 and a second speaker driver 13 when operating in at least one of several operating modes (e.g., public mode or private mode). In public mode, output device 3 is configured to drive the two speaker drivers in a manner that makes them in phase with each other. Specifically, output device 3 drives the two speakers using in-phase driver signals. In one aspect, the driver signals may contain the same audio content for synchronous playback via the two speaker drivers. In another aspect, the two speaker drivers may be driven using the same driver signal (which may be an input audio signal, such as the left audio channel of a musical piece). Thus, driving the two speaker drivers in phase causes the forward-radiated sound 15 to interfere constructively, thereby producing an omnidirectional sound pattern (or a monopole sound source) containing audio content. In one aspect, at least one of the driver signals may be (e.g., slightly) out of phase with the other driver signal to account for the distance between the two speakers. For example, output device 3 (e.g., the processor of the output device) may apply a phase shift to a first driver signal (e.g., at least a portion thereof) used to drive a first speaker driver, but not to a second driver signal (which may be the same as (or different from) the original first driver signal) used to drive a second speaker driver. Further details regarding the application of phase shifting are described herein.

[0047] In private mode, output device 3 is configured to drive the two speaker drivers in a manner that makes them out of phase with each other. Specifically, the output device uses driver signals that are out of phase with each other to drive the two speaker drivers. In one aspect, the two driver signals may be 180° out of phase with each other (or less than 180°). Therefore, the phrase "out of phase" described below can refer to two signals that are out of phase by 0°–180°. For example, the output device may process an audio signal (e.g., by applying one or more audio processing filters) to produce driver signals that are out of phase. When used to drive two driver signals that are out of phase with each other, the output device may produce a dipole sound pattern having a first lobe (or "main" lobe) with audio content and a second lobe (or "back" lobe) containing out-of-phase audio content relative to the audio content contained within the main lobe. In this case, the user of the output device can primarily hear the audio content within the main lobe. However, others located further away from the output device than the user (e.g., beyond a threshold distance) may not hear the audio content due to destructive interference caused by the back lobe. In one respect, the frequency response of a dipole may have a sound pressure level that is 15 dB to 40 dB lower than that of a monopole (e.g., at a given (threshold) distance from the output device) than that of a unipole (e.g., when produced in common mode).

[0048] In one aspect, the output device can operate in both private and public modes (e.g., simultaneously). In this case, the driver signals can be (at least) partially in-phase and (at least) partially out-of-phase. Specifically, the spectral content contained within the driver signals can be partially in-phase and / or partially out-of-phase. For example, the high-frequency content within each driver signal can be partially (or fully) in-phase, while the low-frequency content contained within the driver can be at least partially out-of-phase. Further details regarding operation in both modes are described herein.

[0049] As described herein, applying one or more signal processing operations (e.g., spatial filters) to an audio signal produces one or more sound patterns that can be used to selectively direct sound toward a specific location in space (e.g., a user's ear) and away from another location (e.g., the location of a potential eavesdropper). Further details on the generation of sound patterns are described herein.

[0050] Return to Figure 2Regardless of the operating mode of the device, the volumetric air confinement within the rear volume 14 can negatively impact the performance of the output device 3. In one aspect, the output device may exhibit low low-frequency efficiency, meaning it lacks an extended low-frequency range based on one or more physical characteristics. For example, a smaller housing 11 of the output device can increase its resonant frequency, contrasting with larger output devices, which may also have greater low-frequency efficiency. Furthermore, the volumetric air confinement acts as a “stiff” spring, reducing the potential displacement of the speaker driver diaphragm. This reduction can also be attributed to the increased resonant frequency. In another aspect, the output device may exhibit reduced low-frequency efficiency when operating in privacy mode due to destructive interference at low frequencies.

[0051] Figures 3 to 5 An output device 3 having one or more physical characteristics (or features) is shown, and the output device is shown to be adjacent to a user's ear. Specifically, when the output device is worn (or in use) by the user, the output device (e.g., at least a portion of the output device) can be positioned within a threshold distance from the user's ear.

[0052] Figure 3 An output device 3 with an exhaust port is shown according to one aspect. Specifically, the output device includes an elongated tube (or component) 21 connected to and extending away from a first wall 17. Specifically, the elongated tube has a first open end connected to a common rear volume 14 within a housing 11, such that the interior of the elongated tube (e.g., through the first wall 17) is fluidly connected to the rear volume 14 of the housing 11. The elongated tube also has a second open end (or exhaust port 22) leading to the surrounding environment. Thus, the tube fluidly connects the rear volume to the surrounding environment, such that... Figure 2 A certain volume of air confined within the rear volume of the housing is now able to flow between the common rear volume and the surrounding environment. Therefore, the change in sound pressure within the housing caused by the rearward radiation of sound from the speaker driver (shown as being emitted from exhaust port 22) causes air to move into and out of the exhaust port.

[0053] In one aspect, the elongated tube can have any size, shape, and length. In another aspect, the length of the tube can be set such that the sound level at the exhaust port is less than the sound level at one or more of the speaker drivers 12 and 13. For example, the sound output level (as measured or sensed) of the rearward-radiated sound produced by the first (and / or second) speaker driver at the exhaust port 22 is at least 10 dB SPL less than the sound output level of the forward-radiated sound produced by the same speaker driver. Therefore, the sound output at the exhaust port may not adversely affect the user's sound experience of the output device. In another aspect, the sound output level at the user's ear can be at least a certain threshold less than the sound output level at the exhaust port. For example, the location of the exhaust portion can make the sound output level at the user's ear (which is closest to the exhaust port) at least 10 dB SPL less than the sound output level at the port itself. In some aspects, the elongated tube can be shaped to reduce the audibility of the rearward-radiated sound exhausted from port 22. For example, the elongated tube can be shaped such that the exhaust port (at least partially) is behind the user's ear, so that the user's ear can block at least a portion of the sound produced by the port. In another respect, the pipe can be shaped and / or positioned differently. In some respects, the sound projected by the exhaust port may be inaudible to the user of the output device.

[0054] In one respect, for example, an exhaust port can provide better low-frequency efficiency than an output device without an exhaust port, such as... Figure 2 As shown. Specifically, since the air in the housing is no longer confined and can therefore enter and exit, low-frequency efficiency is improved when the output device drives at least one speaker driver in the speaker drivers.

[0055] Figure 4 An output device with a rear chamber is shown according to one aspect. Specifically, the figure shows that the housing 11 of the output device 3 forms a rear chamber 41 (or an open housing), which is outside a common rear volume 14 and surrounds a second speaker driver 13 (e.g., the front of the second speaker driver). Thus, as shown, the common rear volume contains confined air, such as... Figure 2 As shown, it has a rear chamber formed around the second speaker. In one aspect, the rear chamber may be part of the housing, thus forming an integrated unit. In another aspect, the rear chamber may be removably coupled to (the remainder of) the housing, such that the rear chamber can be attached to and / or removed from the housing.

[0056] The rear chamber 41 includes one or more rear ports 42. This chamber is designed to open to the surrounding environment through these ports, through which the second speaker driver 13 projects forward-radiated sound. In one aspect, each of these ports is positioned such that the forward-radiated sound from the second speaker driver radiates at one or more frequencies. Specifically, each port can simulate a monopole source, thereby generating multiple dipoles when the output device operates in private mode (e.g., when the two speaker drivers output audio content that is at least partially out of phase with each other). In one aspect, each monopole source in the rear ports has a different spectral content depending on its position relative to the second speaker driver. For example, the rear port located furthest from the second speaker driver (e.g., along a longitudinal axis extending through the center of the speaker driver) can (primarily) output low-frequency audio content. As ports are closer to the second speaker driver (and further away from the furthest rear port), these ports can output higher-frequency audio content than ports further away from the second speaker driver.

[0057] In one respect, the output device can control how the audio content output from the rear port is controlled by adjusting the way the second speaker driver is driven. Therefore, based on the adaptive behavior of the second speaker driver (e.g., the speaker's output spectral content), the rear chamber can provide the output device with better low-frequency efficiency and less distortion. This article describes further details regarding controlling the output of the rear port.

[0058] In one aspect, the rear chamber 41 may be positioned such that the sound level of the forward-radiated sound projected from the rear port 42 at the user's position (e.g., the user's ear) is less than the sound level of the forward-radiated sound from the first speaker driver 12 (and / or the second speaker driver 13). For example, the forward-radiated sound projected from the rear port may be at least 6 dB lower than the forward-radiated sound from the first speaker driver.

[0059] Figure 5 An output device having both an exhaust port and a rear chamber is shown according to one aspect. Therefore, in this figure, the output device is... Figure 3 and Figure 4 The output devices can be a combination of various types. Therefore, the output devices may include those that offer performance advantages due to their elongated tubes and rear chambers. For example, when operating in public mode, although the device may not provide sufficient privacy, the release of internal air pressure due to the exhaust port provides good low-frequency efficiency and minimal distortion. When operating in private mode, the output devices can control the performance of the second speaker driver to produce multiple dipoles, thereby increasing low-frequency efficiency (due to less destructive interference) and reducing distortion (due to less required speaker driver offset).

[0060] Figure 6 A block diagram of a system 1 operating in one or more operating modes according to one aspect is shown. Specifically, the diagram illustrates system 1, which includes a controller 51, at least one (e.g., external) microphone 55, a first (extra-auricular) speaker driver 12, and a second (extra-auricular) speaker driver 13. In one aspect, each of these components may be part of an output device 3 (e.g., integrated into the housing of the output device). In another aspect, at least some of these components may be... Figure 2 The diagram shows a portion of the output device and source device 2. For example, a speaker driver may be integrated into the output device (e.g., the housing of the output device), while a controller may be integrated into the source device. In this case, the controller may perform audio privacy operations as described herein to generate one or more driver signals, which are transmitted to the output device (e.g., via a connection, such as...). Figure 2 Wireless connection 4) to drive speaker driver to produce sound.

[0061] Controller 51 may be a dedicated processor such as an application-specific integrated circuit (ASIC), a general-purpose microprocessor, a field-programmable gate array (FPGA), a digital signal controller, or a set of hardware logic structures (e.g., filters, arithmetic logic units, and dedicated state machines). The controller is configured to perform audio signal processing operations, such as audio privacy operations and networking operations as described herein. Further details regarding the operations performed by the controller are described herein. In one aspect, the operations performed by the controller may be implemented in software (e.g., as instructions stored in the memory of the source device (and / or the controller's memory) and executed by the controller), and / or may be implemented by hardware logic structures. In another aspect, the output device may include further elements such as memory elements, one or more display screens, and one or more sensors (e.g., one or more microphones, one or more cameras, etc.). For example, one or more of these elements may be part of the source device, the output device, or may be part of a separate electronic device (not shown).

[0062] As shown in the figure, the controller 51 may have one or more operation blocks, which may include a context engine and decision logic 52 (hereinafter referred to as the context engine), a rendering processor 53, and an environment masking estimator 54.

[0063] An ambient masking estimator 54 is configured to determine an ambient masking threshold (or masking threshold) for ambient sounds within the surrounding environment. Specifically, the estimator is configured to receive a microphone signal generated by microphone 55, wherein the microphone signal corresponds to (or contains) ambient sounds captured by the microphone. The estimator is also configured to use the microphone signal to determine the noise level of the ambient sounds as a masking threshold. Auditory masking occurs when the perception of one sound is affected by the presence of another sound. In one aspect, the estimator determines the frequency response of the ambient sounds as a threshold. Specifically, the estimator determines the magnitude (e.g., dB) of the spectral content contained within the microphone signal. In some aspects, system 1 uses the masking threshold to determine how to process the audio signal, as described herein.

[0064] In one aspect, context engine 52 is configured to determine (or decide) whether output device 3 will operate in one or more operating modes (e.g., public mode or private mode). Specifically, context engine is configured to determine whether (e.g., most) sound output by the first speaker driver and the second speaker driver is only heard by the user (or wearer) of the output device. For example, context engine determines whether a person is within a threshold distance of the output device. In one aspect, in response to determining that a person is within the threshold distance, context engine selects private mode as the mode selection, and in response to determining that a person is not within the threshold distance, context engine selects public mode as the mode selection. Specifically, to make this determination, context engine receives sensor data from one or more sensors (not shown) of system 1. For example, the system (e.g., the system's output device) may include one or more cameras arranged to capture image data of the camera's field of view. Context engine is configured to receive image data (as sensor data) from the cameras and is configured to perform an image recognition algorithm on the image data to detect a person therein. Once a person is detected therein, context engine determines the person's position relative to a reference point (e.g., the position of the output device, the position of the camera, etc.). For example, when the camera is part of the output device, the context engine can receive sensor data indicating the position and / or orientation of the output device (e.g., from an inertial measurement unit (IMU) integrated within the output device). Once the position of the output device (which may correspond to the position of the camera) is determined, the context engine determines the position of the person relative to the output device by analyzing image data (e.g., pixel height and width).

[0065] In one respect, this determination can be based on whether a specific object (or location) is within a threshold distance for the user. For example, the context engine 52 can determine whether another output source (e.g., television, radio, etc.) is within a threshold distance. As another example, the engine can determine whether the user's location is a place where audio content will only be heard by the user (e.g., a library).

[0066] In another aspect, the context engine can acquire additional sensor data to determine whether a person (object or place) is within a threshold distance. For example, the context engine can acquire proximity sensor data (e.g., from one or more proximity sensors of the output device). In some aspects, the context engine can acquire sensor data from another electronic device. For example, controller 51 can acquire data from one or more electronic devices near the output device that indicates the device's location.

[0067] In some respects, the context engine can acquire user input data (as sensor data) indicating a user selection for any mode. For example, the source device's (e.g., touch-sensitive) display screen can receive user selections on graphical user interface (GUI) items displayed on the display screen to initiate (or activate) a public mode (and / or a private mode). Once received, the source device can transmit the user selections as sensor data to controller 51.

[0068] In one aspect, the context engine 52 can determine which operating mode to use based on content analysis of the audio signal. Specifically, the context engine can analyze the (user-expected) audio content contained within the audio signal to determine whether the audio content is private. For example, the context engine can determine whether the audio content contains words that indicate the audio content will be private. In another aspect, the engine can analyze the type of audio content, such as the source of the audio signal. For example, the engine can determine whether the audio signal is a downlink signal received during a telephone call. If so, the context engine can consider the audio signal to be private.

[0069] In one respect, the context engine 52 can determine which mode of operation to use based on system data. In other respects, system data may include user preferences. For example, when a specific type of audio content is output through a speaker driver, the system can determine whether the user of the output device prefers a particular operating mode. For example, when the audio content is a musical work and the user has listened to this type of content in public mode in the past, the context engine can determine which mode to operate in. Therefore, the context engine can execute machine learning algorithms to determine which mode of operation to use based on how the user has listened to audio content in the past.

[0070] In another aspect, system data can indicate system operating parameters (e.g., "overall system health"). Specifically, system data may relate to operating parameters of the output device, such as the battery level of the output device's internal battery, internal temperature (e.g., the temperature of one or more components of the output device), etc. In one aspect, the context engine can determine to operate in public mode in response to an operating parameter falling below a threshold. As described herein, when operating in private mode, distortion can increase due to high driver offset. This increased offset is due to additional power being supplied to the speaker driver (or more power than would normally be required when operating in public mode). Therefore, in response to a battery level falling below a threshold, the context engine can determine to operate in public mode to conserve power. Similarly, high driver offset can cause an increase in the internal temperature of the output device (or more specifically, the driver temperature). If the temperature exceeds a threshold, the context engine can select public mode. In one aspect, the context engine can select public mode in response to an operating parameter (or at least one operating parameter) exceeding a threshold.

[0071] In another aspect, the context engine may rely on one or more conditions to determine which operating mode to operate in, as described herein. Specifically, the context engine may select a particular operating mode based on a confidence score associated with the conditions described herein. In one aspect, the more conditions met, the higher the confidence score. For example, the context engine may assign a high confidence score (e.g., above the confidence threshold) when it detects a person within a threshold and detects that the user is operating in a location where the system is in private mode. When the confidence threshold is exceeded, the context engine selects private mode. In some aspects, the context engine will operate in public mode (e.g., the default) until it is determined to switch to private mode, as described herein.

[0072] In one aspect, the context engine can select one of several operating modes based on ambient noise within the environment. Specifically, the context engine can select a mode based on the spectral content (e.g., the magnitude of the spectral content) of an estimated ambient masking threshold. For example, the context engine can select a public mode in response to an ambient masking threshold having significant low-frequency content (e.g., by determining that at least one frequency band has a magnitude higher than the threshold value than another higher-frequency band). Conversely, the context engine can select a private mode in response to an ambient masking threshold having significant high-frequency content. As described herein, an output device can render an audio signal such that the output spectral content of the audio signal matches the spectral content of the ambient masking threshold in order to mask sounds from other people.

[0073] As described above, the context engine can select one of several operating modes based on one or more parameters, such as ambient noise within the environment. In another aspect, the context engine can select one or more operating modes (e.g., both public and private operating modes) that the system (or output device 3) can operate simultaneously (e.g., to maximize privacy when the output device generates audio content) based on ambient noise. In one aspect, this could be the selection of a third operating mode. Specifically, the context engine can select a "public-private" (or third) operating mode, where the controller applies audio signal processing operations to the audio signal based on the operations described herein related to both public and private operating modes. In this case, system 1 (e.g., the system's rendering processor 53) can generate driver signals for audio signals with some spectral content in phase and others (at least partially) out of phase, as described herein. Specifically, the context engine can determine whether different portions of the spectral content of the audio signal will be processed differently according to different operating modes based on the spectral content (e.g., magnitude) of the ambient noise. For example, the context engine can determine whether a portion of the spectral content of the ambient noise (e.g., signal level) (e.g., spanning one or more frequency bands) exceeds a threshold (e.g., magnitude value). In one respect, the threshold can be a predefined threshold. In another respect, the threshold can be based on the audio signal. Specifically, the threshold can be the signal level of the corresponding spectral content of the audio signal. In this case, the context engine can determine whether ambient noise (at least a portion of it) will mask the audio signal (e.g., the corresponding portion of the audio signal). For example, the context engine can compare the signal level of the ambient noise with the signal level of the audio signal and determine whether the spectral content of the ambient noise (e.g., low-frequency content) is loud enough to mask the corresponding (e.g., low-frequency) content of the audio signal.

[0074] If the ambient noise does exceed a threshold, the context engine may select the corresponding spectral portion of the audio signal (e.g., spanning the same one or more frequency bands) to operate according to a common mode, since the ambient noise sufficiently masks that spectral content of the audio signal. Conversely, if a (e.g., another) portion of the spectral content of the ambient noise does not exceed the threshold (e.g., meaning the audio content of the audio signal might be louder than the ambient noise), the context engine may select another corresponding spectral portion of the audio content to operate according to a private mode. In this case, once both modes are selected, the rendering processor can process the corresponding spectral portion of the audio content according to the selected mode. Specifically, the rendering processor may generate driver signals based on the audio signal according to the selection made by the context engine, where at least some corresponding portions of the driver signals are in-phase, while at least some other corresponding operations of the driver signals are generated out of phase. More details about the rendering processor are described in this paper.

[0075] In one aspect, once it is determined which operating mode the output device will operate in, the context engine can transmit one or more control signals to the rendering processor 53, thereby indicating the selection of one(s) operating modes (such as a public mode or a private mode). The rendering processor 53 is configured to receive the control signals and is configured to process the audio signal to generate (or produce) driver signals for each of the speaker drivers according to the selected mode. As described herein, in response to the selection of a public mode, the rendering processor 53 can generate a first driver signal and a second driver signal that contain audio content and are in phase with each other. In one aspect, the rendering processor can use the audio signal to drive both speaker drivers 12 and 13 such that the two driver signals have the same phase and / or amplitude. In one aspect, the rendering processor can perform one or more audio signal processing operations (e.g., equalization, spectral shaping) on ​​the audio signal.

[0076] In response to the selection of a private mode, the rendering processor can generate two driver signals, one of which is out of phase with the other. In one aspect, the processor can apply one or more linear filters (e.g., low-pass, band-pass, high-pass, etc.) to the audio signal such that one driver signal is out of phase (e.g., 180° out of phase) relative to the other (this may be similar to or identical to the audio signal). In another aspect, the rendering processor can generate driver signals that are at least partially in phase (e.g., between 0° and 180°). In yet another aspect, the rendering processor can perform other audio signal processing operations, such as applying one or more scalar (or vector) gains to give the signals different amplitudes. In some aspects, the rendering processor can perform different spectral shaping on the signals such that at least some frequency bands shared between the signals have the same (or different) amplitudes.

[0077] In response to a choice between a public mode and a private mode (or a public-private mode), the rendering processor can generate two driver signals, wherein a first portion of the corresponding spectral content of these signals is in phase and a second portion of the corresponding spectral content of these signals is (e.g., at least partially) out of phase. In this case, control signals from the context engine can indicate which spectral content (e.g., frequency bands) will be in phase (based on the choice of the public mode) and / or indicate which spectral content will be out of phase.

[0078] In one aspect, output device 3 is configured to generate a beam pattern. For example, when operating in public mode, both speaker drivers 12 and 13 are driven using in-phase driver signals to generate an omnidirectional beam pattern, making the sound produced by the speakers perceptible to the user of the output device and others in the vicinity of the output device. As described herein, both speaker drivers are driven using out-of-phase driver signals to generate dipoles. Specifically, the output device generates a beam pattern having a main lobe containing audio content including an audio signal. In one aspect, a rendering processor is configured to direct the main lobe toward the user of the output device (e.g., the user's ear) by applying one or more (e.g., spatial) filters. For example, the rendering processor is configured to apply one or more spatial filters (e.g., time delay, phase shift, amplitude adjustment, etc.) to the audio signal to generate a directional beam pattern. In one aspect, the direction toward which the main lobe is directed can be a predefined direction. In another aspect, the direction can be based on sensor data (e.g., image data captured by the output device's camera indicating the position of the user's ear relative to the output device). In one aspect, the rendering processor can determine the direction of the beam pattern and / or the location of the null point of the pattern based on the location of a potential eavesdropper in the surrounding environment. For example, the context engine can transmit the location information of one or more people in the surrounding environment to the rendering processor, which can filter the audio signal so that the main lobe points toward the user and at least one zero point is away from the user (e.g., making the zero point point toward another person in the environment).

[0079] In some aspects, the rendering processor can orient the main lobe toward the user of the output device and / or orient one or more nulls toward another person (e.g., in private and / or public-private modes). In other aspects, the rendering processor can orient the nulls and / or lobes differently. For example, the rendering processor can be configured to generate one or more main lobes, each oriented toward someone in the environment other than the user of the output device (or the intended listener). In addition to (or instead of) orienting the main lobes toward others, the rendering processor can orient one or more nulls toward the user of the output device. Thus, the system can direct some sound away from the user of the device, so that the user is not aware of the audio content in the surrounding environment (or is aware of less audio content than others in the surrounding environment). This type of beam pattern configuration can provide privacy to the user of the audio content when the beam pattern includes (masking) noise. Figures 8 to 11 More details are described in the text regarding the generation of noisy beam patterns.

[0080] In one aspect, the rendering processor 53 processes the audio signal based on an ambient masking threshold received from the estimator 54. As described herein, the context engine can select one or more operating modes based on the spectral content of ambient noise within the environment. Furthermore, the rendering processor can process the audio signal according to the spectral content of the ambient noise. For example, as described herein, the context engine can select a common mode in response to significant low-frequency ambient noise spectral content. In one aspect, the rendering processor can render the audio signal in the selected mode to output (corresponding to) low-frequency spectral content. In this way, the spectral content of the ambient noise can help mask the output audio content so that it is not heard by others nearby, while the user of the output device can still experience the audio content.

[0081] Furthermore, the rendering processor 53 can process the audio signal according to one or more operating modes selected by the context engine. For example, upon receiving an instruction from the context engine to select both a private mode and a public mode, the rendering processor can generate (or produce) a driver signal based on audio signals that are at least partially in phase and at least partially out of phase with each other. In one aspect, to operate simultaneously in both modes such that the driver signal is both in phase and out of phase, the rendering processor can process the audio signal based on ambient noise within the environment. Specifically, the rendering processor can determine whether (or which spectral contents) of the ambient noise will mask the audio content desired by the user to be output by the speaker driver. For example, the rendering processor can compare the audio signal (e.g., its signal level) with an ambient masking threshold. A first portion of the spectral content of the audio signal below (or equal to) the threshold can be determined to be masked by the ambient content, while a second portion of the spectral content of the audio signal above the threshold can be determined to be heard by an eavesdropper. Therefore, when generating driver signals, the rendering processor can process a first portion of the spectral content according to a common mode operation, wherein the spectral content of the driver signal corresponding to the first portion can be in phase; and the processor can process a second portion of the spectral content according to a private mode operation, wherein the spectral content of the driver signal corresponding to the second portion can be at least partially out of phase. In some aspects, the rendering processor can perform the determination of which spectral content (or more precisely, which bands of one or more frequency bands) to process according to either mode, as described above. In another aspect, the context engine can provide (e.g., in addition to the operation mode selection) an indication of what spectral content of the audio signal to process according to one or more of the selected operation modes.

[0082] In another aspect, the rendering processor may process audio signals (and / or driver signals) based on ambient noise (e.g., perform one or more audio signal processing operations). Specifically, the rendering processor may determine whether ambient noise will mask user-expected audio content to be output by a speaker driver, making it impossible for the user of the output device to hear the content. For example, the processor may compare the audio signal (e.g., its signal level) with an ambient masking threshold. In one aspect, the rendering processor compares the sound output level of a speaker driver (at least one speaker driver among the speaker drivers) with the ambient masking threshold to determine whether the user of the output device will hear user-expected audio content that is greater than the ambient noise in the surrounding environment. In response to a sound output level below the ambient masking threshold, the rendering processor may increase the sound output level of at least one speaker driver among the speaker drivers to exceed the noise level. For example, the processor may apply one or more scalar gains and / or one or more filters (e.g., low-pass filters, band-pass filters, etc.) to the audio signal (and / or individual driver signals). In some aspects, the processor may estimate the noise level at the detected location of a person in the environment based on the person's location and the ambient masking threshold to produce a revised ambient masking threshold representing an estimate of the noise level at the person's location. The rendering processor can be configured to process audio signals such that the sound output level exceeds the ambient masking threshold but is below the revised ambient masking threshold, so that potential eavesdroppers cannot experience the increase in sound.

[0083] In one aspect, the rendering processor 53 is configured to provide the user of the output device with the minimum privacy required to prevent others from listening (e.g., when operating in private mode), while minimizing the output device resources (e.g., battery power, etc.) required to output the audio content desired by the user. Specifically, the rendering processor determines whether an ambient masking threshold (or the noise level of ambient sound) exceeds the maximum sound output level of the output device. In one aspect, the maximum sound output level may be the maximum rated power of at least one of the first speaker driver 12 and the second speaker driver 13. In another aspect, the maximum sound output level may be the maximum rated power of at least one amplifier (e.g., Class D) driving at least one of the speaker drivers. In yet another aspect, the maximum sound output level may be based on the maximum amount of power available from the output device to drive the speaker drivers. For example, if the ambient masking threshold is higher than the maximum sound output level (e.g., higher than at least a predefined threshold), the rendering processor may not output an audio signal because more power than is available is required to overcome the masking threshold so that the user can hear the audio content. In one aspect, if it is determined that the sound output by the output device cannot overcome the noise level when operating in private mode, the rendering processor can be reconfigured to output the audio content desired by the user in public mode. In other aspects, the output device can output a notification (e.g., an audible notification) requesting the user's authorization to output audio content in public mode. Once authorization is received (e.g., via a voice command), the output device can begin outputting sound.

[0084] In one respect, the rendering processor can adjust audio playback based on an ambient masking threshold as a function of frequency (and signal-to-noise ratio). Specifically, the rendering processor can compare the spectral content of the ambient masking threshold with the audio signal. For example, the rendering processor can compare the magnitude of the low-frequency band of the masking threshold with the magnitude of the same low-frequency band of the audio signal. The rendering processor can determine whether the magnitude of the masking threshold is a threshold greater than the magnitude of the audio signal. In one respect, the threshold can be associated with the maximum rated power, as described herein. In another respect, the threshold can be based on a predefined SNR. In response to the masking threshold magnitude (of one or more frequency bands) being a threshold greater than the magnitude of the same frequency band of the audio signal (or exceeding the threshold by a certain amount), the rendering processor can apply gain to the audio signal to reduce the magnitude of the same frequency band of the audio signal. In other words, the rendering processor can attenuate the low-frequency spectral content of the audio signal to reduce (or eliminate) the output of the speaker driver to the spectral content because the low-frequency spectral content of the masking threshold is too high for the rendering processor to overcome ambient noise. For example, a rendering processor can apply a (first) gain to an audio signal to reduce the magnitude of low-frequency spectral content. Therefore, by attenuating spectral content that cannot overcome ambient noise, the output device can save power and prevent distortion.

[0085] In response to a masking threshold being less than a threshold value (or not exceeding a threshold value in the same frequency band of the audio signal), the rendering processor may apply a (second) gain to the audio signal to increase the magnitude. Continuing the previous example, the rendering processor may boost the low-frequency content of the audio signal above the masking threshold to overcome ambient noise. In one aspect, in response to an audio signal exceeding the masking threshold, the rendering processor may not apply a gain (e.g., cross-band).

[0086] Figure 7 This is a flowchart of one aspect of a process for determining which of the two operating modes a system device will operate in. In one aspect, process 60 is executed by controller 51 of system 1 (e.g., source device 2 and / or output device 3 of the system).

[0087] Process 60 begins with controller 51 receiving an audio signal (at box 61). Specifically, controller 51 may obtain the audio signal from an audio source (e.g., from internal memory or a remote device). In one aspect, the audio signal may include audio content desired by the user, such as musical works, movie soundtracks, etc. In another aspect, the audio signal may include other types of audio, such as a downlink audio signal of a telephone call including the sound of a telephone call (e.g., voice). The controller determines one or more current operating modes for the output device (at box 62). Specifically, the controller determines one or more operating modes for the output device to operate, such as public mode, private mode, or combinations thereof, as described herein. For example, controller 51 may determine whether a person is within a threshold distance of the output device. In this case, when a detected person (e.g., other than the user) is not determined to be within the threshold distance, the controller may determine that the output device will operate in public mode, while when a detected person is determined to be within the threshold distance, the controller may determine that the output device will operate in private mode. In another aspect, the controller may determine the mode to operate based on the presence of ambient noise in the environment. For example, the controller can determine whether ambient noise (e.g., the spectral content of the ambient noise) masks an audio signal across one or more frequency bands (e.g., by a magnitude greater than the spectral content of the audio signal). In response to ambient noise masking a first set of frequency bands (e.g., low-frequency bands), the controller can select a common operating mode for those bands, and / or in response to ambient noise not masking a second set of frequency bands (e.g., high-frequency bands) (or not masking bands above a threshold), the controller can also select a private operating mode for those bands. In one aspect, the controller can select one operating mode. In another aspect, the controller can select two operating modes based on whether a portion of the ambient noise masks or does not mask a corresponding portion of the audio signal. For example, when the first and second frequency bands are non-overlapping (or at least do not overlap beyond a threshold frequency range), the controller can select two modes, allowing the output device to operate simultaneously in both a common mode and a private mode.

[0088] Controller 51 generates a first speaker driver signal and a second speaker driver signal (at block 63) based on an audio signal, according to the determined current operating mode of one or more output devices. Specifically, the controller generates the first and second driver signals based on the audio signal, where the current operating mode corresponds to whether at least a portion of the first and second driver signals are generated as either in-phase or out-of-phase. For example, if the output device will operate in a common mode, the controller processes the audio signal to generate the first and second driver signals, where the two driver signals are in-phase. For example, in response to determining that a person is not within a threshold distance of the output device, the first and second speaker drivers may be generated as in-phase. In one aspect, rendering processor 53 may use (e.g., raw) audio signals as driver signals. In another aspect, the rendering processor may perform any audio signal processing operation (e.g., equalization) on the audio signal while still maintaining the phase between the two driver signals. In some aspects, at least some portions of the first and second driver signals may be generated as in-phase across a frequency band (e.g., a first set of frequency bands) where the output device will operate in a common mode.

[0089] However, if the output device will operate in private mode, controller 51 processes the audio signal to generate a first driver signal and a second driver signal, wherein these two driver signals are out of phase with each other. For example, portions of the first driver signal and the second driver signal may be generated out of phase across the frequency bands (e.g., a second set of frequency bands) in which the output device will operate in private mode. Therefore, when the first driver signal and the second driver signal are generated as in phase across some frequency bands and out of phase across other frequency bands, the output device can operate in both operating modes simultaneously. In one aspect, when operating in private mode, the controller may be configured to process only the portion of the driver signal corresponding to the portion of the audio signal that is not masked by ambient noise as out of phase, without processing the remaining portions (e.g., across other frequency bands) (e.g., where the phase of those portions is not adjusted). The controller uses the first driver signal to drive the first speaker driver and uses the second driver signal to drive the second speaker driver (at block 64).

[0090] Some aspects are feasible Figure 7Variations of process 60. For example, at least some of the specific operations in these processes may not be performed in the exact order shown and described. The specific operation may not be performed in a consecutive series of operations, and different specific operations may be performed in different aspects. In one aspect, although shown as selecting one of two operating modes, the controller may select both modes such that the output device operates simultaneously in (at least) both modes, as described herein. For example, when determining which operating mode to select, the controller may determine whether ambient noise will mask at least a portion of the audio signal. In response to determining that ambient noise will mask a portion (of the spectral content) of the audio signal, the controller may select a common mode to process the audio signal such that the corresponding portion of the driver signal is in phase, while in response to determining that ambient noise will not mask another portion of the audio signal, the controller may select a private mode to process the audio signal such that the corresponding portion of the driver signal is at least partially out of phase, as described herein.

[0091] In some respects, controller 51 may continuously (or periodically) perform at least some of the operations in process 60 while outputting an audio signal. For example, the controller may determine that the output device will operate in private mode based on the detection of a person within a threshold distance. However, when it is determined that the person is no longer within the threshold distance (e.g., the person has left), controller 51 may switch to public mode. As another example, the controller may switch between the two modes based on operating parameters. Specifically, in some cases, the controller may switch from private mode to public mode regardless of whether the output device will be in that mode based on operating parameters. For example, when it is determined that the battery level is below a threshold, controller 51 may switch from private mode to public mode to ensure that audio output is maintained.

[0092] As described herein, System 1 can operate in one or more operating modes: a non-private (or public) mode in which the system generates sound that is heard by the system's user (e.g., the intended listener) and one or more third-party listeners (e.g., eavesdroppers); and a private mode in which the system generates sound that is heard only (or primarily by) the intended listener, and that is not perceived (or heard) by others. To operate in private mode, the system can drive two or more speaker drivers out of phase (or disphase) such that the sound waves generated by the drivers can cancel each other out, making the sound imperceptible to third-party listeners (e.g., listeners at or beyond a threshold distance from the speaker drivers), while the intended listener can still hear the sound. In another aspect, the system can mask private content (or sound intended only for the intended listener) by generating one or more beam patterns that are away from the intended listener (e.g., and towards third-party listeners), including noise to mask the private content. Therefore, audio content (e.g., voice such as a telephone call) can be directed (or transmitted) to one area of ​​space (e.g., directed toward the intended listener), while the audio content is masked in one or more other areas of space, making the noise perceptible (e.g., only or primarily) to people in those other areas. Further details on the use of noise beam patterns are described in this paper.

[0093] Figure 8 A system 1 with two or more speaker drivers is illustrated according to one aspect, which are used to generate noise beam patterns to mask audio content perceived by an intended listener. The figure illustrates system 1 including (at least) a controller 51 and speaker drivers 12 and 13. As described herein, the system may be part of an output device 3, such that the controller and speaker drivers are integrated into the housing of the output device. In another aspect, the speaker drivers may be part of the output device, while the controller may be part of a source device communicatively coupled to the output device. Furthermore, as shown, the system uses the speaker drivers to generate a noise (directional) beam pattern 86 and an audio (directional) beam pattern 87. In one aspect, the system may generate more or fewer beam patterns, wherein each beam pattern may be directed toward a different location within the surrounding environment in which the system is located and include similar (or different) audio content. Further details regarding these beam patterns are described herein.

[0094] Controller 51 includes a signal beamformer 84 and a null (or notch) beamformer 85, each configured to generate one or more (e.g., directional) beam patterns, such as a speaker driver. In one aspect, the controller may include other operating blocks, such as Figure 6The block shown. In this case, the beamformer can be part of the rendering processor 53.

[0095] In some aspects, the null beamformer 85 receives one or more (audio) noise signals (e.g., a first audio signal), which may include any type of noise (e.g., white noise, brown noise, pink noise, etc.). In other aspects, the noise signal may include any type of audio content. In one aspect, the noise signal may be generated by a system (e.g., by an environment masking estimator 54 of controller 51). In this case, the noise signal may be generated based on ambient sound (or noise) in the surrounding environment in which the system is located. Specifically, the masking estimator may define the spectral content of the noise signal based on the magnitude of the spectral content contained in the microphone signal generated by microphone 55. For example, the estimator may apply one or more scalar gains (or vector gains) to the microphone signal such that the magnitude of one or more frequency bands of the signal exceeds (e.g., a predefined) threshold. In another aspect, the estimator may generate the noise signal based on the audio signal and / or ambient noise in the environment. Specifically, the estimator may generate a noise signal such that the noise generated by the system masks the sound of the user-desired audio content generated by the system (e.g., at a threshold distance from the system). A noise beamformer generates (or produces) one or more individual driver signals for one or more speaker drivers in order to “render” the audio content of the one or more noise signals as one or more noise (directional) beam patterns generated (or emitted) by the drivers.

[0096] In one aspect, the signal beamformer receives one or more audio signals (e.g., a second audio signal), which may include audio content desired by the user, such as speech (e.g., the sound of a telephone call), music, podcasts, movie soundtracks, etc., in any audio format (e.g., stereo format, 5.1 surround sound format, etc.). In one aspect, the audio signal may be received (or retrieved) from local memory (e.g., the memory of a controller). In another aspect, the audio signal may be received from a remote source (e.g., streaming from a separate electronic device (such as a server) via a computer network). The signal beamformer may perform operations similar to those of a noise beamformer, such as generating one or more individual driver signals to render the audio content into one or more desired audio (directional) beam patterns.

[0097] Each beamformer in the beamformer generates a driver signal for each speaker driver, wherein the driver signal for each speaker driver is summed by controller 51. The controller uses the summed driver signals to drive the speaker drivers to generate (e.g., primarily) a noise beam pattern 86 that includes noise from the noise signal and to generate (e.g., primarily) an audio beam pattern 87 that includes audio content from the audio signal. The figure also shows a top view (e.g., in the XY plane) of the system that generates beam patterns 86 and 87 pointing towards (or away from) several listeners 80-82. Specifically, the main lobe 88b of the audio beam pattern 87 is oriented towards the intended listener 80 (e.g., the user of the system), while the null point 89b of the pattern is away from the intended listener (e.g., and at least towards a third-party listener 82). Furthermore, the main lobe 88a of the noise beam pattern 86 is oriented towards third-party listeners 81 and 82 (and away from the intended listener 80), while the null point 89a of the pattern is oriented towards the intended listener. Therefore, listeners are expected to experience less (or no) noise in the noise beam pattern while experiencing the audio content contained within the audio beam pattern. Conversely, third-party listeners will experience only (or primarily) the noise in the noise beam pattern 86.

[0098] In one aspect, the beamformer can be configured to shape and manipulate its respective generated beam pattern based on the location of the intended listener 80 and / or the locations of one or more third-party listeners 81 and 82. Specifically, the system can determine whether a person is detected in the surrounding environment and, in response, determine the position of that person relative to a reference point (e.g., the system's location). For example, the system can make these determinations based on sensor data (e.g., image data), as described herein. Once the location of the intended listener is determined, the signal beamformer 84 can manipulate (e.g., by applying one or more vector weights to the audio signal to generate) the audio beam pattern 87 such that the audio beam pattern is directed toward the intended listener. Similarly, upon determining the location of one or more third-party listeners, the null beamformer 85 accordingly directs the noise beam pattern 86. In one aspect, when several third-party listeners are detected, the null beamformer 85 can direct the noise beam pattern such that an optimal amount of noise is directed toward all of these listeners. On the other hand, null beamformers can manipulate the noise pattern by taking into account the location of the intended listener (e.g., so that the null point is always facing the intended listener).

[0099] In one respect, beamformers 84 and 85 can execute any type of (e.g., adaptive) beamformer algorithm to generate the one or more driver signals. For example, any of the beamformers can perform phase-shift beamformer operation, minimum variance distortionless response (MVDR) beamformer operation, and / or linearly constrained minimum variance (LCMV) beamformer operation.

[0100] In one aspect, the beam patterns 86 and 87 generated by the system can create different regions or areas with different (or similar) signal-to-noise ratios (SNRs) within the surrounding environment. For example, the intended listener 80 may be located in an area with a first SNR, while third-party listeners 81 and 82 may be located in areas (or more areas) with a second SNR lower than the first SNR. Therefore, the intended listener can more easily understand the audio content expected by the user in the audio beam pattern 87 compared to a third-party listener who cannot hear the audio content due to the masking characteristics of noise. To illustrate, Figure 9 The graph 90 shows the signal strength of the audio content and noise relative to one or more areas around the system, based on several aspects.

[0101] Specifically, graph 90 shows the sound output level as the signal strength (e.g., in dB) of the noise beam pattern 86 and the audio beam pattern 87 relative to the angle around an axis extending through the system (e.g., the Z-axis). In one aspect, this axis can be the central Z-axis of the region (or part of the system) that includes the speaker drivers. For example, as... Figure 8 As shown, the central axis can be positioned between the first speaker driver and the second speaker driver.

[0102] As shown in graph 90, the beam pattern generated by the system (e.g., around the central Z-axis) creates several zones. Specifically, the graph shows three types of zones: a masking zone 91, a transition zone 92, and a target zone 93. In one aspect, each zone may have a different SNR. For example, masking zone 91 is the zone surrounding the system where the SNR is below (e.g., a first) threshold. In one aspect, this zone is a masking zone such that when positioned within it, the noise generated by the system masks the audio content the user expects, making it impossible for listeners within the zone to perceive (or understand) the audio content they expect. In some aspects, Figure 8 Third-party listeners 81 and 82 can be located within this cover area.

[0103] The target region 93 is the region surrounding the system where the SNR is higher than (e.g., a second) threshold. In one aspect, the second threshold may be greater than the first threshold. In another aspect, the two thresholds may be the same. In some aspects, this region is the target region such that when a listener is positioned within this region, the audio content of the audio beam pattern 87 is intelligible and not overwhelmed (or masked) by noise. In some aspects, it is expected that a listener 80 can be positioned within this region. The figure also shows a transition region 92, which separates the target region from the masking region 91 at any size of the target region. In one aspect, the transition region may have an SNR that transitions from the first threshold to the second threshold. Therefore, the SNR of this region may be between the two thresholds. In one aspect, the system may shape and manipulate the beam pattern to minimize the transition region 92.

[0104] As described above, the system (or more specifically, the output device 3 including the speaker drivers) can generate several beam patterns that can be oriented towards different locations within the surrounding environment to create different zones in order to provide privacy for the intended listener. In one aspect, the output device can be located anywhere within the surrounding environment. For example, the output device can be a standalone electronic device, such as a smart speaker. In another aspect, the output device can be a head-mounted device, such as a pair of smart glasses or a headset. In this case, when the output device is a head-mounted device, the zones can be optimized based on the location (and / or orientation) of one or more speaker drivers of the device to maximize the audio privacy of the intended listener. Figure 10 and Figure 11 An example of a beam pattern generated by an output device is shown, with the listener expected to be very close to the device's speaker driver.

[0105] For example, Figure 10 A top view of a radiation beam pattern 101 with a zero point 100 at the ear of the intended listener is shown, according to some aspects. By placing the zero point 100 close to the ear of the intended listener, while the beam pattern radiates outward and away from the intended listener, radiated sound (e.g., noise) is allowed to propagate within the environment without being heard by the intended listener (or at least the intended listener does not hear sounds above a sound output level threshold).

[0106] As shown in the figure, the output device is positioned close to the intended listener 80. For example, the output device may be within a threshold distance of the listener. Specifically, the output device may be within a threshold distance to the listener's ear (e.g., the right ear). Furthermore, one or more speaker drivers in the output device may be closer to the intended listener than one or more other speaker drivers. As shown, the first speaker driver 12 is closer to the listener's ear (e.g., the right ear) (e.g., within a threshold distance from that ear), while the second speaker driver 13 is further away from the right ear (e.g., beyond a threshold distance from the right ear). In other words, a particular part of the output device may be closer to the user's ear than other parts. For example, and as described herein, the wall (e.g., wall 17, such as the wall to which the first speaker driver 12 is attached (mounted or positioned) may be closer to the user's ear. Figure 2 (As shown) is another wall (e.g., wall 18) to which the second speaker driver 13 is connected. Figure 2 (As shown) closer to the user's ear. In one aspect, the speaker driver can be positioned accordingly when the output device is in use by the intended listener. Specifically, when the output device (e.g., a head-mounted device) is worn on the user's head, the first speaker driver can be closer to the user's ear than the second speaker driver.

[0107] In another aspect, besides (or instead of) being close to the intended listener, the speaker drivers can also be oriented such that they project sound toward the intended listener. Specifically, as shown, the first and second speaker drivers are arranged to project forward-radiating sound toward or in the direction of the user's ear. In one aspect, both (or all) speaker drivers of the output device can be arranged to project sound in the same direction. In another aspect, at least some of the speaker drivers can be arranged to project sound differently. For example, the second speaker driver can be oriented to project sound at an angle different from the angle at which the first speaker driver projects sound (e.g., around the central Z-axis).

[0108] As shown in the figure, the first speaker driver 12 and the second speaker driver 13 generate a directional beam pattern 101 that radiates away from the intended listener (e.g., and to all other locations in the surrounding environment), as indicated by the thinning of the beam pattern as it moves away from the output device. This beam pattern may include masking noise, as described herein. The beam pattern 101 includes a null point 100, which is the location in space where sound from the beam pattern 101 is absent (or very little, below a threshold). In one aspect, this null point may be generated based on the sound output of the first and second speaker drivers. For example, to generate a null point, the output device may drive the first speaker driver 12 with a first driver signal having a first signal level, while driving the second speaker driver 13 with a second driver signal having a second signal level higher than the first signal level. In one aspect, the first driver signal may be out of phase with respect to the second driver signal (e.g., at least partially). Thus, the first speaker driver 12 may generate sound that cancels out the masking noise generated by the second speaker driver 13, wherein the sound output level of the second driver is greater than the sound output level of the first speaker driver. The difference in sound output levels is illustrated by the following: only two curves are positioned in front of the first speaker driver, showing the sound output, while three lines radiate from the second speaker driver 13. Due to the reduced sound output from the first speaker driver that eliminates sound, the listener is expected to experience less masking noise.

[0109] In one aspect, in addition to masking noise, the radiation beam pattern 101 may include the audio content desired by the user. For example, controller 51 may receive audio signals and noise signals, as described herein. The controller may process the audio signals to generate a first driver signal for driving a first speaker driver and a second driver signal for driving a second driver signal. In one aspect, the first driver signal may include more spectral content of the user-desired audio content than the second driver signal. For example, the second driver signal may not include any spectral content of the user-desired audio content. In this case, when the signal is used to drive its respective speaker driver, the sound output of the first speaker driver cancels the masking noise generated by the second speaker driver and produces the sound of the user-desired audio content. In this case, the listener is expected to hear the user-desired audio content while the sound of the content is masked by the masking noise generated by the second speaker driver. In one aspect, this may occur in a “private” operating mode. In this mode, non-users will hear mostly the masking noise generated by the second speaker driver, which masks at least a portion of the user-desired audio content generated by the first speaker driver.

[0110] Figure 11Another radiation beam pattern 102, according to one aspect, is shown to direct sound to the intended listener's ear. In this example, the radiation beam pattern 102 maximizes the SNR at the listener's ear while minimizing the SNR beyond a threshold distance from the listener (e.g., the listener's ear). This is illustrated by the fact that the radiation beam pattern becomes thinner as it radiates away from the intended listener. In this example, two speaker drivers can produce the radiation beam pattern, wherein the two speaker drivers are driven using in-phase driver signals, as described herein. In one aspect, the two speaker drivers can output sound with the same (or different) sound output levels.

[0111] In one respect, the beam pattern described herein can be generated independently by the output device, such as... Figure 10 and Figure 11 As shown. In another aspect, multiple beam patterns can be generated. For example, the output device can generate both radiating beam patterns 101 and 102. In this case, beam pattern 101 can radiate masking noise, while beam pattern 102 includes the audio content desired by the user. Therefore, the sound of the audio content desired by the user can be directed to the user's ears while masking the sound so that it is not heard by others near the intended listener.

[0112] Another aspect of this disclosure is a method performed by a dual-speaker system (e.g., a programmable processor of the dual-speaker system) including a first speaker driver and a second speaker driver. The system receives an audio signal containing audio content (e.g., a musical piece) desired by a user. The system determines that the dual-speaker system will operate in one of a first (“non-private”) operating mode or a second (“private”) operating mode. The system processes the audio signal to generate a first driver signal driving the first speaker driver and a second driver signal driving the second speaker driver. In the first mode, the two signals are in phase with each other. However, in the second mode, the two signals are out of phase with each other. For example, the two signals may be 180° (or less) out of phase. In one aspect, the system uses the corresponding driver signals of different phases to drive the speaker drivers to generate a beam pattern with a main lobe in the direction of the user of the dual-speaker system. In some aspects, the generated beam pattern may have at least one null point away from the user of the output device. For example, the null point may be directed towards another person in the environment.

[0113] In one aspect, both speaker drivers are integrated within a housing, wherein determining whether a person is within a threshold distance of the housing, selecting a second operating mode in response to determining that a person is within the threshold distance, and selecting a first operating mode in response to determining that a person is not within the threshold distance. In another aspect, determining whether a person is within the threshold distance includes receiving image data from a camera and performing an image recognition algorithm on the image data to detect a person therein.

[0114] In some aspects, the system also receives microphone signals generated by microphones arranged to sense ambient sounds in the surrounding environment, uses the microphone signals to determine the noise level of the ambient sounds, and increases the sound output levels of the first speaker driver and the second speaker driver to exceed the noise level. In one aspect, the system determines, for each of several frequency bands of an audio signal, whether the magnitude of a corresponding frequency band of the ambient sound exceeds a threshold value for that frequency band, wherein the increase includes: applying a first gain to the audio signal to reduce the magnitude of the frequency band in response to the magnitude of the corresponding frequency band exceeding the threshold value; and applying a second gain to the audio signal to increase the magnitude of the frequency band in response to the magnitude of the corresponding frequency band not exceeding the threshold value.

[0115] In some aspects, both speaker drivers are integrated within the housing, wherein determining whether a person is within a threshold distance from the housing, selecting a second operating mode in response to determining that the person is within the threshold distance, and selecting a first operating mode in response to determining that the person is not within the threshold distance. In some aspects, determining whether a person is within the threshold distance includes receiving image data from a camera (e.g., which may be integrated within the housing or may be integrated into a separate device), and performing an image recognition algorithm on the image data to detect the person contained therein.

[0116] In one aspect, the method further includes: when in a second operating mode, using a first driver signal and a second driver signal respectively to drive a first speaker driver and a second speaker driver to output an audio signal with a beam pattern having a main lobe in the direction of the user of the system. In another aspect, the main lobe may be pointed in other directions (e.g., away from the user).

[0117] In some aspects, the method further includes: receiving a microphone signal generated by a microphone arranged to sense ambient sound in the surrounding environment; using the microphone signal to determine a noise level of the ambient sound; and increasing the sound output levels of a first speaker driver and a second speaker driver to exceed the noise level. In another aspect, the method further includes: determining, for each of several frequency bands of an audio signal, whether the magnitude of a corresponding frequency band of the ambient sound exceeds a threshold value for that frequency band, wherein increasing includes: applying a first gain (or attenuation) to the audio signal to reduce the magnitude of the frequency band in response to the magnitude of the corresponding frequency band exceeding the threshold value; and applying a second gain to the audio signal to increase the magnitude of the frequency band in response to the magnitude of the corresponding frequency band not exceeding the threshold value.

[0118] In another aspect, when in the second operating mode, at least a portion of the first driver signal is out of phase (at least) 180° with at least a portion of the second driver signal. In some aspects, the first speaker driver and the second speaker driver are integrated within the head-mounted device.

[0119] Personal information used should comply with practices and privacy policies that are generally recognized as meeting (and / or exceeding) government and / or industry requirements for protecting user privacy. For example, any information should be managed to mitigate the risk of unauthorized or unintentional access or use, and users should be clearly informed of the nature of any authorized use.

[0120] As previously described, one aspect of this disclosure may be a non-transitory machine-readable medium (such as microelectronic memory) storing instructions thereon that program one or more data processing units (generally referred to herein as a "processor") to perform network operations and audio signal processing operations, as described herein. In other aspects, some of these operations may be performed by specific hardware components containing hard-wired logic. Alternatively, those operations may be performed by any combination of programmed data processing units and fixed hard-wired circuit components.

[0121] While certain aspects have been described and illustrated in the accompanying drawings, it should be understood that such aspects are merely illustrative of the broad disclosure and not limiting, and that this disclosure is not limited to the specific structures and arrangements shown and described, as various other modifications will be apparent to those skilled in the art. Therefore, the description is to be regarded as exemplary and not restrictive.

[0122] In some aspects, this disclosure may include the language "[element A] and [element B] at least one". This language may refer to one or more of these elements. For example, "at least one of A and B" may refer to "A", "B", or "A and B". Specifically, "at least one of A and B" may refer to "at least one of A and at least one of B" or "at least either A or B". In some aspects, this disclosure may include the language "[element A], [element B], and / or [element C]". This language may refer to any of these elements or any combination thereof. For example, "A, B, and / or C" may refer to "A", "B", "C", "A and B", "A and C", "B and C", or "A, B, and C".

Claims

1. A wearable device, comprising: shell; A first speaker driver and a second speaker driver, both integrated within the housing and arranged to project sound into the surrounding environment, wherein the first speaker driver is positioned closer to the wall of the housing than the second speaker driver, and the first speaker driver and the second speaker driver share a common rear volume within the housing; The controller is configured to: Determine the operating mode of the wearable device; Based on the determined operating mode, an audio signal is used to generate at least one of a first driver signal and a second driver signal that are either in phase or out of phase with each other; as well as The first speaker driver and the second speaker driver are driven by the first driver signal and the second driver signal, respectively. as well as A slender tube having a first open end and a second open end, the first open end being connected to the common rear volume within the housing, and the second open end opening to the surrounding environment, wherein the slender tube is shaped to reduce the audibility of rearwardly radiated sound emitted from the second open end.

2. The wearable device of claim 1, wherein the housing forms an open housing outside the common rear volume and surrounds the front of the second speaker driver.

3. The wearable device of claim 2, wherein the open housing opens to the surrounding environment through a plurality of ports, and the second speaker driver projects forward-radiating sound into the surrounding environment through the plurality of ports.

4. The wearable device according to any one of claims 1 to 3, wherein the first speaker driver is a speaker driver of the same type as the second speaker driver.

5. The wearable device according to any one of claims 1 to 3, wherein the first speaker driver is a speaker of a different type from the second speaker driver.

6. The wearable device of claim 1, wherein the front of the first speaker driver faces a first direction, and the front of the second speaker driver faces a second direction different from the first direction.

7. The wearable device of claim 6, wherein the first direction and the second direction are opposite directions along the same axis.

8. The wearable device according to claim 1, wherein determining the operating mode of the wearable device includes: The audio signal is analyzed to determine whether the audio content of the audio signal is private.

9. The wearable device of claim 1, wherein the controller determines the operating mode by determining whether a person other than the user of the wearable device is within a threshold distance from the wearable device.

10. The wearable device according to claim 9, In response to determining that no one other than the user is within the threshold distance, the determined operating mode is a common mode, in which the first driver signal and the second driver signal are generated in phase with each other. In response to determining that someone other than the user is within the threshold distance, the determined operating mode is a private mode, in which the first driver signal and the second driver signal are generated at least partially out of phase with each other.

11. The wearable device of claim 1, wherein the wall of the housing is a first wall to which the first speaker driver is attached, wherein the housing has a second wall to which the second speaker driver is attached, wherein when the wearable device is worn by a user, the first wall is closer to the user's ear than the second wall.

12. The wearable device of claim 1, wherein the controller is configured to determine system operating parameters of the wearable device, wherein, Determining the operating mode includes: Based on the fact that the system operating parameters are less than a threshold, it is determined that the wearable device is operating in a common mode, in which the first drive signal and the second drive signal are in phase; and Based on the fact that the system operating parameters are greater than the threshold, it is determined that the wearable device is operating in private mode, in which the first drive signal and the second drive signal are out of phase.

13. The wearable device of claim 1, wherein the first speaker driver and the second speaker driver are external speaker drivers arranged to project sound in different directions, wherein, When the first speaker driver and the second speaker driver are out of phase, driving the speaker driver includes: generating a dipole sound pattern having a main lobe oriented toward a person in the surrounding environment other than the user of the wearable device.

14. The wearable device of claim 1, wherein the wearable device is a head-mounted device, such that when a user wears the head-mounted device, the first speaker driver is closer to the user's ear than the second speaker driver.

15. An output device, the output device comprising: An outer casing, the outer casing including an internal volume; A first external ear speaker driver and a second external ear speaker driver, both external ear speaker drivers are integrated within the housing and share the internal volume as the rear volume; A slender tube having a first open end and a second open end, the first open end being connected to the internal volume and the second open end being open to the surrounding environment, wherein the slender tube is shaped to reduce the audibility of rearward radiated sound emitted from the second open end. processor; and A memory having instructions stored therein, which, when executed by the processor, cause the output device to: Receive audio signals; Determine the current operating mode for the output device; Based on the audio signal, a first driver signal and a second driver signal are generated, wherein the current operating mode corresponds to whether at least a portion of the first driver signal and the second driver signal are generated as either in-phase or out-of-phase; and The first driver signal is used to drive the first external ear speaker driver; and The second driver signal is used to drive the second external ear speaker driver.

16. The output device of claim 15, wherein the instruction for determining the current operating mode includes an instruction for determining whether a person is within a threshold distance of the output device, wherein in response to determining that the person is within the threshold distance, the first driver signal and the second driver signal are generated to be at least partially out of phase with each other.

17. The output device of claim 16, wherein in response to determining that the person is not within the threshold distance, the first driver signal and the second driver signal are generated to be in phase with each other.

18. The output device of claim 15, wherein the memory further comprises instructions for driving the first external ear speaker driver and the second external ear speaker driver respectively using the first driver signal and the second driver signal, including instructions for generating a beam pattern having a main lobe in the direction of a user of the output device.

19. The output device of claim 18, wherein the generated beam pattern has at least one null point away from the user of the output device.

20. The output device of claim 15, further comprising receiving a microphone signal generated by a microphone of the output device, the microphone signal including ambient noise of the surrounding environment in which the output device is located, wherein the current operating mode is determined based on the ambient noise.

21. The output device of claim 20, wherein the instructions for determining the current operating mode of the output device include instructions for performing the following operations: Determine whether the ambient noise masks the audio signal across one or more frequency bands; In response to the ambient noise masking a first group of frequency bands in one or more frequency bands, a first operating mode is selected, in which portions of the first driver signal and the second driver signal are generated in phase across the first group of frequency bands; as well as In response to the ambient noise not masking a second set of frequency bands in the one or more frequency bands, a second operating mode is selected, in which portions of the first driver signal and the second driver signal are generated as out of phase across the second set of frequency bands.

22. The output device of claim 21, wherein the first set of frequency bands and the second set of frequency bands are non-overlapping frequency bands, such that the output device operates simultaneously in both the first operating mode and the second operating mode.

23. A head-mounted device, the head-mounted device comprising: A first external ear speaker driver and a second external ear speaker driver, wherein when the head-mounted device is worn on a user's head, the first external ear speaker driver is closer to the user's ear than the second external ear speaker driver; A slender tube having a first open end and a second open end, the first open end being connected to an internal space shared by a first external ear speaker driver and a second external ear speaker driver, the second open end being open to the surrounding environment, wherein the slender tube is shaped to reduce the audibility of rearward radiated sound emitted from the second open end; processor; and A memory having instructions stored therein, which, when executed by the processor, cause the device to: Receives audio signals including noise; The first external ear speaker driver and the second external ear speaker driver are used to generate a directional beam pattern, the directional beam pattern including: 1) having the noise and being away from the user's main lobe, and 2) being oriented toward the user's null point, wherein the sound output level of the second external ear speaker driver is greater than the sound output level of the first external ear speaker driver.

24. The head-mounted device of claim 23, wherein the audio signal is a first audio signal and the directional beam pattern is a first directional beam pattern, wherein the memory further comprises instructions for performing the following operations: Receive a second audio signal that includes the audio content desired by the user; The first external ear speaker driver and the second external ear speaker driver are used to generate a second directional beam pattern, the second directional beam pattern including: 1) It has the audio content desired by the user and is oriented toward the user's main lobe, and 2) It is away from the user's zero point.

25. The head-mounted device of claim 23, wherein the first external ear speaker driver and the second external ear speaker driver project forward-radiating sound toward or in the direction of the user's ear.

26. The head-mounted device of claim 23, wherein the audio signal is a first audio signal, and wherein the memory further comprises instructions for performing the following operations: Receive a second audio signal containing the audio content desired by the user; The first audio signal and the second audio signal are processed to generate a first driver signal and a second driver signal, which generate the directional beam pattern when used to drive the first external ear speaker driver and the second external ear speaker driver, respectively.

27. The head-mounted device of claim 26, wherein the first driver signal includes more spectral content of the user-desired audio content than the second driver signal.