Apparatus, system, and method for balancing voice audio captured by eyewear devices

The eyewear device balances audio volume levels using transducers and circuitry to isolate and modify audio signals, addressing volume imbalances and improving listening experiences in artificial reality systems.

WO2025244960A1PCT designated stage Publication Date: 2025-11-27META PLATFORMS TECHNOLOGIES LLC
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
PCT/US2025/029921
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-05-18
Filing Date
2025-05-18
Publication Date
2025-11-27

AI Technical Summary

Technical Problem

Existing eyewear devices, such as head-mounted displays, struggle with imbalances in audio volume levels between a user's voice and other audio sources, leading to suboptimal listening experiences when capturing and playing back audio data.

Method used

The apparatus and method employ multiple audio transducers and circuitry to identify, isolate, and modify audio signals to balance volume levels, using beamformers and signal modifiers to achieve a desired signal-to-noise ratio, ensuring balanced playback of user and additional audio signals.

Benefits of technology

This approach results in a more optimal and pleasant listening experience by equalizing volume levels, providing a balanced audio presentation that enhances user interaction and immersion in artificial reality environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure US2025029921_27112025_PF_FP_ABST
    Figure US2025029921_27112025_PF_FP_ABST
Patent Text Reader

Abstract

An apparatus for balancing voice audio captured by eyewear devices comprising (1) an eyewear frame (810) dimensioned to be worn by a user, (2) a plurality of audio transducers (820) secured to the frame and configured to capture audio data representative of an environment occupied by the user, and (3) circuitry configured to (A) identify, within the audio data, an audio signal corresponding to a voice of the user and at least one additional audio signal, (B) isolate the audio signal from the at least one additional audio signal, and / or (C) modify the audio signal or the at least one additional audio signal to compensate for an imbalance between a volume level of the voice of the user and a volume level of the at least one additional audio signal. Various other apparatuses, systems, and methods are also disclosed.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] APPARATUS, SYSTEM, AND METHOD FOR BALANCING VOICE AUDIO CAPTURED BY EYEWEAR DEVICES

[0002] CROSS-REFERENCE TO RELATED APPLICATIONS

[0003] This application claims the benefit of U.S. Provisional Application No. 63 / 649,336 filed May 18, 2024.

[0004] SUMMARY

[0005] According to an aspect of the present invention, there is provided an apparatus comprising: a frame dimensioned to be worn by a user; a plurality of audio transducers secured to the frame and configured to capture audio data representative of an environment occupied by the user; and circuitry configured to: identify, within the audio data, an audio signal corresponding to a voice of the user and at least one additional audio signal; isolate the audio signal from the at least one additional audio signal; and modify the audio signal or the at least one additional audio signal to compensate for an imbalance between a volume level of the voice of the user and a volume level of the at least one additional audio signal.

[0006] Optionally, the circuitry is further configured to generate modified audio data for playback of the voice of the user and the additional audio signal by combining the audio signal and the additional audio signal after completion of the modification.

[0007] Optionally, the circuitry is further configured to transmit the modified audio data to a playback device that plays the audio data as an audio presentation.

[0008] Optionally, the circuitry is further configured to: generate a beamformer for isolating the audio signal based at least in part on an arrangement of the audio transducers secured to the frame; and apply the beamformer to the audio data to: identify the audio signal and the at least one additional audio signal; and isolate the audio signal and the at least one additional audio signal from one another.

[0009] Optionally, the audio transducers comprise: a first microphone configured to capture a first instance of the audio data for a first input channel; and a second microphone configured to capture a second instance of the audio data for a second input channel.

[0010] Optionally, the circuitry is further configured to: generate a first beamformer for the first input channel based at least in part on a position of the first microphone relative to the user whose voice captured in the audio signal; generate a second beamformer for the second input channel based at least in part on a position of the second microphone relative to the user whose voice captured in the audio signal; apply the first beamformer to the first instance of the audio data to isolate the audio signal and the at least one additional audio signal from one another for the first input channel; and apply the second beamformer to the second instance of the audio data to isolate the audio signal and the at least one additional audio signal from one another for the second input channel.

[0011] Optionally, the circuitry is further configured to: modify the audio signal or the at least one additional audio signal in the first instance of the audio data to compensate for a first imbalance between the volume level of the voice of the user and the volume level of the at least one additional audio signal in the first input channel; and modify the audio signal or the at least one additional audio signal in the second instance of the audio data to compensate for a second imbalance between the volume level of the voice of the user and the volume level of the at least one additional audio signal in the second input channel.

[0012] Optionally, the circuitry is further configured to generate modified audio data for playback of the voice of the user and the additional audio signal by combining the first instance of the audio data and the second instance of the audio data after completion of the modifications. Optionally, the circuitry is further configured to attenuate the audio signal to a certain decibel level based at least in part on the volume level of the at least one additional audio signal.

[0013] Optionally, the circuitry is further configured to amplify the at least one additional audio signal to a certain decibel level based at least in part on the volume level of the audio signal. Optionally, the circuitry is further configured to balance, based at least in part on the modification, the audio signal and the at least one additional audio signal to a signal-to-noise ratio with a standard deviation between 5 decibels and 7 decibels.

[0014] Optionally, the signal-to-noise ratio of the audio signal and the at least one additional audio signal is approximately 8.5 decibels.

[0015] Optionally, the circuitry is further configured to: detect the imbalance by comparing the volume level of the audio signal and the volume level of the at least one additional audio signal; determine, based at least in part on the comparison, that the imbalance satisfies a certain threshold; and modify, in response to the determination, the audio signal or the at least one additional audio signal to compensate for the imbalance.

[0016] According to a further aspect of the present invention, there is provided a system comprising: an eyewear device that: is dimensioned to be worn by a user; comprises a plurality of audio transducers configured to capture audio data representative of an environment occupied by the user; and circuitry configured to: isolate, within the audio data, an audio signal corresponding to a voice of the user and at least one additional audio signal; and modify the audio signal or the at least one additional audio signal to compensate for an imbalance between a volume level of the voice of the user and a volume level of the at least one additional audio signal; and a playback device configured to play the audio data as an audio presentation after completion of the modification.

[0017] Optionally, the circuitry is further configured to modify the audio data for playback of the voice of the user and the additional audio signal by combining the audio signal and the additional audio signal.

[0018] Optionally, the circuitry is further configured to transmit the audio data to the playback device via a network.

[0019] Optionally, the circuitry is further configured to: generate a beamformer for isolating the audio signal based at least in part on an arrangement of the audio transducers secured to the frame; and apply the beamformer to the audio data to: identify the audio signal and the at least one additional audio signal; and isolate the audio signal and the at least one additional audio signal from one another.

[0020] Optionally, the audio transducers comprise: a first microphone configured to capture a first instance of the audio data for a first input channel; and a second microphone configured to capture a second instance of the audio data for a first input channel.

[0021] Optionally, the circuitry is further configured to: generate a first beamformer for the first input channel based at least in part on a position of the first microphone relative to the user whose voice captured in the audio signal; generate a second beamformer for the second input channel based at least in part on a position of the first microphone relative to the user whose voice captured in the audio signal; apply the first beamformer to the first instance of the audio data to isolate the audio signal and the at least one additional audio signal from one another for the first input channel; apply the second beamformer to the second instance of the audio data to isolate the audio signal and the at least one additional audio signal from one another for the second input channel; modify the audio signal or the at least one additional audio signal in the first instance of the audio data to compensate for a first imbalance between the volume level of the voice of the user and the volume level of the at least one additional audio signal in the first input channel; modify the audio signal or the at least one additional audio signal in the second instance of the audio data to compensate for a second imbalance between the volume level of the voice of the user and the volume level of the at least one additional audio signal in the second input channel; and generate modified audio data for playback of the voice of the user and the additional audio signal by combining the first instance of the audio data and the second instance of the audio data after completion of the modifications.

[0022] According to a further aspect of the present invention, there is provided a method comprising: capturing, by a plurality of audio transducers, audio data representative of an environment occupied by a user; identifying, by circuitry, an audio signal corresponding to a voice of the user and at least one additional audio signal within the audio data; isolating, by the circuitry, the audio signal from the at least one additional audio signal; and modifying, by the circuitry, the audio signal or the at least one additional audio signal to compensate for an imbalance between a volume level of the voice of the user and a volume level of the at least one additional audio signal.

[0023] BRIEF DESCRIPTION OF DRAWINGS

[0024] The accompanying drawings illustrate a number of example embodiments and are a part of the specification. Together with the following description, these drawings demonstrate and explain various principles of the instant disclosure.

[0025] FIG. 1 is an illustration of an exemplary apparatus for balancing voice audio captured by eyewear devices according to one or more implementations of this disclosure.

[0026] FIG. 2 is an illustration of an exemplary eyewear device for balancing voice audio captured by eyewear devices according to one or more implementations of this disclosure.

[0027] FIG. 3 is an illustration of an exemplary implementation of a system for balancing voice audio captured by eyewear devices according to one or more implementations of this disclosure.

[0028] FIG. 4 is an illustration of an exemplary apparatus for balancing voice audio captured by eyewear devices according to one or more implementations of this disclosure.

[0029] FIG. 5 is an illustration of exemplary circuitry for balancing voice audio captured by eyewear devices according to one or more implementations of this disclosure.

[0030] FIG. 6 is an illustration of an exemplary system for balancing voice audio captured by eyewear devices according to one or more implementations of this disclosure.

[0031] FIG. 7 is a flow diagram of an exemplary method for balancing voice audio captured by eyewear devices according to one or more implementations of this disclosure.

[0032] FIG. 8 is an illustration of exemplary augmented-reality glasses that may be used in connection with one or more implementations of this disclosure.

[0033] FIG. 9 is an illustration of an exemplary virtual-reality headset that may be used in connection with one or more implementations of this disclosure.

[0034] While the exemplary embodiments described herein are susceptible to various modifications and alternative forms, specific embodiments have been shown by way of example in the appendices and will be described in detail herein. However, the exemplary embodiments described herein are not intended to be limited to the particular forms disclosed. Rather, the instant disclosure covers all modifications, combinations, equivalents, and alternatives falling within this disclosure.

[0035] DETAILED DESCRIPTION OF EXEMPLARY EMBODIMENTS

[0036] The present disclosure is generally directed to apparatuses, systems, and methods for balancing voice audio captured by eyewear devices. As will be explained in greater detail below, these apparatuses, systems, and methods may provide numerous features and benefits.

[0037] In some examples, eyewear devices like head-mounted displays (HMDs) have revolutionized the way people experience various kinds of digital media. Forexample, HMDs may allow users of artificial reality to experience realistic, immersive virtual and / or augmented environments. Artificial reality may provide users with opportunities to interact with virtual objects and / or environments in one way or another. In this context, artificial reality may constitute a form of reality that has been altered by virtual objects for presentation to a user. Such artificial reality may include and / or represent virtual reality (VR), augmented reality (AR), mixed reality, hybrid reality, or some combination and / or variation of one or more of the same.

[0038] Although artificial-reality systems are commonly implemented for gaming and other entertainment purposes, such systems are also implemented for purposes outside of recreation. For example, governments may use them for military training simulations, pilots may use them for flight simulations, doctors may use them to practice surgery, engineers may use them as visualization aids, and co-workers may use them to facilitate inter-personal interactions and collaboration from across the globe.

[0039] In some examples, HMDs may be equipped with input audio transducers, such as sensors and / or microphones (e.g., condenser microphones, dynamic microphones, ribbon microphones, etc.). In one example, a user may wear an HMD equipped with multiple input audio transducers. In this example, the input audio transducers of the HMD may capture, record, and / or collect audio data representative of the environment occupied by the user. For example, the input audio transducers may capture, record, and / or collect audio data representative of the user's voice and at least one additional person and / or feature (e.g., ambient noise, breathing sounds, dogs barking, sporting activities, etc.) audible in the environment. In this example, the data representative of the user's voice and at least one additional person and / or feature may be modified, manipulated, and / or compensated to balance the user's voice relative to the at least one additional person and / or feature.

[0040] As a specific example, the input audio transducers of the HMD may capture audio data representative of a conversation between the user wearing the HMD and another person present in the environment. In one example, the raw audio data may represent the user's voice at a substantially louder volume and / or level than the other person's voice. As a result, when played back by and / or streamed to a remote device, the raw audio data may provide the listener with a suboptimal, nonideal, and / or somewhat unpleasant listening experience. In other words, unless the raw audio data is modified to compensate for and / or mitigate the imbalance between the user's voice and the other person's voice, the listener may be unable to enjoy the listening experience as much as would otherwise be possible.

[0041] In some examples, the raw audio data may be modified and / or compensated in a variety of different ways and / or contexts. In one example, the HMD may include and / or represent circuitry that modifies the audio data to balance the user's voice with the other person's voice. In another example, the HMD may deliver, transmit, and / or provide the audio data to a remote device that modifies the audio data to balance the user's voice with the other person's voice. Either way, the modified audio data may represent the user's voice and the other person's voice at similar and / or more balanced volumes or levels. As a result, when played back by and / or streamed to output audio transducers (e.g., speakers), the modified audio data may provide the listener with a more optimal, ideal, and / or pleasant listening experience.

[0042] In some examples, various algorithms may be used, applied, and / or implemented to detect and / or correct imbalances between the user's voice and the at one additional person and / or feature in the audio data. In one example, the HMD and / or the remote device may perform one or more single-channel beamformer estimations. Additionally or alternatively, the HMD and / or the remote device may compute a separate beamformer filter for each channel of audio captured by the input audio transducers. For example, the HMD and / or the remote device may generate, apply, and / or implement one or more personalized and / or bespoke filters for the user and / or the at least one additional person and / or feature to account for their respective characteristics and / or attributes based at least in part on the beamformer estimation(s). By doing so, the HMD and / or the remote device may be able to separate and / or isolate the audio signals captured from the different audio sources (e.g., the user, the other person, etc.) that are audible in the environment.

[0043] In some examples, the HMD and / or the remote device may assess the level of imbalance between the volumes of the different audio sources based at least in part on those separate audio signals. Additionally or alternatively, the HMD and / or remote device may modify the audio signals to equalize, mitigate, and / or compensate for the imbalance assessed between the volumes of the different audio sources.

[0044] In some examples, the HMD and / or the remote device may be able to modify and / or change the volume (e.g., amplitude and / or decibel levels) of the individual audio signals separated and / or extracted from one another after they have been captured in the audio data. For example, the HMD and / or the remote device may modify the audio signals corresponding to the user's voice and the other person's voice to equalize, match, and / or balance the relative volumes of the user's voice and the other person's voice. In one example, the HMD and / or the remote device may modify the audio signal corresponding to the user's voice to decrease and / or reduce the volume of the user's voice (e.g., to coincide with and / or match the lower volume of the other person's voice). In this example, upon modifying and / or changing the volume in this way, the HMD and / or the remote device may recombine the individual audio signals to form the modified audio data. As a result, when played back and / or streamed to output audio transducers, the modified audio data may provide and / or represent a more balanced, realistic, and / or pleasant audio presentation of the conversation between the user and the other person.

[0045] The following will provide, with reference to FIGS. 1-6, detailed descriptions of exemplary apparatuses, devices, systems, components, and corresponding configurations or implementations for balancing voice audio captured by eyewear devices. In addition, detailed descriptions of methods for balancing voice audio captured by eyewear devices will be provided in connection with FIG. 7. The discussion corresponding to FIGS. 8 and 9 will provide detailed descriptions of types of exemplary artificial-reality devices, wearables, and / or associated systems capable of balancing voice audio captured by eyewear devices. FIG. 1 illustrates an exemplary apparatus 100 for balancing voice audio captured by eyewear devices. As illustrated in FIG. 1, apparatus 100 may include and / or represent an eyewear frame 102 dimensioned to be worn by a user. In some examples, eyewear frame 102 may include and / or be equipped with one or more transducers 104( l)-(N) and / or circuitry 106. In one example, transducers 104(l)-(N) and / or circuitry 106 may be physically coupled and / or secured to eyewear frame 102. In certain implementations, transducers 104(l)-(N) and circuitry 106 may be electrically and / or communicatively coupled to one another.

[0046] In some examples, transducers 104(l)-(N) may capture, generate, and / or produce audio data 108 representative of an environment occupied by the user. In one example, transducers 104(l)-(N) may transform and / or convert voices, sounds, or noises present in the environment to audio data 108. In this example, transducers 104(l)-(N) may pass, deliver, and / or transmit audio data 108 to circuitry 106.

[0047] In some examples, circuitry 106 may receive, obtain, and / or receive audio data 108 from transducers 104(l)-(N). In one example, circuitry 106 may identify, detect, and / or discern an audio signal 114 corresponding to the user's voice. In this example, circuitry 106 may isolate, extract, and / or separate audio signal 114 from an audio signal 116 included in and / or represented by audio data 108.

[0048] Circuitry 106 may identify and / or isolate audio signal 114 in a variety of different ways. For example, circuitry 106 may generate, implement, and / or provide a beamformer 112 that facilitates and / or supports isolating audio signal 114 from audio signal 116. In one example, circuitry 106 may generate and / or implement beamformer 112 based at least in part on an arrangement of audio transducers secured to eyewear frame 102. In this example, circuitry 106 may rely on and / or leverage knowledge and / or information about the positions and / or locations of transducers 104(l)-(N) relative to one another and / or relative to the user wearing eyewear frame 102.

[0049] As a specific example, circuitry 106 may calibrate, customize, adapt, and / or personalize beamformer 112 to the user. Beamformer 112 may adapt to the characteristics of the user's voice to facilitate and / or support distortionless separation of the user's voice from other sounds and / or noises represented in audio data 108. In one example, circuitry 106 may obtain and / or access knowledge and / or information about the geometry of transducers 104(l)-(N) as arranged on eyewear frame 102. In this example, circuity 106 may also obtain and / or access one or more samples of the user's voice captured in an otherwise quiet environment (e.g., an empty and / or quiet room). Additionally or alternatively, circuitry 106 may generate and / or implement beamformer 112 based at least in part on the geometry of transducers 104(l)-(N) and / or such samples of the user's voice.

[0050] In some examples, circuitry 106 may apply beamformer 112 to audio data 108. In one example, by applying beamformer 112 to audio data 108, circuitry 106 may identify and / or detect audio signal 114 and / or audio signal 116. Additionally or alternatively, by applying beamformer 112 to audio data 108, circuitry 106 may isolate and / or separate audio signal 114 and audio signal 116 from one another.

[0051] In some examples, circuitry 106 may modify, manipulate, and / or alter audio signal 114 and / or audio signal 116 to compensate for an imbalance between the volume level of the user's voice and the volume level of the audio signal 116. For example, circuitry 106 may amplify audio signal 116 to increase its volume level (e.g., amplitude and / or decibel level) upon playback by a playback device. Additionally or alternatively, circuitry 106 may attenuate audio signal 114 to decrease its volume level (e.g., amplitude and / or decibel level) upon playback by the playback device.

[0052] As another example, circuitry 106 may amplify audio signal 114 to increase its volume level upon playback by the playback device. Additionally or alternatively, circuitry 106 may attenuate audio signal 116 to decrease its volume level upon playback by the playback device. In certain implementations, circuitry 106 may modify, manipulate, and / or alter audio signal 114 and / or audio signal 116 by applying one or more audio filters to audio signal 114 and / or audio signal 116. Examples of such audio filters include, without limitation, low-pass filters, high-pass filters, band-pass filters, shelf filters, notch filters, formant filters, comb filters, combinations or variations of one or more of the same, and / or any other suitable audio filters. Circuitry 106 may identify and / or isolate audio signal 114 in a variety of different ways. For example, circuitry 106 may apply a signal modifier 118 to audio signal 114 and / or audio signal 116. In one example, signal modifier 118 may include and / or represent an amplifier that amplifies audio signal 114 or audio signal 116 to a certain amplitude and / or decibel level based at least in part on the other's volume level, amplitude, and / or decibel level. Additionally or alternatively, signal modifier 118 may include and / or represent an attenuator that attenuates audio signal 114 or audio signal 116 to a certain amplitude and / or decibel level based at least in part on the other's volume level, amplitude, and / or decibel level.

[0053] In some examples, circuitry 106 my detect, identify, determine, and / or characterize the imbalance by com aring the volume levels, amplitudes, and / or decibel levels of audio signals 114 and 116 with one another. In one example, circuitry 106 may determine, based at least in part on the comparison, that the imbalance satisfies and / or exceeds a certain threshold. For example, circuitry 106 may be configured and / or programmed to modify one or more of audio signals 114 and 116 only if their respective volume levels, amplitudes, and / or decibel levels deviate from one another by more than 6 decibels. In other words, circuitry 106 may be configured and / or programmed to forego modifying any of audio signals 114 and 116 when their respective volume levels, amplitudes, and / or decibel levels fail to deviate from one another by more than 6 decibels. In this example, circuitry 106 may determine that the imbalance between audio signals 114 and 116 satisfies and / or exceeds the 6 decibels of deviation.

[0054] In some examples, circuitry 106 may substantially balance and / or equalize audio signals 114 and 116 based at least in part on the modification(s) made to one or more of the same. For example, circuitry 106 may balance and / or equalize audio signals 114 and 116 to a signal-to- noise ratio (SNR) with a standard deviation between 5 decibels and 7 decibels (e.g., 6 decibels). In one example, the SNR of audio signal 114 and 116 may be measured at approximately 8.5 or 8.6 decibels. In certain implementations, circuitry 106 may modify, manipulate, and / or alter one or more of audio signals 114 and 116 to ensure that the combination and / or integration of audio signals 114 and 116 results in an SNR between 8 and 9 decibels (e.g., 8.5 or 8.6 decibels).

[0055] In some examples, circuitry 106 may generate, create, and / or produce modified audio data 120 for playback of audio signals 114 and 116. In one example, circuitry 106 may generate modified audio data 120 by combining and / or integrating audio signals 114 and 116 after completion of the modification(s). For example, signal modifier 118 may combine and / or integrate audio signals 114 and 116 with one another and then output modified audio data 120. In this example, circuitry 106 may subsequently cause one or more output audio transducers (e.g., audio speakers) to play modified audio data 120 as an audio presentation. Additionally or alternatively, circuitry 106 may send, transmit, and / or transfer modified audio data 120 to a remote playback device (e.g., via a network) for playback.

[0056] In some examples, transducers 104(l)-(N) may include and / or represent input and / or output devices implemented and / or incorporated in eyewear frame 102. For example, transducers 104(l)-(N) may include and / or represent one or more audio speakers and / or microphones. Examples of transducers 104(l)-(N) include, without limitation, voice coil speakers, ribbon speakers, electrostatic speakers, piezoelectric speakers, bone conduction transducers, cartilage conduction transducers, tragus-vibration transducers, tissue transducers, condenser microphones, dynamic microphones, ribbon microphones, combinations or variations of one or more of the same, and / or any other suitable transducers.

[0057] In some examples, circuitry 106 may include and / or represent one or more electrical and / or electronic circuits capable of processing, applying, modifying, transforming, displaying, transmitting, receiving, and / or executing data for apparatus 100. In one example, circuitry 106 may process and / or analyze audio data 108 received from transducers 104(l)-(4). Additionally or alternatively, circuitry 106 may implement, apply, and / or modify certain audio or visual features presented to the user wearing eyewear frame 102. In certain implementations, circuitry 106 may provide this audio or visual content for presentation on a display device and / or transducers 104(l)-(N) such that the audio or visual content is sensed, consumed, and / or experienced by the user.

[0058] In some examples, circuitry 106 may launch, perform, and / or execute certain executable files, code snippets, and / or computer-readable instructions to facilitate and / or support balancing voice audio captured by eyewear devices. Although illustrated as a single unit in FIG. 1, circuitry 106 may include and / or represent a collection of multiple processing units and / or electrical or electronic components that work and / or operate in conjunction with one another. In one example, circuitry 106 may include and / or represent an application-specific integrated circuit (ASIC). In another example, circuitry 106 may include and / or represent a central processing unit (CPU).

[0059] Examples of circuitry 106 include, without limitation, processing devices, microprocessors, microcontrollers, graphics processing units (GPUs), field-programmable gate arrays (FPGAs), systems on chips (SoCs), parallel accelerated processors, tensor cores, integrated circuits, chiplets, optical modules, receivers, transmitters, transceivers, optical modules, memory devices, transistors, antennas, resistors, capacitors, diodes, inductors, switches, registers, flipflops, digital logic, connections, traces, buses, semiconductor (e.g., silicon) devices and / or structures, storage devices, audio controllers, portions of one or more of the same, variations or combinations of one or more of the same, and / or any other suitable circuitry. In certain implementations, circuitry 106 may execute and / or implement software and / or firmware that performs one or more of the steps and / or features described herein for balancing audio signals 114 and 116.

[0060] In some examples, eyewear frame 102 may include and / or represent any type or form of structure and / or assembly capable of securing and / or mounting transducers 104(l)-(N), circuitry 106, and / or sensors 108(l)-(N) to the user's head or face. In one example, eyewear frame 102 may be sized, dimensioned, and / or shaped in any suitable way to facilitate securing and / or mounting an artificial-reality device to the user's head or face. In one example, eyewear frame 102 may include and / or contain a variety of different materials. Additional examples of such materials include, without limitation, plastics, acrylics, polyesters, metals (e.g., aluminum, magnesium, etc.), nylons, conductive materials, rubbers, neoprene, carbon fibers, composites, combinations or variations of one or more of the same, and / or any other suitable materials.

[0061] In some examples, apparatus 100 may include and / or represent a head-mounted display (HMD). In one example, the term "head-mounted display" and / or the abbreviation "HMD" may refer to any type or form of display device or system that is worn on or about a user's face and displays virtual content, such as computer-generated objects and / or AR content, to the user. HMDs may present and / or display content in any suitable way, including via a display screen, a liquid crystal display (LCD), a light-emitting diode (LED), a microLED display, a plasma display, a projector, a cathode ray tube, an optical mixer, combinations or variations of one or more of the same, and / or any other suitable HMDs. HMDs may present and / or display content in one or more media formats. For example, HMDs may display videos, photos, computer-generated imagery (CGI), and / or variations or combinations of one or more of the same. Additionally or alternatively, HMDs may include and / or incorporate see-through lenses that enable the user to see the user's surroundings in addition to such computer-generated content.

[0062] HMDs may provide diverse and distinctive user experiences. Some HMDs may provide virtual reality experiences (i.e., they may display computer-generated or pre-recorded content), while other HMDs may provide real-world experiences (i.e., they may display live imagery from the physical world). HMDs may also provide any mixture of live and virtual content. For example, virtual content may be projected onto the physical world (e.g., via optical or video see-through lenses), which may result in AR and / or mixed reality experiences.

[0063] FIG. 2 illustrates an exemplary eyewear device 200 for balancing voice audio. In some examples, eyewear device 200 may include and / or represent certain devices, components, and / or features that perform and / or provide functionalities that are similar and / or identical to those described above in connection with FIG. 1. As illustrated in FIG. 2, eyewear device 200 may include and / or represent eyewear frame 102 dimensioned to be worn by a user. In one example, eyewear frame 102 may include and / or represent a front frame 202, temples 204(1) and 204(2), optical elements 206(1) and 206(2), endpieces 208(1) and 208(2), nose pads 210, and / or a bridge 212. Additionally or alternatively, eyewear frame 102 may include, implement, and / or incorporate transducers 104(l)-(N) and / or circuitry 106— some of which are not necessarily illustrated, visible, and / or labelled in FIG. 2.

[0064] In some examples, optical elements 206(1) and 206(2) may be inserted and / or installed in front frame 202. In other words, optical elements 206(1) and 206(2) may be coupled to, incorporated in, and / or held by eyewear frame 102. In one example, optical elements 206(1) and 206(2) may be configured and / or arranged to provide one or more virtual visual features for presentation to a user wearing eyewear device 200. These virtual visual features may be driven, influenced, and / or controlled by one or more wireless technologies supported by eyewear device 200.

[0065] In some examples, optical elements 206(1) and 206(2) may each include and / or represent optical stacks, lenses, and / or films. In one example, optical elements 206(1) and 206(2) may each include and / or represent various layers that facilitate and / or support the presentation of virtual features and / or elements that overlay real-world features and / or elements. Additionally or alternatively, optical elements 206(1) and 206(2) may each include and / or represent one or more screens, lenses, and / or fully or partially see-through components. Examples of optical elements 206(1) and 206(2) include, without limitation, electrochromic layers, dimming stacks, transparent conductive layers (such as indium tin oxide films), metal meshes, antennas, transparent resin layers, lenses, films, combinations or variations of one or more of the same, and / or any other suitable optical elements.

[0066] FIG. 3 illustrates an exemplary implementation 300 of a system for balancing voice audio captured by eyewear devices. In some examples, implementation 300 may include, involve, and / or represent certain devices, components, and / or features that perform and / or provide functionalities that are similar and / or identical to those described above in connection with either FIG. 1 or FIG. 2. As illustrated in FIG. 3, a user 302 may wear and / or don eyewear device 200. In one example, implementation 300 may involve eyewear device 200 capture audio data representative of a voice 306 of user 302 and / or a sound 308 originating from a sound source 304. In certain implementations, sound 308 may include and / or represent the voice of another person present in the same room and / or environment as user 302.

[0067] In some examples, user 302 may include and / or represent the near-field source of sound captured by eyewear device 200. In such examples, sound source 304 may include and / or represent the far-field source of sound 308 captured by eyewear device 200. In one example, eyewear device 200 may identify and / or distinguish the audio signal representative of voice 306 within the audio data (e.g., via an adaptive personalized beamformer). Additionally or alternatively, eyewear device 200 may isolate and / or separate the audio signal of voice 306 from the audio signal of sound 308 (e.g., via the adaptive personalized beamformer).

[0068] In some examples, eyewear device 200 may compare and / or characterize the amplitudes, decibel levels, and / or volume levels of the audio signals corresponding to voice 306 and / or sound 308. In one example, eyewear device 200 may detect and / or characterize an imbalance between the audio signals corresponding to voice 306 and / or sound 308. In this example, eyewear device 200 may mitigate and / or compensate for the imbalance by attenuating and / or amplifying one or more of the audio signals corresponding to voice 306 and / or sound 308.

[0069] In some examples, eyewear device 200 may balance and / or equalize those audio signals to an SNR with a standard deviation between 5 decibels and 7 decibels (e.g., 6 decibels). In one example, the SNR of those audio signals may be approximately 8.5 or 8.6 decibels. Accordingly, eyewear device 200 may modify, manipulate, and / or alter one or more of those audio signals to ensure that the subsequent combination of audio signals 114 and 116 results in an SNR between 8 and 9 decibels (e.g., 8.5 or 8.6 decibels). In this example, eyewear device 200 may play back the resulting combination of those audio signals for user 302 to hear and / or may send, transfer, and / or transmit the resulting combination of those audio signals to a remote device for playback.

[0070] FIG. 4 illustrates an exemplary apparatus 400 for balancing voice audio captured by eyewear devices. In some examples, apparatus 400 may include, involve, and / or represent certain devices, components, and / or features that perform and / or provide functionalities that are similar and / or identical to those described above in connection with any of FIGS. 1-3. As illustrated in FIG. 4, apparatus 400 may include and / or represent microphones 404(l)-(N) and / or circuitry 106 secured and / or coupled to eyewear frame 102. In one example, microphones 404(l)-(N) may capture, generate, and / or produce audio data 408(l)-(N) representative of the environment occupied by the user from different perspectives.

[0071] In some examples, microphones 404(l)-(N) may capture and / or record audio data 408(l)-(N), respectively. For example, microphone 404(1) may capture and / or record audio data 408(1) for a first input channel corresponding to the perspective of microphone 404(1). In this example, microphone 404(N) may capture and / or record audio data 408(N) for a first input channel corresponding to the perspective of microphone 404(N).

[0072] In some examples, circuitry 106 may receive, obtain, and / or receive audio data 108 from microphones 404(l)-(N). In one example, circuitry 106 may identify, detect, and / or discern an audio signal 414(1) corresponding to the user's voice within audio data 408(1). In this example, circuitry 106 may isolate, extract, and / or separate audio signal 414(1) from an audio signal 416(1) included in and / or represented by audio data 408(1). Additionally or alternatively, circuitry 106 may identify, detect, and / or discern an audio signal 414(N) corresponding to the user's voice within audio data 408(N). In this example, circuitry 106 may isolate, extract, and / or separate audio signal 414(N) from an audio signal 416(N) included in and / or represented by audio data 408(N).

[0073] Circuitry 106 may identify and / or isolate audio signal 114 in a variety of different ways. For example, circuitry 106 may generate, implement, and / or provide a beamformer 412(1) for the first input channel and / or a beamformer 412(N) for the second input channel. In one example, circuitry 106 may generate and / or implement beamformers 412(1)-(N) based at least in part on the positions of microphones 404(l)-(N) relative to one another and / or relative to the user wearing eyewear frame 102. In this example, circuitry 106 may rely on and / or leverage knowledge and / or information about the positions and / or locations of microphones 404(l)-(N) relative to the user wearing eyewear frame 102.

[0074] As a specific example, circuitry 106 may calibrate, customize, adapt, and / or personalize beamformers 412(1)-(N) to the user. Beamformers 412(1)-(N) may adapt to the characteristics of the user's voice to facilitate and / or support distortionless separation of different instances of the user's voice across the different input channels. In one example, circuitry 106 may obtain and / or access knowledge and / or information about the geometry of microphones 404(l)-(N) as arranged on eyewear frame 102. In this example, circuity 106 may also obtain and / or access one or more samples of the user's voice captured in an otherwise quiet environment (e.g., an empty and / or quiet room). Additionally or alternatively, circuitry 106 may generate and / or implement beamformers 412( l)-(N) based at least in part on the geometry of microphones 404(l)-(N) and / or such samples of the user's voice.

[0075] In some examples, circuitry 106 may apply beamformers 412( l)-(N) to audio data 408( l)-(N), respectively. In one example, by applying beamformers 412(1)-(N) to audio data 408(l)-(N), circuitry 106 may identify and / or detect audio signals 414(1)-(N) and / or audio signals 416(1)- (N). Additionally or alternatively, by applying beamformers 412(1)-(N) to audio data 408(1)- (N), circuitry 106 may isolate and / or separate audio signals 414(1)-(N) from audio signals 416(1)-(N), respectively.

[0076] In some examples, circuitry 106 may modify, manipulate, and / or alter audio signals 414(1)- (N) and / or audio signals 416(1)-(N) to compensate for any excessive imbalance(s) between the volume levels of audio signals 414(1)-(N) and the volume levels of audio signals 416(1)- (N) across the different input channels. For example, circuitry 106 may amplify audio signals 416(1)-(N) to increase their volume levels (e.g., amplitudes and / or decibel levels) upon playback by a playback device. Additionally or alternatively, circuitry 106 may attenuate audio signals 414(1)-(N) to decrease their volume levels (e.g., amplitudes and / or decibel levels) upon playback by the playback device.

[0077] In some examples, circuitry 106 may substantially balance and / or equalize audio signals 414(1)-(N) and audio signals 416(1)-(N), respectively, based at least in part on the modification(s) made to one or more of the same. For example, circuitry 106 may balance and / or equalize audio signals 414(1) and 416(1) to an SNR with a standard deviation between 5 decibels and 7 decibels (e.g., 6 decibels). In one example, the SNR of audio signals 414(1) and 416(1) may be measured at approximately 8.5 or 8.6 decibels. In certain implementations, circuitry 106 may modify, manipulate, and / or alter one or more of audio signals 414(1)-(N) and 416(1)-(N) to ensure that the combination and / or integration of audio signals 414(1)-(N) and 416(1)-(N) results in an SNR between 8 and 9 decibels (e.g., 8.5 or 8.6 decibels).

[0078] In some examples, circuitry 106 may generate, create, and / or produce modified audio data 420 for playback of audio signals 414(1)-(N) and 416(1)-(N). In one example, circuitry 106 may generate modified audio data 420 by combining and / or integrating audio signals 414(1)-(N) and 416(1)-(N) after completion of the modification(s). For example, signal modifier 118 may combine and / or integrate audio signals 414(1)-(N) and 416(1)-(N) with one another and then output modified audio data 420. In this example, circuitry 106 may subsequently cause one or more output audio transducers (e.g., audio speakers) to play modified audio data 420 as an audio presentation. Additionally or alternatively, circuitry 106 may send, transmit, and / or transfer modified audio data 420 to a remote playback device (e.g., via a network) for playback.

[0079] FIG. 5 illustrates an exemplary implementation of circuitry 106 for balancing voice audio captured by eyewear devices. In some examples, circuitry 106 may include, involve, and / or represent certain devices, components, and / or features that perform and / or provide functionalities that are similar and / or identical to those described above in connection with any of FIGS. 1-4. As illustrated in FIG. 5, circuitry 106 may include and / or represent a mixing block 502 and / or a beamformer block 504. In one example, beamformer block 504 may include, represent, and / or implement an array transfer function (ATF) estimator 520 and / or a beamformer estimator 522. Additionally or alternatively, mixing block 502 may include, represent, and / or implement a reference beamformer 512, a reference nullformer 514, a SNR estimator 516, and / or an own-voice adjustor 518.

[0080] In some examples, circuitry 106 may receive, retrieve, and / or obtain audio data 408(l)-(N). In one example, audio data 408(l)-(N) may include and / or represent audio captured from various microphones (e.g., 2, 3, 4, 5 or more microphones) secured to eyewear frame 102. In this example, audio data 408(l)-(N) may correspond to and / or represent various audio input channels (e.g., 2, 3, 4, 5 or more channels).

[0081] In some examples, circuitry 106 may pass and / or deliver audio data 408(l)-(N) to mixing block 502 and / or beamformer block 504. For example, ATF estimator 520 and / or beamformer estimator 522 in beamformer block 504 may receive and / or obtain audio data 408(l)-(N). Additionally or alternatively, reference beamformer 512, reference nullformer 514, and / or own-voice adjustor 518 in mixing block 502 may receive and / or obtain audio data 408( l)-(N) . In some examples, ATF estimator 520 may generate and / or formulate an ATF estimate for mixing block 502 based at least in part on audio data 408(l)-(N). Additionally or alternatively, beamformer estimator 522 may generate and / or formulate beamformer coefficients for mixing block 502 based at least in part on audio data 408(l)-(N). In one example, beamformer block 504 may also generate a single-channel unreferenced beamformer signal for mixing block 502.

[0082] In some examples, mixing block 502 may use the ATF estimate and / or beamformer coefficients to compute, generate, and / or formulate reference beamformer 512 and / or reference nullformer 514 for each of the audio input channels. In one example, the beamformed signal BFCreferenced to an input channel C may be represented as: BFC— HCBF°, where Hccorresponds to the ATF and BF° corresponds to an unreferenced singlechannel beamformer stream. Additionally or alternatively, the nullformed signal NFCreferenced to an input channel C may be represented as: NFC= XC— BFC, where Xccorresponds to the multichannel input audio data and BFCcorresponds to the beamformed signal referenced to input channel C.

[0083] In some examples, mixing block 502 may compute the beamformed signal and the nullformed signal for each of the input audio channels at and / or via reference beamformer 512 and / or reference nullformer 514, respectively. In one example, mixing block 502 may use these beamformed and nullformed signals to estimate a gain level (whether positive or negative) for achieving a target SNR at and / or via SNR estimator 516. Additionally or alternatively, mixing block 502 may generate and / or output modified audio data 420 by applying the estimated gain level to the user's voice to achieve the target SNR at and / or via own-voice adjustor 518. In certain implementations, the gain level may be applied by amplifying and / or attenuating the user's voice signal and / or the ambient noise signal.

[0084] FIG. 6 illustrates an exemplary system 600 for balancing voice audio captured by eyewear devices. In some examples, system 600 may include, involve, and / or represent certain devices, components, and / or features that perform and / or provide functionalities that are similar and / or identical to those described above in connection with any of FIGS. 1-5. As illustrated in FIG. 6, system 600 may include and / or represent eyewear device 200 and / or a playback device 602 that a communicatively coupled to one another via a network 604. In one example, eyewear device 200 may send, transmit, and / or transfer modified audio data 120 to playback device 602 via network 604. In this example, playback device 602 may then play back modified audio data 120 as an audio presentation via an audio transducer 606 (e.g., a speaker).

[0085] Network 604 generally represents any medium or architecture capable of facilitating communication or data transfer. In one example, network 604 may include one or more network devices and / or connections. In this example, network 604 may facilitate communication or data transfer using wireless and / or wired connections. Examples of network 604 include, without limitation, an intranet, a Wide Area Network (WAN), a Local Area Network (LAN), a Personal Area Network (PAN), the Internet, Power Line Communications (PLC), a cellular network (e.g., a Global System for Mobile Communications (GSM) network), portions of one or more of the same, variations or combinations of one or more of the same, and / or any other suitable network.

[0086] In some examples, the various apparatuses, devices, and systems described in connection with FIGS. 1-6 may include and / or represent one or more additional circuits, components, and / or features that are not necessarily illustrated and / or labeled in FIGS. 1-6. For example, the apparatuses, devices, and systems illustrated in FIGS. 1-6 may also include and / or represent additional analog and / or digital circuitry, onboard logic, transistors, radiofrequency (RF) transmitters, RF receivers, RF transceivers, antennas, resistors, capacitors, diodes, inductors, switches, registers, flipflops, digital logic, connections, traces, buses, semiconductor (e.g., silicon) devices and / or structures, processing devices, storage devices, circuit boards, sensors, packages, substrates, housings, combinations or variations of one or more of the same, and / or any other suitable components. In certain implementations, one or more of these additional circuits, components, and / or features may be inserted and / or applied between any of the existing circuits, components, and / or features illustrated in FIGS. 1-6 consistent with the aims and / or objectives described herein. Accordingly, the couplings and / or connections described with reference to FIGS. 1-6 may be direct connections with no intermediate components, devices, and / or nodes or indirect connections with one or more intermediate components, devices, and / or nodes.

[0087] In some examples, the phrase "to couple" and / or the term "coupling", as used herein, may refer to a direct connection and / or an indirect connection. For example, a direct coupling between two components may constitute and / or represent a coupling in which those two components are directly connected to each other by a single node that provides continuity from one of those two components to the other. In other words, the direct coupling may exclude and / or omit any additional components between those two components.

[0088] Additionally or alternatively, an indirect coupling between two components may constitute and / or represent a coupling in which those two components are indirectly connected to each other by multiple nodes that fail to provide continuity from one of those two components to the other. In other words, the indirect coupling may include and / or incorporate at least one additional component between those two components. In some examples, one or more components and / or features illustrated in FIGS. 1-6 may be excluded and / or omitted from the various apparatuses, devices, and / or systems described in connection with FIGS. 1-6.

[0089] FIG. 7 is a flow diagram of an exemplary method 700 for balancing voice audio captured by eyewear devices. In one example, the steps shown in FIG. 7 may be achieved and / or accomplished by an AR / VR HMD worn by a user. Additionally or alternatively, the steps shown in FIG. 7 may incorporate and / or involve certain sub-steps and / or variations consistent with the descriptions provided above in connection with FIGS. 1-6.

[0090] As illustrated in FIG. 7, method 700 may include the step of capturing, by a plurality of audio transducers, audio data representative of an environment occupied by the user (710). Step 710 may be performed in a variety of ways, including any of those described above in connection with FIGS. 1-6. For example, microphones incorporated in an AR / VR HMD may capture, by a plurality of audio transducers, audio data representative of an environment occupied by the user.

[0091] Method 700 may also include the step of identifying, by circuitry, an audio signal corresponding to a voice of the user and at least one additional audio signal within the audio data (720). Step 720 may be performed in a variety of ways, including any of those described above in connection with FIGS. 1-6. For example, circuitry incorporated in the AR / VR HMD may identify an audio signal corresponding to a voice of the user and at least one additional audio signal within the audio data.

[0092] Method 700 may further include the step of isolate, by the circuitry, the audio signal from the at least one additional signal (730). Step 730 may be performed in a variety of ways, including any of those described above in connection with FIGS. 1-6. For example, the circuitry incorporated in the AR / VR HMD may isolate the audio signal from the at least one additional signal.

[0093] Method 700 may further include the step of modify, by the circuitry, the audio signal from the at least one additional signal to compensate for an imbalance between a volume level of the voice of the user and a volume level of the at least one additional audio signal (740). Step 740 may be performed in a variety of ways, including any of those described above in connection with FIGS. 1-6. For example, the circuitry incorporated in the AR / VR HMD may modify the audio signal from the at least one additional signal to compensate for an imbalance between a volume level of the voice of the user and a volume level of the at least one additional audio signal.

[0094] Example Embodiments

[0095] Example 1: An apparatus comprising (1) an eyewear frame dimensioned to be worn by a user, (2) a plurality of audio transducers secured to the frame and configured to capture audio data representative of an environment occupied by the user, and (3) circuitry configured to (A) identify, within the audio data, an audio signal corresponding to a voice of the user and at least one additional audio signal, (B) isolate the audio signal from the at least one additional audio signal, and / or (C) modify the audio signal or the at least one additional audio signal to compensate for an imbalance between a volume level of the voice of the user and a volume level of the at least one additional audio signal.

[0096] Example 2: The apparatus of Example 1, wherein the circuitry is further configured to generate modified audio data for playback of the voice of the user and the additional audio signal by combining the audio signal and the additional audio signal after completion of the modification.

[0097] Example 3: The apparatus of either Example 1 or Example 2, wherein the circuitry is further configured to transmit the modified audio data to a playback device that plays the audio data as an audio presentation.

[0098] Example 4: The apparatus of any of Examples 1-3, wherein the circuitry is further configured to (1) generate a beamformer for isolating the audio signal based at least in part on an arrangement of the audio transducers secured to the frame and (2) apply the beamformer to the audio data to (A) identify the audio signal and the at least one additional audio signal and (B) isolate the audio signal and the at least one additional audio signal from one another.

[0099] Example 5: The apparatus of any of Examples 1-4, wherein the audio transducers comprise

[0100] (1) a first microphone configured to capture a first instance of the audio data for a first input channel and (2) a second microphone configured to capture a second instance of the audio data for a first input channel.

[0101] Example 6: The apparatus of any of Examples 1-5, wherein the circuitry is further configured to (1) generate a first beamformer for the first input channel based at least in part on a position of the first microphone relative to the user whose voice captured in the audio signal,

[0102] (2) generate a second beamformer for the second input channel based at least in part on a position of the first microphone relative to the user whose voice captured in the audio signal,.

[0103] (3) apply the first beamformerto the first instance of the audio data to isolate the audio signal and the at least one additional audio signal from one another for the first input channel, and / or (4) apply the second beamformer to the second instance of the audio data to isolate the audio signal and the at least one additional audio signal from one another for the second input channel. Example 7: The apparatus of any of Examples 1-6, wherein the circuitry is further configured to (1) modify the audio signal or the at least one additional audio signal in the first instance of the audio data to compensate for a first imbalance between the volume level of the voice of the user and the volume level of the at least one additional audio signal in the first input channel and / or (2) modify the audio signal or the at least one additional audio signal in the second instance of the audio data to compensate fora second imbalance between the volume level of the voice of the user and the volume level of the at least one additional audio signal in the second input channel.

[0104] Example 8: The apparatus of any of Examples 1-7, wherein the circuitry is further configured to generate modified audio data for playback of the voice of the user and the additional audio signal by combining the first instance of the audio data and the second instance of the audio data after completion of the modifications.

[0105] Example 9: The apparatus of any of Examples 1-8, wherein the circuitry is further configured to attenuate the audio signal to a certain decibel level based at least in part on the volume level of the at least one additional audio signal.

[0106] Example 10: The apparatus of any of Examples 1-9, wherein the circuitry is further configured to amplify the at least one additional audio signal to a certain decibel level based at least in part on the volume level of the audio signal.

[0107] Example 11: The apparatus of any of Examples 1-10, wherein the circuitry is further configured to balance, based at least in part on the modification, the audio signal and the at least one additional audio signal to a signal-to-noise ratio with a standard deviation between 5 decibels and 7 decibels.

[0108] Example 12: The apparatus of any of Examples 1-11, wherein the signal-to-noise ratio of the audio signal and the at least one additional audio signal is approximately 8.5 decibels.

[0109] Example 13: The apparatus of any of Examples 1-12, wherein the circuitry is further configured to (1) detect the imbalance by comparing the volume level of the audio signal and the volume level of the at least one additional audio signal, (2) determine, based at least in part on the comparison, that the imbalance satisfies a certain threshold, and / or (3) modify, in response to the determination, the audio signal or the at least one additional audio signal to compensate for the imbalance.

[0110] Example 14: A system comprising (1) an eyewear device that (A) is dimensioned to be worn by a user and (B) comprises a plurality of audio transducers secured to the frame and configured to capture audio data representative of an environment occupied by the user, (2) circuitry configured to (A) isolate, within the audio data, an audio signal corresponding to a voice of the user and at least one additional audio signal and (B) modify the audio signal or the at least one additional audio signal to compensate for an imbalance between a volume level of the voice of the user and a volume level of the at least one additional audio signal, and (3) a playback device configured to play the audio data as an audio presentation after completion of the modification.

[0111] Example 15: The system of Example 14, wherein the circuitry is further configured to modify the audio data for playback of the voice of the user and the additional audio signal by combining the audio signal and the additional audio signal.

[0112] Example 16: The system of Example 14 or 15, wherein the circuitry is further configured to transmit the audio data to the playback device via a network.

[0113] Example 17: The system of any of Examples 14-16, wherein the circuitry is further configured to (1) generate a beamformer for isolating the audio signal based at least in part on an arrangement of the audio transducers secured to the frame and (2) apply the beamformer to the audio data to (A) identify the audio signal and the at least one additional audio signal and (B) isolate the audio signal and the at least one additional audio signal from one another.

[0114] Example 18: The system of any of Examples 14-17, wherein the audio transducers comprise

[0115] (1) a first microphone configured to capture a first instance of the audio data for a first input channel and (2) a second microphone configured to capture a second instance of the audio data for a first input channel.

[0116] Example 19: The system of any of Examples 14-18, wherein the circuitry is further configured to (1) generate a first beamformer for the first input channel based at least in part on a position of the first microphone relative to the user whose voice captured in the audio signal,

[0117] (2) generate a second beamformer for the second input channel based at least in part on a position of the first microphone relative to the user whose voice captured in the audio signal,.

[0118] (3) apply the first beamformerto the first instance of the audio data to isolate the audio signal and the at least one additional audio signal from one another for the first input channel, (4) apply the second beamformer to the second instance of the audio data to isolate the audio signal and the at least one additional audio signal from one another for the second input channel, (5) modify the audio signal or the at least one additional audio signal in the first instance of the audio data to compensate for a first imbalance between the volume level of the voice of the user and the volume level of the at least one additional audio signal in the first input channel, and / or (6) modify the audio signal or the at least one additional audio signal in the second instance of the audio data to compensate for a second imbalance between the volume level of the voice of the user and the volume level of the at least one additional audio signal in the second input channel.

[0119] Example 20: A method comprising (1) capturing, by a plurality of audio transducers, audio data representative of an environment occupied by a user, (2) identifying, by circuitry, an audio signal corresponding to a voice of the user and at least one additional audio signal within the audio data, (3) isolating, by the circuitry, the audio signal from the at least one additional audio signal, and / or (4) modifying, by the circuitry, the audio signal or the at least one additional audio signal to compensate for an imbalance between a volume level of the voice of the user and a volume level of the at least one additional audio signal.

[0120] Embodiments of the present disclosure may include or be implemented in conjunction with various types of artificial-reality systems. Artificial reality is a form of reality that has been adjusted in some manner before presentation to a user, which may include, for example, a VR, an AR, a mixed reality, a hybrid reality, or some combination and / or derivative thereof. Artificial-reality content may include completely computer-generated content or computer-generated content combined with captured (e.g., real-world) content. The artificial-reality content may include video, audio, haptic feedback, or some combination thereof, any of which may be presented in a single channel or in multiple channels (such as stereo video that produces a 3D effect to the viewer). Additionally, in some embodiments, artificial reality may also be associated with applications, products, accessories, services, or some combination thereof, that are used to, for example, create content in an artificial reality and / or are otherwise used in (e.g., to perform activities in) an artificial reality.

[0121] Artificial-reality systems may be implemented in a variety of different form factors and configurations. Some artificial-reality systems may be designed to work without near-eye displays (NEDs). Other artificial-reality systems may include an NED that also provides visibility into the real world (such as, e.g., augmented-reality system 800 in FIG. 8) or that visually immerses a user in an artificial reality (such as, e.g., virtual-reality system 900 in FIG. 9). While some artificial-reality devices may be self-contained systems, other artificial-reality devices may communicate and / or coordinate with external devices to provide an artificial-reality experience to a user. Examples of such external devices include handheld controllers, mobile devices, desktop computers, devices worn by a user, devices worn by one or more other users, and / or any other suitable external system.

[0122] Turning to FIG. 8, augmented-reality system 800 may include an eyewear device 802 with a frame 810 configured to hold a left display device 815(A) and a right display device 815(B) in front of a user's eyes. Display devices 815(A) and 815(B) may act together or independently to present an image or series of images to a user. While augmented-reality system 800 includes two displays, embodiments of this disclosure may be implemented in augmented- reality systems with a single NED or more than two NEDs.

[0123] In some embodiments, augmented-reality system 800 may include one or more sensors, such as sensor 840. Sensor 840 may generate measurement signals in response to motion of augmented-reality system 800 and may be located on substantially any portion of frame 810. Sensor 840 may represent one or more of a variety of different sensing mechanisms, such as a position sensor, an inertial measurement unit (IMU), a depth camera assembly, a structured light emitter and / or detector, or any combination thereof. In some embodiments, augmented-reality system 800 may or may not include sensor 840 or may include more than one sensor. In embodiments in which sensor 840 includes an IMU, the IMU may generate calibration data based on measurement signals from sensor 840. Examples of sensor 840 may include, without limitation, accelerometers, gyroscopes, magnetometers, other suitable types of sensors that detect motion, sensors used for error correction of the IMU, or some combination thereof.

[0124] In some examples, augmented-reality system 800 may also include a microphone array with a plurality of acoustic transducers 820(A)-820(J), referred to collectively as acoustic transducers 820. Acoustic transducers 820 may represent transducers that detect air pressure variations induced by sound waves. Each acoustic transducer820 may be configured to detect sound and convert the detected sound into an electronic format (e.g., an analog or digital format). The microphone array in FIG. 8 may include, for example, ten acoustic transducers: 820(A) and 820(B), which may be designed to be placed inside a corresponding ear of the user, acoustic transducers 820(C), 820(D), 820(E), 820(F), 820(G), and 820(H), which may be positioned at various locations on frame 810, and / or acoustic transducers 820(1) and 820(J), which may be positioned on a corresponding neckband 805.

[0125] In some embodiments, one or more of acoustic transducers 820(A)-(J) may be used as output transducers (e.g., speakers). For example, acoustic transducers 820(A) and / or 820(B) may be earbuds or any other suitable type of headphone or speaker.

[0126] The configuration of acoustic transducers 820 of the microphone array may vary. While augmented-reality system 800 is shown in FIG. 8 as having ten acoustic transducers 820, the number of acoustic transducers 820 may be greater or less than ten. In some embodiments, using higher numbers of acoustic transducers 820 may increase the amount of audio information collected and / or the sensitivity and accuracy of the audio information. In contrast, using a lower number of acoustic transducers 820 may decrease the computing power required by an associated controller 850 to process the collected audio information. In addition, the position of each acoustic transducer 820 of the microphone array may vary. For example, the position of an acoustic transducer 820 may include a defined position on the user, a defined coordinate on frame 810, an orientation associated with each acoustic transducer 820, or some combination thereof.

[0127] Acoustic transducers 820(A) and 820(B) may be positioned on different parts of the user's ear, such as behind the pinna, behind the tragus, and / or within the auricle or fossa. Or, there may be additional acoustic transducers 820 on or surrounding the ear in addition to acoustic transducers 820 inside the ear canal. Having an acoustic transducer 820 positioned next to an ear canal of a user may enable the microphone array to collect information on how sounds arrive at the ear canal. By positioning at least two of acoustic transducers 820 on either side of a user's head (e.g., as binaural microphones), AR system 800 may simulate binaural hearing and capture a 3D stereo sound field around about a user's head. In some embodiments, acoustic transducers 820(A) and 820(B) may be connected to augmented-reality system 800 via a wired connection 830, and in other embodiments acoustic transducers 820(A) and 820(B) may be connected to augmented-reality system 800 via a wireless connection (e.g., a BLUETOOTH connection). In still other embodiments, acoustic transducers 820(A) and 820(B) may not be used at all in conjunction with augmented-reality system 800.

[0128] Acoustic transducers 820 on frame 810 may be positioned in a variety of different ways, including along the length of the temples, across the bridge, above or below display devices 815(A) and 815(B), or some combination thereof. Acoustic transducers 820 may also be oriented such that the microphone array is able to detect sounds in a wide range of directions surrounding the user wearing the augmented-reality system 800. In some embodiments, an optimization process may be performed during manufacturing of augmented-reality system 800 to determine relative positioning of each acoustic transducer 820 in the microphone array.

[0129] In some examples, augmented-reality system 800 may include or be connected to an external device (e.g., a paired device), such as neckband 805. Neckband 805 generally represents any type or form of paired device. Thus, the following discussion of neckband 805 may also apply to various other paired devices, such as charging cases, smart watches, smart phones, wrist bands, other wearable devices, hand-held controllers, tablet computers, laptop computers, other external compute devices, etc.

[0130] As shown, neckband 805 may be coupled to eyewear device 802 via one or more connectors. The connectors may be wired or wireless and may include electrical and / or non-electrical (e.g., structural) components. In some cases, eyewear device 802 and neckband 805 may operate independently without any wired or wireless connection between them. While FIG. 8 illustrates the components of eyewear device 802 and neckband 805 in example locations on eyewear device 802 and neckband 805, the components may be located elsewhere and / or distributed differently on eyewear device 802 and / or neckband 805. In some embodiments, the components of eyewear device 802 and neckband 805 may be located on one or more additional peripheral devices paired with eyewear device 802, neckband 805, or some combination thereof.

[0131] Pairing external devices, such as neckband 805, with augmented-reality eyewear devices may enable the eyewear devices to achieve the form factor of a pair of glasses while still providing sufficient battery and computation power for expanded capabilities. Some or all of the battery power, computational resources, and / or additional features of augmented-reality system 800 may be provided by a paired device or shared between a paired device and an eyewear device, thus reducing the weight, heat profile, and form factor of the eyewear device overall while still retaining desired functionality. For example, neckband 805 may allow components that would otherwise be included on an eyewear device to be included in neckband 805 since users may tolerate a heavier weight load on their shoulders than they would tolerate on their heads. Neckband 805 may also have a larger surface area over which to diffuse and disperse heat to the ambient environment. Thus, neckband 805 may allow for greater battery and computation capacity than might otherwise have been possible on a stand-alone eyewear device. Since weight carried in neckband 805 may be less invasive to a user than weight carried in eyewear device 802, a user may tolerate wearing a lighter eyewear device and carrying or wearing the paired device for greater lengths of time than a user would tolerate wearing a heavy standalone eyewear device, thereby enabling users to more fully incorporate artificial-reality environments into their day-to-day activities.

[0132] Neckband 805 may be communicatively coupled with eyewear device 802 and / or to other devices. These other devices may provide certain functions (e.g., tracking, localizing, depth mapping, processing, storage, etc.) to augmented-reality system 800. In the embodiment of FIG. 8, neckband 805 may include two acoustic transducers (e.g., 820(1) and 820(J)) that are part of the microphone array (or potentially form their own microphone subarray). Neckband 805 may also include a controller 825 and a power source 835.

[0133] Acoustic transducers 820(1) and 820(J) of neckband 805 may be configured to detect sound and convert the detected sound into an electronic format (analog or digital). In the embodiment of FIG. 8, acoustic transducers 820(1) and 820(J) may be positioned on neckband 805, thereby increasing the distance between the neckband acoustic transducers 820(1) and 820(J) and other acoustic transducers 820 positioned on eyewear device 802. In some cases, increasing the distance between acoustic transducers 820 of the microphone array may improve the accuracy of beamforming performed via the microphone array. For example, if a sound is detected by acoustic transducers 820(C) and 820(D) and the distance between acoustic transducers 820(C) and 820(D) is greater than, e.g., the distance between acoustic transducers 820(D) and 820(E), the determined source location of the detected sound may be more accurate than if the sound had been detected by acoustic transducers 820(D) and 820(E).

[0134] Controller 825 of neckband 805 may process information generated by the sensors on neckband 805 and / or augmented-reality system 800. For example, controller 825 may process information from the microphone array that describes sounds detected by the microphone array. Foreach detected sound, controller 825 may perform a direction-of-arrival (DOA) estimation to estimate a direction from which the detected sound arrived at the microphone array. As the microphone array detects sounds, controller 825 may populate an audio data set with the information. In embodiments in which augmented-reality system 800 includes an inertial measurement unit, controller 825 may compute all inertial and spatial calculations from the IMU located on eyewear device 802. A connector may convey information between augmented-reality system 800 and neckband 805 and between augmented-reality system 800 and controller 825. The information may be in the form of optical data, electrical data, wireless data, or any other transmittable data form. Moving the processing of information generated by augmented-reality system 800 to neckband 805 may reduce weight and heat in eyewear device 802, making it more comfortable to the user.

[0135] Power source 835 in neckband 805 may provide power to eyewear device 802 and / or to neckband 805. Power source 835 may include, without limitation, lithium-ion batteries, lithium-polymer batteries, primary lithium batteries, alkaline batteries, or any other form of power storage. In some cases, power source 835 may be a wired power source. Including power source 835 on neckband 805 instead of on eyewear device 802 may help better distribute the weight and heat generated by power source 835.

[0136] As noted, some artificial-reality systems may, instead of blending an artificial reality with actual reality, substantially replace one or more of a user's sensory perceptions of the real world with a virtual experience. One example of this type of system is a head-worn display system, such as virtual-reality system 900 in FIG. 9, that mostly or completely covers a user's field of view. Virtual-reality system 900 may include a front rigid body 902 and a band 904 shaped to fit around a user's head. Virtual-reality system 900 may also include output audio transducers 906(A) and 906(B). Furthermore, while not shown in FIG. 9, front rigid body 902 may include one or more electronic elements, including one or more electronic displays, one or more inertial measurement units (IMUs), one or more tracking emitters or detectors, and / or any other suitable device or system for creating an artificial-reality experience.

[0137] Artificial-reality systems may include a variety of types of visual feedback mechanisms. For example, display devices in augmented-reality system 800 and / or virtual-reality system 900 may include one or more liquid crystal displays (LCDs), light emitting diode (LED) displays, microLED displays, organic LED (OLED) displays, digital light project (DLP) micro-displays, liquid crystal on silicon (LCoS) micro-displays, and / or any other suitable type of display screen. These artificial-reality systems may include a single display screen for both eyes or may provide a display screen for each eye, which may allow for additional flexibility for varifocal adjustments or for correcting a user's refractive error. Some of these artificial-reality systems may also include optical subsystems having one or more lenses (e.g., concave or convex lenses, Fresnel lenses, adjustable liquid lenses, etc.) through which a user may view a display screen. These optical subsystems may serve a variety of purposes, including to collimate (e.g., make an object appear at a greater distance than its physical distance), to magnify (e.g., make an object appear larger than its actual size), and / or to relay (to, e.g., the viewer's eyes) light. These optical subsystems may be used in a non-pupil-forming architecture (such as a single lens configuration that directly collimates light but results in so-called pincushion distortion) and / or a pupil-forming architecture (such as a multi-lens configuration that produces so- called barrel distortion to nullify pincushion distortion).

[0138] In addition to or instead of using display screens, some of the artificial-reality systems described herein may include one or more projection systems. For example, display devices in augmented-reality system 800 and / or virtual-reality system 900 may include micro-LED projectors that project light (using, e.g., a waveguide) into display devices, such as clear combiner lenses that allow ambient light to pass through. The display devices may refract the projected light toward a user's pupil and may enable a user to simultaneously view both artificial-reality content and the real world. The display devices may accomplish this using any of a variety of different optical components, including waveguide components (e.g., holographic, planar, diffractive, polarized, and / or reflective waveguide elements), lightmanipulation surfaces and elements (such as diffractive, reflective, and refractive elements and gratings), coupling elements, etc. Artificial-reality systems may also be configured with any other suitable type or form of image projection system, such as retinal projectors used in virtual retina displays.

[0139] The artificial-reality systems described herein may also include various types of computer vision components and subsystems. For example, augmented-reality system 800 and / or virtual-reality system 900 may include one or more optical sensors, such as two-dimensional (2D) or 3D cameras, structured light transmitters and detectors, time-of-flight depth sensors, single-beam orsweeping laser rangefinders, 3D LiDAR sensors, and / or any othersuitable type or form of optical sensor. An artificial-reality system may process data from one or more of these sensors to identify a location of a user, to map the real world, to provide a user with context about real-world surroundings, and / or to perform a variety of other functions.

[0140] The artificial-reality systems described herein may also include one or more input and / or output audio transducers. Output audio transducers may include voice coil speakers, ribbon speakers, electrostatic speakers, piezoelectric speakers, bone conduction transducers, cartilage conduction transducers, tragus-vibration transducers, and / or any other suitable type or form of audio transducer. Similarly, input audio transducers may include condenser microphones, dynamic microphones, ribbon microphones, and / or any other type or form of input transducer. In some embodiments, a single transducer may be used for both audio input and audio output. In some embodiments, the artificial-reality systems described herein may also include tactile (i.e., haptic) feedback systems, which may be incorporated into headwear, gloves, body suits, handheld controllers, environmental devices (e.g., chairs, floormats, etc.), and / or any other type of device or system. Haptic feedback systems may provide various types of cutaneous feedback, including vibration, force, traction, texture, and / or temperature. Haptic feedback systems may also provide various types of kinesthetic feedback, such as motion and compliance. Haptic feedback may be implemented using motors, piezoelectric actuators, fluidic systems, and / or a variety of other types of feedback mechanisms. Haptic feedback systems may be implemented independent of other artificial-reality devices, within other artificial-reality devices, and / or in conjunction with other artificial-reality devices.

[0141] By providing haptic sensations, audible content, and / or visual content, artificial-reality systems may create an entire virtual experience or enhance a user's real-world experience in a variety of contexts and environments. For instance, artificial-reality systems may assist or extend a user's perception, memory, or cognition within a particular environment. Some systems may enhance a user's interactions with other people in the real world or may enable more immersive interactions with other people in a virtual world. Artificial-reality systems may also be used for educational purposes (e.g., for teaching or training in schools, hospitals, government organizations, military organizations, business enterprises, etc.), entertainment purposes (e.g., for playing video games, listening to music, watching video content, etc.), and / or for accessibility purposes (e.g., as hearing aids, visual aids, etc.). The embodiments disclosed herein may enable or enhance a user's artificial-reality experience in one or more of these contexts and environments and / or in other contexts and environments.

[0142] In some embodiments, one or more objects (e.g., content or other types of objects) of a computing system may be associated with one or more privacy settings. The one or more objects may be stored on or otherwise associated with any suitable computing system or application, such as, for example, a social-networking system, a client system, a third-party system, a social-networking application, a messaging application, a photo-sharing application, or any other suitable computing system or application.

[0143] Privacy settings (or "access settings") for an object may be stored in any suitable manner, such as, for example, in association with the object, in an index on an authorization server, in another suitable manner, or any suitable combination thereof. A privacy setting for an object may specify how the object (or particular information associated with the object) can be accessed, stored, or otherwise used (e.g., viewed, shared, modified, copied, executed, surfaced, or identified) within the online social network. When privacy settings for an object allow a particular user or other entity to access that object, the object may be described as being "visible" with respect to that user or other entity. As an example and not by way of limitation, a user of an online social network may specify privacy settings for a user-profile page that identify a set of users that may access work-experience information on the userprofile page, thus excluding other users from accessing that information.

[0144] In some embodiments, privacy settings for an object may specify a "blocked list" of users or other entities that should not be allowed to access certain information associated with the object. In some cases, the blocked list may include third-party entities. The blocked list may specify one or more users or entities for which an object is not visible. As an example and not by way of limitation, a user may specify a set of users who may not access photo albums associated with the user, thus excluding those users from accessing the photo albums (while also possibly allowing certain users not within the specified set of users to access the photo albums). In some embodiments, privacy settings may be associated with particular socialgraph elements. Privacy settings of a social-graph element, such as a node or an edge, may specify how the social-graph element, information associated with the social-graph element, or objects associated with the social-graph element can be accessed using the online social network. As an example and not by way of limitation, a particular concept node corresponding to a particular photo may have a privacy setting specifying that the photo may be accessed only by users tagged in the photo and friends of the users tagged in the photo. In some embodiments, privacy settings may allow users to opt in to or opt out of having their content, information, or actions stored / logged by a social-networking system or shared with other systems (e.g., a third-party system). Although this disclosure describes using particular privacy settings in a particular manner, this disclosure contemplates using any suitable privacy settings in any suitable manner.

[0145] In some embodiments, privacy settings may be based on one or more nodes or edges of a social graph. A privacy setting may be specified for one or more edges or edge-types of the social graph, or with respect to one or more nodes or node-types of the social graph. The privacy settings applied to a particular edge connecting two nodes may control whether the relationship between the two entities corresponding to the nodes is visible to other users of the online social network. Similarly, the privacy settings applied to a particular node may control whether the user or concept corresponding to the node is visible to other users of the online social network. As an example and not by way of limitation, a first user may share an object to the socialnetworking system. The object may be associated with a concept node connected to a user node of the first user by an edge. The first user may specify privacy settings that apply to a particular edge connecting to the concept node of the object, or may specify privacy settings that apply to all edges connecting to the concept node. As another example and not by way of limitation, the first user may share a set of objects of a particular object-type (e.g., a set of images). The first user may specify privacy settings with respect to all objects associated with the first user of that particular object-type as having a particular privacy setting (e.g., specifying that all images posted by the first user are visible only to friends of the first user and / or users tagged in the images).

[0146] In some embodiments, a social-networking system may present a "privacy wizard" (e.g., within a webpage, a module, one or more dialog boxes, or any other suitable interface) to the first user to assist the first user in specifying one or more privacy settings. The privacy wizard may display instructions, suitable privacy-related information, current privacy settings, one or more input fields for accepting one or more inputs from the first user specifying a change or confirmation of privacy settings, or any suitable combination thereof. In some embodiments, the social-networking system may offer a "dashboard" functionality to the first user that may display, to the first user, current privacy settings of the first user. The dashboard functionality may be displayed to the first user at any appropriate time (e.g., following an input from the first user summoning the dashboard functionality, following the occurrence of a particular event or trigger action). The dashboard functionality may allow the first user to modify one or more of the first user's current privacy settings at any time, in any suitable manner (e.g., redirecting the first user to the privacy wizard).

[0147] Privacy settings associated with an object may specify any suitable granularity of permitted access or denial of access. As an example and not by way of limitation, access or denial of access may be specified for particular users (e.g., only me, my roommates, my boss), users within a particular degree-of-separation (e.g., friends, friends-of-friends), user groups (e.g., the gaming club, my family), user networks (e.g., employees of particular employers, students or alumni of particular university), all users ("public"), no users ("private"), users of third-party systems, particular applications (e.g., third-party applications, external websites), other suitable entities, or any suitable combination thereof. Although this disclosure describes particular granularities of permitted access or denial of access, this disclosure contemplates any suitable granularities of permitted access or denial of access.

[0148] In some embodiments, one or more servers may be authorization / privacy servers for enforcing privacy settings. In response to a request from a user (or other entity) for a particular object stored in a data store, the social-networking system may send a request to the data store for the object. The request may identify the user associated with the request and the object may be sent only to the user (or a client system of the user) if the authorization server determines that the user is authorized to access the object based on the privacy settings associated with the object. If the requesting user is not authorized to access the object, the authorization server may prevent the requested object from being retrieved from the data store or may prevent the requested object from being sent to the user. In the searchquery context, an object may be provided as a search result only if the querying user is authorized to access the object, e.g., if the privacy settings for the object allow it to be surfaced to, discovered by, or otherwise visible to the querying user. In some embodiments, an object may represent content that is visible to a user through a newsfeed of the user. As an example and not by way of limitation, one or more objects may be visible to a user's "Trending" page. In some embodiments, an object may correspond to a particular user. The object may be content associated with the particular user or may be the particular user's account or information stored on the social-networking system or other computing system. As an example and not by way of limitation, a first user may view one or more second users of an online social network through a "People You May Know" function of the online social network, or by viewing a list of friends of the first user. As an example and not by way of limitation, a first user may specify that they do not wish to see objects associated with a particular second user in their newsfeed or friends list. If the privacy settings for the object do not allow it to be surfaced to, discovered by, or visible to the user, the object may be excluded from the search results. Although this disclosure describes enforcing privacy settings in a particular manner, this disclosure contemplates enforcing privacy settings in any suitable manner.

[0149] In some embodiments, different objects of the same type associated with a user may have different privacy settings. Different types of objects associated with a user may have different types of privacy settings. As an example and not by way of limitation, a first user may specify that the first user's status updates are public, but any images shared by the first user are visible only to the first user's friends on the online social network. As another example and not by way of limitation, a user may specify different privacy settings for different types of entities, such as individual users, friends-of-friends, followers, user groups, or corporate entities. As another example and not by way of limitation, a first user may specify a group of users that may view videos posted by the first user, while keeping the videos from being visible to the first user's employer. In some embodiments, different privacy settings may be provided for different user groups or user demographics. As an example and not by way of limitation, a first user may specify that other users who attend the same university as the first user may view the first user's pictures, but that other users who are family members of the first user may not view those same pictures.

[0150] In some embodiments, the social-networking system may provide one or more default privacy settings for each object of a particular object-type. A privacy setting for an object that is set to a default may be changed by a user associated with that object. As an example and not by way of limitation, all images posted by a first user may have a default privacy setting of being visible only to friends of the first user and, for a particular image, the first user may change the privacy setting for the image to be visible to friends and friends-of-friends.

[0151] In some embodiments, privacy settings may allow a first user to specify (e.g., by opting out, by not opting in) whether the social-networking system may receive, collect, log, or store particular objects or information associated with the user for any purpose. In some embodiments, privacy settings may allow the first user to specify whether particular applications or processes may access, store, or use particular objects or information associated with the user. The privacy settings may allow the first user to opt in or opt out of having objects or information accessed, stored, or used by specific applications or processes. The social-networking system may access such information in order to provide a particular function or service to the first user, without the social-networking system having access to that information for any other purposes. Before accessing, storing, or using such objects or information, the social-networking system may prompt the user to provide privacy settings specifying which applications or processes, if any, may access, store, or use the object or information prior to allowing any such action. As an example and not by way of limitation, a first user may transmit a message to a second user via an application related to the online social network (e.g., a messaging app), and may specify privacy settings that such messages should not be stored by the social-networking system.

[0152] In some embodiments, a user may specify whether particular types of objects or information associated with the first user may be accessed, stored, or used by the social-networking system. As an example and not by way of limitation, the first user may specify that images sent by the first user through the social-networking system may not be stored by the socialnetworking system. As another example and not by way of limitation, a first user may specify that messages sent from the first user to a particular second user may not be stored by the social-networking system. As yet another example and not by way of limitation, a first user may specify that all objects sent via a particular application may be saved by the socialnetworking system.

[0153] In some embodiments, privacy settings may allow a first user to specify whether particular objects or information associated with the first user may be accessed from particular client systems or third-party systems. The privacy settings may allow the first user to opt in or opt out of having objects or information accessed from a particular device (e.g., the phone book on a user's smart phone), from a particular application (e.g., a messaging app), or from a particular system (e.g., an email server). The social-networking system may provide default privacy settings with respect to each device, system, or application, and / or the first user may be prompted to specify a particular privacy setting for each context. As an example and not by way of limitation, the first user may utilize a location-services feature of the socialnetworking system to provide recommendations for restaurants or other places in proximity to the user. The first user's default privacy settings may specify that the social-networking system may use location information provided from a client device of the first user to provide the location-based services, but that the social-networking system may not store the location information of the first user or provide it to any third-party system. The first user may then update the privacy settings to allow location information to be used by a third-party imagesharing application in order to geo-tag photos.

[0154] Privacy Settings for Mood, Emotion, or Sentiment Information

[0155] In some embodiments, privacy settings may allow a user to specify whether current, past, or projected mood, emotion, or sentiment information associated with the user may be determined, and whether particular applications or processes may access, store, or use such information. The privacy settings may allow users to opt in or opt out of having mood, emotion, or sentiment information accessed, stored, or used by specific applications or processes. For example, a social-networking system may predict or determine a mood, emotion, or sentiment associated with a user based on, for example, inputs provided by the user and interactions with particular objects, such as pages or content viewed by the user, posts or other content uploaded by the user, and interactions with other content of the online social network.

[0156] In some embodiments, the social-networking system may use a user's previous activities and calculated moods, emotions, or sentiments to determine a present mood, emotion, or sentiment. A user who wishes to enable this functionality may indicate in their privacy settings that they opt in to the social-networking system receiving the inputs necessary to determine the mood, emotion, or sentiment. As an example, the social-networking system may determine that a default privacy setting is to not receive any information necessary for determining mood, emotion, or sentiment until there is an express indication from a user that the social-networking system may do so. By contrast, if a user does not opt in to the socialnetworking system receiving these inputs (or affirmatively opts out of the social-networking system receiving these inputs), the social-networking system may be prevented from receiving, collecting, logging, or storing these inputs or any information associated with these inputs. In some embodiments, the social-networking system may use the predicted mood, emotion, or sentiment to provide recommendations or advertisements to the user.

[0157] In some embodiments, if a user desires to make use of this function for specific purposes or applications, additional privacy settings may be specified by the user to opt in to using the mood, emotion, or sentiment information for the specific purposes or applications. As an example, the social-networking system may use the user's mood, emotion, or sentiment to provide newsfeed items, pages, friends, or advertisements to a user. The user may specify in their privacy settings that the social-networking system may determine the user's mood, emotion, or sentiment. The user may then be asked to provide additional privacy settings to indicate the purposes for which the user's mood, emotion, or sentiment may be used. The user may indicate that the social-networking system may use his or her mood, emotion, or sentiment to provide newsfeed content and recommend pages, but not for recommending friends or advertisements. The social-networking system may then only provide newsfeed content or pages based on user mood, emotion, or sentiment, and may not use that information for any other purpose, even if not expressly prohibited by the privacy settings.

[0158] Privacy Settings for Ephemeral Sharing In some embodiments, privacy settings may allow a user to engage in the ephemeral sharing of objects on an online social network. Ephemeral sharing refers to the sharing of objects (e.g., posts, photos) or information for a finite period of time. Access or denial of access to the objects or information may be specified by time or date. As an example and not by way of limitation, a user may specify that a particular image uploaded by the user is visible to the user's friends for the next week, after which time the image may no longer be accessible to other users. As another example and not by way of limitation, a company may post content related to a product release ahead of the official launch, and specify that the content may not be visible to other users until after the product launch.

[0159] In some embodiments, for particular objects or information having privacy settings specifying that they are ephemeral, the social-networking system may be restricted in its access, storage, or use of the objects or information. The social-networking system may temporarily access, store, or use these particular objects or information in order to facilitate particular actions of a user associated with the objects or information, and may subsequently delete the objects or information, as specified by the respective privacy settings. As an example and not by way of limitation, a first user may transmit a message to a second user, and the socialnetworking system may temporarily store the message in a data store until the second user has viewed or downloaded the message, at which point the social-networking system may delete the message from the data store. As another example and not by way of limitation, continuing with the prior example, the message may be stored for a specified period of time (e.g., 2 weeks), after which point the social-networking system may delete the message from the data store.

[0160] Privacy Settings Based on Location

[0161] In some embodiments, privacy settings may allow a user to specify one or more geographic locations from which objects can be accessed. Access or denial of access to the objects may depend on the geographic location of a user who is attempting to access the objects. As an example and not by way of limitation, a user may share an object and specify that only users in the same city may access or view the object. As another example and not by way of limitation, a first user may share an object and specify that the object is visible to second users only while the first user is in a particular location. If the first user leaves the particular location, the object may no longer be visible to the second users. As another example and not by way of limitation, a first user may specify that an object is visible only to second users within a threshold distance from the first user. If the first user subsequently changes location, the original second users with access to the object may lose access, while a new group of second users may gain access as they come within the threshold distance of the first user.

[0162] Privacy Settings for User-Authentication and Experience-Personalization Information

[0163] In some embodiments, a social-networking system may have functionalities that may use, as inputs, personal or biometric information of a user for user-authentication or experiencepersonalization purposes. A user may opt to make use of these functionalities to enhance their experience on the online social network. As an example and not by way of limitation, a user may provide personal or biometric information to the social-networking system. The user's privacy settings may specify that such information may be used only for particular processes, such as authentication, and further specify that such information may not be shared with any third-party system or used for other processes or applications associated with the social-networking system. As another example and not by way of limitation, the social-networking system may provide a functionality for a user to provide voice-print recordings to the online social network. As an example and not by way of limitation, if a user wishes to utilize this function of the online social network, the user may provide a voice recording of his or her own voice to provide a status update on the online social network. The recording of the voice-input may be compared to a voice print of the user to determine what words were spoken by the user. The user's privacy setting may specify that such voice recording may be used only for voice-input purposes (e.g., to authenticate the user, to send voice messages, to improve voice recognition in order to use voice-operated features of the online social network), and further specify that such voice recording may not be shared with any third-party system or used by other processes or applications associated with the socialnetworking system. As another example and not by way of limitation, the social-networking system may provide a functionality for a user to provide a reference image (e.g., a facial profile, a retinal scan) to the online social network. The online social network may compare the reference image against a later-received image input (e.g., to authenticate the user, to tag the user in photos). The user's privacy setting may specify that such voice recording may be used only for a limited purpose (e.g., authentication, tagging the user in photos), and further specify that such voice recording may not be shared with any third-party system or used by other processes or applications associated with the social-networking system. User-Initiated Changes to Privacy Settings

[0164] In some embodiments, changes to privacy settings may take effect retroactively, affecting the visibility of objects and content shared prior to the change. As an example and not by way of limitation, a first user may share a first image and specify that the first image is to be public to all other users. At a later time, the first user may specify that any images shared by the first user should be made visible only to a first user group. A social-networking system may determine that this privacy setting also applies to the first image and make the first image visible only to the first user group. In some embodiments, the change in privacy settings may take effect only going forward. Continuing the example above, if the first user changes privacy settings and then shares a second image, the second image may be visible only to the first user group, but the first image may remain visible to all users. In some embodiments, in response to a user action to change a privacy setting, the social-networking system may further prompt the user to indicate whether the user wants to apply the changes to the privacy setting retroactively. In some embodiments, a user change to privacy settings may be a one-off change specific to one object. In some embodiments, a user change to privacy may be a global change for all objects associated with the user.

[0165] In some embodiments, the social-networking system may determine that a first user may want to change one or more privacy settings in response to a trigger action associated with the first user. The trigger action may be any suitable action on the online social network. As an example and not by way of limitation, a trigger action may be a change in the relationship between a first and second user of the online social network (e.g., "un-friending" a user, changing the relationship status between the users). In some embodiments, upon determining that a trigger action has occurred, the social-networking system may prompt the first user to change the privacy settings regarding the visibility of objects associated with the first user. The prompt may redirect the first user to a workflow process for editing privacy settings with respect to one or more entities associated with the trigger action. The privacy settings associated with the first user may be changed only in response to an explicit input from the first user, and may not be changed without the approval of the first user. As an example and not by way of limitation, the workflow process may include providing the first user with the current privacy settings with respect to the second user or to a group of users (e.g., un-tagging the first user or second user from particular objects, changing the visibility of particular objects with respect to the second user or group of users), and receiving an indication from the first user to change the privacy settings based on any of the methods described herein, or to keep the existing privacy settings.

[0166] In some embodiments, a user may need to provide verification of a privacy setting before allowing the user to perform particular actions on the online social network, or to provide verification before changing a particular privacy setting. When performing particular actions or changing a particular privacy setting, a prompt may be presented to the user to remind the user of his or her current privacy settings and to ask the user to verify the privacy settings with respect to the particular action. Furthermore, a user may need to provide confirmation, double-confirmation, authentication, or other suitable types of verification before proceeding with the particular action, and the action may not be complete until such verification is provided. As an example and not by way of limitation, a user's default privacy settings may indicate that a person's relationship status is visible to all users (i.e., "public"). However, if the user changes his or her relationship status, the social-networking system may determine that such action may be sensitive and may prompt the userto confirm that his or her relationship status should remain public before proceeding. As another example and not by way of limitation, a user's privacy settings may specify that the user's posts are visible only to friends of the user. However, if the user changes the privacy setting for his or her posts to being public, the social-networking system may prompt the user with a reminder of the user's current privacy settings of posts being visible only to friends, and a warning that this change will make all of the user's past posts visible to the public. The user may then be required to provide a second verification, input authentication credentials, or provide other types of verification before proceeding with the change in privacy settings. In some embodiments, a user may need to provide verification of a privacy setting on a periodic basis. A prompt or reminder may be periodically sent to the user based either on time elapsed or a number of user actions. As an example and not by way of limitation, the social-networking system may send a reminder to the user to confirm his or her privacy settings every six months or after every ten photo posts. In some embodiments, privacy settings may also allow users to control access to the objects or information on a per-request basis. As an example and not by way of limitation, the social-networking system may notify the user whenever a third-party system attempts to access information associated with the user, and require the user to provide verification that access should be allowed before proceeding.

[0167] The preceding description has been provided to enable others skilled in the art to best utilize various aspects of the exemplary embodiments disclosed herein. This exemplary description is not intended to be exhaustive or to be limited to any precise form disclosed. Many modifications and variations are possible without departing from the scope of the present disclosure. The embodiments disclosed herein should be considered in all respects illustrative and not restrictive. Reference may be made to any claims appended hereto and their equivalents in determining the scope of the present disclosure.

[0168] Unless otherwise noted, the terms "connected to" and "coupled to" (and their derivatives), as used in the specification and / or claims, are to be construed as permitting both direct and indirect (i.e., via other elements or components) connection. In addition, the terms "a" or "an," as used in the specification and / or claims, are to be construed as meaning "at least one of." Finally, for ease of use, the terms "including" and "having" (and their derivatives), as used in the specification and / or claims, are interchangeable with and have the same meaning as the word "comprising."

Claims

WHAT IS CLAIMED IS:

1. An apparatus comprising: a frame dimensioned to be worn by a user; a plurality of audio transducers secured to the frame and configured to capture audio data representative of an environment occupied by the user; and circuitry configured to: identify, within the audio data, an audio signal corresponding to a voice of the user and at least one additional audio signal; isolate the audio signal from the at least one additional audio signal; and modify the audio signal or the at least one additional audio signal to compensate for an imbalance between a volume level of the voice of the user and a volume level of the at least one additional audio signal.

2. The apparatus of claim 1, wherein the circuitry is further configured to generate modified audio data for playback of the voice of the user and the additional audio signal by combining the audio signal and the additional audio signal after completion of the modification, in which case optionally wherein the circuitry is further configured to transmit the modified audio data to a playback device that plays the audio data as an audio presentation.

3. The apparatus of claim 1 or 2, wherein the circuitry is further configured to: generate a beamformer for isolating the audio signal based at least in part on an arrangement of the audio transducers secured to the frame; and apply the beamformer to the audio data to: identify the audio signal and the at least one additional audio signal; and isolate the audio signal and the at least one additional audio signal from one another.

4. The apparatus of any preceding claim, wherein the audio transducers comprise: a first microphone configured to capture a first instance of the audio data for a first input channel; and a second microphone configured to capture a second instance of the audio data for a second input channel.

5. The apparatus of claim 4, wherein the circuitry is further configured to:generate a first beamformer for the first input channel based at least in part on a position of the first microphone relative to the user whose voice captured in the audio signal; generate a second beamformer for the second input channel based at least in part on a position of the second microphone relative to the user whose voice captured in the audio signal; apply the first beamformer to the first instance of the audio data to isolate the audio signal and the at least one additional audio signal from one anotherforthe first input channel; and apply the second beamformer to the second instance of the audio data to isolate the audio signal and the at least one additional audio signal from one another for the second input channel.

6. The apparatus of claim 5, wherein the circuitry is further configured to: modify the audio signal or the at least one additional audio signal in the first instance of the audio data to compensate for a first imbalance between the volume level of the voice of the user and the volume level of the at least one additional audio signal in the first input channel; and modify the audio signal or the at least one additional audio signal in the second instance of the audio data to compensate for a second imbalance between the volume level of the voice of the user and the volume level of the at least one additional audio signal in the second input channel, in which case optionally wherein the circuitry is further configured to generate modified audio data for playback of the voice of the user and the additional audio signal by combining the first instance of the audio data and the second instance of the audio data after completion of the modifications.

7. The apparatus of any preceding claim, wherein the circuitry is further configured to attenuate the audio signal to a certain decibel level based at least in part on the volume level of the at least one additional audio signal.

8. The apparatus of any preceding claim, wherein the circuitry is further configured to amplify the at least one additional audio signal to a certain decibel level based at least in part on the volume level of the audio signal.

9. The apparatus of any preceding claim, wherein the circuitry is further configured to balance, based at least in part on the modification, the audio signal and the at least one additional audio signal to a signal-to-noise ratio with a standard deviation between5 decibels and 7 decibels, in which case optionally wherein the signal-to-noise ratio of the audio signal and the at least one additional audio signal is approximately 8.5 decibels.

10. The apparatus of any preceding claim, wherein the circuitry is further configured to: detect the imbalance by comparing the volume level of the audio signal and the volume level of the at least one additional audio signal; determine, based at least in part on the comparison, that the imbalance satisfies a certain threshold; and modify, in response to the determination, the audio signal or the at least one additional audio signal to compensate for the imbalance.

11. A system comprising: an eyewear device that: is dimensioned to be worn by a user; comprises a plurality of audio transducers configured to capture audio data representative of an environment occupied by the user; and circuitry configured to: isolate, within the audio data, an audio signal corresponding to a voice of the user and at least one additional audio signal; and modify the audio signal or the at least one additional audio signal to compensate for an imbalance between a volume level of the voice of the user and a volume level of the at least one additional audio signal; and a playback device configured to play the audio data as an audio presentation after completion of the modification.

12. The system of claim 11, wherein the circuitry is further configured to modify the audio data for playback of the voice of the user and the additional audio signal by combining the audio signal and the additional audio signal, in which case optionally wherein the circuitry is further configured to transmit the audio data to the playback device via a network.

13. The system of claim 11 or 12, wherein the circuitry is further configured to: generate a beamformer for isolating the audio signal based at least in part on an arrangement of the audio transducers secured to the frame; and apply the beamformer to the audio data to:identify the audio signal and the at least one additional audio signal; and isolate the audio signal and the at least one additional audio signal from one another.

14. The system of any one of claims 11 to 13, wherein the audio transducers comprise: a first microphone configured to capture a first instance of the audio data for a first input channel; and a second microphone configured to capture a second instance of the audio data for a first input channel, in which case optionallywherein the circuitry is further configured to: generate a first beamformer for the first input channel based at least in part on a position of the first microphone relative to the user whose voice captured in the audio signal; generate a second beamformer for the second input channel based at least in part on a position of the first microphone relative to the user whose voice captured in the audio signal; apply the first beamformer to the first instance of the audio data to isolate the audio signal and the at least one additional audio signal from one anotherforthe first input channel; apply the second beamformer to the second instance of the audio data to isolate the audio signal and the at least one additional audio signal from one another for the second input channel; modify the audio signal or the at least one additional audio signal in the first instance of the audio data to compensate for a first imbalance between the volume level of the voice of the user and the volume level of the at least one additional audio signal in the first input channel; modify the audio signal or the at least one additional audio signal in the second instance of the audio data to compensate for a second imbalance between the volume level of the voice of the user and the volume level of the at least one additional audio signal in the second input channel; and generate modified audio data for playback of the voice of the user and the additional audio signal by combining the first instance of the audio data and the second instance of the audio data after completion of the modifications.

15. A method comprising: capturing, by a plurality of audio transducers, audio data representative of anenvironment occupied by a user; identifying, by circuitry, an audio signal corresponding to a voice of the user and at least one additional audio signal within the audio data; isolating, by the circuitry, the audio signal from the at least one additional audio signal; and modifying, by the circuitry, the audio signal or the at least one additional audio signal to compensate for an imbalance between a volume level of the voice of the user and a volume level of the at least one additional audio signal.

Citation Information

Patent Citations

  • Headset Interview Mode

    US20150112671A1

  • Hearing assist device employing dynamic processing of voice signals

    US20210345047A1

  • Spatial Capture with Noise Mitigation

    US20240107259A1