Gaze-based audio beamforming
By employing gaze-based beamforming with stored patterns, AR devices effectively enhance audio signals in the user's direction while conserving power and reducing processing demands, addressing the inefficiencies of existing AR audio processing.
Patent Information
- Application Number
- JP2024530011
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2021-11-22
- Filing Date
- 2022-11-17
- Publication Date
- 2025-12-25
- Estimated Expiration
- 2042-11-17
AI Technical Summary
Existing augmented reality (AR) devices face challenges in efficiently processing audio while minimizing power consumption and latency through eye-tracking beamforming, which demands significant processing resources.
Implementing stored beam patterns based on user gaze direction, enabling selective audio beamforming that reduces complexity and conserves power by using eye-tracking to adjust beamforming techniques dynamically.
Enhances audio signals in the user's gaze direction without significantly impacting battery life or processing resources, enabling efficient and responsive audio experiences.
Smart Images

Figure 0007792517000006 
Figure 0007792517000007 
Figure 0007792517000008
Abstract
Description
[Technical Field]
[0001] CROSS-REFERENCE TO RELATED APPLICATIONS This application is a continuation of and claims priority to U.S. Patent Application No. 17 / 456,007, filed November 22, 2021, the disclosure of which is incorporated herein by reference in its entirety.
[0002] The present disclosure relates to augmented reality, and more particularly to an augmented reality device configured to process audio based on eye tracking. [Background technology]
[0003] Head-mounted computing devices (e.g., smart glasses) can be configured with various sensors that enable augmented reality (AR), in which virtual elements are presented along with real elements of the environment. The virtual elements can be presented on a head-up display, so that they appear as if they are in the real world. The head-up display can be implemented with a device similar to glasses (i.e., AR glasses).
[0004] The AR glasses may be configured with eye-tracking sensor(s) that determine the user's gaze direction and / or gaze point as they change over time. The AR glasses may also be configured with multiple microphones operating as a microphone array with a sensitivity pattern having a beam, so that sounds from the beam direction are received with the highest sensitivity of the microphone array. The sound from the microphone array may be processed so that the beams can be steered (i.e., beamformed) in various directions. Summary of the Invention
[0005] In at least one aspect, the present disclosure generally describes a method. The method includes receiving audio channels from a plurality of microphones configured to operate as a microphone array of an augmented reality (AR) device. The method further includes tracking the eyes of a user of the AR device to determine the user's gaze direction. The method further includes selecting a beam pattern for the microphone array from a set of stored beam patterns based on the user's gaze direction. The method further includes generating a beamformed audio signal based on the selected beam pattern, and transmitting the beamformed audio signal to a speaker (e.g., a built-in speaker, a paired speaker) of the AR device to play the beamformed audio signal for the user.
[0006] In another aspect, the present disclosure generally describes an AR device, such as smart glasses. The AR device may include components, specifically a microphone array, an eye tracker, a speaker, and a processor configured to implement the proposed method. For example, the proposed smart glasses include a microphone array having microphones configured to generate audio channels based on sounds from the environment. The smart glasses further include an eye tracker configured to determine a user's gaze direction. The smart glasses further include a speaker and a processor. The processor is configured with software to receive the audio channels from the microphone array and the gaze direction from the eye tracker. The processor is further configured to retrieve weights for the audio channels from a lookup table (or another type of data memory, specifically another type of database, array, or table) based on the gaze direction. The processor is further configured to apply the weights to the channels and sum the channels to generate a beamformed audio signal that amplifies sounds in the environment from the gaze direction. The processor is further configured to transmit the beamformed audio signal to a speaker for playback to the user.
[0007] The foregoing summary, as well as other exemplary objects and / or advantages of the present disclosure, and the manner in which they are achieved, are further explained in the following detailed description and accompanying drawings. [Brief explanation of the drawings]
[0008] [Figure 1] 10 is a sensitivity plot of a microphone array after beamforming, according to a possible embodiment of the present disclosure. [Figure 2] FIG. 1 illustrates a view through AR glasses including a measured viewpoint, according to a possible embodiment of the present disclosure. [Figure 3] 3A and 3B are diagrams illustrating beamforming based on gaze direction according to an embodiment of the present disclosure. [Figure 4] 1 is a perspective view of AR glasses according to a possible embodiment of the present disclosure. FIG. [Figure 5] 1 illustrates a beamforming process according to a possible embodiment of the present disclosure. [Figure 6] 6A and 6B are diagrams illustrating possible processing of beamforming audio according to embodiments of the present disclosure. [Figure 7] 1 is a flowchart of a method for eye-tracking audio beamforming, according to an embodiment of the present disclosure. [Figure 8] FIG. 10 illustrates retrieving a beam pattern from a database based on line of sight, according to an embodiment of the present disclosure. [Figure 9] 1 is a flowchart of a method for selecting a beam pattern from a database based on a line of sight, according to an embodiment of the present disclosure. [Figure 10] FIG. 1 illustrates a flowchart of a method for generating a beam pattern according to an embodiment of the present disclosure. [Figure 11] 1 is a flowchart for detecting line of sight for beamforming, according to an embodiment of the present disclosure. [Figure 12] FIG. 1 illustrates beamforming zoom-in and zoom-out according to an embodiment of the present disclosure. DETAILED DESCRIPTION OF THE INVENTION
[0009] The elements in the drawings are not necessarily drawn to scale relative to each other. Like reference numerals indicate corresponding parts throughout the several views.
[0010] The present disclosure describes audio beamforming (i.e., beamforming) of a microphone array in an augmented reality (AR) device, such as smart glasses (e.g., AR glasses), where the beamforming is based at least in part on the position(s) of a user's eye(s) (i.e., eye-tracking beamforming). A technical problem with eye-tracking beamforming relates to the demands placed on the power / processing resources of the AR glasses. To be effective, eye tracking and beamforming need not consume too much power (i.e., extend battery life) and be responsive (i.e., avoid significant latency). The present disclosure provides systems and methods for eye-tracking beamforming that are based on techniques that reduce complexity and increase processing / power efficiency. The disclosed techniques may have the technical effect of automatically enhancing signals in the user's gaze direction without significantly impacting the battery life or processing resources of the AR glasses.
[0011] The power / processing efficiency of the disclosed eye-tracking beamforming techniques may result from several different aspects. First, the disclosed eye-tracking beamforming may rely, at least in part, on stored beam patterns that may be acquired and applied based on the user's line of sight. Second, the eye-tracking beamforming may be configured to be enabled / disabled in certain situations so that it is not always operational.
[0012] The technical effects of the disclosed eye-tracking beamforming techniques may enable new audio applications. For example, this disclosure further describes implementations in which audio beamforming can zoom in or out in the direction of gaze to enhance a user's audio experience.
[0013] Beamforming (i.e., beamsteering) is a signal processing technique in which multiple audio channels can be processed (e.g., filtered, delayed, phase-shifted) to generate a beamformed audio signal in which audio from various directions can be strengthened (i.e., amplified) or weakened (i.e., attenuated). For example, a first microphone and a second microphone can be spatially separated by a certain distance along the array direction. The spatial separation distance and the direction of sound (relative to the array direction) can result in an interaural delay between a first audio stream at the first microphone and a second audio stream at the second microphone. Beamforming can include further delaying one of the audio streams by a beamforming delay, such that after beamforming, the first audio stream and the second audio stream are phase-shifted by the interaural delay and the beamforming delay. The phase-shifted audio streams are then combined (e.g., summed) to create the beamformed audio. By adjusting the beamforming delay relative to the interaural delay, sounds from particular directions can be adjusted (e.g., canceled, attenuated, or enhanced) by the summation process. For example, a pure sine wave received by the first and second microphones can be completely canceled for a particular direction if, after the interaural delay and beamforming delay, the sine wave versions at the combiner are 180 degrees out of phase. Alternatively, if, after the interaural delay and beamforming delay, the sine wave versions at the combiner are in phase (i.e., phase difference is 0 degrees), the sine wave versions at the combiner can be enhanced.
[0014] Multiple audio channels may be captured (i.e., collected) by an array of microphones (i.e., a microphone array). Each microphone in a microphone array may be of the same type or of a different type. For example, all microphones in a microphone array may be omnidirectional. The microphones may be spaced (e.g., equally spaced) in one, two, or three dimensions. A one-dimensional microphone array may be capable of beam steering in one dimension, while a two-dimensional microphone array may be capable of beam steering in either or both of the two dimensions. The number and spacing of microphones in a microphone array may correspond to the beamwidth (i.e., directivity, focus, angular range) of the beam.
[0015] FIG. 1 shows a polar plot of the sensitivity of a microphone array 101 after beamforming. The pattern of sensitivity is known as the microphone array's "beam pattern." The microphone array's beam pattern has a beam 110 of maximum sensitivity in a beam direction 120, with beamwidths 130 of specific sensitivity (e.g., -3 decibels (dB)) on either side of the beam direction 120. The beamformed sound generated from the beam pattern reinforces (e.g., amplifies) sounds from sound sources in the beam direction 120 and suppresses (e.g., attenuates) sounds from sound sources not in the beam direction. In other words, a first sound 104 coming from a direction aligned with the beam direction 120 will sound louder to a listener than a second sound 105 coming from a direction inconsistent with the beam direction 120.
[0016] The spatially selective enhancement / suppression resulting from beamforming can help users distinguish between sounds (e.g., in noisy environments). Additionally (or alternatively), beamforming can improve the accuracy of other computer-assisted speech applications (e.g., speech recognition, speech-to-text (VTT), language translation, etc.). Furthermore, beamforming can improve privacy because other sounds received from directions other than the speech direction (e.g., nearby conversations) can be amplified much less than the speech sound. Controlling beamforming based on the listener's intent, which can be determined by tracking the listener's eyes, can improve the versatility of these applications.
[0017] Eye-tracking beamforming involves adjusting the processing (e.g., filtering, delaying, phase-shifting) of multiple audio channels from a microphone array according to a user's eye(s) to generate a beam with a beam direction that closely matches (e.g., exactly matches) the user's line of sight. The user's line of sight may include the direction in which the user is looking (i.e., gaze direction). Determining the gaze direction (e.g., gaze(θ), gaze(φ,θ)) may include determining a viewpoint within the field of view from which the user is looking (e.g., gaze(x,y)).
[0018] Gaze may be determined by tracking the user's eye(s). One possible method of gaze tracking involves using a camera to measure eye metrics and determine eye position. In one possible implementation, pupil position relative to a light pattern (near-infrared light) projected onto the eye may be measured by analyzing high-resolution images of the eye and the pattern. The eye position may then be applied to a machine learning model to determine gaze. Variations on this method that do not use a projected pattern are also possible. For example, standard glint-based tracking or convolutional neural net methods can convert a two-dimensional (2D) infrared image captured by a camera pointed at the eye (or an image of the eye reflected from a mirror) into coordinates (x, y) within the field of view of the AR glasses, as shown in FIG. 2.
[0019] 2 illustrates a field of view 205 through AR glasses 201, including a measured viewpoint 210. The viewpoint may correspond to a gaze direction relative to the AR glasses. A viewpoint remaining stationary for a certain period of time may indicate user interest in the corresponding viewing region. This interest may trigger beamforming of a gaze direction corresponding to the viewpoint. For example, the gaze direction may be based on a projection, with respect to the coordinate system of the AR glasses, intersecting the viewpoint on the viewing surface.
[0020] 3A and 3B illustrate beamforming based on gaze direction. As indicated by the dotted arrows, the gaze direction can be determined based on the user's eye position. FIG. 3A shows a first beam pattern 321 along a first gaze direction 331 of a user 301, and FIG. 3B shows a second beam pattern 322 along a second gaze direction 332 of a user 302. The user's gaze can change over time. Thus, the first beam pattern 321 can be used at a first time, and the second beam pattern 322 can be used at a second time. The user's gaze can be correlated or uncorrelated with the user's head position. In some implementations, the user's head position and gaze direction can be combined when determining the beam direction.
[0021] FIG. 4 is a perspective view of AR glasses according to a possible embodiment of the present disclosure. The AR glasses 400 are configured to be worn on a user's head and face. The AR glasses 400 include a right earpiece 401 and a left earpiece 402 that are supported by the user's ears. The AR glasses further include a bridge portion 403 that is supported by the user's nose so that a left lens 404 and a right lens 405 can be positioned in front of the user's left eye and right eye, respectively. The portions of the AR glasses may collectively be referred to as the frame of the AR glasses. The frame of the AR glasses may include electronics that enable functionality. For example, the frame may include a battery, a processor, memory (e.g., a non-transitory computer-readable medium), and electronics that support sensors (e.g., a camera, a depth sensor, etc.), and an interface device (e.g., a speaker, a display, a network adapter, etc.).
[0022] The AR glasses 400 may include an FOV camera 410 (e.g., an RGB camera) oriented with a camera field of view that overlaps the natural field of view of the user's eyes when the glasses are worn. In possible implementations, the AR glasses may also include a depth sensor 411 (e.g., a LIDAR camera, structured light camera, time-of-flight camera, depth camera) oriented with a depth sensor field of view that overlaps the natural field of view of the user's eyes when the glasses are worn. Data from the depth sensor 411 and / or the FOV camera 410 may be used to measure depth within the user's (i.e., wearer's) field of view (i.e., region of interest). In possible implementations, the camera field of view and the depth sensor field of view may be calibrated so that the depth (i.e., range) of objects in images from the FOV camera 410 can be determined, with the depth measured between the object and the AR glasses.
[0023] The AR glasses 400 may further include a display 415. The display may present AR data (e.g., images, graphics, text, icons, etc.) on a portion of the lens(es) of the AR glasses such that the user can view the AR data when looking through the lenses of the AR glasses. In this way, the AR data may be overlaid with the user's view of the environment.
[0024] The AR glasses 400 may further include an eye tracking sensor. The eye tracking sensor may include a right-eye camera 420 and a left-eye camera 421. The right-eye camera 420 and the left-eye camera 421 may be positioned in the lens portion of the frame such that, when the AR glasses are worn, the right FOV 422 of the right-eye camera includes the user's right eye, and the left FOV 423 of the left-eye camera includes the user's left eye. The viewpoint (x, y) may be measured at the frequency of the video feed of the camera (e.g., right-eye camera 420, left-eye camera 421). For example, the gaze coordinates (x, y) may be measured at a frame rate (e.g., 15 frames / second) or less of the camera.
[0025] The AR glasses 400 may further include multiple microphones (i.e., two or more microphones). The multiple microphones may be arranged at intervals on the frame of the AR glasses. As shown in FIG. 4 , the multiple microphones may include a first microphone 431, a second microphone 432, a third microphone 433, a fourth microphone 434, and a fifth microphone 435. The multiple microphones may be configured to operate together as a microphone array that directs beams in a specific direction with respect to a coordinate system 430 of the AR glasses 400. Alternatively, the microphones may be configured to operate in groups (i.e., subarrays), with each group configured to operate as a microphone array. In one example, the third microphone 433 and the fourth microphone 434 may be configured to operate as a (horizontal) microphone array along the X direction of the coordinate system 430 of the AR glasses 400. In other words, the third microphone 433 and the fourth microphone 434 may be used for horizontal beamforming. Furthermore, the third microphone 433 and the fifth microphone 435 may be configured to operate as a (vertical) microphone array along the Y direction of the coordinate system 430 of the AR glasses 400. In other words, the third microphone 433 and the fifth microphone 435 may be used for vertical beamforming. Furthermore, the first microphone 431 and the second microphone 432 may be configured to operate as a (horizontal) microphone array along the Z direction of the coordinate system 430 of the AR glasses 400. The number of microphones used for each subarray may be two or more. Increasing the number of microphones in a subarray can narrow the beamwidth of the beam pattern resulting from the subarray. Because the beamforming process of the subarray can be parallelized, two or more subarrays may share one or more microphones.
[0026] The AR glasses 400 may further include a left speaker 441 and a right speaker 442 configured to transmit sound (e.g., beamformed sound) to the user. Additionally or alternatively, transmitting sound to the user may include transmitting sound to a listening device (e.g., hearing aid, earphones, etc.) via a wireless communication link 445. For example, the AR glasses may transmit sound (e.g., beamformed sound) to a left wireless earphone 446 and a right wireless earphone 447. Wireless The AR device may transmit the audio to the earphones 447. In the case of beamforming audio that tracks the user's viewpoint (x, y), sounds in the audio from a field of view that includes the viewpoint may be amplified, while sounds from other field of view may not be amplified or may be attenuated. In other words, the speaker of the AR device may include a speaker (e.g., earphones) communicatively connected (i.e., paired) with the AR glasses 400 or a speaker integrated (i.e., built-in) into the AR glasses 400.
[0027]
number
[0028]
number
[0029] Beamforming may be parallelized by orientation, such that arrays positioned horizontally with respect to the coordinate system (i.e., horizontal arrays) have a first set of weights, while arrays positioned vertically with respect to the coordinate system (i.e., vertical arrays) have a second set of weights. Audio from each directional array may be processed separately to generate horizontal and vertical beamforming signals. The horizontal and vertical beamforming signals may be averaged to form directional beamforming audio that includes a horizontal component (e.g., x) and a vertical component (e.g., y). While this parallel processing approach may have the advantage of simplicity, other approaches may be possible. For example, it may be possible to determine beamforming weights in both the horizontal and vertical directions so that an additional averaging step is not required. Furthermore, adding a third dimension (e.g., z) to the aforementioned steps may enable three-dimensional (3D) beamforming.
[0030]
number
[0031]
number
[0032] 7 is a flowchart of a method for eye-tracking audio beamforming according to an embodiment of the present disclosure. Method 700 may be implemented as a computer program product tangibly embodied on a non-transitory computer-readable medium. In other words, the steps of method 700 may be included as part of a computer program (i.e., module, software, application, code) that may be implemented in a programming language or machine language. When executed, the computer program may configure at least one processor to perform the method for eye-tracking audio beamforming.
[0033] The method 700 for eye-tracking audio beamforming includes capturing 705 audio from a plurality of microphones (i.e., a microphone array). In one possible implementation, each microphone in the microphone array has an omnidirectional sensitivity pattern. In another possible implementation, one or more of the microphones in the microphone array have a directional sensitivity pattern.
[0034] The microphones may be integrated with the AR glasses. In a possible implementation, the AR glasses may be configured in a beamforming mode (i.e., with beamforming) or a normal mode (i.e., without beamforming). In the beamforming mode, audio from the microphone array may be processed to steer the sensitivity of the microphone array in a direction corresponding to the user's gaze. The selection of the mode may depend on various factors. For example, whether to perform beamforming may be determined based on the processing and power resources available in the AR glasses. Specifically, when the AR glasses are in a low power mode (e.g., a power level below 25%), eye tracking may be avoided. Therefore, the method 700 may optionally include determining (710) whether the device is in a beamforming mode. When the AR device is not in a beamforming mode, audio from one or more of the microphones may be provided to the user (745). However, when the AR glasses are in a beamforming mode, steps may be performed to perform eye tracking audio beamforming. In some implementations, the AR glasses are configured to automatically beamform audio when the user's line of sight meets the criteria(s). In these implementations, determining whether the device is in beamforming mode (710) may be omitted.
[0035] The method 700 includes tracking 715 the user's eye(s). Results of the eye tracking may be used to detect gaze 720. If gaze is detected, audio may be beamformed and provided to the user as beamformed audio; otherwise, audio may be provided to the user without beamforming 745. Details regarding determining when to beamform and when not to beamform based on gaze are discussed further below (see, e.g., FIG. 11).
[0036] After detecting the gaze, a gaze direction may be determined (725). As described above (e.g., FIG. 2), determining the user's gaze may include determining a gaze point from the measured position of one or both eyes. The position may be measured from a captured image of the eye(s). In some implementations, the method may further include collecting additional information to facilitate confirmation and / or improvement (730) (i.e., adjustment) of the gaze direction. For example, an image 727 of the user's field of view may be captured and analyzed at the determined gaze direction. The analysis may include searching for known sound sources within the image 727. For example, if a person is speaking in the gaze direction, the gaze direction may be confirmed to be toward the person speaking. If the person speaking is in a direction close to the determined gaze direction (e.g., within ±10 degrees) but does not exactly match the gaze direction, an adjustment may be made to align the gaze direction with the person speaking.
[0037] After the gaze direction is determined, method 700 includes selecting 735 a beam pattern according to the gaze direction. The beam pattern may be selected from a plurality of beam patterns stored in memory. The memory may be local memory of the AR glasses or may be memory available over a network communicatively connected to the AR glasses. For example, the beam patterns may be stored in a lookup table or database 737 that can be queried using (at least) the gaze direction.
[0038] 8 illustrates selecting and retrieving a beam pattern from a database based on the line of sight, according to an embodiment of the present disclosure. A database or lookup table may be queried using the line of sight direction to return a set of weights (w1, w2, ... wn) to provide a beam in the line of sight direction (or a direction close to the line of sight direction). As shown, the database includes multiple beam patterns (BP1, BP2, BP3, BPn) in multiple directions (d1, d2, d3, ... dn), each with a beamwidth (bw1, bw2, bw3, ... bwn).
[0039] The returned set of weights (w1, w2, ... wn) may have each weight corresponding to a microphone. The beamwidth of the beampattern may correspond to the number of microphones in the array. Thus, the stored beampatterns may include different numbers of weights to provide different beamwidths. For example, two beampatterns with the same direction but different beamwidths may have different numbers of weights. Alternatively, two beampatterns with the same direction but different beamwidths may have the same number of weights, but one of the beampatterns may have some of the weights with zero values. If a weight has a zero value, the microphone corresponding to the weight may actually be turned off.
[0040] Optionally, the selection (query) of the database or lookup table of beam patterns may further include the mode / metric of the device. For example, a device may be capable of various microphone configurations, and the selection may be made based on the device's particular microphone configuration. Specifically, based on the device's mode / metric, only horizontal beam patterns may be selected. Alternatively, some microphones may be disabled (e.g., based on power state). This disabling may result in only beam patterns having beamwidths (i.e., number of weights) corresponding to the number of enabled microphones based on the device's mode / metric. Further details regarding the selection of beam patterns based on line of sight are discussed further below (e.g., see FIG. 9).
[0041] Returning to FIG. 7, after the beam pattern is selected (735), the method 700 further includes generating beamformed audio based on the selected beam pattern (740) and providing the beamformed audio to a user (745). Generating the beamformed audio may be performed as described above (e.g., see FIG. 5). For example, generating the beamformed audio signal based on the selected beam pattern may include obtaining a weight set corresponding to the selected beam pattern. Each audio channel is then applied to a corresponding weight in the weight set to generate weighted audio channels, which are summed to generate the beamformed audio signal.
[0042] FIG. 9 is a flowchart of a method for selecting (i.e., obtaining) a beam pattern from the database of FIG. 8 based on a line of sight, according to an embodiment of the present disclosure. Method 900 includes comparing a determined line of sight (905) with beam directions of a stored set of beam patterns. If a beam pattern for the line of sight is found (910), weights for that beam pattern may be obtained (915) and used to generate beamformed audio (see, e.g., FIG. 5). If a beam pattern for the line of sight is not found, the search of the database (or lookup table) may be expanded to include beam patterns for directions within an angular range (e.g., ±10 degrees) around the line of sight. If multiple beam patterns are found within the range (920), the beam pattern for the direction closest to the line of sight may be selected. Weights for the closest beam pattern may be obtained (925) and used for beamforming. Although the closest weights may result in a mismatch between the beam and the object of interest, the resulting beamforming may still provide enhanced (i.e., amplified) audio from the object of interest because the object of interest may still be within the beamwidth of the beam (see, e.g., FIG. 1). If no beam pattern is found in the line of sight direction and in a range of directions surrounding the line of sight, a default beam pattern may be obtained (930) and used for beamforming. The default beam pattern may have weights that steer the beam in a direction (e.g., z direction) in front of the user (e.g., azimuth angle 0 degrees, elevation angle 0 degrees with respect to the AR glasses' coordinate system 430). Alternatively, the default pattern may have weight values of zero for all microphones except one or two of the microphones in the array (e.g., stereo L / R). If the weight values are zero, the corresponding microphones may actually be disabled (see, e.g., FIG. 5), while if the weight values are non-zero, the gain is one and no phase shift may occur. In one possible implementation, the default pattern may allow the left and right microphones to operate as a stereo pair of microphones with no phase shift other than the interaural delay caused by the spacing of the microphones on the AR glasses.In another possible implementation, the default pattern may average the audio from the microphones by applying equal weights (eg, 1) to each channel.
[0043] The stored beam patterns may be generated based on training that is performed before (i.e., offline) beamforming is used in operation (i.e., online, at runtime). Figure 10 shows a flowchart of a method for generating beam patterns according to an embodiment of the present disclosure. The method includes determining a line of sight direction (1005).
[0044] A gaze direction may be determined based on the popularity of the gaze direction over time. A popular gaze direction may be determined based on gazes monitored over time for one or more users. For example, a user's eyes may be tracked over time to determine the probability of various gaze directions or gaze points (see, for example, FIG. 2). For example, a gaze point may be determined when the gaze stays at the gaze point for longer than a certain period of time. A gaze probability map may be generated based on the gaze points collected during a training period. For example, each possible gaze point (x, y) may have a large number of gaze points collected during the training period. The probability (i.e., likelihood) of that gaze point may be the number of gaze points of the gaze point collected during the training period divided by the total number of gaze points detected during the training period. The probability map may be implemented as a heat map image in which the intensity of more popular gaze points is higher than the intensity of less popular gaze points. The probability map may be analyzed to determine a set of gaze directions. For example, one or more gaze points (i.e., pixels) in an area of the heat map that have an intensity higher than a threshold may be highlighted (i.e., selected) as popular gaze points, and gaze directions toward the popular gaze points may be determined.
[0045] The method further includes selecting 1010 a first gaze direction from the (set of) gaze directions. A target beam pattern for the selected gaze direction may be determined 1015. Determining the target beam pattern may include determining a beam width appropriate for the particular gaze direction. For example, a wide beam width may be selected so that a single beam pattern can cover a range of popular gaze directions (i.e., a region of popular viewpoints). Once the target beam pattern is selected, weights for that beam pattern may be calculated. Calculating the weights may include collecting 1020 audio from multiple directions and optimizing the following equation according to a least-squares optimization process:
[0046]
number
[0047] In the above equation, y is the target beam pattern (e.g., a two-dimensional matrix with sensitivity values corresponding to the beam pattern), X is the audio from multiple angles (e.g., a full-rank pseudo-inverse data matrix), and w is the weights to solve for (e.g., a vector of weights corresponding to the number of audio channels). Inversion to learn weights for a particular gaze direction may be possible when the forward matrix (X) is full rank.
[0048] A practical (offline) setup for collecting audio may involve moving a sound source around the AR glasses while recording audio from each channel so that the same audio data can be collected from the sound source at various angles. An optimization may then be performed that tries different weights for each channel until the audio has a spatial sensitivity pattern that corresponds to that of the target beam pattern. This results in a set of weights that approximates the target beam pattern. The quality of the approximation may be based on the number of weights (i.e., microphones). For example, increasing the number of weights so that the least-squares optimization process is minimized closer to zero may result in a better match with the target beam pattern.
[0049] Returning to FIG. 10 , after the optimal weights for the target beam patterns are calculated (1025), they may be stored (1030) in a database (or lookup table). The weights may be indexed in the database by their corresponding beam direction and / or beam width. The method may be repeated for other gaze directions by selecting (1035) the next gaze direction from the gaze directions, determining the target beam pattern for the next gaze direction, and repeating the optimization process to calculate / store the weights for the next beam pattern in the database. After all beam patterns for all determined gaze directions are stored in the database, they may be downloaded to the local memory of the AR glasses or stored in cloud memory accessible online by the AR glasses.
[0050] The stored beam patterns and lookup approach to beamforming is very computationally efficient, power efficient, and fast because no optimization needs to be done on the AR glasses while the user is using them (i.e., at runtime). At runtime, beamforming can be done by simply recalling weights from a database. While the weights may not provide a beam pattern that perfectly matches the user's line of sight, they can often enhance the audio sufficiently to allow the user to better hear what they are looking at.
[0051] To determine when beamforming should occur (e.g., item 720 of FIG. 7 ), identifying that a user's gaze remains stationary from a viewpoint can be utilized offline (e.g., through training as described above) and / or online. FIG. 11 shows a flowchart for detecting gaze for beamforming according to an embodiment of the present disclosure. The method includes identifying (1105) a viewpoint (i.e., gaze coordinates) (x, y). Identifying gaze can be challenging due to rapid eye movements (i.e., saccades). Therefore, the method further includes temporally filtering (1110) the gaze coordinates. For example, gaze coordinates obtained from real-time eye tracking can be low-pass filtered to generate a time-varying signal corresponding to gaze with little change over time. In possible embodiments, eye-tracking coordinates can be measured and averaged over time to obtain average eye-tracking coordinates. If the average eye-tracking coordinates meet a dwell time criterion, the user's gaze direction can be identified. For example, average eye-tracking coordinates that remain within a range (e.g., area) for longer than a threshold time can indicate a stable gaze. Beamforming may be determined based on line of sight stability. For example, if the line of sight is stable (1115), Beamforming may be performed in a direction corresponding to the gaze determined from the filtered gaze coordinates and may terminate in a direction corresponding to the average eye-tracking coordinates. A stable line of sight (1115) may trigger the AR glasses to enter beamforming mode (1120), while an unsteady line of sight (i.e., an unstable line of sight) will not trigger beamforming (i.e., no beamforming (1125)). For example, in possible implementations, the AR glasses will not be set into beamforming mode unless a stable line of sight is detected.
[0052] The computational efficiency, power efficiency, and speed of stored beam patterns and lookup techniques can enable new beamforming applications. For example, beamforming can be gradually focused (zoomed in) or defocused (zoomed out) in a direction. In other words, beamforming can be changed over time. Rather than enabling beamforming on / off like a switch, zooming beamforming corresponds to increasing / decreasing beamforming over time, creating an "audio zoom" experience on the object the user is visually focusing on. Combined with dwell time gaze detection as described above, case, Stepwise beamforming (i.e., beamforming zoom) To or from focused beamforming Smooth audio transitions may be possible.
[0053] 12 illustrates zooming in and out of beamforming according to an embodiment of the present disclosure. The figure shows three beam patterns. A first beam pattern (BP1) has a first beam width, a second beam pattern (BP2) has a second beam width smaller than the first beam width, and a third beam pattern (BP3) has a beam width smaller than the second beam width. The first beam pattern (BP1), the second beam pattern (BP2), and the third beam pattern (BP3) are oriented in the same direction (d1). The beam patterns can be stored in a database and applied sequentially (i.e., one at a time) to gradually change the beamforming (i.e., zoom in or out).
[0054] FIG. 12 includes a time graph 1210 of possible zoom-in and zoom-out sequences, displayed below the beam patterns. At a first time point 1215, audio is provided to the user without beamforming (i.e., no beamforming). Beamforming is triggered at a second time point 1220. Once beamforming is triggered (e.g., by a sustained line of sight), beamforming can be zoomed in by applying a sequence that transitions from no beamforming to full beamforming, for example, by applying a sequence of at least two beam patterns with decreasing beamwidths for zooming in. FIG. 12 illustrates one possible sequence: At the second time point 1220, a first beam pattern (BP1) is obtained (e.g., from a database) and applied to the audio channel. Then, at a third time point 1230, a second beam pattern (BP2) is obtained and applied to the audio channel. Then, at a fourth time point 1240, a third beam pattern (BP3) is obtained and applied to the audio channel. At a fourth point in time 1240, beamforming may be fully on and remain in this state based on the user's line of sight. At a fifth point in time 1250, beamforming may end (e.g., the user's line of sight changes). At the fifth point in time 1250, beamforming may zoom out by transitioning from the third beam pattern (BP3) to the second beam pattern (BP2). Next, at a sixth point in time 1260, beamforming transitions from the second beam pattern (BP2) to the first beam pattern (BP1). Finally, at a seventh point in time 1270, beamforming transitions back to no beamforming (no BF). The duration and beam width of the sequences may be adjustable (e.g., by the user). Additionally, some sequences may include more (or fewer) beam patterns.
[0055] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art. Methods and materials similar or equivalent to those described herein can be used in the practice or testing of the present disclosure. As used in this specification and the appended claims, the singular forms "a," "an," and "the" include plural referents unless the context clearly dictates otherwise. As used herein, the term "comprising" and variations thereof are used synonymously with the term "including" and variations thereof and are open-ended terms. As used herein, the terms "optional" or "optionally" mean that the subsequently described feature, event, or circumstance may or may not occur, and that the description includes both instances where the feature, event, or circumstance occurs. Ranges may be expressed herein, such as from "about" one particular value and / or to "about" another particular value. When such a range is expressed, the aspect includes from the one particular value and / or to the other particular value. Similarly, when values are expressed as approximations, by use of the antecedent "about," it will be understood that the particular value forms another aspect. It will be further understood that the endpoints of each of the ranges are understood both relative to the other endpoint, and independent of the other endpoint.
[0056] As described herein, while certain features of the described embodiments have been illustrated, numerous modifications, substitutions, changes, and equivalents will occur to those skilled in the art. It is therefore to be understood that the appended claims are intended to cover all such modifications and variations that fall within the scope of the embodiments. It should be understood that these have been presented by way of example only, and not limitation, and that various changes in form and detail may be made. Any portion of the apparatus and / or methods described herein may be combined in any combination except in mutually exclusive combinations. The embodiments described herein may include various combinations and / or subcombinations of the functions, components, and / or features of the different described embodiments.
[0057] When the foregoing description refers to an element being on, connected to, electrically connected to, coupled to, or electrically coupled to another element, it will be understood that the element may be directly on, directly connected to, or directly coupled to another element, or that one or more intervening elements may be present. In contrast, when an element is referred to as being directly on, directly connected to, or directly coupled to another element, there are no intervening elements present. Throughout the detailed description, elements that are shown to be directly on, directly connected, or directly coupled may be so referenced, even if the terms directly on, directly connected, or directly coupled are not used. The claims of this application may be amended to recite the exemplary relationships, if any, described in the specification or shown in the drawings.
[0058] As used herein, the singular can include the plural unless the context clearly indicates otherwise. Spatially relative terms (e.g., above, above, top of, below, below, below, and below) are intended to encompass various orientations of the device in use or operation in addition to the orientation shown in the figures. In some embodiments, the relative terms "above" and "below" can include vertically above and vertically below, respectively. In some embodiments, the term "adjacent" can include laterally adjacent or horizontally adjacent.
Claims
1. receiving audio channels from a plurality of microphones of an augmented reality (AR) device, the plurality of microphones being configured to operate as a microphone array; capturing an eye position of a user of the AR device to identify a gaze direction of the user; selecting a beam pattern for the microphone array, the beam pattern being selected from a set of stored beam patterns based on the gaze direction of the user and a power state of the AR device; generating a beamformed audio signal based on the selected beam pattern; transmitting the beamformed audio signal to a speaker of the AR device for playing the beamformed audio signal to the user; A method comprising:
2. The method described in claim 1, wherein the generation of the beamforming audio signal is triggered by the user's viewpoint remaining stationary for a specific period of time.
3. generating the beamformed audio signal based on the selected beam pattern, obtaining a weight set corresponding to the selected beam pattern; applying a corresponding weight from the set of weights to each audio channel to generate weighted audio channels; summing the weighted audio channels to generate the beamformed audio signal, the audio being amplified according to the beam pattern; The method of claim 1 , comprising:
4. The method of any one of claims 1 to 3, wherein each beam pattern in the set of stored beam patterns has a beam direction and a beam width.
5. Selecting the beam pattern from the set of stored beam patterns based on the line of sight direction of the user includes: comparing the line of sight direction to a beam direction of each beam pattern in the set of stored beam patterns; obtaining a beam pattern from the set of stored beam patterns having a beam direction closest to the line of sight direction; The method of claim 4, comprising:
6. receiving audio channels from a plurality of microphones of an augmented reality (AR) device, the plurality of microphones configured to operate as a microphone array; capturing an eye position of a user of the AR device to identify a gaze direction of the user; selecting a beam pattern for the microphone array, the beam pattern being selected from a set of stored beam patterns based on the gaze direction of the user; generating a beamformed audio signal based on the selected beam pattern; transmitting the beamformed audio signal to a speaker of the AR device for playing the beamformed audio signal to the user; Including, each beam pattern in the set of stored beam patterns has a beam direction and a beam width; Selecting the beam pattern from the set of stored beam patterns based on the line of sight direction of the user includes: comparing the line of sight direction to a beam direction of each beam pattern in the set of stored beam patterns; obtaining a beam pattern from the set of stored beam patterns having a beam direction closest to the line of sight direction; Including, selecting a beam pattern from the set of stored beam patterns based on the beam width if a plurality of beam patterns have the beam direction closest to the line of sight direction; The method further comprises:
7. determining a power state of the AR device; selecting a beam pattern having a beamwidth based on the power condition; The method of claim 6 further comprising:
8. zooming in the beam patterns in the line of sight direction by generating a beamforming audio signal based on a sequence of selected beam patterns having the beam direction, wherein each successive beam pattern in the sequence has a smaller beamwidth; The method of claim 6 further comprising:
9. zooming out the beam patterns in the line of sight direction by generating a beamforming audio signal based on a sequence of selected beam patterns having the beam direction, wherein each successive beam pattern in the sequence has a larger beamwidth; The method of claim 6 further comprising:
10. transmitting the beamformed audio signal to a speaker of the AR device for playback to the user; Splitting the beamformed audio signal into a left channel and a right channel; adjusting a phase and amplitude difference between the left and right channels based on the gaze direction of the user; The method according to any one of claims 1 to 3, comprising:
11. determining a target beam pattern based on a training experiment; calculating weights for a beam pattern that approximates the target beam pattern; storing the weights of the beam pattern in a memory as a first beam pattern having a first line of sight direction and a first beamwidth; repeating the determining, calculating, and storing for other target beam patterns to generate the set of stored beam patterns; The method of any one of claims 1 to 3, further comprising:
12. The method of claim 11 , wherein the training experiments include gaze direction likelihoods.
13. Calculating weights for a beam pattern that approximates the target beam pattern includes: performing a least squares optimization; The method of claim 11 , comprising:
14. Capturing an eye position of a user of the AR device to identify a gaze direction of the user includes: measuring target coordinates over time; averaging the target coordinates over time to obtain average target coordinates; determining the gaze direction of the user based on the average target coordinates if the average target coordinates are within range of each other during a dwell time; The method according to any one of claims 1 to 3, comprising:
15. acquiring an image of the user's field of view; analyzing the image within a region based on the gaze direction; ascertaining the gaze direction based on the analysis; and triggering beamforming when the gaze direction is determined, the beamforming including selecting a beam pattern from a set of stored beam patterns based on the gaze direction of the user; 15. The method of claim 14, further comprising:
16. 1. An augmented reality (AR) device, comprising: a microphone array including microphones configured to generate audio channels based on sounds from the environment; an eye position meter configured to determine a gaze direction of a user of the AR device; A speaker and a software-configured processor; The processor, by the software, receiving the audio channel from the microphone array; receiving the gaze direction from the eye position measurement device; selecting a beam pattern for the microphone array from a set of stored beam patterns based on the gaze direction of the user and a power state of the AR device; generating a beamformed audio signal based on the selected beam pattern; transmitting the beamformed audio signal to the speaker for playback of the beamformed audio signal to the user; The AR device is configured to:
17. The AR device of claim 16, wherein generation of the beamforming audio signal is triggered by the user's viewpoint remaining stationary for a specific period of time.
18. Smart glasses, a microphone array including microphones configured to generate audio channels based on sounds from the environment; an eye position meter configured to determine a gaze direction of a user of the smart glasses; A speaker and a processor configured by software; The processor, by the software, receiving the audio channel from the microphone array; receiving the gaze direction from the eye position measurement device; Obtaining weights for the audio channels from a lookup table based on the gaze direction and a power state of the smart glasses; applying the weights to the audio channels and summing the audio channels to generate a beamformed audio signal that amplifies sounds in the environment from the line of sight; transmitting the beamformed audio signal to the speaker for playback to the user; Smart glasses configured as follows.
19. The smart glasses of claim 18, wherein the generation of the beamforming audio signal is triggered by the user's gaze remaining stationary for a specified period of time.
20. To obtain the weights for the audio channels from the lookup table, the processor further comprises: accessing the lookup table containing a set of weights for a plurality of predetermined gaze directions; selecting a weight for one predetermined gaze direction that is closest to said gaze direction; The smart glasses of claim 18 configured to:
21. The smart glasses of claim 20 , wherein the plurality of predetermined gaze directions are based on likelihoods of gaze directions determined by a training process.
22. 22. The smart glasses of claim 20 or 21, wherein the look-up table is stored on the smart glasses.
23. 22. The smart glasses of claim 20 or 21, wherein the lookup table is stored on a network connected to the smart glasses.
24. The smart glasses of any one of claims 18 to 21, wherein the microphone array includes a vertical array generating a vertical audio channel and a horizontal array generating a horizontal audio channel.
25. 25. The smart glasses of claim 24, wherein the gaze direction includes a horizontal gaze direction and a vertical gaze direction, and the processor is configured to determine a horizontal weight for the horizontal audio channel based on the horizontal gaze direction and to determine a vertical weight for the vertical audio channel based on the vertical gaze direction.
26. A computer program product which, when executed by at least one processor of a computer, causes said at least one processor to carry out the method of any one of claims 1 to 3 or 6.
Citation Information
Patent Citations
Sound input device
JP1995177595A
Design and implementation of super-directional beamforming
JP2003514481A
Sound collecting device, hearing aid, and sound collecting device set
JP2019054385A
Gaze-based audio direction
US20160080874A1
Hearing device or system adapted for navigation
US20190174237A1