Audio system and method for determining an audio filter based on a device location

By determining the relative position of the speakers of the external audio device and the user's anatomical characteristics, a compensation filter is generated, which solves the problem of inaccurate sound perception in the external audio device, and accurately spatial audio reproduction on the external audio device is achieved.

CN115150716BActive Publication Date: 2025-08-05APPLE INC
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210342536.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2021-03-31
Filing Date
2022-03-31
Publication Date
2025-08-05
Estimated Expiration
2042-03-31

AI Technical Summary

Technical Problem

In the prior art, when using extra-ear audio equipment, spatialized sounds are inaccurate in the sound perception due to the deviation between the entrance of the ear canal and the speaker position, and it is impossible to accurately simulate the sound position in the real world.

Method used

By receiving images of the audio device and monitoring the anatomical characteristics of the user, determining the relative position between the speaker and the anatomical characteristics, a compensation filter is generated to correct the deviation, applied to the audio input signal for accurate spatial audio reproduction.

Benefits of technology

It realizes accurate reproduction of spatial audio on extra-ear audio devices, compensates for the deviation of anatomical characteristics and speaker position, and provides a realistic spatial sound experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115150716B_ABST
    Figure CN115150716B_ABST
Patent Text Reader

Abstract

The present application relates to an audio system and method for determining audio filters based on device position. An audio system and method for determining audio filters based on the position of an audio device of the audio system are described. The audio system receives an image of an audio device worn by a user and determines the relative position between the electroacoustic transducer and an anatomical feature of the user based on the image and a known geometric relationship between a reference on the audio device and the electroacoustic transducer of the audio device. An audio filter is determined based on the relative position. The audio filter can be applied to an audio input signal to present spatialized sound to the user through the electroacoustic transducer, or the audio filter can be applied to a microphone input signal to capture the user's speech by the electroacoustic transducer. Other aspects are also described and claimed.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] This application claims the benefit of priority to U.S. Provisional Patent Application No. 63 / 169,004, filed on March 31, 2021, which is incorporated herein by reference in its entirety. Technical Field

[0002] The present invention discloses aspects related to devices with audio capabilities, and more particularly, aspects related to devices for presenting spatial audio. Background Art

[0003] Spatial audio can be rendered using audio devices worn by the user. For example, headphones can reproduce spatial audio signals that simulate the soundscape surrounding the user. Effective spatial sound reproduction can render sounds so that the user perceives them as coming from locations within the soundscape outside the user's head, just as the user would experience sounds in the real world.

[0004] When sound propagates from the real-world surroundings to a listener, it travels along a direct path, such as through the air to the listener's ear canal entrance, and along one or more indirect paths, such as through reflection and diffraction off the listener's head or shoulders. When sound propagates along indirect paths, artifacts may be introduced into the acoustic signal received at the ear canal entrance. These artifacts are anatomically dependent and therefore user-specific. Consequently, the user perceives the artifacts as natural.

[0005] User-specific artifacts can be incorporated into binaural audio through signal processing algorithms that use spatial audio filters. For example, a head-related transfer function (HRTF) is a filter that contains all the acoustic information needed to describe how sound reflects or diffracts around the listener's head before entering their auditory system at the entrance to their ear canal. HRTFs can be measured for a specific user in a laboratory. HRTFs can be applied to an audio input signal to shape the signal so that the reproduction of the shaped signal realistically simulates the sound propagating from the surrounding environment to the user. Thus, a listener can use simple stereo headphones to create the illusion of a sound source somewhere in the listening environment by applying an HRTF to the audio input signal. Summary of the Invention

[0006] Existing methods for generating and applying head-related transfer functions (HRTFs) assume that headphones emit spatialized sound directly into the listener's ear canal entrance. However, this assumption can be incorrect. For example, when a listener wears an audio device with speakers located far from the ear canal entrance, such as with over-the-ear headphones, the spatialized sound may experience additional artifacts before entering the ear canal entrance. As a result, the user may perceive the spatialized sound as a flawed representation of sound, as would normally be experienced.

[0007] An audio system and a method of using the audio system to determine an audio filter that compensates for the relative positioning between an electroacoustic transducer (e.g., a speaker) and an anatomical feature (e.g., an ear canal entrance) are described. By compensating for the relative positioning, spatialized sound output to a user can accurately represent the sound as the user would normally experience it. In one aspect, the method includes receiving an image of an audio device worn on a user's head. A monitoring device (e.g., a wearable device) can output one or more of a visual cue, an audio cue, or a tactile cue to guide the user to move a remote device relative to the audio device for image capture. Thus, a camera of the remote device can capture an image that includes a datum of the audio device and an anatomical feature of the user.

[0008] In one aspect, one or more processors of the audio system determine the relative position between an anatomical feature and an electroacoustic transducer of the audio device. The determination can be made based on an image and also based on a known geometric relationship between a reference and the electroacoustic transducer. For example, the electroacoustic transducer may not be visible in the image, however, the geometric relationship between the reference visible in the image and the hidden electroacoustic transducer can be used to determine the position of the electroacoustic transducer. The relative position between the hidden electroacoustic transducer (e.g., a speaker or microphone of the audio device) and a visible anatomical feature (e.g., the entrance to the ear canal or mouth of the user) can then be determined.

[0009] In one aspect, an audio filter can be determined based on relative position. The audio filter can compensate for the relative position between the electroacoustic transducer and the anatomical feature. For example, artifacts can be introduced by the separation between the entrance of the user's ear canal and the external speaker of the wearable device. The audio filter can compensate for those artifacts and can therefore be selected based on the determined separation. Thus, the audio filter can be applied to the audio input signal to generate a spatial input signal, and the external speaker can be driven with the spatial input signal to present realistic spatialized sound to the user.

[0010] In one aspect, a device includes a memory and one or more processors configured to perform the above method. For example, the memory may store an image of an audio device and instructions executable by the processor to cause the device to perform a method comprising determining a relative position based on the image, and determining an audio filter based on the relative position.

[0011] The above summary does not include an exhaustive list of all aspects of the present invention. It is contemplated that the present invention includes all systems and methods that can be practiced from all suitable combinations of the various aspects summarized above and disclosed in the following detailed description and particularly pointed out in the claims filed with this patent application. Such combinations have particular advantages not specifically recited in the above summary. BRIEF DESCRIPTION OF THE DRAWINGS

[0012] Figure 1 is a pictorial view of a user wearing an audio device and holding a remote device according to one aspect.

[0013] Figure 2 is a block diagram of an audio system according to one aspect.

[0014] Figure 3 is a perspective view of an audio device according to one aspect.

[0015] Figure 4 is a perspective view of an audio device according to one aspect.

[0016] Figure 5 is a flowchart of a method of determining an audio filter according to one aspect.

[0017] Figure 6 is a pictorial illustration of a user capturing an image of an audio device worn on the user's head according to an aspect.

[0018] Figure 7 is a flow chart of a method of guiding a user to capture an image of an audio device worn on the user's head, according to an aspect.

[0019] Figure 8 is a pictorial representation of an image of an audio device worn on a user's head according to one aspect.

[0020] Figure 9 is a flow chart of a method of using an audio filter for audio playback according to one aspect.

[0021] Figure 10 is a pictorial illustration of a method for audio playback of spatialized sound using audio filters according to an aspect.

[0022] Figure 11 is a flow chart of a method for using an audio filter for audio pickup according to one aspect.

[0023] Figure 12 is a pictorial diagram of a method for audio pickup using an audio filter according to one aspect. DETAILED DESCRIPTION

[0024] Various aspects describe an audio system and a method for determining an audio filter based on a position of an audio device relative to an anatomical feature of a listener, and using the audio filter to implement audio playback or audio pickup by the audio system. The audio system may include an audio device, and the audio filter may be applied to an audio input signal to generate a spatial input signal for playback by the audio device. For example, the audio device may be a wearable device such as an over-the-ear earphone, a head-mounted device having an over-the-ear earphone, etc. However, the audio device may be another wearable device such as an in-ear headphone or a telephone headset, to name a few possible applications.

[0025] In various aspects, description is made with reference to the accompanying drawings. However, certain aspects may be practiced without one or more of these specific details or without being combined with other known methods and configurations. In the following description, a number of specific details such as specific configurations, dimensions, and processes are set forth in order to provide a thorough understanding of the aspects. In other instances, well-known processes and manufacturing techniques are not described in particular detail so as not to unnecessarily obscure the description. References to "one aspect," "aspect," etc. throughout this specification mean that the specific features, structures, configurations, or characteristics being described are included in at least one aspect. Therefore, the phrases "one aspect," "aspect," etc. appearing in various places throughout this specification do not necessarily refer to the same aspect. In addition, specific features, structures, configurations, or characteristics may be combined in one or more aspects in any suitable manner.

[0026] Relative terms are used throughout the description to refer to relative positions or directions. For example, "in front of" may indicate a first direction away from a reference point. Similarly, "behind" may indicate a position in a second direction away from the reference point that is opposite to the first direction. However, such terms are provided to establish a relative frame of reference and are not intended to limit the use or orientation of an audio system or system component (e.g., an audio device) to the specific configurations described in the various aspects below.

[0027] In one aspect, an audio system includes an audio device worn by a user, and a remote device that can image the audio device while it is being worn. Based on the image captured by the remote device, the relative position between an electroacoustic transducer of the audio device (e.g., a speaker or microphone) and an anatomical feature of the user (e.g., an ear canal entrance or mouth) can be determined. The electroacoustic transducer may not be visible in the image, and therefore, a known geometric relationship between the electroacoustic transducer and a visible reference of the audio device can be used for determination. An audio filter can be determined based on the relative position. The audio filter can compensate for the spatial offset between the anatomical feature and the electroacoustic transducer, and thus can generate spatialized audio that is more realistic for the user, or can generate a microphone pickup signal that more accurately captures external sounds (such as the user's voice).

[0028] refer to Figure 1 , shows a pictorial view of a user wearing an audio device and holding a remote device, according to one aspect. An audio system 100 may include devices, such as a remote device 102 (such as a smartphone, laptop, portable speaker, etc.), which communicates with an audio device 104 worn on a user's head 106. As shown, the user 108 may wear several audio devices 104. For example, the audio device 104 may be a wearable device such as over-the-ear headphones 110, a head-mounted display for applications such as virtual reality or augmented reality videos or games, or another device with speakers and / or microphones spaced apart from the user's ears or mouth. More specifically, the wearable device 110 may include over-the-ear speakers, a microphone, and optionally a display, as described below. Alternatively, the audio device 104 may be in-ear headphones 112. The in-ear headphones 112 may include speakers that transmit sound directly into the ears of the user 108. Thus, the user 108 can listen to audio played by the audio device 104, such as music, movie or game content, binaural audio reproduction, phone calls, and the like. In an aspect, the remote device 102 may drive the audio device 104 to present spatial audio to the user 108 .

[0029] In one aspect, the audio device 104 may include a microphone. The microphone may be built into the wearable device 110 or the in-ear headphones 112 to detect sounds inside and / or outside the audio device 104. For example, the microphone may be mounted on the audio device 104 in a position facing the surrounding environment. Thus, the microphone may detect input signals corresponding to sounds received from the surrounding environment. For example, the microphone may be pointed toward the mouth 120 of the user 108 to pick up the user's 108 speech and generate a corresponding microphone output signal.

[0030] In one aspect, the remote device 102 includes a camera 114 to capture images of the audio device 104 worn on the head 106 of the user 108 as the remote device 102 moves around the head 106. For example, the remote device 102 may capture several images, for example via the camera 114, as the remote device 102 is continuously moved around the head 106. The images may be used to determine audio filters to implement output of a speaker or microphone of the audio device 104, as described below. In addition, the remote device 102 may include circuitry to connect to the audio device 104 wirelessly or via a wired connection to transmit signals for audio rendering (e.g., binaural audio reproduction).

[0031] refer to Figure 2, shows a block diagram of an audio system according to one aspect. The audio system 100 may include a remote device 102, which may be any of several types of portable devices or apparatuses having circuitry suitable for a particular function. Similarly, the audio system 100 may include a first audio device 104 (e.g., a wearable device 110) and / or a second audio device 104 (e.g., an in-ear headphone 112). More specifically, the audio device 104 may include any of several types of wearable devices or apparatuses having circuitry suitable for a particular function. The wearable device may be head-mounted, wrist-mounted, or worn on any other part of the body of a user 108. The illustrated circuitry is provided by way of example and not limitation.

[0032] The audio system 100 may include one or more processors 202 to execute instructions to perform the various functions and capabilities described below. Instructions executed by the processor 202 may be retrieved from a memory 204, which may include non-transitory machine-readable media. The instructions may be in the form of an operating system program having a device driver and / or an audio rendering engine for rendering music playback, two-channel audio playback, etc. according to the methods described below. The processor 202 may retrieve data from the memory 204 for various purposes, including: for image processing; for audio filter selection, generation, or application; or for any other operations, including those involved in the methods described below.

[0033] One or more processors 202 may be distributed throughout the audio system 100. For example, the processor 202 may be incorporated into the remote device 102 or the audio device 104. The processors 202 of the audio system 100 may communicate with each other. For example, the processor 202 of the remote device 102 and the processor 202 of the audio device 104 may wirelessly transmit signals to each other via corresponding RF circuits 205, as shown by the arrows, or transmit signals to each other via a wired connection. The processor 202 of the audio system 100 may also communicate with one or more device components within the audio system 100. For example, the processor 202 of the audio device 104 may communicate with the electroacoustic transducer 208 (e.g., speaker 210 or microphone 212) of the audio device 104.

[0034] In one aspect, the processor 202 may access and retrieve audio data stored in the memory 204. The audio data may be an audio input signal provided by one or more audio sources 206. The audio source may include a phone and / or music playback function controlled by a phone or audio application running on top of the operating system. Similarly, the audio source may include an augmented reality (AR) or virtual reality (VR) application running on top of the operating system. In one aspect, the AR application may generate a spatial input signal to be output to the electroacoustic transducer 208 (e.g., speaker 210) of the audio device 104. For example, the remote device 102 and the audio device 104 (e.g., wearable device 110 or in-ear headphones 112) may transmit the signal wirelessly. Thus, the audio device 104 may present spatial audio to the user 108 based on the spatial input signal from the audio source.

[0035] In one aspect, memory 204 stores audio filter data for use by processor 202. For example, memory 204 may store audio filters that can be applied to an audio input signal from an audio source to generate a spatial input signal. As used herein, audio filters can be implemented in digital signal processing code or computer software as digital filters that perform equalization or filtering of the audio input signal. For example, a dataset may include measured or estimated HRTFs corresponding to user 108. A single HRTF in the dataset may be a pair of acoustic filters (one for each ear) that characterize the acoustic transmission from a specific location in an anechoic environment to the entrance of the user's 108 ear canal. Individual equalization can also be performed for each ear. Ears and their position relative to the head are asymmetrical, and the audio device 104 may be worn so that the relative position varies between ears. Therefore, the acoustic filters selected for each ear can be personalized for that ear, rather than being selected as a fixed pair. The dataset of HRTFs summarizes the fundamental tones of user 108's spatial hearing. The dataset may also include audio filters that compensate for the separation between the entrance of the user's 108 ear canal and the speaker 210 of the audio device 104. Such audio filters can be applied directly to the audio input signal, or applied to the audio input signal filtered by an HRTF-related audio filter, as described below. Thus, the processor 202 can select one or more audio filters from the database in the memory 204 to apply to the audio input signal to generate the spatial input signal. The audio filters in the memory 204 can also be used to influence the microphone input signal of the microphone 212, as described below.

[0036] The memory 204 may also store data generated by the imaging system of the remote device 102. For example, the structured light scanner or RGB camera 114 of the remote device 102 may capture an image of the audio device 104 being worn on the head 106 of the user 108, and the image may be stored in the memory 204. The image may be accessed and processed by the processor 202 to determine the relative position between the anatomical features of the user 108 and the electroacoustic transducer of the audio device 104.

[0037] To perform various functions, the processor 202 may directly or indirectly implement a control loop and receive input signals from other electronic components and / or provide output signals to other electronic components. For example, the processor 202 may receive input signals from a microphone or an input control (such as a menu button of the remote device 102). The input control may be displayed as a user interface element on a display of the remote device 102 or the audio device 104, and may be selected by inputting an input selection of a user interface element displayed on the display 211, for example, when the wearable device 110 is a head-mounted display.

[0038] refer to Figure 3 , shows a perspective view of an audio device according to one aspect. The audio device 104 may be a wearable device 110 and may have features germane to and typically associated with that type of device. For example, when the wearable device 110 is a head-mounted display, the device may have a housing that incorporates a display 211 for the user to view video content while wearing the audio device 104. The portion of the housing that holds the display 211 may rest on the nose of the user 108, and the audio device 104 may include other features to support the housing on the head 106 of the user 108. For example, the head-mounted display may include temples or a headband to support the housing on the head 106 of the user 108. Similarly, when the wearable device 110 includes over-the-ear headphones, such as Figure 3 As shown, the headset may include temples 302 to support the device on the head 106 of the user 108.

[0039] The wearable device 110 may include an electroacoustic transducer 208 to output sound or receive sound from the user 108. For example, the electroacoustic transducer 208 may include a speaker 210, which may be an extra-ear speaker integrated into the temple 302 of the wearable device 110. The wearable device 110 may include other features, such as embossing or hinges on the temple 302, markings on the temple 302, a headband, a housing, etc.

[0040] The overall geometry of the wearable device 110 can be designed and modeled using computer-aided design. More specifically, the audio device 104 can be represented by a computer-aided design (CAD) model, which can be a virtual representation of the physical object of the audio device 104. Figure 3 The view may be a view of a CAD model. The CAD model may have the same properties as a physical object, and therefore, the geometric relationships between features of the audio device 104 may be represented by the CAD model.

[0041] In one aspect, several features of the audio device 104 can be related via geometric relationships 304. Geometric relationships 304 can be distinguished from relative positions because geometric relationships are known or determined relative to a predetermined model of the audio device 104, as opposed to the actual relative positions between audio device components as they may exist in free space. The audio device 104 has a predetermined geometry that is known based on a CAD model, and thus any two physical features of the device can have a relative orientation or position that can be determined based on the CAD model. By way of example, the audio device 104 can include a reference 306. The reference 306 can be any feature of the audio device 104 that is identifiable and / or imageable and can be used as a basis for determining the position of another feature of the audio device 104. For example, the reference 306 can be a marking on the temple 302, an embossing of the temple 302, a cover or hinge, or any other feature that can be imaged. The marking can be a diamond, a rectangle, or any other shape that can be identified using image processing techniques.

[0042] As shown, a fiducial 306 (in this case, the embossing of the temple) may have a geometric relationship 304 with the electro-acoustic transducer 208. More specifically, a point on the fiducial 306 may be spaced apart from the electro-acoustic transducer 208, and the relative position between the features may be the geometric relationship 304. The geometric relationship of the features may be modeled in the CAD model. The geometric relationship 304 may be the difference in coordinates of the features within a Cartesian coordinate system, or any other system for representing features in the CAD model.

[0043] refer to Figure 4 , shows a perspective view of an audio device according to one aspect. The audio device 104 may be an in-ear headphone 112 and may have features germane to and typically associated with that type of device. For example, the in-ear headphone 112 may have a housing that incorporates a speaker 210 and a microphone 212. The in-ear headphone 112 may be fitted into the outer ear of the user 108 such that the speaker 210 may output sound into the entrance to the ear canal of the user 108. Similarly, the in-ear headphone 112 may have a microphone 212, spaced apart from the speaker 210, for example at the distal end of the body 402, to receive sound when the user 108 speaks.

[0044] Like the wearable device 110, the in-ear headphone 112 may have one or more fiducials 306 that are represented by a CAD model and that can be identified in an image of the audio device 104. Like the wearable device 110, the in-ear headphone 112 may be designed and modeled using CAD, and features of the in-ear headphone 112 may be related to each other through the resulting CAD model. For example, the geometric relationship 304 between a rectangular mark on the body 402 and the speaker 210 may be known and used to determine the spatial position of the speaker 210 when only the fiducial 306 is visible. Similarly, the geometric relationship 304 between a rectangular mark on the body 402 and the microphone 212 may be known and used to determine the spatial position of the microphone 212 when only the fiducial 306 is visible. The fiducial 306 may be any identifiable physical feature, such as a bump, groove, color change, or any other feature of the audio device 104 that can be imaged.

[0045] The geometric relationship 304 between the fiducial 306 and the electroacoustic transducer 208 (e.g., the speaker 210 or the microphone 212) may allow the location of one feature to be determined based on the known location of the other feature. Even if only one feature (e.g., the fiducial 306) is identifiable in the image, other features (e.g., hidden in the image) may be detected. Figure 3 The position of the speaker 210 behind the temple 302 in the audio device 104 can also be determined based on the predetermined geometry of the audio device 104 known based on the CAD model. More specifically, based on the CAD model, the visible portion of the audio device 104 can be associated with the hidden portion of the audio device 104.

[0046] refer to Figure 5 , shows a flow chart of a method for determining an audio filter according to one aspect. The method can be used to determine an audio filter based on a relationship between an electroacoustic transducer 208 (e.g., a speaker 210 or a microphone 212) of an audio device 104 and an anatomical feature of a user 108 (e.g., an ear canal entrance or mouth 120). More specifically, an audio filter can be determined that compensates for artifacts introduced due to a separation between the anatomical feature and the electroacoustic transducer 208. For example, applying an audio filter to an audio input signal can provide acoustic compensation for the way the user 108 is wearing the audio device 104. The operation of the method is Figures 6 and 7 and therefore, the operation of the method will be similar to those of Figure 1 Start description.

[0047] refer to Figure 6, shows a pictorial view of a user capturing an image of an audio device worn on the user's head, according to one aspect. At operation 502, an image of an audio device 104 may be received by one or more processors 202 of the audio system 100. The image may be received from the camera 114 of the remote device 102. More specifically, during the enrollment process, the user 108 may move the remote device 102 in an arcuate path around the user's 108 head 106, with the remote device 102's forward-facing camera 114 facing the user's 108 head 106. As the remote device 102 is moved around the head 106, the forward-facing camera 114 may capture and record one or more images of a known device (e.g., the audio device 104) being worn on the user's 108 head 106. For example, when the user 108 wears the wearable device 110 or the in-ear headphones 112, the remote device 102 may record anatomical features of the audio device 104 and the head 106, such as the user's 108 mouth 120 or ears. The one or more images may be multiple images. More specifically, the input data may be several images instead of just one image.

[0048] The images from the enrollment process may be used to determine an appropriate HRTF for the user 108. More specifically, the method provides for mapping the anatomy of the user 108 to a particular HRTF that is stored, for example, in a database of the remote device 102 and selected to be applied to the audio input signal. The method of determining the HRTF will not be described in detail, but it will be understood that the image capture used to map the anatomy of the user 108 to a particular HRTF may also be used to determine an audio filter that compensates for the separation between the electro-acoustic transducer 208 and the anatomical feature. Alternatively, the anatomy of the user 108 may be scanned a first time to determine the complete anatomy of the user 108, for example when the user 108 is not wearing the audio device 104, and the anatomy of the user 108 may be scanned a second time to determine the relative positioning of the anatomy and the electro-acoustic transducer 208, for example when the user 108 is wearing the audio device 104.

[0049] The goal of the enrollment process is to capture an image that shows the relative position of the audio device 104 and the anatomy of the user 108. The relative position can be the relative positioning of the audio device 104 (or a portion thereof) to the anatomy of the user 108 in the environment in which the image is captured (e.g., in free space where the user is located). For example, the image can show how the in-ear earphone 112 fits within the ear, the direction in which the body 402 of the in-ear earphone 112 extends away from the ear or toward the mouth 120, how the wearable device 110 sits on the ear or face of the user 108, how the headband of the wearable device 110 is positioned around the head 106 of the user 108, and so on. This information about the fit (and more specifically, the relative position of the audio device 104 and the user's anatomy) can be used to determine information such as whether the user 108 has long hair that can affect the user's HRTF, from which direction sound will be received at the microphone 212 when the user 108 is speaking, in which direction and how far the sound must travel from the speaker 210 to the ear canal entrance, and so on. More specifically, when the captured image shows the relative position between the electro-acoustic transducer 208 and the user's anatomy, or as described below, the relative position between the user's anatomy and the fiducial 306 (which may be related to the electro-acoustic transducer 208), the audio signal can be appropriately adjusted to maintain realistic spatial audio performance and accurate audio pickup.

[0050] Correctly positioning the remote device 102 relative to the head-mounted device can allow the camera 114 to capture an image of the audio device 104 being worn on the head 106 of the user 108 at an angle that provides information about the relative position between the audio device 104 and the user's anatomy. However, sometimes, it can be difficult for the user 108 to determine whether the remote device 102 is correctly positioned from the display 211 of the remote device 102 (which can display the image captured by the camera 114). More specifically, because the remote device 102 may be scanning one side of the head 106, the user 108 may not be able to see the display 211 of the remote device 102 and, therefore, may not be able to rely on the display 211 for guidance in positioning the remote device 102.

[0051] refer to Figure 7, shows a flow chart of a method for guiding a user to capture an image of an audio device worn on the user's head according to one aspect. At operation 702, the camera 114 of the remote device 102 may capture an image of the audio device 104 worn on the head 106 of the user 108. In one aspect, feedback may be provided to the user 108 by the auxiliary device to guide the user 108 to move the remote device 102 to an appropriate position for image capture. More specifically, at operation 704, the auxiliary device may output one or more of a visual prompt, an audio prompt, or a tactile prompt to guide the user 108 to move the remote device 102 relative to the audio device 104. The auxiliary device may be a monitoring device 602 ( Figure 6 ), which is a device other than the remote device 102, and may output prompts to the user 108. The prompts may prompt the user 108 to move the remote device 102 to an appropriate position for image capture.

[0052] The monitoring device 602 can be a phone, computer, or another device with a visual display, speaker, haptic motor, or any other component capable of providing guidance cues to the user 108 to help the user 108 properly position the camera 114 of the remote device 102. The monitoring device 602 can visually display, audibly describe, tactilely stimulate, or otherwise feed information back to the user 108 regarding the progress of the scan or the position of the remote device 102 relative to the audio device 104. Feedback provides for more efficient and accurate imaging operations during the enrollment process.

[0053] In one aspect, the monitoring device 602 is a wearable device. More specifically, the user 108 can wear the monitoring device 602 while performing the enrollment process, including the imaging operation. The wearable device can be a device other than the remote device 102. For example, the monitoring device 602 can be an audio device 104, such as the wearable device 110 or the in-ear headphones 112, worn on the head 106 of the user 108. The ability to wear the monitoring device 602 ensures that whenever the user 108 wants to perform acoustic adjustments based on the fit of the audio device 104, the device is present and easily visible.

[0054] The wearable device may be a device other than the remote device 102 and the audio device 104. For example, the monitoring device 602 may be a smartwatch worn on the wrist of the user 108. The smartwatch may have a computer architecture similar to that of the remote device 102. The smartwatch may include a display for presenting visual cues, a speaker for presenting audio cues, or a vibration motor or other actuator for providing tactile cues. When the smartwatch is worn on the wrist, it can be easily positioned in the field of view of the user 108, while the remote device 102 remains at a position outside the field of view of the user 108. The remote device 102 may stream images or other position information (e.g., inertial measurement unit (IMU) data) to the monitoring device 602. The monitoring device 602 may use the position information to determine guidance instructions and present the guidance instructions to the user 108 in a visual, audio, or tactile form. Thus, monitoring device 602 may be a third device in audio system 100 in addition to remote device 102 and audio device 104 to allow user 108 to register and determine audio filters that may compensate for separation between electroacoustic transducer 208 and anatomical features.

[0055] In one aspect, the monitoring device 602 provides visual cues to guide the user 108. The remote device 102 may stream images captured by the camera 114 to the audio device 104 for presentation on the display 211. For example, the user 108 may be viewing an image of the side of their head 106 on the audio device display 211. The image may be provided by the remote device 102, which the user is holding with their arms straight and extended to their sides. The user 108 may move the remote device 102 based on the streamed image until the remote device 102 is in the desired position. In addition to the image of the audio device 104 worn on the head 106 of the user 108, the audio device 104 may also display text instructions, icons, indicators, or other information that directs the user 108 to move the remote device 102 in a specific manner. For example, the monitoring device 602 may determine the current position and orientation of the remote device 102 based on the image or position information provided by the remote device 102. A flashing arrow may be displayed to indicate the direction in which the remote device 102 should be moved to best capture the relative position between the audio device 104 and the user's anatomy. For example, the arrow may guide the user 108 to move the remote device 102 from its current position to the optimal position. Thus, the monitoring device 602 provides prompts to guide the user 108 to position the phone at a specific location with a specific orientation (pitch, yaw, and roll) relative to the gravity vector or the audio device 104, or at a specific distance from the audio device 104.

[0056] In one aspect, the monitoring device 602 provides audio prompts to guide the user 108. For example, the speaker 210 of a wearable device (e.g., a smartwatch or audio device 104) can provide a descriptive version of the visual prompts described above. More specifically, audio instructions (such as "tilt your head to the left," "rotate your head," "move your phone to the left," "tilt your phone away from you") or other instructions can be provided to guide the user 108 to correctly position the remote device 102 relative to the audio device 104. No spoken instructions are required. For example, a tone can be periodically output in the form of a radar beep. As the remote device 102 approaches the optimal position, the frequency of the beeps can increase. Thus, when the user 108 has moved the remote device 102 with the purpose of reaching the optimal position based on feedback from the increased frequency of the beeps, the remote device 102 will become correctly positioned. When correctly positioned, the remote device 102 can capture an image representing the relative position between the audio device 104 and the anatomical feature.

[0057] In one aspect, the monitoring device 602 provides tactile cues to guide the user 108. For example, a vibration motor or other actuator of a wearable device (e.g., a smartwatch or audio device 104) can provide tactile feedback, such as vibration, in a manner similar to the above-mentioned audio cues. More specifically, vibration pulses can be output periodically in the manner of radar beeps. As the remote device 102 approaches the optimal position, the frequency of the pulses can increase. Thus, when the user 108 has moved the remote device 102 with the purpose of reaching the optimal position based on the feedback of the increased frequency of the pulses, the remote device 102 will become correctly positioned. When correctly positioned, the remote device 102 can capture an image representing the relative position between the audio device 104 and the anatomical feature.

[0058] refer to Figure 8 , showing a pictorial view of an image of an audio device worn on a user's head according to one aspect. In operation 504 ( Figure 5), a relative position 808 between the anatomical feature 804 and the electroacoustic transducer 208 is determined based on the image 802. When the user 108 is holding the remote device 102 near the above-mentioned optimal position, the image 802 is shown on the display 211 of the remote device 102. It should be understood that for illustrative purposes, the image 802 is shown on the display 211, but the image 802 can be received as an image file representing the shown view. Therefore, the image 802 can be processed to identify certain image features. For example, the image 802 may include a reference 306 of the audio device 104 and one or more anatomical features 804 of the user 108. The reference 306 can be a marking on the temple 302 of the wearable device 110, as described above. The reference can also be a feature such as an edge, structure, or any feature of the audio device 104 that can be identified in the image 802. The anatomical feature 804 can be the entrance to the ear canal 806 of the user 108 or the upper edge of the pinna, as shown. The anatomical feature 804 may also be the mouth 120 of the user 108 , an earlobe of the user 108 , or any other anatomical feature identifiable in the image 802 .

[0059] In one aspect, image 802 does not include electro-acoustic transducer 208. More specifically, electro-acoustic transducer 208 may be hidden in image 802. For example, electro-acoustic transducer 208 may be speaker 210 mounted on an inner surface of temple 302, which is hidden behind temple 302. Therefore, the relative position 808 between anatomical feature 804 and electro-acoustic transducer 208 may not be directly identifiable from image 802.

[0060] To determine relative position 808, geometric relationship 304 between identifiable fiducial 306 and electro-acoustic transducer 208 can be used. More specifically, the geometry of audio device 104 can be known and stored, for example, as a CAD model of audio device 104. Thus, the geometry can be used to associate any identifiable point on audio device 104 with another point on audio device 104, regardless of whether the other point is visible in image 802. In one aspect, when electro-acoustic transducer 208 is hidden from view, the position of fiducial 306 can be identified and then associated with electro-acoustic transducer 208. More specifically, CAD model-based geometric relationship 304 can be used to mathematically determine the unknown position of electro-acoustic transducer 208 based on the known position of fiducial 306.

[0061] When the position of the electroacoustic transducer 208 is known, it can be used to determine the relative position 808 between the electroacoustic transducer 208 and the anatomical feature 804. For example, the relative position 808 between the electroacoustic transducer 208 and the anatomical feature 804 can be determined based on the known geometric relationship 304. Figure 8The relative position 808 between the speaker 210 and the ear canal entrance 806 can be determined based on the image 802. Alternatively, when the image 802 includes the in-ear earphone body 402 positioned relative to the mouth 120, the relative position between the microphone and the mouth of the user 108 can be determined. Thus, the relative position 808 between the anatomical feature 804 and the electroacoustic transducer 208 of the audio device 104 can be determined based on the image 802 and the geometric relationship 304 between the reference 306 and the electroacoustic transducer 208.

[0062] In operation 506 ( Figure 5 ), an audio filter is determined based on relative position 808. By determining the relative position and / or orientation of electroacoustic transducer 208 to anatomical feature 804, a personalized audio filter (e.g., a personalized equalizer) can be generated or selected to compensate for the separation. Relative position 808 can be used to reference a lookup table, for example, or otherwise identify an audio filter stored in memory 204 that corresponds to the separation between electroacoustic transducer 208 and anatomical feature 804.

[0063] In the case of audio output, audio filters can be used in conjunction with HRTFs to account not only for anatomy but also for how the audio device 104 is mounted on the user 108 when providing spatial audio. In the case of audio input, audio filters can be used to filter the input based on how the orientation of the audio device 104 (e.g., the body 402 of the in-ear headphones 112) positions and directs the microphone 212 relative to the sound source (e.g., the mouth 120). Thus, as described below, the determined audio filters can be used for audio playback to adjust how the speaker 210 outputs sound, or the determined audio filters can be used for audio pickup to adjust how the microphone 212 picks up sound. In either case, the audio filters can compensate for artifacts introduced by the relative position 808.

[0064] refer to Figure 9 , a flowchart of a method for using an audio filter for audio playback according to one aspect is shown. Figure 10 The operations of the method are illustrated in FIG, and therefore, these operations are described below with reference to that figure.

[0065] refer to Figure 10, shows a pictorial view of a method for audio playback of spatialized sound using audio filters, according to one aspect. At operation 902, an audio filter 1002 is applied to an audio input signal 1004 to generate a spatial input signal 1008. The audio input signal 1004 may be audio data provided by one or more audio sources 206 of a remote device 102. The audio filter 1002 may be applied directly or indirectly to the audio input signal 1004. For example, the audio filter 1002 may be applied to the audio input signal 1004 before or after the audio input signal 1004 is modified by the HRTF 1006. In one aspect, the HRTF 1006 is applied to the audio input signal 1004 to modify the audio input signal 1004 so that it is spatialized based on the specific anatomy of the user 108. The specific anatomy of the region of interest (such as the user's pinna) can have a significant impact on how sound reflects or diffracts around the listener's head before entering the listener's auditory system, and the HRTF 1006 can be applied to the audio input signal 1004 to shape the signal so that the reproduction of the shaped signal realistically simulates the sound propagating from the surrounding environment to the user. As described above, the HRTF 1006 can be selected as part of the registration process. The audio filter 1002 can then be applied to the modified signal to adjust the HRTF 1006 based not only on the anatomy but also on the position of the speaker 210 relative to the ear canal entrance 806.

[0066] The result of modifying the audio input signal 1004 with both the HRTF 1006 and the audio filter 1002 is a spatial input signal 1008. The spatial input signal 1008 is the audio input signal 1004 filtered by the HRTF 1006 and the audio filter 1002, such that the input sound recording is altered to simulate the diffraction and reflection characteristics of the anatomy of the user 108 and to compensate for artifacts introduced by separating the speaker 210 from the ear canal entrance 806. The spatial input signal 1008 may be transmitted by the processor 202 to the speaker 210. At operation 904, the speaker 210 is driven with the spatial input signal 1008 to present spatialized sound 1010 to the user 108. The spatialized sound 1010 may simulate the sound (e.g., speech) generated by a spatialized sound source 1012 (e.g., a person speaking) in a virtual environment surrounding the user 108. More specifically, by driving the speaker 210 with the spatial input signal 1008, the spatialized sound 1010 may be accurately and clearly presented to the user 108.

[0067] In addition to improving sound spatialization, personalized equalization of playback using audio filter 1002 can improve playback consistency from user to user. Personalized equalization can make the sound entering the ear canal consistent for all users. More specifically, the timbre of stereo playback can be perceived as the same across a user population. Such consistency can be beneficial in homogenizing the user experience.

[0068] refer to Figure 11 , a flowchart of a method for using an audio filter for audio pickup according to one aspect is shown. Figure 12 The operations of the method are illustrated in FIG, and therefore, these operations are described below with reference to that figure.

[0069] refer to Figure 12 , shows a pictorial view of a method for using an audio filter for audio pickup according to one aspect. As described above, the determined audio filter 1002 can be used for audio pickup. At operation 1102, an audio filter 1202 is applied to a microphone input signal 1204 of a microphone 212. For example, the microphone 212 can generate the microphone input signal 1204 based on an incident sound wave, and the audio filter 1202 can be applied to the microphone input signal 1204 to generate a pickup output signal 1206. Thus, the audio filter 1202 can adjust the microphone input signal 1204 based on the relative position 808 between the microphone 212 and the mouth 120 of the user 108 (or another sound source). The adjustment can produce a more accurate pickup output signal 1204. For example, the audio filter 1202 can be derived to improve voice pickup, intelligibility, active noise control, or other microphone pickup functions.

[0070] It is understood that the use of personally identifiable information should be subject to privacy policies and practices that are generally recognized to meet or exceed industry or government requirements for maintaining user privacy. Specifically, personally identifiable information data should be managed and processed to minimize the risk of unintentional or unauthorized access or use, and the nature of authorized use should be clearly stated to users.

[0071] In the foregoing description, the present invention has been described with reference to specific exemplary aspects thereof. It will be apparent that various modifications may be made to the specific exemplary aspects without departing from the broader spirit and scope of the present invention as set forth in the following claims. Accordingly, the description and drawings are to be regarded in an illustrative rather than a restrictive sense.

Claims

1. A method for determining an audio filter, comprising: receiving, by one or more processors, an image of an audio device worn on a head of a user, wherein the image includes a fiducial of the audio device and an anatomical feature of the user; determining, by the one or more processors, a relative position between the anatomical feature and the electroacoustic transducer of the audio device based on the image and a geometric relationship between the fiducial and the electroacoustic transducer; as well as An audio filter is determined, by the one or more processors, based on the relative position. The method of claim 1 , wherein the image excludes the electroacoustic transducer. The method of claim 1 , wherein the geometric relationship is based on a computer-aided design model of the audio device.

4. The method of any one of claims 1 to 3, wherein the electroacoustic transducer is a loudspeaker, and wherein the anatomical feature is an entrance to the user's ear canal.

5. The method according to claim 4, further comprising: applying, by the one or more processors, the audio filter to an audio input signal to generate a spatial input signal; as well as The speakers are driven by the one or more processors with the spatial input signal to render spatialized sound.

6. The method of any one of claims 1 to 3, wherein the electroacoustic transducer is a microphone, and wherein the anatomical feature is the user's mouth.

7. The method according to claim 6, further comprising: The audio filter is applied by the one or more processors to a microphone input signal of the microphone.

8. The method according to any one of claims 1 to 3, further comprising: capturing, by a camera of a remote device, the image of the audio device worn on the head of the user; as well as One or more of a visual prompt, an audio prompt, or a tactile prompt is output by a monitoring device to guide the user to move the remote device relative to the audio device. The method of claim 8 , wherein the monitoring device is a wearable device.

10. The method of claim 9, wherein the wearable device is the audio device.

11. An audio system comprising: a memory configured to store an image of an audio device worn on a user's head, wherein the image includes a fiducial of the audio device and an anatomical feature of the user; as well as One or more processors configured to: determining a relative position between the anatomical feature and the electroacoustic transducer of the audio device based on the image and a geometric relationship between the fiducial and the electroacoustic transducer; and An audio filter is determined based on the relative position.

12. The audio system of claim 11, wherein the image excludes the electroacoustic transducer.

13. An audio system according to any one of claims 11 to 12, wherein the electroacoustic transducer is a loudspeaker, and wherein the anatomical feature is an entrance to the user's ear canal.

14. The audio system of claim 13, wherein the one or more processors are configured to: applying the audio filter to an audio input signal to generate a spatial input signal; and The loudspeaker is driven with the spatial input signal to render spatialized sound.

15. The audio system of any one of claims 11 to 12, wherein the electroacoustic transducer is a microphone, and wherein the anatomical feature is the user's mouth.

16. A non-transitory machine-readable medium storing instructions executable by one or more processors of an audio system to cause the audio system to perform a method comprising: receiving an image of an audio device worn on a user's head, wherein the image includes a fiducial of the audio device and an anatomical feature of the user; determining a relative position between the anatomical feature and the electroacoustic transducer of the audio device based on the image and a geometric relationship between the fiducial and the electroacoustic transducer; as well as An audio filter is determined based on the relative position. The non-transitory machine-readable medium of claim 16 , wherein the image excludes the electroacoustic transducer.

18. The non-transitory machine-readable medium of any one of claims 16 to 17, wherein the electroacoustic transducer is a speaker, and wherein the anatomical feature is an entrance to the user's ear canal.

19. The non-transitory machine-readable medium of claim 18, wherein the method comprises: applying the audio filter to an audio input signal to generate a spatial input signal; as well as The loudspeaker is driven with the spatial input signal to render spatialized sound.

20. The non-transitory machine-readable medium of any one of claims 16 to 17, wherein the electroacoustic transducer is a microphone, and wherein the anatomical feature is a mouth of the user.

Citation Information

Patent Citations

  • In-ear speaker hybrid audio transparency system

    CN106982400A

  • System and method for operating a wearable loudspeaker device

    CN107690110A