Virtual microphone calibration based on displacement of the outer ear

By using a virtual microphone calibration method based on ear displacement, the sound pressure at the ear canal entrance is estimated using transducers and displacement sensors, and a filter is generated to adjust the audio content. This solves the problems of impracticality and discomfort of traditional methods and achieves a high-quality spatial audio experience.

CN116195269BActive Publication Date: 2025-11-25CTRL-LABS CORP
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202180064339.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2020-07-21
Filing Date
2021-07-20
Publication Date
2025-11-25
Estimated Expiration
2041-07-20

AI Technical Summary

Technical Problem

Traditional methods for calibrating in-ear microphones in existing head-mounted devices are impractical or uncomfortable, making it difficult to achieve a high-quality spatial audio experience.

Method used

By using a virtual microphone calibration method based on the user's ear displacement, a calibration signal is generated using a transducer, combined with a displacement sensor to measure the ear displacement, and a model is used to estimate the sound pressure at the ear canal entrance to generate a filter to adjust the audio content.

Benefits of technology

It achieves a high-quality audio experience without the need to place a physical microphone at the entrance of the ear canal, improves the auditory effect and spatialization of audio content, and enhances user comfort.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116195269B_ABST
    Figure CN116195269B_ABST
Patent Text Reader

Abstract

An audio system calibrates virtual microphones using displacement of a user's outer ear (207). Transducers (220, 230) present audio content to the user. One or more sensors (240) monitor displacement of a portion of the user's pinna (215). The displacement is caused in part by the presented audio content. The audio system estimates sound pressure at the user's ear canal entrance (210) based on the monitored displacement of the portion of the pinna (215), generates a sound filter accordingly, and adjusts the audio content using the sound filter. The transducers present the adjusted audio content to the user, thereby improving the user's auditory experience.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure generally relates to audio systems in headsets, and more specifically, to virtual microphone calibration based on the displacement of the user's outer ear in a headset. Background Technology

[0002] Headsets deliver audio content to users. Typically, to calibrate a headset to provide spatialized sound, a microphone is placed in the user's ear canal (usually at the ear canal entrance). The sound captured by the microphone is used to calibrate and equalize the system's output, and then a head-related transfer function (HRTF) is used to deliver the 3D spatialized sound. The device can use the HRTF to generate audio content that is presented via one or more speakers to provide spatialized audio. To ensure high reproduction quality, one or more speakers can be equalized at the same point where the HRTF is captured. However, using an in-ear microphone to calibrate a headset is not always practical or ideal. Summary of the Invention

[0003] Therefore, the present invention relates to methods, systems, and computer-readable non-transitory storage media.

[0004] This document describes an audio system. The audio system is configured to calibrate a virtual microphone based on the displacement of one or both of a user's ears. For the user's ears, the audio system generates a calibration signal (e.g., via a transducer) and measures (e.g., via a displacement sensor) the displacement of a portion of the ear that may be partially caused by the calibration signal. The audio system provides the displacement information as input to a model configured to output the estimated sound pressure level at the entrance to the ear canal. Thus, the audio system can simulate how a virtual microphone at the entrance to the ear canal would detect audio content. In some embodiments, one or more displacement sensors of the audio system are integrated into one or more transducers of the audio system. For example, a cartilage conduction transducer incorporated into the user's ear (e.g., configured to present audio content via cartilage conduction) may include a displacement sensor and / or a displacement sensor coupled to a displacement sensor that measures the displacement of the ear when the cartilage conduction transducer vibrates the ear.

[0005] Audio content is presented to a user via one or more transducers. One or more sensors monitor displacement of at least a portion of the user's auricle, said displacement being partially caused by the presented audio content. The sound pressure level at the user's ear canal entrance is estimated based on the displacement of the portion of the auricle. A sound filter for the transducer is generated using the remotely estimated sound pressure level at the ear canal entrance, and the generated filter is used to adjust the audio content. The transducer then presents the adjusted audio content to the user.

[0006] In some embodiments, an audio system for calibrating a virtual microphone is disclosed. The audio system includes a transducer, one or more sensors, and a controller. The transducer is configured to measure displacement of a portion of a user's auricle caused by presented audio content. The controller is configured to estimate sound pressure at the user's ear canal entrance based on the monitored displacement of the auricle portion, use the estimated sound pressure to generate a sound filter for the transducer, and use the generated filter to adjust the audio content. The controller instructs the transducer to present the adjusted audio content to the user.

[0007] In some embodiments of the audio system according to the invention, the transducer may be a cartilage conduction transducer configured to present audio content. One or more sensors may be the cartilage conduction transducer. Additionally, the system may be configured to monitor displacement of a portion of the user's auricle by measuring the preload of the cartilage conduction transducer.

[0008] In some embodiments of the audio system according to the present invention, one or more sensors may further include at least one of an accelerometer and an optical displacement sensor.

[0009] In some embodiments of the audio system according to the invention, the transducer may be a speaker configured to present audio content to a user via air conduction.

[0010] In some embodiments of the audio system according to the invention, the controller is configured to provide a model as input the displacement of a monitored portion of the auricle, the model being configured to output the sound pressure at the entrance of the ear canal based on the displacement of the auricle. Alternatively, the model may include at least one of a convolutional neural network, a linear model, and a numerical simulation. Alternatively or additionally, the model may be configured to receive the geometry of a user's ear as input, the geometry including measurements determined based on one or more images of the user's ear.

[0011] In some embodiments of the audio system according to the present invention, the adjusted audio content has a frequency response of a target amplitude.

[0012] The present invention also relates to a method for calibrating a virtual microphone based on the displacement of one or both of a user's ears. Audio content is presented to the user via one or more transducers. One or more sensors monitor the displacement of at least a portion of the user's auricle, said displacement being partially caused by the presented audio content. The sound pressure level at the user's ear canal entrance is estimated based on the displacement of the portion of the auricle. A sound filter for the transducer is generated using the remotely estimated sound pressure level at the ear canal entrance, and the generated filter is used to adjust the audio content. The transducer then presents the adjusted audio content to the user.

[0013] In some embodiments of the method according to the invention, the transducer may be a cartilage conduction transducer configured to present audio content. One or more sensors may be a cartilage conduction transducer. Furthermore, the method may also include monitoring displacement of a portion of the user's auricle by measuring the preload of the cartilage conduction transducer.

[0014] In some embodiments of the method according to the invention, one or more sensors may further include at least one of an acceleration sensor and an optical displacement sensor.

[0015] In some embodiments of the method according to the invention, the transducer may be a speaker configured to present audio content to a user via air conduction.

[0016] In some embodiments of the method according to the invention, estimating the sound pressure at the ear canal entrance may further include: providing a model as input a displacement of a monitored portion of the auricle, the model being configured to output the sound pressure at the ear canal entrance based on the auricle displacement. Additionally, the model may include at least one of a convolutional neural network, a linear model, and numerical simulation. Alternatively or additionally, the model may be configured to receive the geometry of a user's ear as input, the geometry including measurements determined based on one or more images of the user's ear.

[0017] In some embodiments of the method according to the invention, the adjusted audio content has a frequency response of a target amplitude. Attached Figure Description

[0018] Figure 1A A perspective view of a head-mounted device implemented as an eyeglasses device according to one or more embodiments, the head-mounted device being configured to calibrate a virtual microphone.

[0019] Figure 1B A perspective view of a head-mounted device implemented as a head-mounted display according to one or more embodiments, the head-mounted device being configured to calibrate a virtual microphone.

[0020] Figure 2A side view of a head-mounted device configured as part of a virtual microphone according to one or more embodiments.

[0021] Figure 3A The diagram shows a cartilage conduction transducer according to one or more embodiments, which is configured to monitor the displacement of a user's ear using a capacitive displacement sensor.

[0022] Figure 3B The diagram shows a cartilage conduction transducer according to one or more embodiments, which is configured to use an optical encoder to monitor the displacement of a user's ear.

[0023] Figure 4 This is a block diagram of an audio system according to one or more embodiments.

[0024] Figure 5 This is a flowchart of a process for calibrating a virtual microphone according to one or more embodiments.

[0025] Figure 6 This is a block diagram of an example artificial reality system environment according to one or more embodiments.

[0026] These accompanying drawings depict various embodiments for illustrative purposes only. Those skilled in the art will readily recognize from the following discussion that alternative embodiments of the structures and methods illustrated herein can be employed without departing from the principles described herein. Detailed Implementation

[0027] The audio system calibrates a "virtual microphone" located at the entrance to the ear canal of the user's ear. In practice, the virtual microphone simulates the presence of a microphone at the entrance to the ear canal by characterizing how sound is detected there. The audio system plays a calibration signal via a transducer and then measures the displacement of at least a portion of the user's ear caused by the calibration signal. The audio system provides this displacement information as input to a model, which outputs an estimated sound pressure level at the entrance to the ear canal of the user's ear. In some embodiments, the audio system calibrates the virtual microphones for both of the user's ears. The audio system can use the estimated sound pressure level at the entrance to the ear canal to generate a sound filter and use the sound filter to adjust the user's audio content.

[0028] Head-mounted devices present audio content to users. To improve the user's auditory experience, conventional audio systems require a targeted microphone to be placed at the entrance of the ear canal. Therefore, the audio system characterizes how sound is perceived at the entrance of the ear canal. However, this conventional calibration technique is often impractical or uncomfortable for users. In contrast, the audio system described in this paper eliminates the need for conventional calibration techniques, namely calibrating virtual microphones at the entrance of the ear canal in one or both ears of the user.

[0029] Embodiments of the present invention may include an artificial reality system or a combination thereof. Artificial reality is a form of reality that has been adjusted in some way before being presented to a user. This artificial reality may include, for example, virtual reality (VR), augmented reality (AR), mixed reality (MR), hybrid reality, or some combination and / or derivative thereof. Artificial reality content may include fully generated content or content generated by combining it with captured (e.g., real-world) content. Artificial reality content may include video, audio, haptic feedback, or some combination thereof, and any of the above may be presented in a single-channel or multi-channel manner (e.g., stereoscopic video providing a three-dimensional effect to the viewer). Furthermore, in some embodiments, artificial reality may also be associated with applications, products, accessories, services, or some combination thereof for creating content in artificial reality and / or otherwise used in artificial reality (e.g., performing activities in artificial reality). Artificial reality systems that deliver artificial reality content can be implemented on a variety of platforms, including head-mounted devices (e.g., head-mounted displays (HMDs) and / or near-eye displays (NEDs)) connected to a host computer system, stand-alone head-mounted devices, mobile devices or computing systems, or any other hardware platform capable of delivering artificial reality content to one or more viewers.

[0030] System Overview

[0031] Figure 1AThis is a perspective view of a head-mounted device 100 implemented as an eyewear device according to one or more embodiments, configured to calibrate a virtual microphone. In some embodiments, the eyewear device is a near-eye display (NED). Typically, the head-mounted device 100 can be worn on a user's face such that content (e.g., media content) is presented using a display component and / or an audio system. However, the head-mounted device 100 can also be used to present media content to a user in a different manner. Examples of media content presented by the head-mounted device 100 include one or more images, videos, audio, or some combination thereof. The head-mounted device 100 includes a frame and may include a display component (which includes one or more display elements 120), a depth camera assembly (DCA), and other components such as an audio system. Although Figure 1A The illustration shows example locations of components of the head-mounted device 100 on the head-mounted device 100, but these components may be located at other locations on the head-mounted device 100, on peripheral devices paired with the head-mounted device 100, or some combination thereof. Similarly, there may be more components on the head-mounted device 100 than... Figure 1A The number of components shown may be more or less.

[0032] The frame 110 holds other components of the head-mounted device 100. The frame 110 includes a front portion that holds one or more display elements 120, and end pieces (e.g., temples) that attach to the user's head. The front portion of the frame 110 extends across the top of the user's nose. The length of the end pieces may be adjustable (e.g., adjustable temple length) to fit different users. The end pieces may also include curved portions behind the user's ears (e.g., temple covers, eyeglass temples).

[0033] One or more display elements 120 provide light to a user wearing a head-mounted device 100. As shown, the head-mounted device includes a display element 120 for each of the user's eyes. In some embodiments, the display element 120 generates image light that is provided to the eyebox of the head-mounted device 100. The eyebox is the location in the space occupied by the user's eyes when wearing the head-mounted device 100. For example, the display element 120 may be a waveguide display. A waveguide display includes a light source (e.g., a two-dimensional source, one or more line sources, one or more point sources, etc.) and one or more waveguides. Light from the light source is internally coupled into one or more waveguides, which output light in a manner that creates a pupil replication in the eyebox of the head-mounted device 100. The internal coupling of light and / or the external coupling of light from one or more waveguides may be accomplished using one or more diffraction gratings. In some embodiments, the waveguide display includes a scanning element (e.g., a waveguide, a mirror, etc.) that scans the light as it is internally coupled into one or more waveguides. Note that in some embodiments, one or both of the two display elements 120 are opaque and do not transmit light from a local area surrounding the head-mounted device 100. This local area is the area surrounding the head-mounted device 100. For example, this local area could be a room where a user wearing the head-mounted device 100 is inside, or an outdoor area where the user wearing the head-mounted device 100 may be outside. In this context, the head-mounted device 100 generates VR content. Alternatively, in some embodiments, one or both of the two display elements 120 are at least partially transparent, such that light from the local area can be combined with light from the one or more display elements to generate AR and / or MR content.

[0034] In some embodiments, display element 120 does not generate image light, but instead uses a lens to transmit light from a localized area to the eye-friendly area. For example, one or both of the two display elements 120 may be an uncorrected (over-the-counter) lens or a prescription lens (e.g., a single-vision lens, bifocal and trifocal lens, or graduated lens) to help correct the user's visual impairment. In some embodiments, display element 120 may be polarized and / or tinted to protect the user's eyes from the sun.

[0035] In some embodiments, display element 120 may include an additional optics block (not shown). The optics block may include one or more optical elements (e.g., lenses, Fresnel lenses, etc.) that guide light from display element 120 to the eye-friendly area. The optics block may, for example, correct aberrations in some or all of the image content, magnify some or all of the image, or some combination thereof.

[0036] DCA determines depth information for a portion of a local area surrounding the head-mounted device 100. DCA includes one or more imaging devices 130 and a DCA controller. Figure 1A (Not shown in the image), and may also include an illuminator 140. In some embodiments, the illuminator 140 uses light to illuminate a portion of a local area. This light may be, for example, infrared (IR) structured light (e.g., dot-patterned structured light, strip structured light, etc.), an IR flash for time-of-flight, etc. In some embodiments, the one or more imaging devices 130 acquire an image of a portion of the local area including light from the illuminator 140. As shown, Figure 1A A single illuminator 140 and two imaging devices 130 are shown. In an alternative embodiment, the illuminator 140 is absent and at least two imaging devices 130 are present.

[0037] The DCA controller uses acquired images and one or more depth determination techniques to calculate depth information for a portion of a local region. Depth determination techniques may include, for example, direct time-of-flight (ToF) depth sensing, indirect ToF depth sensing, structured light, passive stereo analysis, active stereo analysis (using textures added to the scene by light from illuminator 140), some other technique for determining the depth of the scene, or some combination thereof. In some embodiments, the head-mounted device 100 may provide simultaneous localization and mapping (SLAM) for the position of the head-mounted device 100 and for updating the model of the local region. For example, the head-mounted device 100 may include a passive camera assembly (PCA) that generates color image data. The PCA may include one or more RGB cameras that acquire images of some or all portions of the local region. In some embodiments, some or all of the imaging devices 130 in the DCA may also be used as the PCA. The images acquired by the PCA and the depth information determined by the DCA may be used to determine parameters of the local region, generate a model of the local region, update the model of the local region, or some combination thereof. In some embodiments, a sensor array (as discussed below) generates measurement signals in response to movement of the head-mounted device 100 and tracks the positioning (e.g., position and orientation) of the head-mounted device 100 within a room.

[0038] An audio system presents audio content to a user. The audio system includes a transducer array, a sensor array, and an audio controller 160. However, in other embodiments, the audio system may include different components and / or additional components. Similarly, in some cases, the functions described for the components in the reference audio system may be distributed among these components in a manner different from that described herein. For example, some or all of the controller's functions may be performed by a remote server.

[0039] A transducer array presents sound to the user. The transducer array includes one or more transducers, including one or more tissue-conducting transducers 170 and one or more air-conducting transducers 180. In some embodiments, one or more transducers in the transducer array are housed within a frame 110. In some embodiments, the head-mounted device 100 includes one or more transducers along each temple of the frame 110 and / or at the end of each temple of the frame 110. Therefore, multiple transducers can improve the directionality of the presented audio content.

[0040] One or more tissue conduction transducers 170 generate sound via tissue conduction. Each tissue conduction transducer 170 may be, for example, a cartilage conduction transducer and / or a bone conduction transducer. The tissue conduction transducer 170 is incorporated into the user's tissue (e.g., bone and / or cartilage) and directly vibrates the user's tissue to generate sound waves perceived by at least one of the user's inner ears. The user thus perceives the sound waves as sound. Each tissue conduction transducer 170 is positioned close to and / or in contact with tissue of the user's ear (e.g., on the back of the auricle). In some embodiments, the head-mounted device 100 includes at least one tissue conduction transducer 170 at each of the user's ears. The number and / or location of the tissue conduction transducers 170 may vary. Figure 1A The quantities and / or positions shown are different.

[0041] One or more air conduction transducers 180 generate sound via air conduction. The air conduction transducer 180 may be, for example, a loudspeaker that generates sound waves perceived as sound by at least one inner ear of a user. In some embodiments, a plurality of air conduction transducers 180 are positioned on and / or along the frame 110 of the head-mounted device 100. The number and / or position of the air conduction transducers 180 may vary with... Figure 1A The quantities and / or positions shown are different.

[0042] The head-mounted device 100 has a sensor array that measures various parameters. The sensor array includes one or more acoustic sensors 185 and one or more displacement sensors 190. In some embodiments, the sensor array may include additional sensors in addition to and / or in lieu of those sensors described herein.

[0043] One or more acoustic sensors 185 detect sound within a localized area of ​​the head-mounted device 100. The acoustic sensors 185 acquire sound emitted from one or more sound sources, including an array of transducers, within the localized area (e.g., a room). The sound detected by the acoustic sensors 185 is used to calibrate a “virtual” microphone at the user’s ear canal entrance. This virtual microphone is not a physical device, but rather a virtual device that simulates the presence of a physical microphone and can therefore be used by the head-mounted device 100 to characterize how sound is perceived at the simulated location of the virtual microphone. For example, the head-mounted device 100 can better represent spatialized audio content based on the calibrated virtual microphone at the ear canal entrance.

[0044] Each acoustic sensor is configured to detect sound and convert the detected sound into an electronic format (analog or digital). Acoustic sensor 185 may be a sound wave sensor, microphone, sound transducer, or similar sensor suitable for detecting sound. In some embodiments, acoustic sensor 185 may be placed on the outer surface of head-mounted device 100, placed on the inner surface of head-mounted device 100, separate from head-mounted device 100 (e.g., part of some other device), or some combination thereof. The number and / or location of acoustic sensors 185 may vary. Figure 1A The number and / or locations shown may differ. For example, the number of acoustic detection locations can be increased to increase the amount of audio information collected and the sensitivity and / or accuracy of that information. The acoustic detection locations can be oriented such that the virtual microphone is calibrated to take into account sound from a wide range of directions around the user wearing the headset 100.

[0045] One or more displacement sensors 190 measure the displacement of multiple parts of a user's ear. For example, each displacement sensor 190 may be individually coupled to one of the user's ears. After audio content is generated by one or more transducers, each part of the ear may vibrate partially due to that audio content. Therefore, the displacement sensor 190 measures the displacement caused partially by the vibration of a part of the ear. In some embodiments, the head-mounted device 100 may assign more than one displacement sensor 190 to each of the user's ears. Each displacement sensor 190 may be configured to measure the displacement of a different part of the user's ear. The displacement sensor 190 may be an optical displacement sensor, an inertial measurement unit, an accelerometer, a velocity meter, a gyroscope, or another suitable type of sensor for detecting motion, or some combination thereof.

[0046] Displacement sensor 190 can be positioned in addition to Figure 1A In locations other than those shown in the diagram. In some embodiments, displacement sensors 190 measure displacement of a portion of the user's facial tissue due to vibration. For example, displacement sensors 190 may measure displacement of a portion of the user's temples, forehead, etc. In some embodiments, one or more of the aforementioned displacement sensors 190 may be coupled to a portion of the head-mounted device 100 that contacts the user's nose. Thus, these displacement sensors 190 measure displacement of facial tissue caused by bone conduction of the user's voice.

[0047] In some embodiments, at least one displacement sensor 190 is part of a tissue conduction transducer and is located within the tissue conduction transducer. For example, displacement sensor 190 can measure the displacement of a portion of the user's ear that is attached to the tissue conduction transducer. Regarding Figure 3A and Figure 3B An embodiment of a cartilage conduction transducer including a displacement sensor is described in more detail.

[0048] Audio controller 160 processes information from the sensor array and instructs the transducer array to present audio content. In some embodiments, audio controller 160 calibrates a virtual microphone at the user's ear canal entrance based on measurements from displacement sensor 190. Audio controller 160 can calibrate the virtual microphone for one or both ears of the user (e.g., at each entrance of the ear canal). For a given ear, audio controller 160 takes as input measurements of the displacement of at least a portion of that ear. A model executed by audio controller 160 correlates the measured displacement information with the estimated sound pressure at the ear canal entrance using a functional mapping from the measured displacement information to the estimated sound pressure. Therefore, audio controller 160 outputs the estimated sound pressure at the ear canal entrance, and based on the estimated sound pressure, audio controller 160 can generate a sound filter and apply the sound filter to the audio content. For example, the sound filter can better spatialize the audio content, prevent sound leakage (e.g., by amplifying and / or attenuating some or all frequencies of the audio content), and improve the clarity of the audio content (e.g., by amplifying frequencies that the user might otherwise mishear). Additionally, users can experience improved audio quality and perceive audio content more naturally. The audio controller 160 instructs the transducer array to present the resulting filtered audio content. In some embodiments, the audio controller 160 may include a processor and a computer-readable storage medium. Furthermore, the audio controller 160 may be configured to generate direction of arrival (DOA) estimates, generate acoustic transfer functions (e.g., array transfer function and / or head-related transfer function), track the location of sound sources, form beams in the direction of sound sources, classify sound sources, or some combination thereof.

[0049] Figure 1B This is a perspective view of a head-mounted display (HMD) implemented according to one or more embodiments, configured to calibrate a virtual microphone. In embodiments describing AR and / or MR systems, multiple portions of the front of the HMD are at least partially transparent to the visible wavelength range (~380 nm to 750 nm), and multiple portions of the HMD located between the front of the HMD and the user's eyes are at least partially transparent (e.g., partially transparent electronic displays). The HMD includes a front rigid body 115 and a band 195. The head-mounted device 105 includes the components referenced above. Figure 1A The components described are some of the same, but these components are modified to be combined with the HMD form factor. For example, an HMD includes a display component, DCA, and audio system. Figure 1BMultiple imaging devices 130, an illuminator 140, an audio controller 160, a tissue transducer 170, an air transducer 180, an acoustic sensor 185, and a displacement sensor 190 are shown. The various components can be located in various positions, for example, coupled to a band 195 (as shown), coupled to a front rigid body 115, or configured to be inserted into a user's ear canal.

[0050] Headset for calibrating virtual microphones

[0051] Figure 2 This is a side view of a head-mounted device 205 configured to calibrate a virtual microphone 207, according to one or more embodiments. The head-mounted device 205 simulates the presence of the virtual microphone 207 at the ear canal entrance 210 of a user's ear. The virtual microphone 207 is used to characterize how audio content is perceived at the ear canal entrance 210. Figure 2 A portion of the head-mounted device 205 shown includes an air conduction transducer 220, a cartilage conduction transducer 230, and one or more displacement sensors 240. The head-mounted device 205 may be... Figure 1A The head-mounted device 100 is an embodiment of the device, and therefore may also include components other than those shown herein. For example, the head-mounted device 205 may include a controller, a display assembly, etc.

[0052] Air conduction transducer 220 can present audio content to a user. Air conduction transducer 220 can be a speaker that presents audio content via air conduction. In some embodiments, air conduction transducer 220 is a component in a transducer array of head-mounted device 205. In some embodiments, the transducer array includes a plurality of air conduction transducers configured to provide audio content to one or both ears of a user.

[0053] The cartilage conduction transducer 230 can present audio content to the user via cartilage conduction. The cartilage conduction transducer 230 can be positioned directly and / or indirectly in contact with and / or near the ear tissues. The cartilage conduction transducer 230 causes the contacting portion in the ear to vibrate, thereby generating a series of sound pressure waves that are transmitted to the cochlea of ​​the user's inner ear. Figure 2 (Not shown in the image) was detected as sound. Figure 2In this embodiment, when a user wears the head-mounted device 205, the cartilage conduction transducer 230 contacts the auricle 215. In other embodiments, the cartilage conduction transducer 230 may be positioned to contact the tragus of the ear, the earlobe of the ear, another part of the ear, or a combination thereof. In some embodiments, the cartilage conduction transducer 230 is a component of a transducer array of the head-mounted device 205. In some embodiments, the transducer array includes a plurality of cartilage conduction transducers configured to provide audio content to one or both ears of a user.

[0054] Displacement sensor 240 measures displacement of a portion of auricle 215. In some embodiments, one of the plurality of displacement sensors 240 is incorporated into a portion of the back of auricle 215. In other embodiments, at least one of the plurality of displacement sensors 240 is incorporated into the top of auricle 215. Displacement sensor 240 measures displacement of auricle 215 when auricle 215 vibrates due to audio content generated by air conduction transducer 220 and / or cartilage conduction by cartilage conduction transducer 230. Displacement sensor 240 may be an accelerometer, an optical displacement sensor, or some combination thereof. In some embodiments, displacement sensor 240 is integrated into and / or coupled to cartilage conduction transducer 230. For example, displacement sensor 240 may measure displacement of a portion of a user's ear incorporated into a tissue conduction transducer. Figures 3A to 3B To provide a more detailed description.

[0055] In some embodiments (not shown), one or more displacement sensors measure displacement of the user's face and / or other parts of the ear that move and / or vibrate in response to audio content generated by the air conduction transducer 220 and / or the cartilage conduction transducer 230. For example, the displacement sensors may contact and measure the displacement of a portion of the user's face, such as the temples or forehead.

[0056] Monitoring ear displacement via cartilage conduction transducer

[0057] Figure 3AThe block diagram illustrates a cartilage conduction transducer 300 according to one or more embodiments, configured to monitor the displacement of a user's ear using a capacitive displacement sensor 310. The cartilage conduction transducer 300 presents audio content to the user via cartilage conduction and is configured to measure the displacement of a portion of the user's ear. The cartilage conduction transducer 300 may be a component of a head-mounted device (e.g., head-mounted device 205) and is an embodiment of a cartilage conduction transducer 230 coupled to a displacement sensor 240. The cartilage conduction transducer 300 includes magnets 320A and 320B (collectively referred to as magnet 320), a motion coil 330, a contact pad 340, a preload spring 350, and a capacitive displacement sensor 310. The cartilage conduction transducer 300 may include, in addition to... Figure 3A Other components besides those shown.

[0058] Magnet 320 generates a magnetic field that causes the movable coil 330 to vibrate. Magnet 320 includes soft and / or hard magnets. For example, magnet 320A may be a soft magnet and magnet 320B may be a hard magnet. Soft magnets may be made of steel and / or nickel-plated, while hard magnets may be neodymium magnets and / or zinc-plated. Cartilage conduction transducer 300 may include, in addition to Figure 3A and Figure 3B Magnets other than those shown.

[0059] The moving coil 330 responds to an input signal and vibrates due to the magnetic field generated by the magnet 320. When current flows through it, the moving coil 330 experiences a Lorentz force, which causes it to vibrate at a frequency specified in the input signal. The moving coil 330 may be a printed circuit board (PCB) or another structure with sufficient rigidity to receive the Lorentz force. In some embodiments, the moving coil 330 may include a flexible printed circuit.

[0060] The contact pad 340 is attached to the tissue of the user's ear. In order to present audio content via cartilage conduction, the cartilage conduction transducer 300 causes the tissue at and / or near the user's ear (e.g., the auricle) to vibrate. The contact pad 340 is in direct and / or indirect contact with this tissue, which vibrates as the moving coil 330 moves.

[0061] A preloaded spring 350 positions the cartilage conduction transducer 300 in contact with the tissue of the user's ear. When the user wears a head-mounted device including the cartilage conduction transducer 300, the preloaded spring 350 is configured to position the cartilage conduction transducer 300 in contact with the user's ear at a nominal position. When the cartilage conduction transducer 300 is in the nominal position, the preloaded spring 350 can have a predictable response when vibrating against the tissue of the user's ear. When measured, the displacement of the preloaded spring 350 characterizes the contact force from the cartilage conduction transducer 300 to the tissue of the ear. In some embodiments, the preloaded spring 350 can be used as an error detection mechanism. For example, when the preload amount of the preloaded spring 350 exceeds a threshold amount, such that the cartilage conduction transducer 300 may not be in contact with the tissue of the user's ear, the user can be notified that the head-mounted device needs to be repositioned.

[0062] A capacitive displacement sensor 310 measures the displacement of the contact pad 340. When the cartilage conduction transducer 300 presents audio content via cartilage conduction, the displacement can be partly attributed to the vibration of the movable coil 330. The capacitive displacement sensor 310 accordingly measures the displacement of the contact pad 340 in the ear, which is to which, and therefore to which, the cartilage conduction transducer 300 is to which. For example, the cartilage conduction transducer 300 may receive instructions (e.g., from a controller of a head-mounted device) to present audio content. The movable coil 330 vibrates when presenting audio content, causing it to shift from its rest position. The displacement of the movable coil 330 causes a change in capacitance detected by the capacitive displacement sensor 310. In some embodiments, the capacitive displacement sensor 310 determines the displacement of the auricle 215 caused by audio content presented by an air conduction transducer (e.g., air conduction transducer 220). In some embodiments, the capacitive displacement sensor 310 is constructed in such a way that it includes two electrodes spaced apart to form a capacitor. When the moving coil 330 vibrates, the distance between the two electrodes of the capacitive displacement sensor 310 changes, thereby altering the capacitance. In other embodiments, the capacitive displacement sensor 310 is configured differently from the capacitive displacement sensor described herein.

[0063] Figure 3BThe diagram illustrates a cartilage conduction transducer 360 according to one or more embodiments, configured to monitor displacement of a user's ear using an optical encoder 370. The cartilage conduction transducer 360 is structurally and functionally similar to the cartilage conduction transducer 300, except that it includes an optical encoder 370 for measuring displacement of the user's ear instead of a capacitive displacement sensor 310. The optical encoder 370 can have higher sensitivity and can therefore be used to measure smaller displacement values ​​of the user's ear. In some embodiments, the cartilage conduction transducer 360 differs from the cartilage conduction transducer 300 structurally and / or functionally.

[0064] When the cartilage conduction transducer 360 presents audio content via cartilage conduction, the optical encoder 370 measures the displacement of the contact pad 340 caused by the vibration of the movable coil 330. Similar to the capacitive displacement sensor 310, the optical encoder 370 can be used to determine the displacement of a portion of the tissue in the user's ear caused by the cartilage conduction transducer 360 and / or the air conduction transducer. For example, the optical encoder 370 can be used to determine the displacement of the auricle 215. In some embodiments, the optical encoder 370 includes a light source (e.g., an LED) and a mechanism (e.g., a shaft) that moves the light emitted by the light source when the movable coil 330 is vibrating. Thus, the optical encoder 370 monitors the position of the light caused by the movable coil 330, thereby measuring the displacement of the portion of the ear to which the cartilage conduction transducer 360 is attached.

[0065] Audio System Overview

[0066] Figure 4 This is a block diagram of an audio system 400 according to one or more embodiments. The audio system 400 provides audio content to a user. In some embodiments, the audio system 400 calibrates: (1) a virtual microphone positioned at the entrance to the ear canal of the user's left ear (e.g., ear canal entrance 210); (2) a virtual microphone positioned at the entrance to the ear canal of the user's right ear; or (3) virtual microphones positioned at the respective entrances to the ear canals of the right and left ears. The audio system may adjust the user's audio content in part based on the calibrated virtual microphones. The audio system 400 may be a component in and / or coupled to a head-mounted device (e.g., head-mounted device 100, 105). The audio system 400 includes a transducer array 410, a sensor array 420, and a controller 430. In some embodiments, the audio system 400 includes additional components.

[0067] Transducer array 410 presents audio content to a user according to instructions from controller 430. Transducer array 410 includes one or more transducers that present audio content via air conduction (e.g., air conduction transducer 220) and / or one or more transducers that present audio content via tissue conduction (e.g., cartilage conduction transducer 230). Transducer array 410 can be configured to present audio content in a frequency range such as 20 Hz (Hertz) to 20 kHz (kilohertz), which is generally near the average human hearing range. In some embodiments, transducer array 410 presents adjusted (e.g., filtered, enhanced, amplified, or attenuated) audio content.

[0068] Sensor array 420 measures various parameters related to the head-mounted device. Sensor array 420 includes one or more acoustic sensors (e.g., acoustic sensor 185) and / or one or more displacement sensors (e.g., displacement sensor 240). The acoustic sensors detect sound from a localized area, which is used to generate one or more virtual microphones for one or both ears of the user. The displacement sensors calibrate the generated virtual microphones by measuring the displacement of a portion of the user's ear (e.g., auricle 215). This portion of the user's ear may vibrate and / or displace due to the audio content generated by transducer array 410, and this vibration and / or displacement is measured by the displacement sensors in sensor array 420. In some embodiments, at least one of the plurality of displacement sensors is integrated into a cartilage conduction transducer of transducer array 410. The displacement sensors may be optical displacement sensors, inertial measurement units, accelerometers, gyroscopes, or another suitable type of sensor for detecting motion, or some combination thereof.

[0069] In some embodiments, sensor array 420 further includes one or more acoustic sensors (e.g., acoustic sensor 185) configured to detect sound. The acoustic sensors may be configured to detect sound pressure waves from a localized area surrounding the user and convert the detected sound pressure waves into analog and / or digital formats. The acoustic sensors may be, for example, microphones, accelerometers, another sensor for detecting sound pressure waves, or some combination thereof.

[0070] Controller 430 processes the data received from sensor array 420 and instructs transducer array 410 to present audio content, thereby enabling audio system 400 to calibrate a virtual microphone at the user's ear canal entrance. The audio controller 160 of Figure 1 is an embodiment of controller 430. Controller 430 includes a data store 435, a direction of arrival (DOA) estimation module 440, a transfer function module 450, a tracking module 460, a beamforming module 470, a sound pressure level estimation module 480, and a sound filter module 490. In some embodiments, controller 430 includes other modules and / or components besides those described herein.

[0071] Data storage area 435 stores data related to audio system 400. This data may include, for example, measured displacement information of one or both ears of a user, calibration signals, training model data, generated sound filters, other information about audio system 400, or some combination thereof. Furthermore, the data in data storage area 435 may include sounds recorded in local areas of audio system 400, audio content, head-related transfer functions (HRTF), transfer functions of one or more sensors, array transfer functions (ATF) of one or more acoustic sensors, sound source locations, virtual models of local areas, direction-of-arrival estimation results, sound filters, and other data related to the use of audio system 400, or any combination thereof.

[0072] The DOA estimation module 440 is configured to locate sound sources in a local area, in part based on information from the sensor array 420. Localization is the process of determining the location of a sound source relative to the user of the audio system 400. The DOA estimation module 440 performs DOA analysis to locate one or more sound sources within the local area. DOA analysis may include analyzing the intensity, spectrum, and / or time of arrival of each sound at the sensor array 420 to determine the direction of origin of the sound. In some cases, DOA analysis may include any suitable algorithm used to analyze the surrounding acoustic environment in which the audio system 400 is located.

[0073] For example, DOA analysis can be designed to receive input signals from sensor array 420 and apply digital signal processing algorithms to these input signals to estimate the direction of arrival (DOA). These algorithms may include, for example, a delay summation algorithm, in which the input signal is sampled and the weighted and delayed versions of the resulting sampled signals are averaged together to determine the DOA. A least mean squared (LMS) algorithm can also be implemented to create an adaptive filter. This adaptive filter can then be used to identify, for example, differences in signal strength or differences in arrival time. These differences can then be used to estimate the DOA. In another embodiment, the DOA can be determined by converting the input signal to the frequency domain and selecting specific frequency bins within the time-frequency (TF) domain for processing. Each selected TF frequency bin can be processed to determine whether the frequency bin includes a portion of the audio spectrum containing a direct-path audio signal. Those frequency bins containing the direct-path signal portion can then be analyzed to identify the angle at which sensor array 420 receives the direct-path audio signal. The determined angle can then be used to identify the DOA of the received input signal. Other algorithms not listed above can be used alone or in combination with the algorithms above to determine DOA.

[0074] In some embodiments, the DOA estimation module 440 can also determine the DOA related to the absolute position of the audio system 400 within a local area. The position of the sensor array 420 can be received from an external system (e.g., another component of the head-mounted device, an AI console, a mapping server, a position sensor, etc.). The external system can create a virtual model of the local area, mapping the local area and position of the audio system 400 within that model. The received position information may include the position and / or orientation of some or all of the audio system 400 (e.g., the sensor array 420). The DOA estimation module 440 can update the estimated DOA based on the received position information.

[0075] The transfer function module 450 is configured to generate one or more acoustic transfer functions. Typically, a transfer function is a mathematical function that gives a corresponding output value for each possible input value. Based on the parameters of the detected sound, the transfer function module 450 generates one or more acoustic transfer functions associated with the audio system. The acoustic transfer function can be an array transfer function (TF), a head-related transfer function (HRTF), other types of acoustic transfer functions, or some combination thereof. The ATF characterizes how a microphone receives sound from a point in space.

[0076] The ATF comprises multiple transfer functions that characterize the relationship between a sound source and the corresponding sound received by the multiple acoustic sensors in the sensor array 420. Therefore, for a sound source, there exists a corresponding transfer function for each acoustic sensor in the sensor array 420. This set of transfer functions is collectively referred to as the ATF. Thus, for each sound source, there exists a corresponding ATF. It should be noted that the sound source can be, for example, a person or object generating sound in a local area, a user, or one or more transducers in the transducer array 410. Because human physiology (e.g., ear shape, shoulders, etc.) affects the sound as it travels towards the ear, the ATF relative to a specific sound source location in the sensor array 420 may vary from user to user. Therefore, these ATFs of the sensor array 420 are personalized for each user of the audio system 200.

[0077] In some embodiments, the transfer function module 450 determines one or more HRTFs for a user of the audio system 400. HRTFs characterize how an ear receives sound from a point in space. Because a person's physiological structure (e.g., ear shape, shoulders, etc.) affects the sound as it travels towards the ear, the HRTF relative to a particular sound source location is unique for each ear of that person (and thus unique to that person). In some embodiments, the transfer function module 450 may use a calibration process to determine the user's HRTF. In some embodiments, the transfer function module 450 may provide information about the user to a remote system. The user can adjust privacy settings to allow or prevent the transfer function module 450 from providing information about the user to any remote system. The remote system uses, for example, machine learning to determine a set of HRTFs tailored to the user and provides this tailored set of HRTFs to the audio system 400.

[0078] Tracking module 460 is configured to track the location of one or more sound sources. Tracking module 460 can compare multiple current DOA estimates and compare them to a stored history of previous DOA estimates. In some embodiments, audio system 400 can recalculate the DOA estimates according to a periodic schedule (e.g., once per second or once per millisecond). The tracking module can compare current DOA estimates with previous DOA estimates, and in response to changes in the DOA estimates of a sound source, tracking module 460 can determine that the sound source has moved. In some embodiments, tracking module 460 can detect changes in location based on visual information received from a head-mounted device or some other external source. Tracking module 460 can track the movement of one or more sound sources over time. Tracking module 460 can store the number of sound sources and the location of each sound source at each time point. In response to changes in the number or location of sound sources, tracking module 460 can determine that the sound source has moved. Tracking module 460 can calculate an estimate of the localization variance. The localization variance can be used as the confidence level for each determination of movement change.

[0079] Beamforming module 470 is configured to process one or more ATFs to selectively emphasize sound from a sound source within a certain region while disregarding sound from other regions. When analyzing sound detected by sensor array 420, beamforming module 470 can combine information from different acoustic sensors to emphasize sound associated with a specific area of ​​the local region while disregarding sound from outside that area. Beamforming module 470 can isolate audio signals associated with sound from a specific sound source from other sound sources in the local region based on, for example, different DOA estimation results from DOA estimation module 440 and tracking module 460. Beamforming module 470 can therefore selectively analyze discrete sound sources in the local region. In some embodiments, beamforming module 470 can amplify the signal from the sound source. For example, beamforming module 470 can apply a sound filter that eliminates signals above certain frequencies, below certain frequencies, or between certain frequencies. Signal enhancement is used to amplify the sound associated with a given identified sound source relative to other sounds detected by sensor array 420.

[0080] When audio content is played, the sound pressure estimation module 480 estimates the sound pressure at the entrance to the ear canal. The sound pressure estimation module 480 uses a model to characterize how the audio content is perceived at the entrance to the ear canal (e.g., predicting what the binaural microphones at the entrance to the ear canal will detect in response to the audio content), thereby simulating a virtual microphone. The sound pressure estimation module 480 may instruct the transducer array 410 to play a calibration signal via air conduction and / or tissue conduction. The calibration signal may be audio content that generates sound waves perceptible to the user, such as playing a tone for a period of time, a piece of music, etc. The sound pressure estimation module 480 uses displacement information of a portion of the ear (e.g., the auricle) from the sensor array 420 (e.g., as measured by one or more displacement sensors in the sensor array 420). The displacement of said portion of the ear is at least partially due to the calibration signal.

[0081] The sound pressure estimation module 480 uses a model to estimate the sound pressure at the entrance of the ear canal. This model can be configured to take measured displacement information of a portion of the ear as input and output the estimated sound pressure at the entrance of the ear canal accordingly. In some embodiments, when outputting the estimated sound pressure at the entrance of the ear canal, the model is configured to take into account the geometry of the user's ear (e.g., measurements of features of the user's ear). The geometry of the user's ear can be determined based on images and / or videos of the user. The model can be, for example, a machine learning model, such as a convolutional neural network, a linear model, numerical simulation, or some combination thereof. The model can be trained and / or built on a dataset that includes data from multiple other users. For each of the multiple other users, the data correlates the measured displacement information of a portion of the ear with the sound pressure at the entrance of the ear canal (e.g., measured by binaural microphones). In some embodiments, the model can correlate the displacement information of a portion of the user's ear with the sound pressure at the entrance of the ear canal based on the following formula.

[0082]

[0083] In equation (1) shown above, p represents sound pressure, a represents acceleration, and F represents the functional mapping between p and a. This model assumes a strong linear relationship between acceleration and sound pressure if there is high coherence between p and a in the time domain or between P (e.g., the complex frequency response of p) and A (e.g., the complex frequency response of a) in the frequency domain. If p is considered a time-invariant function of a, then p can be described by a according to the following equation:

[0084] p(t)=a(t)*h(t) (2)

[0085] The above formula describes temporal convolution, or:

[0086] P(f)=A(f)H(f) (3)

[0087] Equation (3) describes the spectral multiplication in the frequency domain, where h(t) or H(f) characterizes the transfer function between the outer ear vibration and the corresponding sound pressure. Considering the linear time-invariant (LTI) relationship between the outer ear acceleration a and the sound pressure p at the ear canal entrance, the pressure at the ear canal entrance can then be calibrated by calibrating the right side of Equations (2) and (3).

[0088] The sound pressure level estimation module 280 also distinguishes between the displacement of the auricle caused by the audio content presented by the transducer array 410 and the displacement caused by other noise. The sound pressure level estimation module 280 uses a correlation model to measure the correlation between the audio content output by the transducer array 410 and the displacement measured by the displacement sensors in the sensor array 420. A high correlation indicates that the displacement is largely due to the audio content presented by the transducer array 410.

[0089] The sound filter module 490 generates one or more sound filters for the user based on the estimated sound pressure at the entrance of the ear canal. The estimated sound pressure at the entrance of the ear canal indicates how the user perceives the audio content at the entrance of the ear canal, and the sound filter module 490 generates sound filters to adjust the audio content accordingly. Examples of sound filters include low-pass filters, high-pass filters, band-pass filters, etc. When applied to audio content, the sound filter adjusts the audio content to improve the user's auditory experience. For example, the user may perceive the adjusted audio content as filtered, enhanced, amplified, attenuated, or some combination thereof. In some embodiments, the sound filter results in adjusted audio content having a frequency response (e.g., a flat frequency response) of a target magnitude. In other embodiments, the sound filter may target a specific frequency range to help a hearing-impaired user hear more clearly within those frequency ranges. After adjusting the audio content using the sound filter, the sound filter module 490 instructs the transducer array 410 to present the adjusted audio content to the user. In some embodiments, a user provides feedback to the audio system 400 regarding the adjusted audio content, which can be incorporated into a dataset used by the sound pressure estimation module 480 to train and / or build a model.

[0090] In some embodiments, the audio system 400 may use a sound filter to spatialize audio content so that the audio content sounds as originating from a target area within a local region. The sound filter module 490 may use HRTF and / or acoustic parameters to generate the sound filter. The acoustic parameters describe the acoustic characteristics of the local region. Acoustic parameters may include, for example, reverberation time, reverberation level, room impulse response, etc. In some embodiments, the sound filter module 490 calculates one or more of these acoustic parameters. In some embodiments, the sound filter module 490 obtains data from a map building server (e.g., as described below regarding...). Figure 6 (Description) Request acoustic parameters.

[0091] Figure 5 This is a flowchart of a process 500 for calibrating a virtual microphone according to one or more embodiments. The process 500 may be performed by a component of an audio system (e.g., audio system 400). In some embodiments, the audio system is a component of a head-mounted device (e.g., head-mounted device 205) configured to calibrate a virtual microphone at the entrance to the ear canal of a user by performing the process 500. In some embodiments, the audio system performs the process 500 for one or both ears of the user. In other embodiments, other entities may perform the process. Figure 5 Some or all of the steps in the process. Implementations may include different steps and / or additional steps, or perform these steps in a different order.

[0092] The audio system presents audio content 510 to a user via one or more transducers (e.g., transducers in transducer array 410). The transducers may generate the audio content based on instructions from a controller (e.g., controller 430) of the audio system 400. The audio content may be a calibration signal. The transducers may be air conduction transducers (e.g., air conduction transducer 180), tissue conduction transducers (e.g., tissue conduction transducer 170), or some combination thereof.

[0093] The audio system monitors displacement of the auricles (e.g., auricle 215) of one or both ears of a user 520 via one or more sensors (e.g., sensors in sensor array 420). The displacement of one or both auricles may be partly due to vibrations caused by the audio content. In some embodiments, the sensor monitoring the displacement of one or both auricles may be integrated with and / or coupled to one or more transducers of the audio system. For example, for a given auricle, the sensor may monitor displacement of the auricle caused by a cartilaginous conductive transducer attached to the auricle. One or more of the aforementioned sensors may be a displacement sensor (e.g., displacement sensor 240) and / or an optical microphone.

[0094] The audio system estimates the sound pressure level at the ear canal entrance (e.g., ear canal entrance 210) of 530 ears based on the monitored displacement. For example, the audio system can provide the monitored displacement as input to a model configured to output the estimated sound pressure level at the ear canal entrance. Based on the estimated sound pressure level at the ear canal entrance, the audio system characterizes how audio content is perceived at the ear canal entrance, thereby calibrating a virtual microphone at the ear canal entrance. In some embodiments, the audio system estimates the sound pressure level at both ear canal entrances.

[0095] The audio system generates one or more sound filters for the transducer based on the estimated sound pressure level. The sound filters can amplify, attenuate, and / or enhance certain frequencies. In some embodiments, the sound filters are configured to spatialize sound detected by the audio system's sensors from a localized region.

[0096] The audio system uses the generated sound filters to adjust the 550 audio content. In some embodiments, adjusting the audio content using the generated sound filters includes applying gain, filtering out certain frequencies, etc. In some embodiments, the adjusted audio content is sound from a local area. In other embodiments, the adjusted audio content is configured to be part of an artificial reality and / or mixed reality experience.

[0097] The audio system presents 560° adjusted audio content to the user via one or more transducers. This adjusted audio content can provide an improved auditory experience for the user. For example, it can preserve spatial cues, amplify certain frequencies for hearing-impaired users, and enhance sound from localized areas around the user for use in artificial reality and / or mixed reality applications.

[0098] Artificial Reality System Environment

[0099] Figure 6 This is a block diagram of an example artificial reality system environment 600 according to one or more embodiments. System 600 can operate in an artificial reality environment (e.g., a virtual reality environment, an augmented reality environment, a mixed reality environment, or some combination thereof). Figure 6 The system 600 shown includes a head-mounted device 605, an input / output (I / O) interface 610 coupled to a console 615, a network 620, and a map building server 625. In some embodiments, the head-mounted device 605 may be... Figure 1A Head-mounted devices 100 or Figure 1B The head-mounted device 105 is configured to calibrate a virtual microphone at the entrance of the user's ear canal.

[0100] although Figure 6The illustrated example system 600 includes a head-mounted device 605 and an I / O interface 610; however, in other embodiments, system 600 may include any number of these components. For example, multiple head-mounted devices may be present, each with an associated I / O interface 610, wherein each head-mounted device and I / O interface 610 communicates with a console 615. In alternative configurations, system 600 may include different and / or additional components. Additionally, in some embodiments, [the following is a continuation of the previous sentence, but the translation is incomplete]. Figure 6 The functions described in one or more components shown can be combined with Figure 6 The different ways in which they are described are distributed among the components. For example, some or all of the functions of the console 615 may be provided by the head-mounted device 605.

[0101] Head-mounted device 605 includes a display assembly 630, an optical component block 635, one or more position sensors 640, and a DCA 645. Some embodiments of head-mounted device 605 have [integration / combination / etc.]. Figure 6 These components are described as different parts. Additionally, they are combined. Figure 6 The functions provided by the various components described may be distributed in different ways among the components of the head-mounted device 605 in other embodiments, or embodied in separate components remote from the head-mounted device 605.

[0102] Display component 630 displays content to the user based on data received from console 615. Display component 630 uses one or more display elements (e.g., display element 120) to display content. The display element may be, for example, an electronic display. In various embodiments, display component 630 includes a single display element or multiple display elements (e.g., one display for each of the user's eyes). Examples of electronic displays include: liquid crystal display (LCD), organic light emitting diode (OLED) display, active-matrix organic light-emitting diode (AMOLED) display, waveguide display, some other display, or some combination thereof. It should be noted that in some embodiments, display element 120 may also include some or all of the functions of optical component block 635.

[0103] Optical element block 635 can amplify image light received from an electronic display, correct optical errors associated with the image light, and present corrected image light to one or both eye-correcting zones of head-mounted device 605. In various embodiments, optical element block 635 includes one or more optical elements. Example optical elements included in optical element block 635 include: apertures, Fresnel lenses, convex lenses, concave lenses, filters, reflective surfaces, or any other suitable optical elements that affect image light. Furthermore, optical element block 635 can include combinations of different optical elements. In some embodiments, one or more optical elements in optical element block 635 may have one or more coatings, such as a partial reflective coating or an anti-reflective coating.

[0104] The amplification and focusing of image light by the optical element block 635 allows the electronic display to be physically smaller, lighter, and consume less power compared to larger displays. Additionally, the amplification increases the field of view of the content presented on the electronic display. For example, the field of view of the displayed content is such that the displayed content is presented using almost the entire user's field of view (e.g., approximately 110 degrees diagonally), and in some cases, the displayed content is presented using the entire user's field of view. Furthermore, in some embodiments, the amplification amount can be adjusted by adding or removing optical elements.

[0105] In some embodiments, the optical element block 635 may be designed to correct one or more types of optical errors. Examples of optical errors include barrel or pincushion distortion, longitudinal or lateral chromatic aberration. Other types of optical errors may include spherical aberration; chromatic aberration; or errors due to lens field curvature, astigmatism; or any other type of optical error. In some embodiments, the content provided to the electronic display for display is pre-distorted, and the optical element block 635 corrects this distortion when it receives image light from the electronic display (which is generated based on the content).

[0106] Position sensor 640 is an electronic device that generates data indicating the position of head-mounted device 605. Position sensor 640 generates one or more measurement signals in response to movement of head-mounted device 605. Displacement sensor 190 is an embodiment of position sensor 640. Examples of position sensor 640 include one or more IMUs, one or more accelerometers, one or more gyroscopes, one or more magnetometers, another suitable type of sensor for detecting motion, or some combination thereof. Position sensor 640 may include multiple accelerometers for measuring translational motion (forward / backward, up / down, left / right) and multiple gyroscopes for measuring rotational motion (e.g., pitch, yaw, roll). In some embodiments, the IMU rapidly samples the measurement signals and calculates an estimated position of head-mounted device 605 based on the sampled data. For example, the IMU integrates the measurement signals received from the accelerometers over time to estimate a velocity vector, and integrates the velocity vector over time to determine the estimated position of a reference point on head-mounted device 605. A reference point is a point that can be used to describe the position of the head-mounted device 605. Although a reference point can generally be defined as a point in space, it is actually defined as a point within the head-mounted device 605.

[0107] The DCA645 generates depth information for a portion of a local area. The DCA includes one or more imaging devices and a DCA controller. The DCA645 may also include an illuminator. (See above reference.) Figure 1A The operation and structure of DCA645 are described.

[0108] Audio system 400 provides audio content to a user of head-mounted device 605. Audio system 400 calibrates a virtual microphone positioned at the entrance to the ear canal of the user's ear and adjusts the audio content accordingly. In some embodiments, the audio system uses a machine learning model to calibrate the virtual microphone. The audio system provides the model with displacement information about a portion of the user's ear as input, and the model outputs an estimated sound pressure level at the entrance to the ear canal. Therefore, the audio system can predict how to perceive the audio content generated by the transducer array at the entrance to the ear canal. In some embodiments, the audio system calibrates the virtual microphone for each of the user's ears. (As stated above regarding...) Figure 4 As described, the audio system 400 may include a transducer array 410, a sensor array 420, and a controller 430. The audio system 400 may include other components besides those described herein.

[0109] In addition to calibrating the virtual microphone at the user's ear canal entrance, the audio system 400 can perform other functions. In some embodiments, the audio system 400 can request acoustic parameters from the map-building server 625 via network 620. The acoustic parameters describe one or more acoustic characteristics of a local area (e.g., room impulse response, reverberation time, reverberation level, etc.). The audio system 400 can provide, for example, information describing at least a portion of the local area from the DCA 645 and / or location information of the head-mounted device 605 from the position sensor 640. The audio system 400 can use one or more acoustic parameters received from the map-building server 625 to generate one or more sound filters and use these sound filters to provide audio content to the user.

[0110] I / O interface 610 is a device that allows a user to send action requests to console 615 and receive responses from console 615. An action request is a request to perform a specific action. For example, an action request may be an instruction to start or stop acquiring image or video data, or an instruction to perform a specific action within an application. I / O interface 610 may include one or more input devices. Example input devices include a keyboard, mouse, game controller, or any other suitable device for receiving action requests and transmitting them to console 615. Action requests received by I / O interface 610 are transmitted to console 615, which performs the action corresponding to the action request. In some embodiments, I / O interface 610 includes an IMU that acquires calibration data indicating an estimated position of I / O interface 610 relative to its initial position. In some embodiments, I / O interface 610 may provide haptic feedback to a user based on instructions received from console 615. For example, haptic feedback can be provided when a motion request is received, or the console 615 can send instructions to the I / O interface 610 when the console 615 performs an action, thereby causing the I / O interface 610 to generate haptic feedback.

[0111] The console 615 provides content to the head-mounted device 605 for processing based on information received from one or more of the following: DCA 645, head-mounted device 605, and I / O interface 610. Figure 6 In the example shown, console 615 includes an application store 655, a tracking module 660, and an engine 665. Some embodiments of console 615 have a combination with... Figure 6 These modules or components are described as different modules or components. Similarly, the functions further described below can be arranged in combination with... Figure 6The different ways in which they are described are distributed among the components of console 615. In some embodiments, the functions of console 615 discussed herein may be implemented in head-mounted device 605 or a remote system.

[0112] Application store 655 stores one or more applications for execution by console 615. An application is a set of instructions that, when executed by a processor, generate content to be presented to a user. The content generated by the application may respond to input received from the user via movement of head-mounted device 605 or I / O interface 610. Examples of applications include: game applications, conferencing applications, video playback applications, or other suitable applications.

[0113] Tracking module 660 uses information from DCA 645, one or more position sensors 640, or some combination thereof, to track the movement of head-mounted device 605 or I / O interface 610. For example, tracking module 660 determines the position of a reference point of head-mounted device 605 in a mapping of a local region based on information from head-mounted device 605. Tracking module 660 can also determine the position of an object or virtual object. Additionally, in some embodiments, tracking module 660 can use data portions from position sensors 640 indicating the position of head-mounted device 605 and a representation of a local region from DCA 645 to predict the future position of head-mounted device 605. Tracking module 660 provides engine 665 with the estimated or predicted future position of head-mounted device 605 or I / O interface 610.

[0114] Engine 665 executes the application and receives position information, acceleration information, velocity information, predicted future position, or a combination thereof from tracking module 660 of head-mounted device 605. Based on the received information, engine 665 determines the content to be presented to the user on head-mounted device 605. For example, if the received information indicates that the user has looked to the left, engine 665 generates content for head-mounted device 605 that is a mirror image of the user's movement in a virtual local area or a local area (enhanced with additional content). Additionally, in response to an action request received from I / O interface 610, engine 665 executes an action within the application running on console 615 and provides feedback to the user that the action has been performed. The feedback provided can be visual or auditory feedback via head-mounted device 605, or haptic feedback via I / O interface 610.

[0115] Network 620 couples the head-mounted device 605 and / or console 615 to the map building server 625. Network 620 may include any combination of local area networks and / or wide area networks using wireless communication systems and / or wired communication systems. For example, network 620 may include the Internet and mobile phone networks. In one embodiment, network 620 uses standard communication technologies and / or protocols. Therefore, network 620 may include links using technologies such as Ethernet, 802.11, worldwide interoperability for microwave access (WiMAX), 2G / 3G / 4G mobile communication protocols, digital subscriber line (DSL), asynchronous transfer mode (ATM), InfiniBand, PCI Express Advanced Switching, etc. Similarly, networking protocols used on Network 620 may include Multiprotocol Label Switching (MPLS), Transmission Control Protocol / Internet Protocol (TCP / IP), User Datagram Protocol (UDP), Hypertext Transfer Protocol (HTTP), Simple Mail Transfer Protocol (SMTP), File Transfer Protocol (FTP), etc. Data exchanged through Network 620 may be represented using technologies and / or formats including binary image data (e.g., Portable Network Graphic (PNG)), Hypertext Markup Language (HTML), Extensible Markup Language (XML), etc.In addition, conventional encryption techniques can be used to encrypt all or some links, such as Secure Sockets Layer (SSL), Transport Layer Security (TLS), Virtual Private Network (VPN), and Internet Protocol Security (IPsec).

[0116] Map building server 625 may include a database storing virtual models describing multiple spaces, wherein a location in the virtual model corresponds to the current configuration of a local area of ​​head-mounted device 605. Map building server 625 receives information describing at least a portion of the local area and / or location information of the local area from head-mounted device 605 via network 620. A user may adjust privacy settings to allow or prevent head-mounted device 605 from sending information to map building server 625. Map building server 625 determines the location in the virtual model associated with the local area of ​​head-mounted device 605 based on the received information and / or location information. Map building server 625 determines (e.g., retrieves) one or more acoustic parameters associated with the local area, in part based on the determined location in the virtual model and any acoustic parameters associated with the determined location. Map building server 625 may send the location of the local area and any acoustic parameter values ​​associated with the local area to head-mounted device 605.

[0117] One or more components in system 600 may include a privacy module that stores one or more privacy settings for user data elements. The user data elements describe a user or head-mounted device 605. For example, a user data element may describe the user's physical characteristics, actions performed by the user, the user's location on head-mounted device 605, the location of head-mounted device 605, the user's HRTF, etc. Privacy settings (or "access settings") for user data elements may be stored in any suitable manner, such as being stored in association with the user data element, stored in an index on an authorization server, stored in another suitable manner, or any suitable combination thereof.

[0118] Privacy settings for user data elements specify how user data elements (or specific information associated with user data elements) can be accessed, stored, or otherwise used (e.g., viewed, shared, modified, copied, performed, displayed, or identified). In some embodiments, privacy settings for user data elements may specify a “blacklist” of entities that may be denied access to certain information associated with the user data element. Privacy settings associated with user data elements may specify any appropriate granularity for granting or denying access. For example, some entities may have permission to ascertain the existence of a specific user data element, some entities may have permission to view the content of a specific user data element, and some entities may have permission to modify a specific user data element. Privacy settings may allow a user to permit other entities to access or store user data elements for a limited period of time.

[0119] Privacy settings allow users to specify one or more geographic locations from which they can access user data elements. Access to or denial of access to user data elements can depend on the geographic location of the entity attempting to access the user data element. For example, a user can allow access to a user data element and specify that the user data element is only accessible to an entity while the user is in a specific location. If the user leaves that specific location, the user data element may no longer be accessible to that entity. As another example, a user can specify that a user data element is only accessible to entities within a threshold distance of the user (e.g., another user of a headset in the same local area as the user). If the user subsequently changes location, the entity with access to the user data element may lose access, while a new set of entities may gain access when they come within the user's threshold distance.

[0120] System 600 may include one or more authorization / privacy servers for implementing privacy settings. A request from an entity for a specific user data element can identify the entity associated with the request, and if the authorization server determines, based on the privacy settings associated with the user data element, that the entity is authorized to access the user data element, it can send the user data element only to that entity. If the requesting entity is not authorized to access the user data element, the authorization server can prevent the requested user data element from being retrieved or from being sent to the entity. Although this disclosure describes implementing privacy settings in a particular manner, this disclosure contemplates implementing privacy settings in any suitable manner.

[0121] Additional configuration information

[0122] The above description of embodiments has been presented for illustrative purposes and is not intended to be exhaustive, nor is it intended to limit the patent right to the precise form disclosed. Those skilled in the art will understand that many modifications and variations are possible in light of the foregoing disclosure.

[0123] Some portions of this description describe embodiments of algorithms and symbolic representations for manipulating information. These algorithmic descriptions and representations are commonly used by those skilled in the art of data processing to effectively communicate the substance of their work to others skilled in the art. Although these operations are described functionally, computationally, or logically, they are understood to be implemented by computer programs or equivalent circuits or microcode, etc. Furthermore, it has been proven that, without loss of generality, the arrangement of these operations is sometimes referred to as a module for convenience. The described operations and their associated modules can be implemented in software, firmware, hardware, or any combination thereof.

[0124] Any step, operation, or process described herein may be performed or implemented using one or more hardware or software modules, individually or in combination with other devices. In one embodiment, a software module is implemented using a computer program product comprising a computer-readable medium containing computer program code that can be executed by a computer processor to perform any or all of the steps, operations, or processes described herein.

[0125] The embodiments may also relate to an apparatus for performing the operations described herein. This apparatus may be specifically constructed for the desired purpose, and / or the apparatus may include a general-purpose computing device selectively activated or reconfigured by a computer program stored in a computer. Such a computer program may be stored in a non-transitory tangible computer-readable storage medium or any type of medium suitable for storing electronic instructions, which may be coupled to a computer system bus. Furthermore, any computing system mentioned in this specification may include a single processor, or may be an architecture employing a multiple-processor design to increase computing power.

[0126] The embodiments may also relate to a product generated by the computational process described herein. Such a product may include information generated according to the computational process, wherein the information is stored on a non-transitory tangible computer-readable storage medium and may include any embodiment of a computer program product or other combinations of data described herein.

[0127] Finally, the language used in this specification has been chosen primarily for readability and guidance purposes, and may not have been chosen to define or limit patent rights. Therefore, the scope of patent rights is not limited to this specific embodiment, but rather to any claims published in the application based on this document. Thus, the disclosure of the embodiments is intended to exemplify, not limit, the scope of patent rights, which is set forth in the appended claims.

Claims

1. A method comprising: presenting audio content to a user via a transducer; monitoring, via one or more sensors, displacement of a portion of a pinna of the user, the displacement being caused in part by the presented audio content; distinguishing between the displacement of the portion of the pinna due to the presented audio content and displacement of a portion of the pinna due to noise by providing, as input to a correlation model, the displacement of the portion of the pinna monitored via the sensors to measure a correlation between the audio content presented via the transducer and the monitored displacement; in response to determining that the measured correlation indicates that the monitored displacement is due to the presented audio content, providing the monitored displacement as input to a sound pressure estimation model that relates sound pressure at an entrance to an ear canal of the user as a function of outer ear acceleration; estimating the sound pressure at the entrance to the ear canal of the user based on output from the sound pressure estimation model; using the estimated sound pressure at the entrance to the ear canal to generate a sound filter for the transducer; adjusting audio content using the generated sound filter; and presenting the adjusted audio content to the user via the transducer, wherein the sound pressure estimation model is configured to receive, as input, a geometry of an ear of the user, the geometry comprising measurements determined from one or more images of the ear of the user. the transducer is an ossicular conduction transducer configured to present the audio content.

2. The method of claim 1, wherein, one of the one or more sensors is the ossicular conduction transducer.

3. The method of claim 2, wherein, the method further comprises:

4. The method of claim 3, wherein, monitoring the displacement of the portion of the pinna of the user by measuring an amount of preload of the ossicular conduction transducer. the one or more sensors comprise at least one of an acceleration sensor and an optical displacement sensor.

5. The method of claim 1, wherein, the transducer is a loudspeaker configured to present the audio content to the user via air conduction.

6. The method of claim 1, wherein, the sound pressure estimation model comprises at least one of a convolutional neural network, a linear model, and a numerical simulation.

7. The method of claim 1, wherein, the adjusted audio content has a frequency response of a target magnitude.

8. The method of claim 1, wherein, 9. An audio system comprising: a transducer configured to present audio content to a user; one or more sensors configured to monitor displacement of a portion of a pinna of the user, the displacement being caused by presented audio content; and a controller configured to: distinguish between the displacement of the portion of the pinna due to the presented audio content and displacement of a portion of the pinna due to noise by providing, as input to a correlation model, the displacement of the portion of the pinna monitored via the sensors to measure a correlation between the audio content presented via the transducer and the monitored displacement; ​ ​ in response to determining that the measured correlation indicates that the monitored displacement is due to the rendered audio content, providing the monitored displacement as input to a sound pressure estimation model that relates sound pressure at an entrance of an ear canal of the user as a function of an outer ear acceleration; based on output from the sound pressure estimation model, estimating the sound pressure at the entrance of the ear canal of the user; using the estimated sound pressure at the entrance of the ear canal, generating a sound filter for the transducer; using the generated sound filter to adjust audio content; and indicating to the transducer to render adjusted audio content to the user, wherein the sound pressure estimation model is configured to receive as input a geometry of an ear of the user, the geometry comprising measurements determined from one or more images of the ear of the user.

10. The audio system of claim 9, wherein, the transducer is a cartilage conduction transducer configured to render the audio content.

11. The audio system of claim 10, wherein, one of the one or more sensors is the cartilage conduction transducer.

12. The audio system of claim 11, wherein, the controller is further configured to monitor the displacement of the portion of the pinna of the user by measuring a preload amount of the cartilage conduction transducer.

13. The audio system of claim 9, wherein, the one or more sensors comprise at least one of an acceleration sensor and an optical displacement sensor.

14. The audio system of claim 9, wherein, the transducer is a loudspeaker configured to render the audio content to the user via air conduction.

15. The audio system of claim 9, wherein, the sound pressure estimation model comprises at least one of a convolutional neural network, a linear model, and a numerical simulation.

16. The audio system of claim 9, wherein, the adjusted audio content has a frequency response of a target magnitude.

17. A computer-readable non-transitory storage medium comprising software that is operable when executed on a computer processor to perform the method of any one of claims 1 to 8.

Citation Information

Patent Citations

  • Bone conduction earphone test method and test system

    CN111065035A

  • Hybrid audio system for eyewear devices

    US20190342647A1