Method and storage medium for generating equalization filter for a user
By receiving images of the user's ears and using a machine learning model to generate a personalized equalization filter, the acoustic parameters of the audio output are adjusted, solving the problem of inconsistent audio output caused by anatomical features and fit inconsistencies in head-mounted devices, thus improving the user experience.
Patent Information
- Application Number
- CN202080059986.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2019-09-04
- Filing Date
- 2020-08-15
- Publication Date
- 2026-01-09
- Estimated Expiration
- 2040-08-15
AI Technical Summary
Existing head-mounted devices, such as AR and VR devices, suffer from inconsistent audio output due to differences in the anatomical features of users' ears and heads, as well as inconsistencies in device compatibility, which affects the user experience.
By receiving images of the user's ears, a machine learning model is used to generate a personalized equalization filter, which adjusts the acoustic parameters of the audio output to match the target response of the user's ears. The personalized equalization filter is then attached to the user's profile.
It improves the personalization and adaptability of audio output, reduces audio output variations caused by differences between users and devices, and enhances the user's audio experience.
Smart Images

Figure CN114303388B_ABST
Abstract
Description
Technical Field
[0001] This disclosure generally relates to artificial reality systems, and more specifically to generating personalized equalization filters for users. Background Technology
[0002] Existing head-mounted devices (such as artificial reality (AR) and virtual reality (VR) headsets) typically use multiple transducers to deliver audio content to users. However, sound propagation from the transducers to the user's ears can vary based on the anatomical features of the user's ears and / or head. For example, differences in ear size and shape between users can affect the sound generated by the head-mounted device and perceived by the user, potentially negatively impacting the user's audio experience. Current audio systems are insufficient for delivering high-fidelity audio content because they may not account for variations in anatomical features between users and inconsistencies in head-mounted device fit. Therefore, a method is needed to adjust the audio output to compensate for variations in anatomical features and fit inconsistencies, ensuring that the audio content delivered by the head-mounted device is tailored to the user. Summary of the Invention
[0003] A system and method are disclosed that use a machine learning model to generate personalized equalization filters to enhance a user's audio experience. The system receives one or more images of a portion of a user's head, including at least the user's ears. The images may include images of the user (e.g., the user's head, the user's ears) and / or images of the user wearing a head-mounted device. The head-mounted device may include multiple transducers that provide audio content to the user. Features describing the user's ears are extracted from one or more images and fed into a model. The model is configured to predict how the audio output will sound at the user's ears. An equalization filter is generated for the user based on the difference between the target audio response and the predicted audio output at the user's ears. The equalization filter adjusts one or more acoustic parameters of the audio output (e.g., wavelength, frequency, volume, pitch, balance, etc.) based on the user's ears to generate a target response at the user's ears, such that the user perceives the audio output as the creator of the audio output intended it to be heard. The equalization filter can be used in a head-mounted device to provide audio content to the user. The equalization filter can also be attached to the user's social network profile.
[0004] According to one embodiment of the present invention, a method is provided, comprising: receiving one or more images including a user's ear; identifying one or more features of the user's ear from one or more images; providing one or more features of the user's ear to a model configured to predict audio output at the user's ear based on the identified one or more features; and generating an equalization filter based on the audio output at the user's ear, the equalization filter being configured to adjust one or more acoustic parameters of audio content provided to the user.
[0005] In some embodiments, the method further includes providing the generated equalization filter to a head-mounted device configured to use the equalization filter when providing audio content to a user.
[0006] In some embodiments, when applied to audio content provided to a user, the equalization filter adjusts one or more acoustic parameters of the audio content for the user based on the predicted audio output at the user's ear.
[0007] In some embodiments, the method further includes providing an equalization filter to an online system for storage in association with a user profile, wherein the equalization filter can be retrieved by one or more head-mounted devices associated with the user and authorized to access the user profile for use in providing content to the user.
[0008] In some embodiments, the method further includes: training a model using multiple labeled images, each of which identifies features of the ears of an additional user for whom the audio output at the ear is known.
[0009] In some embodiments, a user in one or more images wears a head-mounted device, and one or more features are identified based at least in part on the position of the head-mounted device relative to the user's ears.
[0010] In some embodiments, the head-mounted device includes an eyeglass frame with two arms, each arm coupled to the eyeglass body, and one or more images include at least a portion of one of the two arms, the at least a portion of one of the two arms including a transducer among a plurality of transducers.
[0011] In some embodiments, the model is configured to determine the audio output response based at least in part on the position of one of the multiple transducers relative to the user's ear.
[0012] In some embodiments, one or more images are depth images captured using a depth camera component.
[0013] In some embodiments, one or more of the identified features are anthropometric features describing the size or shape of the user's ears.
[0014] In some embodiments, the method further includes: comparing the determined audio output at the user's ear with the measured audio output at the user's ear; and updating the model based on the comparison.
[0015] In some embodiments, the measured audio output response is measured by providing audio content to the user via a headset and by analyzing the audio output at the user's ears using one or more microphones placed near the user's ears.
[0016] According to some embodiments of the present invention, a non-transitory computer-readable storage medium is provided thereon storing instructions which, when executed by a processor, cause the processor to perform steps including: receiving one or more images including a user's ear; identifying one or more features of the user's ear based on the one or more images; providing one or more features to a model configured to determine an audio output at the user's ear based on the identified one or more features; and generating an equalization filter based on the audio output at the user's ear, the equalization filter being configured to adjust one or more acoustic parameters of the audio content provided to the user.
[0017] In some embodiments, the instructions, when executed by the processor, also cause the processor to perform the following steps: training a model using a plurality of labeled images, each of which identifies features of the ears of additional users for whom the audio output response is known.
[0018] In some embodiments, when applied to audio content provided to a user, the equalization filter adjusts one or more acoustic parameters of the audio content for the user based on the predicted audio output at the user's ear.
[0019] In some embodiments, one or more images are depth images captured using a depth camera component.
[0020] In some embodiments, a user in one or more images wears a head-mounted device, and one or more features are identified based at least in part on the position of the head-mounted device relative to the user's ears.
[0021] In some embodiments, one or more of the identified features are anthropometric features describing the size or shape of the user's ears.
[0022] In some embodiments, the head-mounted device includes an eyeglass frame with two arms, each arm coupled to the eyeglass body, and one or more images include at least a portion of one of the two arms, the at least a portion of one of the two arms including a transducer among a plurality of transducers.
[0023] In some embodiments, the model is configured to determine the audio output at the user's ear based at least in part on the position of one of the multiple transducers relative to the user's ear. Attached Figure Description
[0024] Figure 1A This is a perspective view of a first embodiment of a head-mounted device according to one or more embodiments.
[0025] Figure 1B This is a perspective view of a second embodiment of a head-mounted device according to one or more embodiments.
[0026] Figure 2 A system environment for providing audio content to a device is illustrated according to one or more embodiments.
[0027] Figure 3 An equalization system according to one or more embodiments is shown.
[0028] Figure 4A This is an example view of an imaging device capturing an image of a user's head according to one or more embodiments.
[0029] Figure 4B According to one or more embodiments Figure 4A An image captured by an imaging device in the image, showing a portion of the user's head.
[0030] Figure 5A This is an example view of an imaging device according to one or more embodiments capturing an image of the head of a user wearing a head-mounted device.
[0031] Figure 5B According to one or more embodiments Figure 5A An image captured by an imaging device in the image, showing a portion of the user's head.
[0032] Figure 6A This is an example view of an imaging device according to one or more embodiments capturing an image of the head of a user wearing a head-mounted device with a visual marker.
[0033] Figure 6B According to one or more embodiments Figure 6A An image captured by an imaging device in the image, showing a portion of the user's head.
[0034] Figure 7 A method for generating personalized equalization filters for users based on simulation, according to one or more embodiments, is illustrated.
[0035] Figure 8A An example flow for generating a representation of a user's ear using a machine learning model is shown according to one or more embodiments.
[0036] Figure 8B It is a flowchart for determining a PCA model according to one or more embodiments.
[0037] Figure 9A A machine learning model for predicting audio output at a user's ear, according to one or more embodiments, is shown.
[0038] Figure 9B A method for generating personalized equalization filters using a machine learning model according to one or more embodiments is illustrated.
[0039] Figure 10 It is a block diagram of an audio system according to one or more embodiments.
[0040] Figure 11 This is a system environment for providing audio content to a user, according to an embodiment.
[0041] The accompanying drawings depict various embodiments for illustrative purposes only. Those skilled in the art will readily recognize from the following discussion that alternative embodiments of the structures and methods shown herein can be employed without departing from the principles described herein. Detailed Implementation
[0042] Overview
[0043] Head-mounted devices, such as artificial reality (AR) headsets, include one or more transducers (e.g., speakers) for providing audio content to a user. However, sound propagation from the transducers to the user's ears can vary between users and between devices. In particular, the audio output at the user's ears can vary based on anthropometry characteristics of the user's ears and / or head. Anthropometry characteristics are the user's physical features (e.g., ear shape, ear size, ear orientation / position on the head, head size, etc.). Furthermore, the fit of the head-mounted device may vary based on anthropometry characteristics, which can also affect the audio output response. Therefore, it may be useful to adjust the audio content provided to the user by the head-mounted device to provide a personalized audio output response, thereby enhancing the user experience and providing the user with high-quality content. Thus, an equalization filter is generated based on the user's ears, which adjusts one or more acoustic parameters of the audio output (e.g., wavelength, frequency, volume, pitch, balance, other spectral content, acoustic time delay, etc.). When applied to audio content, the equalization filter modulates the audio content to a target response at the user's ear, so that the user perceives the audio content as the content creator intends it to be heard. In one embodiment, the target response is associated with a predetermined value (or an acceptable range of values) for each of a set of acoustic parameters. The predetermined value (or acceptable range of values) for each of these acoustic parameters corresponds to a relatively high acceptable sound quality threshold that the content creator wants the audio content to be perceived by the user.
[0044] In one embodiment, an imaging system (e.g., a user's mobile device, etc.) captures one or more images of a user wearing a head-mounted device to collect anthropometry information associated with the user. The imaging system may capture image data (e.g., still image data or video image data) of the user's ears, head, and / or the user wearing the head-mounted device. In one embodiment, each of the one or more images is a frame from a captured video of the user's ears, head, and / or the user wearing the head-mounted device. The head-mounted device may be a virtual reality (VR) head-mounted device, an AR head-mounted device, or some other head-mounted device configured to provide audio content to a user. The head-mounted device may include multiple transducers for providing audio content, and the positions of the transducers may be known. The size of the head-mounted device may also be known. In some embodiments, the head-mounted device includes one or more visual markers for determining its positional information relative to the user's head. For example, the head-mounted device may include markers positioned along its frame (e.g., along each temple arm). The position of each marker relative to other markers and the head-mounted device is known. In some embodiments, each marker has a unique size and / or shape.
[0045] An equalization system (e.g., from an imaging system, from a head-mounted device, etc.) receives one or more images of a user to generate a customized equalization filter for the user. In one embodiment, the imaging system provides one or more images to the head-mounted device, and the head-mounted device provides one or more images to the equalization system. The equalization system identifies features of the user's ear (e.g., shape, size) based on the received images. In some embodiments, the equalization system extracts depth information associated with the images and generates a 3D representation of the user's ear based on the extracted depth information and the identified features. The equalization system may use a machine learning model to generate the 3D representation, and in some embodiments, the 3D representation includes a representation of the head-mounted device. The equalization system performs a simulation of audio propagation from an audio source (e.g., the transducer array of the head-mounted device) to the 3D representation of the user's ear. Based on the simulation, the equalization system can predict the audio output at the user's ear. An equalization filter is generated for the user based on the difference between the target audio response and the predicted audio output at the user's ear. In one embodiment, the equalization filter is generated based on a transfer function that is the ratio between two complex frequency responses (i.e., the target response and the predicted response). An equalization filter adjusts one or more acoustic parameters of the audio output (e.g., wavelength, frequency, volume, pitch, balance, other spectral content, acoustic time delay, etc.) based on the user's ear to generate a target response at the user's ear, so that the user perceives the audio output as the audio output creator intended it to be heard. In some embodiments, the equalization system generates an audio profile for the user based on the equalization filter, which specifies the amount of compensation for one or more acoustic parameters.
[0046] In another embodiment, the equalization system uses a machine learning model to predict the audio output at the user's ear. The equalization system (e.g., from an imaging system) receives one or more images and extracts one or more features describing the user's ear based on these images. The equalization system can use machine learning techniques, imaging techniques, algorithms, or any other model to extract features of the user's ear based on the images. The equalization system uses the machine learning model to determine the audio output at the user's ear based on the extracted one or more features. In one embodiment, the model is trained using images of other users' ears / heads with previously identified features (e.g., those identified by the model, those identified by the person) and the known audio output at each user's ear. An equalization filter is generated for the user based on the difference between the target audio response and the predicted audio output at the user's ear. The equalization filter adjusts one or more acoustic parameters of the audio output (e.g., wavelength, frequency, volume, pitch, balance, other spectral content, acoustic time delay, etc.) based on the user's ear to generate a target response at the user's ear, such that the user perceives the audio output as the creator of the audio output intended it to be heard.
[0047] An equalization system can provide a personalized equalization filter to a headset. Therefore, a personalized equalization filter can modify one or more acoustic parameters of the audio content provided to the user by the headset, thus customizing the audio content for the user. Personalized equalization filters improve the audio experience by reducing audio output variations caused by differences between users and between devices. Furthermore, personalized equalization filters can be attached to a user's profile (e.g., a social network profile), eliminating the need for the user to recalibrate their device during subsequent use.
[0048] Embodiments of the present invention may be implemented using an artificial reality system or in combination with an artificial reality system. Artificial reality is a form of reality that has been adjusted in some way before being presented to a user, and may include, for example, virtual reality (VR), augmented reality (AR), mixed reality (MR), hybrid reality, or some combination and / or derivative thereof. Artificial reality content may include fully generated content or generated content combined with captured (e.g., real-world) content. Artificial reality content may include video, audio, haptic feedback, or some combination thereof, any of which may be presented in a single channel or in multiple channels (e.g., stereoscopic video that produces a three-dimensional effect for the viewer). Furthermore, in some embodiments, artificial reality may also be associated with applications, products, accessories, services, or some combination thereof for creating content in artificial reality and / or being used in artificial reality in other ways. Artificial reality systems that provide artificial reality content can be implemented on a variety of platforms, including wearable devices (e.g., head-mounted devices) connected to a host computer system, stand-alone wearable devices (e.g., head-mounted devices), mobile devices or computing systems, or any other hardware platform capable of providing artificial reality content to one or more viewers.
[0049] Example head-mounted device
[0050] Figure 1A This is a perspective view of a first embodiment of a head-mounted device according to one or more embodiments. In some embodiments, the head-mounted device is a near-eye display (NED) or eyewear device. Typically, the head-mounted device 100 can be worn on a user's face to present content (e.g., media content) using a display component and / or an audio system. However, the head-mounted device 100 can also be used to present media content to the user in different ways. Examples of media content presented by the head-mounted device 100 include one or more images, videos, audio, or some combination thereof. The head-mounted device 100 includes a frame and may include a display component, a depth camera component (DCA), an audio system, and a position sensor 190, as well as other components, the display component including one or more display elements 120. Although Figure 1A The example positioning of components on the head-mounted device 100 illustrates the components of the head-mounted device 100, but these components may be located elsewhere on the head-mounted device 100, on a peripheral device paired with the head-mounted device 100, or some combination of both. Similarly, components on the head-mounted device 100 may be more numerous than those on the head-mounted device 100. Figure 1A More or less as shown.
[0051] Frame 110 holds other components of the head-mounted device 100. Frame 110 includes a front portion that holds one or more display elements 120 and end pieces (e.g., temples) that attach to the user's head. The front portion of frame 110 bridges the top of the user's nose. The length of the end pieces may be adjustable (e.g., adjustable temple length) to fit different users. The end pieces may also include curled portions behind the user's ears (e.g., temple tips, ear loops). In some embodiments, frame 110 includes one or more visual markers, which are described below. Figures 6A-6B To provide a more detailed description.
[0052] One or more display elements 120 provide light to a user wearing a head-mounted device 100. As shown, the head-mounted device includes a display element 120 for each of the user's eyes. In some embodiments, the display elements 120 generate image light provided to an eyebox of the head-mounted device 100. The eyebox is the spatial location occupied by the user's eyes when wearing the head-mounted device 100. For example, the display element 120 may be a waveguide display. A waveguide display includes a light source (e.g., a two-dimensional light source, one or more line light sources, one or more point light sources, etc.) and one or more waveguides. Light from the light source is inwardly coupled into one or more waveguides, which output light in a manner that results in pupil replication in the eyebox of the head-mounted device 100. The inward and / or outward coupling of light from one or more waveguides may be accomplished using one or more diffraction gratings. In some embodiments, the waveguide display includes a scanning element (e.g., a waveguide, a mirror, etc.) that scans the light from the light source as it is coupled inward into one or more waveguides. It should be noted that in some embodiments, one or two of the display elements 120 are opaque and do not transmit light from a local area surrounding the head-mounted device 100. The local area is the area surrounding the head-mounted device 100. For example, the local area could be a room where the user wearing the head-mounted device 100 is located, or the user wearing the head-mounted device 100 could be outside, and the local area is an external area. In this context, the head-mounted device 100 generates VR content. Alternatively, in some embodiments, one or both of the display elements 120 are at least partially transparent, such that light from the local area can be combined with light from one or more display elements to produce AR and / or MR content.
[0053] In some embodiments, the display element 120 does not generate image light; instead, lenses deliver light from localized areas to the viewing window. For example, one or both of the display elements 120 may be uncorrected lenses (non-prescription lenses) or prescription lenses (e.g., monocular lenses, bifocal and trifocal lenses, or progressive lenses) to help correct for the user's vision deficiencies. In some embodiments, the display element 120 may be polarized and / or tinted to protect the user's eyes from sunlight.
[0054] It should be noted that in some embodiments, the display element 120 may include additional optical blocks (not shown). The optical blocks may include one or more optical elements (e.g., lenses, Fresnel lenses, etc.) that guide light from the display element 120 to the viewing window. The optical blocks may, for example, correct some or all aberrations in the image content, magnify some or all of the image, or some combination thereof.
[0055] DCA determines depth information for a portion of a local area surrounding the head-mounted device 100. DCA includes one or more imaging devices 130 and a DCA controller (not included in...). Figure 1A As shown in the figure, it may also include an illuminator 140. In some embodiments, the illuminator 140 illuminates a portion of a local area with light. The light may be, for example, structured light in the infrared (IR) (e.g., dot patterns, stripes, etc.), an IR flash for time-of-flight, etc. In some embodiments, one or more imaging devices 130 capture an image of a portion of the local area including light from the illuminator 140. As shown in the figure, Figure 1A A single illuminator 140 and two imaging devices 130 are shown. In an alternative embodiment, there is no illuminator 140 and at least two imaging devices 130.
[0056] The DCA controller uses captured images and one or more depth determination techniques to calculate partial depth information of local regions. The depth determination techniques can be, for example, direct time-of-flight (ToF) depth sensing, indirect ToF depth sensing, structured light, passive stereo analysis, active stereo analysis (using textures added to the scene via light from illuminator 140), some other technique for determining the depth of the scene, or some combination thereof.
[0057] An audio system provides audio content. The audio system includes a transducer array, a sensor array, and an audio controller 150. However, in other embodiments, the audio system may include different and / or additional components. Similarly, in some cases, the functions described with reference to the components of the audio system may be distributed among the components in a manner different from that described herein. For example, some or all of the functions of the audio controller may be performed by a remote server.
[0058] A transducer array presents sound to the user. The transducer array includes multiple transducers. The transducers may be speakers 160 or tissue transducers 170 (e.g., bone conduction transducers or cartilage conduction transducers). Although speakers 160 are shown outside the frame 110, they may be enclosed within the frame 110. In some embodiments, instead of separate speakers for each ear, the head-mounted device 100 includes a speaker array comprising multiple speakers integrated into the frame 110 to improve the directionality of the presented audio content. Tissue transducers 170 are coupled to the user's head and directly vibrate the user's tissue (e.g., bone or cartilage) to generate sound. The number and / or positioning of the transducers can be related to… Figure 1A The differences are shown.
[0059] A sensor array detects sound within a localized area of the head-mounted device 100. The sensor array includes multiple acoustic sensors 180. Each acoustic sensor 180 captures sound emitted from one or more sound sources within the localized area (e.g., a room). Each acoustic sensor is configured to detect sound and convert the detected sound into an electronic format (analog or digital). The acoustic sensor 180 may be a sound wave sensor, a microphone, a sound transducer, or a similar sensor suitable for detecting sound.
[0060] In some embodiments, one or more acoustic sensors 180 may be placed in the ear canal of each ear (e.g., acting as binaural microphones). In some embodiments, the acoustic sensors 180 may be placed on the outer surface of the head-mounted device 100, on the inner surface of the head-mounted device 100, separate from the head-mounted device 100 (e.g., part of some other device), or some combination thereof. The number and / or positioning of the acoustic sensors 180 may be related to... Figure 1A The differences are illustrated. For example, the number of acoustic detection localizations can be increased to increase the amount of audio information collected, as well as the sensitivity and / or accuracy of the information. Acoustic detection localizations can be oriented such that the microphones can detect sound in a wide range of directions around the user wearing the headset 100.
[0061] Audio controller 150 processes information from a sensor array describing the sound detected by the sensor array. Audio controller 150 may include a processor and a computer-readable storage medium. Audio controller 150 may be configured to generate direction of arrival (DOA) estimates, generate acoustic transfer functions (e.g., array transfer function and / or head-related transfer function), track the localization of sound sources, form beams in the direction of sound sources, classify sound sources, generate sound filters for loudspeaker 160, or some combination thereof.
[0062] Position sensor 190 generates one or more measurement signals in response to motion of head-mounted device 100. Position sensor 190 may be located on a portion of frame 110 of head-mounted device 100. Position sensor 190 may include an inertial measurement unit (IMU). Examples of position sensor 190 include: one or more accelerometers, one or more gyroscopes, one or more magnetometers, another suitable type of sensor for detecting motion, a type of sensor for error correction of the IMU, or some combination thereof. Position sensor 190 may be located external to the IMU, internal to the IMU, or some combination thereof.
[0063] In some embodiments, the head-mounted device 100 can provide real-time localization and mapping (SLAM) of the head-mounted device 100's position and updates to the model of a local region. For example, the head-mounted device 100 may include a passive camera assembly (PCA) that generates color image data. The PCA may include one or more red, green, and blue (RGB) cameras that capture some or all of the images of a local region. In some embodiments, some or all of the imaging devices 130 of the DCA may also be used as the PCA. The images captured by the PCA and the depth information determined by the DCA can be used to determine parameters of the local region, generate a model of the local region, update the model of the local region, or some combination thereof. Furthermore, the position sensor 190 tracks the position (e.g., localization and orientation) of the head-mounted device 100 within a room. The following is in conjunction with... Figure 10 Additional details regarding the components of the head-mounted device 100 are discussed.
[0064] Figure 1B This is a perspective view of a second embodiment of a head-mounted device implemented as a head-mounted display (HMD) according to one or more embodiments. In embodiments describing AR and / or MR systems, the front portion of the HMD is at least partially transparent in the visible light band (approximately 380 nm to 750 nm), and the portion of the HMD between the front portion and the user's eyes is at least partially transparent (e.g., a partially transparent electronic display). The HMD includes a front rigid body 115 and a strap 175. The head-mounted device 105 includes a plurality of components referenced above. Figure 1A The same components are described, but these components are modified to integrate with the HMD's shape factor. For example, the HMD includes a display assembly, DCA, audio system, and position sensor 190. Figure 1B An illuminator 140, multiple speakers 160, multiple imaging devices 130, multiple acoustic sensors 180, and a position sensor 190 are shown.
[0065] System environment for providing personalized audio content
[0066] Figure 2A system environment for providing personalized audio content to a user via a head-mounted device is illustrated according to one or more embodiments. System environment 200 includes a head-mounted device 210 connected via a network 250, an imaging system 220, an equalization system 230, and an online system 240. System environment 200 may include fewer or additional components than those described herein. Furthermore, the structure and / or function of the components may differ from those described herein.
[0067] Headset 210 is a device configured to be worn by a user over the area of their head (e.g., headset 100, headset 105). Headset 210 includes an audio system 215 configured to transmit audio content to the user wearing the headset 210. The audio system 215 may include one or more transducers (e.g., speakers) for providing audio content to the user. The following is about... Figure 10 The audio system 215 is described in more detail below. In some embodiments, the head-mounted device 210 includes additional components (e.g., a display system, a haptic feedback system) for providing the user with other types of content (e.g., digital content, haptic content). Furthermore, the head-mounted device 210 may include one or more visual markers for determining the position of the head-mounted device 210 relative to the user wearing the device. The markers may be positioned along the frame of the head-mounted device 210 (e.g., frame 110). The position of the markers relative to other markers and the head-mounted device 210 is known. See below for more details. Figures 6A-6B Describe the marker in more detail.
[0068] Imaging system 220 includes imaging device 225 configured to capture one or more images of a user's head, head-mounted device 210, and / or at least a portion of the user wearing head-mounted device 210. Imaging device 225 can be any suitable type of sensor, such as a multispectral camera, stereo camera, CCD camera, single-lens camera, hyperspectral imaging system, LIDAR system (light detection and ranging system), DCA, dynamometer, IR camera, some other imaging device, or some combination thereof. Therefore, imaging device 225 can capture RGB images, depth images (e.g., 3D images captured using a structured light camera, stereo camera, etc.), or some other suitable type of image. In one embodiment, imaging device 225 is a user device with image capture capabilities (e.g., a smartphone, tablet, laptop). Imaging device 225 can additionally or alternatively capture video. Although in Figure 2In this embodiment, imaging system 220 is shown separate from head-mounted device 210, but in alternative embodiments, imaging system 220 is included in head-mounted device 210. For example, imaging device 225 may be a camera coupled to head-mounted device 210 or a camera integrated into head-mounted device 210 (e.g., imaging device 130).
[0069] In some embodiments, imaging system 220 may apply one or more imaging techniques (e.g., stereo triangulation, sheet of light triangulation, structured light analysis, time-of-flight analysis, interferometry) to determine depth information associated with an image captured by imaging device 225. In a particular embodiment, imaging system 220 includes a DCA that captures an image of a user, and the DCA uses the captured image to determine depth information of the user's head. The depth information describes the distance between surfaces in the captured image and the DCA. The DCA may use one or more of the following to determine the depth information: stereo vision, photometric stereo, time-of-flight (ToF), and structured light (SL). The DCA may compute depth information based on the captured image or send the captured image to another component (e.g., equalization system 230) to extract the depth information. In embodiments where imaging system 220 does not include a DCA, imaging system 220 may provide the captured image to equalization system 230 or some other device and / or console to determine depth information.
[0070] Equalization system 230 generates an equalization filter for the user, which adjusts one or more acoustic parameters of the audio content provided to the user via headset 210 such that the audio output of headset 210 matches a target response at the user's ear. In one embodiment, the equalization filter is generated based on the difference between the audio output and the target response at the ear canal entrance point (EEP) or tympanic membrane reference point (DRP). In this embodiment, EEP refers to the entrance location of the ear canal, while DRP refers to the location of the eardrum. The target response and its physically defined location may vary depending on the type of audio material presented. In one embodiment, the target response may be a flat frequency response as measured at the EEP. In one embodiment, the equalization filter is generated based on a transfer function as a ratio between two complex frequency responses (i.e., the target response and the predicted response).
[0071] Therefore, the equalization filter adjusts the audio output based on the user's ear, so that the user hears the audio output in the way the content creator wants it to be heard. Although in Figure 2In this embodiment, equalization system 230 is shown separate from head-mounted device 210, but in some embodiments, equalization system 230 may be included in head-mounted device 210. In some embodiments, equalization system 230 generates a representation of at least a portion of a user's head (e.g., ears) based on images and / or video received from imaging system 220. Equalization system 230 may use this representation to simulate audio output at the user's ears (e.g., from head-mounted device 210) and determine an equalization filter for the user based on the difference between the audio output at the user's ears and a target response. The target response is how the content creator wants the sound to be heard by the user, and would be standard if not for differences in ear shape and fit. The target response is then the ideal version of the audio output with the highest achievable sound quality. Therefore, the equalization filter specifies the amount of compensation for one or more acoustic parameters of the audio content provided to the user via head-mounted device 210 to account for differences in ear shape and fit, so that the user hears a version of the audio output that is as reasonably close to the target response as possible. Equalization system 230 is discussed below regarding... Figure 3 To provide a more detailed description.
[0072] Online system 240 maintains user profile information and content to be presented to the user. For example, online system 240 may be a social networking system. In some embodiments, online system 240 stores the user profile of head-mounted device 210. Therefore, equalization system 230 may transmit an audio profile including one or more equalization filters to online system 240, and online system 240 may store the audio profile with equalization filters along with the user's online profile. Online system 240 may store personalized equalization filters corresponding to one or more devices for a single user. For example, online system 240 may store a personalized equalization filter for head-mounted device 100 and another personalized equalization filter for head-mounted device 105 for a user. Therefore, head-mounted device 100 and head-mounted device 105 can retrieve the equalization filter for each device, and head-mounted device 100 and head-mounted device 105 can retrieve and use the equalization filter when providing content to the user. Therefore, the user can use head-mounted device 210 without having to re-execute the process for generating personalized equalization filters.
[0073] Network 250 can be any suitable communication network used for data transmission. Network 250 is typically the Internet, but can be any network, including but not limited to a Local Area Network (LAN), Metropolitan Area Network (MAN), Wide Area Network (WAN), mobile wired or wireless network, private network, or Virtual Private Network. In some example embodiments, network 250 is the Internet and uses standard communication technologies and / or protocols. Therefore, network 250 may include links that use technologies such as Ethernet, 802.11, Global Microwave Access Interoperability (WiMAX), 3G, 4G, Digital Subscriber Line (DSL), Asynchronous Transfer Mode (ATM), unlimited bandwidth, PCIexpress advanced switching, etc. In some example embodiments, entities use custom and / or dedicated data communication technologies instead of or supplementing the aforementioned technologies.
[0074] Equilibrium System
[0075] As described above, the equalization system 230 is configured as a user-generated personalized equalization filter for the head-mounted device 210. The equalization system 230 includes an image analysis module 305 and an audio customization model 325. In other embodiments, the equalization system 230 may include fewer or more components than those described herein. Furthermore, the functional distribution of the components may differ from that described below.
[0076] Image analysis module 305 includes feature extraction module 310, which is configured to extract information from one or more images of a user's head and / or ears captured by the user. Feature extraction module 310 receives images from one or more components of system environment 200 (e.g., head-mounted device 210, imaging system 220). The images may include a portion of the user's head (e.g., ears), which may be present when the user is wearing a head-mounted device (e.g., head-mounted device 100) or a head-mounted display (e.g., head-mounted device 105). Feature extraction module 310 may extract information (e.g., depth information, color information) from the images and apply one or more techniques and / or models to determine features describing the user's ears and / or head (e.g., size, shape). Examples include distance imaging techniques, machine learning models (e.g., feature recognition models), algorithms, etc. In one embodiment, feature extraction module 310 extracts anthropometry features describing the user's body characteristics (e.g., ear size, ear shape, head size, etc.).
[0077] In some embodiments, a machine learning model is used to train the feature extraction module 310. The feature extraction module 310 can be trained using images of other users with previously identified features. For example, multiple images can be labeled (e.g., by a person, by another model) using identifying features of the user's ears and / or head (e.g., earlobe size and shape, ear position on the head, etc.). The image analysis module 305 can use the images and associated features to train the feature extraction module 310.
[0078] Image analysis module 305 further includes a depth map generator 315 configured to generate one or more depth maps based on information extracted by feature extraction module 310 (e.g., depth information). Depth map generator 315 can create depth maps of at least a portion of a user's head and identify the relative positions of the user's features. The depth map indicates the positional or spatial relationship between features of interest (e.g., ears) from an image of the user's head. For example, a depth map can indicate the distance between the user's left and right ears or the position of the user's ears relative to other features such as the eyes and shoulders. Similarly, depth map generator 315 can be used to create a depth map of the head wearing a head-mounted device based on an image of the user's head wearing the device. In some embodiments, depth map generator 315 can be used to create a depth map of the head-mounted device using images received from an isolated (i.e., not worn by the user) head-mounted device.
[0079] The reconstruction module 320 generates a 3D representation of at least a portion of the user's head based on features extracted by the feature extraction module 310 and / or a depth map generated by the depth map generator 315. More specifically, the reconstruction module 320 may generate representations of one or both of the user's ears. In one example, the reconstruction module 320 generates a representation of one ear (e.g., the left ear) and a mirror representation of the other ear (e.g., the right ear). Additionally or alternatively, the reconstruction module 320 may generate a 3D mesh representation of the user's head that describes the location of features such as the user's head (e.g., eyes, ears, neck, and shoulders). The reconstruction module 320 may combine the features of the user's head with features of the head-mounted device 210 to obtain a representation of the user's head wearing the head-mounted device. In some embodiments, the representation of the head-mounted device 210 may be predetermined because the head-mounted device 210 worn by the user may have a unique, known identifier to identify the device. In some embodiments, the head-mounted device 210 worn by the user may be identified from an image of the device captured using the imaging device 225 when the device is worn.
[0080] In some embodiments, reconstruction module 320 generates a PCA-based representation of the user based on images of the heads of human test subjects wearing the test headset and measured audio outputs at each test subject's ears. In the PCA-based representation, the user's head or features of the user's head (e.g., ear shape) are represented as linear combinations of principal components multiplied by corresponding PCA coefficients. For this purpose, reconstruction module 320 receives, for example, images from a database and measured audio outputs at the user's ears from a set of test transducers (e.g., a speaker array). Based on the received images of the test subjects (e.g., 500-215 test subjects), reconstruction module 320 performs principal component analysis (PCA), which uses orthogonal transformations to determine a set of linearly uncorrelated principal components. For example, the orientation of the headset on the test subject's ears may be the focus of the PCA.
[0081] Reconstruction module 320 can generate a PCA model to determine PCA-based geometry, as detailed below. Figures 8A-8B This will be discussed further. Although the PCA model is described as being generated and executed within the equalization system 230, the PCA model can actually be executed on a separate computing device. In this case, the results of the PCA are processed and provided to the reconstruction module 320 to process the user's PCA-based representation.
[0082] The audio customization model 325 is configured to predict the audio output at the user's ear and generate a personalized equalization filter for the user based on the difference between the audio output at the ear and the target response. The audio customization model 325 includes a sound simulation module 330, an audio prediction module 335, and an equalization filter generator 345. In other embodiments, the audio customization model 325 may include additional components not described herein.
[0083] The sound simulation module 330 uses the representation generated by the reconstruction module 320 to simulate the audio output at the user's ears from an audio source (e.g., a speaker, speaker array, transducer of a head-mounted device, etc.). In one example, the sound simulation module 330 generates the simulated audio output at the user's ears based on a representation of at least a portion of the user's head. In another example, the sound simulation module 330 generates the simulated audio output at the user's ears based on a representation of at least a portion of the user's head wearing head-mounted device 210 (e.g., head-mounted device 100, head-mounted device 105). Additionally, the head-mounted device 210 in the representation may include multiple transducers (e.g., speakers), and for each transducer (or a subset thereof) in the representation, the sound simulation module 330 simulates the propagation of sound from the transducer to the user's ears. The sound simulation module 330 may also simulate the audio output at one or both of the user's ears.
[0084] In one embodiment, the sound simulation module 330 is a numerical simulation engine. To obtain the simulated audio output at the ear, the sound simulation module 330 can use various simulation schemes, such as (i) the boundary element method (BEM), which is described, for example, in the following articles: Carlos A. Brebbia et al., “Boundary Element Methods in Acoustics,” Springer; 1st ed., ISBN 1851666796 (1991) and Gumerov NA et al., “A broadband fast multipole accelerated boundary element method for the three dimensional Helmholtz equation,” J. Acoust. Soc. Am, Vol. 125, No. 1, pp. 191-205 (2009); (ii) the finite element method (FEM), which is described, for example, in the following article: Thompson, LL., “A review of finite-element methods for time-harmonic…” “Acoustics”, J. Acoust. Soc. Am, Vol. 119, No. 3, pp. 1315-1330 (2006), (iii) Finite-Difference Time-Domain (FDTD) method, which is described, for example, in: “Computational Electrodynamics: The Finite-Difference Time-Domain Method”, Taflove, A. et al., 3rd ed.; Chapters 1 and 4, Artech House, 2005 and “Numerical solution of initial boundary value problems involving Maxwell’s sequences in isotropic media”, IEEE Transactions on Antennas and Propagations, Vol. 14, No. 3, pp. 302-307 (1966), (iv) Fourier pseudospectral time-domain (PSTD) method, which is described, for example, in: “Numerical analysis of sound propagation in rooms using the finitedifference time domain”, Sakamoto, S. et al. “Method”, J. Acoust. Soc. Am, Vol. 120, No. 5, 3008 (2006) and Sakamoto, S."Calculation of impulse responses and acoustic parameters in a hall by the finite-difference time-domain method," *Acoustic Science and Technology*, Vol. 29, No. 4 (2008).
[0085] Audio prediction module 335 is configured to predict features (e.g., spectral content and acoustic group delay) of the audio output at the user's ear of the head-mounted device 210. Audio prediction module 335 may determine the predicted audio output at the ear based on features extracted by feature extraction module 310, a representation of at least a portion of the user's head generated by reconstruction module 320, and / or a simulation performed by sound simulation module 330. The predicted audio output at the ear may be a simulated audio output at the ear generated by sound simulation module 330. Alternatively, audio prediction module 335 may use a machine learning model to determine the predicted audio output at the ear, as will be discussed below. Figures 9A-9B To provide a more detailed description, for example, the audio prediction module 335 can input the features extracted by the feature extraction module 310 into a machine learning model, which is configured to determine the audio output at the ear based on the features.
[0086] One or more machine learning techniques can be used to train the audio prediction module 335 directly from input data of images and videos of head and ear geometry. In one embodiment, the audio prediction module 335 is retrained periodically at a defined time frequency. Feature vectors from both positive and negative training sets can be used as input to train the audio prediction module 335. Different machine learning techniques—such as linear support vector machines (linear SVMs), boosting for other algorithms (e.g., AdaBoost), neural networks, logistic regression, Naive Bayes, memory-based learning, random forests, bagged trees, decision trees, boosted trees, boosted stumps, nearest neighbors, k-nearest neighbors, kernel machines, probabilistic models, conditional random fields, Markov random fields, manifold learning, generalized linear models, generalized exponential models, kernel regression, or Bayesian regression—can be used in different embodiments. The following is about... Figure 9A The training audio prediction module 335 is described in more detail.
[0087] Equalization filter generator 345 generates an equalization filter customized for the user. In one embodiment, equalization filter generator 345 generates an equalization filter based on predicted audio output at the user's ear as predicted by audio prediction module 335. In another embodiment, equalization filter generator 345 generates an equalization filter based on audio output at the user's ear as simulated by sound simulation module 330. As described elsewhere herein, the equalization filter is configured to adjust one or more acoustic parameters of the audio output for the user when applied to the audio output of head-mounted device 210. For example, the equalization filter may be configured to adjust other acoustic parameters, such as pitch, dynamics, timbre, etc. The equalization filter may be a high-pass filter, a low-pass filter, a parametrically personalized equalization filter, a graphic equalization filter, or any other suitable type of personalized equalization filter. In some embodiments, equalization filter generator 345 selects an equalization filter from a set of existing equalization filters, adjusts the parameters of an existing equalization filter to generate a new equalization filter, or adjusts an equalization filter previously generated by equalization filter generator 345 based on predicted audio output at the user's ear. The equalizer filter generator 345 can provide an equalizer filter to the headset 210, and the headset 210 can use the equalizer filter to provide personalized audio content to the user. Additionally or alternatively, the equalizer filter generator 345 can provide an equalizer filter to the online system 240, storing the equalizer filter in association with the user's profile in the online system 240.
[0088] Example Method
[0089] Figure 4A This is an example view of an image of the head of a user 405 captured by an imaging device 225 according to one or more embodiments. Figure 4AIn some embodiments, imaging device 225 captures images including at least the user's ears. Furthermore, imaging device 225 can capture images of the user's head at different angles and orientations. For example, user 405 (or some other party) can position imaging device 225 at different locations relative to his / her head, such that the captured images cover different portions of user 405's head. Additionally, user 405 can hold imaging device 225 at different angles and / or distances relative to user 405. For example, user 405 can hold imaging device 225 at an arm's length directly in front of user 405's face and use imaging device 225 to capture images of user 405's face. User 405 can also hold imaging device 225 at a distance shorter than arm's length, wherein imaging device 225 is pointed to one side of user 405's head to capture images of user 405's ears and / or shoulders. In some embodiments, imaging device 225 is positioned to capture images of both the user's left and right ears. Alternatively, imaging device 225 can capture a 180-degree panoramic view of the user's head, allowing both ears to be captured in a single image or video.
[0090] In some embodiments, imaging device 225 uses feature recognition software and automatically captures images when features of interest (e.g., ears, shoulders) are identified. Additionally or alternatively, imaging device 225 may prompt the user to capture an image when the feature of interest is within the field of view of imaging device 225. In some embodiments, imaging device 225 includes an application with a graphical user interface (GUI) that guides user 405 to capture multiple images of user 405's head from a specific angle and / or distance relative to user 405. For example, the GUI may request a frontal image of user 405's face, an image of user 405's right ear, and an image of user 405's left ear. Imaging device 225 may also determine, (e.g., based on image quality and features captured in the image), whether an image is suitable for use by equalization system 230.
[0091] Figure 4B The invention illustrates the invention according to one or more embodiments. Figure 4A The image captured by imaging device 225 is a side view of the user 405. The focus of the captured image is the user's ear 407. In some embodiments, the equalization system 230 may use... Figure 4B The images shown are used to determine features associated with the user's ear 407 and / or the user's head. Imaging device 225 may capture additional images to determine additional features associated with the user's head.
[0092] Although Figure 4AThe image shown is of the head of user 405 captured by imaging device 225, but imaging device 225 can also capture images of users wearing head-mounted devices (e.g., head-mounted device 100, head-mounted device 105).
[0093] Figure 5A This is an example view of an image captured by an imaging device 225 according to one or more embodiments of a user 405 wearing a head-mounted device 510. The head-mounted device 510 may be an embodiment of a head-mounted device 210, a near-eye display including audio output (e.g., a speaker), or some other head-mounted display including audio output.
[0094] Figure 5B The invention illustrates the invention according to one or more embodiments. Figure 5A The imaging device 225 captures a side view of an image of a user 405 wearing the head-mounted device 510. The equalization system 230 can determine features associated with the user's ear 407 relative to the position of the head-mounted device 510, as described in more detail below. In one embodiment, the head-mounted device 510 includes one or more transducers, and at least one of the transducers is captured in... Figure 5A As shown in the image, the equalization system 230 can therefore determine the distance between the user's ear 407 and one or more transducers.
[0095] Visual models benefit from scale and orientation information of the operation. There are certain situations where scale or orientation information can be encoded; however, these situations are often important. Therefore, in another embodiment, the head-mounted device (e.g., head-mounted device 210) includes one or more visual markers for determining the position of the head-mounted device relative to the user's ears when the user wears the head-mounted device. As described above and elsewhere herein, machine learning-based prediction engines use images and videos of human heads and ears to predict personalized acoustic transfer functions from the head-mounted device measured at the user's ears. Therefore, accurate information about the size and orientation of visually captured anthropometric features is a key requirement for the images and videos to be useful to the model. Various methods can be designed to provide this information, such as including a reference visual object of known size (e.g., a coin) or markers (e.g., multiple points) drawn on features of interest (e.g., ears and eyeglass frames), where their relative distances are measured using a ruler in the captured images and videos. However, these methods are cumbersome and / or unreliable and unsuitable for product applications.
[0096] One approach to eliminating ambiguity regarding the size and orientation of anthropometry features is to use markers designed into the head-mounted device for the explicit purpose of providing visual references within the data. Thus, in one embodiment, images and / or videos are captured while a user wears the head-mounted device (as is typically fitted to a user). Since the dimensions of these markers are known by design, and their orientation relative to the head and ears is expected to be consistent on every user, the markers can achieve the desired property of providing reliable visual references within the image. Furthermore, the unique marker design associated with each head-mounted device model, with its unique dimensions, can be used to identify a product model in the image from which precise information about its industrial design can be inferred.
[0097] Figure 6A This is an example view of an imaging device 225 according to one or more embodiments capturing an image of a user 405 wearing a head-mounted device 610 including a plurality of markers 615. The head-mounted device 610 may be an embodiment of a head-mounted device 210, a near-eye display including audio output (e.g., a speaker), or a head-mounted display including audio output.
[0098] Figure 6B The invention illustrates the invention according to one or more embodiments. Figure 6A The imaging device 225 captures an image of a portion of the user's head. The head-mounted device 610 captured in the image includes four markings 615a, 615b, 615c, and 615d along its right temple arm 612. The head-mounted device 610 may be symmetrical, such that the corresponding location of the head-mounted device 610 on the left temple arm (not shown) includes the same markings. In other embodiments, the head-mounted device 610 may include any other suitable number of markings (e.g., one, three, ten) along the right temple arm, left temple arm, and / or the front of the frame. Figure 6B In this embodiment, each marker 615 has a unique shape and size, such that each marker 615 can be easily identified by the equalization system 230. Alternatively, the markers 615 may be substantially the same size and / or shape. Furthermore, the size of the head-mounted device 610 and the position of the markers 615 relative to the head-mounted device 610 are known. The equalization system 230 can use... Figure 6B The image shown is used to determine information relating to the user's ear 407 relative to the head-mounted device 610. For example, the equalization system 230 can determine the distance between each marker and a point on the user's ear 407.
[0099] Imaging system 220 can capture one or more images (e.g. Figure 4B , Figure 5B and Figure 6BThe images shown are provided to the equalization system 230 to generate an equalization filter for the user 405. The equalization system 230 can also receive additional images from the imaging device 225, showing other views of the user's ears and / or head. The equalization system 230 can determine the audio output at the user's ears based on the images. Furthermore, the images can be used to train one or more components of the equalization system 230, as described in more detail below.
[0100] Deterministic Equalization Filter Based on Simulation
[0101] A high-fidelity audio experience from head-mounted devices requires their audio output to match a consistent target response at the user's ear in terms of spectral content and acoustic time delay. For output modules built into the device frame, static, non-personalized EQ tuned on a mannequin and / or ear coupler is insufficient to provide such high-fidelity audio output because the audio each user hears is affected by multiple sources of variation, such as their anthropometric characteristics (e.g., pinna size and shape), fit inconsistencies, transducer component sensitivity to environmental factors, manufacturing tolerances, etc. Among these, person-to-person and fit-to-fit variations account for the largest portion of audio output variability and are determined by the shape of the user's head and / or ears and the relative position of the audio output module on the head-mounted device to the user's ears.
[0102] For headsets with open-ear audio output (consisting of a frame with speaker modules embedded in the temple arms), industrial design practices can be employed to ensure that the user's fit is repeatable and stable during normal use of the device, thereby minimizing fit-to-fit variation in audio output. However, eliminating person-to-person variation requires understanding the audio output at the user's ear, which can be used to compensate for this variation by applying a personalized inverse equalization filter, as described herein. One approach to obtaining this understanding is to place a microphone at the entrance of the ear canal to measure the raw response of the audio output. The practical application of this approach presents challenges in terms of comfort and aesthetics in industrial design, and also challenges in terms of ease of use in user experience. Therefore, an alternative method is desired to measure or predict the audio output at the wearer's ear.
[0103] In one embodiment, a method for achieving such a goal includes building a machine learning model capable of reconstructing the 3D geometry of a human head and / or ear wearing a headset based on videos and images. This model is trained using a dataset consisting of images and videos of human subjects wearing the headset, along with high-quality 3D scan meshes of the corresponding user's head and ears. The reconstructed 3D geometry is then used as input to a numerical simulation engine to predict the acoustic propagation of the headset output to the ear, thereby predicting the audio output observed at the user's ear. The predicted response can be used to generate device-specific, personalized equalization filters for the user's audio.
[0104] Figure 7 An example method for generating equalization filters for a user based on a user's ear representation, according to an embodiment, is shown. These steps can be performed by... Figure 2 The system environment 200 shown may be executed by one or more components (e.g., the equalization system 230). In other embodiments, these steps may be performed in a different order than those described herein.
[0105] The equalization system 230 receives one or more images of at least a portion of the user's head. In one embodiment, the equalization system 230 receives one or more images of the user's ears, the user's head, and / or the user wearing the head-mounted device 210. For example, the equalization system 230 receives... Figure 4B The image shown. The image can be captured using an imaging device 225 associated with a user device (e.g., a mobile phone).
[0106] The equalization system 230 generates a representation of at least a portion of a user's head based on one or more images. In some embodiments, the equalization system 230 generates a representation of one or both ears of the user. Alternatively, the equalization system 230 may generate a representation of the user's head (including one or both ears). The generated representation may be a 3D mesh representing the user's ears and / or head, or a PCA-based representation, as described below. Figures 8A-8B To provide a more detailed description.
[0107] The equalization system 230 performs an analog simulation of audio propagation from an audio system included in the head-mounted device to the user's ear based on the user's ear representation. The audio system may be an array of transducers coupled to the left temple arm and / or the right temple arm of the head-mounted device 210. The equalization system 230 determines a predicted audio output response based on the simulation. For example, the equalization system 230 may determine one or more acoustic parameters (e.g., pitch, frequency, volume, balance, etc.) perceived by the user based on the simulation.
[0108] Equalization system 230 generates an equalization filter 740 based on the predicted audio output response. Therefore, a user can experience a customized audio environment provided by head-mounted device 210. For example, due to the user's anthropometric characteristics, the predicted audio output response may have frequencies above average, and equalization system 230 generates an equalization filter that reduces the frequencies of the audio content provided to the user. In some embodiments, equalization system 230 provides the equalization filter to head-mounted device 210, allowing head-mounted device 210 to use the equalization filter to adjust the audio content provided to the user. Additionally, equalization system 230 can provide the equalization filter to online system 240, and online system 240 can store the equalization filter in the profile of the user associated with online system 240 (e.g., a social network profile).
[0109] In some embodiments, a trained model (e.g., a PCA model) is used to generate the aforementioned representation of the user's ear. Employing machine learning techniques allows the reconstruction module 320 to generate a more accurate representation of the user's ear and / or head. Figure 8A This is a block diagram of a trained PCA model 860 according to one or more embodiments. The machine learning process can be used to generate a PCA-based representation of a user's ear and determine an audio output response for the user.
[0110] The reconstruction module 320 receives information (e.g., features from an image of a user's head) from the feature extraction module 310 and / or the depth map generator 315. Based on this information, the reconstruction module 320 generates a PCA-based representation of the user's head using a PCA model 860. In one embodiment, the PCA-based representation also includes a representation of the head-mounted device. Thus, the reconstruction module 320 can use the PCA model 860, trained to generate a PCA-based representation in which the shape of a human head wearing a head-mounted device (e.g., ear shape) is represented as a linear combination of the three-dimensional shapes of the head or head features of a representative test subject wearing the head-mounted device. In other embodiments, the PCA model 860 is trained to generate a PCA-based representation of a head-mounted device (e.g., head-mounted device 210), which is represented as a linear combination of the three-dimensional shapes of representative images of the head-mounted device. PCA model 860 can also be trained to generate PCA-based representations, wherein the shape of a human head or human head feature (e.g., ear shape) is represented as a linear combination of the three-dimensional shapes of the head or head features of a representative test subject. In other embodiments, PCA model 860 may combine a PCA-based representation of the head with a PCA-based representation of the head-mounted device to obtain a PCA-based representation of the head wearing the head-mounted device. Alternatively, PCA model 860 may be trained to generate PCA-based representations, wherein the shape of a human head or human head feature (e.g., ear shape) wearing the head-mounted device (e.g., head-mounted device 210) is represented as a linear combination of the three-dimensional shapes of the head or head features of a representative test subject when wearing the head-mounted device.
[0111] Taking PCA analysis of the ear shape of a head wearing a head-mounted device as an example, the three-dimensional shape of a random ear shape E can be represented as follows:
[0112] E = ∑ (α) i ×ε i (1)
[0113] Where, α i Let ε represent the i-th principal component (i.e., the i-th representative ear shape in 3D). i This represents the PCA coefficient of the i-th principal component. The number of principal components (the number of "i") is chosen such that it is less than the total number of test subjects who provided the measured audio output response. In the example, the number of principal components is between 5 and 10.
[0114] In some embodiments, the PCA-based representation is generated using a representation of the head shape of a test subject wearing the headset and its measured audio output response, such that the PCA-based representation obtained from PCA model 860 can produce a more accurate equalization filter through simulation compared to performing a simulation on the three-dimensional mesh geometry of the same user's head wearing the headset. The test subject referred to herein is a human or a physical model of a human, for which their head shape geometry (or head shape image) and audio output response (i.e., the "measured audio output response") are known. To obtain the audio output response, the test subject may be placed in an anechoic chamber and exposed to sound from one or more transducers, with microphones placed at the test subject's ears. In some embodiments, the audio output response is measured for a test headset (including a test transducer array) worn by the test subject. The test headset is substantially identical to a headset worn by a user.
[0115] like Figure 8A As shown, the PCA model 860 provides a PCA-based representation to the sound simulation module 330, and the sound simulation module 330 uses the PCA-based representation to perform a simulated audio output response. The equalization system 230 can compare the measured audio output response of the test subject with the simulated audio output response to update the PCA model 860, as described below regarding... Figure 8B A more detailed description follows. After the PCA model is determined and / or updated, PCA model 860 is trained using images of the heads of test subjects wearing the head-mounted device and their PCA-based representations, based on PCA model 860. The trained PCA model 860 can predict or infer a PCA-based representation of the user's head from images of the user's head wearing the head-mounted device. In some embodiments, the trained PCA model 860 can predict or infer a PCA-based representation of the user's head wearing the head-mounted device from images of the user's head and other images of the head-mounted device.
[0116] In some embodiments, the generation and training of the PCA model 860 can be performed offline. The trained PCA model 860 can then be deployed in the reconstruction module 320 of the equalization system 230. Using the trained PCA model 860 enables the reconstruction module 320 to generate the user's PCA-based representation in a robust and efficient manner.
[0117] Figure 8B This is a flowchart illustrating the generation and updating of PCA model 860 according to one or more embodiments. In one embodiment, Figure 8BThe process is performed by components of the equalization system 230. In other embodiments, other entities may perform some or all of the steps of the process. Similarly, embodiments may include different steps and / or additional steps, or perform these steps in a different order.
[0118] The equilibration system 230 determines an initial PCA model. In some embodiments, the equilibration system 230 determines the initial PCA model by selecting a subset of the test subject's head (or a portion thereof) as principal components to represent random head shapes or head shape features.
[0119] Equalization system 230 uses the current PCA model to determine a PCA-based representation of the 820 test images. For example, the initial PCA model processes images of the test subject's head wearing a test headset, which includes a test transducer array, to determine a PCA-based representation of the test subject's head or portions of the test subject's head (e.g., ears) when wearing the test headset. That is, the head shape (or the shape of a portion of the head) of all test subjects wearing the headset is represented as a linear combination of subsets of the test subject's head shapes multiplied by the corresponding PCA coefficients, as explained above with reference to equation (1). Note that the test headset is substantially the same as the headset worn by the user.
[0120] The equalization system 230 uses a PCA-based representation to perform one or more analog operations to generate an analog audio output response. (See reference above.) Figure 3 The process involves performing one or more simulations on the PCA-based representation using one or more of the BEM, FEM, FDTD, or PSTD methods. As a result of the simulations, the equalization system 230 obtains the simulated audio output response of the test subject based on the current PCA model.
[0121] The equalization system 230 determines whether the difference between the measured audio output response and the simulated audio output response of the 840 test subjects is greater than a threshold. This difference can be the sum of the magnitudes of the differences between the measured audio output response and the simulated audio output response for each test subject.
[0122] If the difference is greater than a threshold, the equilibrium system 230 updates the PCA model 850 to a new current PCA model. Updating the PCA model may include increasing or decreasing the number of principal components, updating PCA coefficient values, or updating the representative shape. The process then returns to determining a new set of PCA-based representations 820 based on the updated current PCA model, and repeats the subsequent steps.
[0123] If the equilibrium system 230 determines that the difference is less than or equal to the threshold, then the current PCA model is finally determined as the PCA model for deployment (i.e., for the purposes of the above regarding...). Figure 7 The described equalization system 230 is used.
[0124] Use the trained model to determine the audio output response.
[0125] In another embodiment, the equalization system 230 uses a machine learning model to determine the audio output response. The machine learning model can be trained using a dataset consisting of images and videos of human subjects wearing a head-mounted device and audio output responses measured at the ears of the respective subjects. The model will be able to predict the audio output response for new users based on images and videos of their head and ear geometry. Therefore, in this embodiment, the machine learning model computes the equalization filter directly based on anthropometry features visually extracted from the images and videos.
[0126] Figure 9A A machine learning process for predicting an audio output response according to an embodiment is illustrated. Feature extraction module 310 receives images of a user's head, including at least the user's ears. Feature extraction module 310 extracts features describing the user's ears and provides the extracted features to audio prediction module 335. Audio prediction module 335 uses response model 970 (i.e., a machine learning model) to predict the audio output response based on the features of the user's ears. Response model 970 is generated and trained using images of additional users, their associated features, and measured audio response profiles. In some embodiments, response model 970 can be updated by comparing the predicted audio output response of the additional user with the measured audio output response of the additional user. The additional user described herein refers to a human or a physical model of a human, for which their anthropometry features and audio output response are known. Anthropometry features can be determined by a human or another model. To obtain the audio output response, the additional user can be placed in an anechoic chamber and exposed to sound from one or more transducers, with microphones placed at the additional user's ears. In some embodiments, the audio output response is measured against a test headset (including a test transducer array) worn by the additional user. The test headset is essentially the same as the one worn by the user. The trained response model 970 can be used to predict the audio output response, which is discussed below. Figure 9B To provide a more detailed description.
[0127] Figure 9B A method for generating an equalization filter based on the audio output at the user's ear determined using a response model 970, according to one embodiment, is illustrated. These steps can be performed by... Figure 2The system environment 200 shown is executed by one or more components (e.g., the equalization system 230). In one embodiment, the above-mentioned... Figure 9A The described machine learning response model 970 performs this process. Method 900 may include fewer or more steps than those described herein.
[0128] The equalization system 230 receives one or more images of the user's ears and / or head. In one embodiment, the equalization system 230 receives one or more images of the user's ears, the user's head, and / or the user wearing the head-mounted device 210 (e.g., Figure 4B , Figure 5B and Figure 6B (The images shown). These images can be captured using an imaging device 225 (e.g., a mobile phone).
[0129] The equalization system 230 identifies 920 one or more features from an image describing the user's ears. These features may describe anthropometry information (e.g., size, position, shape) related to the user's ears and / or head. Features may be based on information extracted from the image (e.g., depth information, color information). In some embodiments, features may be identified relative to a head-mounted device. For example, in Figure 6B In one embodiment, the equalization system extracts information related to the position of the head-mounted device 510 relative to the user's ear 407 based on marker 615 to determine the characteristics of the ear.
[0130] Equalization system 230 provides features as input to model 930 (e.g., response model 970). This model is configured to determine the audio output response based on the features. The model is trained using images of additional users' ears and features extracted from those images, where the audio output response for each additional user is known. Equalization system 230 can periodically retrain the model and use the trained model to predict audio output responses for users.
[0131] Equalization system 230 generates an equalization filter 940 based on the predicted audio output at the user's ear. This equalization filter is configured to adjust one or more acoustic parameters of the audio content provided to the user by the headset. Equalization system 230 can provide the equalization filter to the headset (e.g., headset 610) so that the headset can use the equalization filter to provide audio content to the user. Additionally, equalization system 230 can provide the equalization filter to online system 240 for associating the equalization filter with the user's online profile.
[0132] The trained response model 970 allows the equalization system 230 to quickly and efficiently predict the audio output at the user's ears based on the user's image. Therefore, the equalization system 230 can generate equalization filters configured to tailor audio content to the user, thereby enhancing the user's audio experience. In some embodiments, the response model 970 can be used for multiple users and multiple devices. Alternatively, the response model 970 can be customized for a specific device to adjust the audio output for a particular user-device combination. For example, the equalization filter generator 345 can generate a model for head-mounted device 100 and another model for head-mounted device 105, the models being generated based on an image of a user wearing the respective device and the audio output measured at the user's ears. The equalization system 230 can thus generate personalized equalization filters specific to each user and device.
[0133] In some embodiments, Figure 7 and Figures 8A-8B The aspects of the process shown can be related to Figures 9A-9B The aspects of the process shown are combined to enhance the user's audio experience. For example, in Figure 9B In some embodiments, the equalization system 230 may additionally generate a 3D representation of the user's ear, such as regarding Figure 7 As described, a 3D representation can be input into the model to generate a prediction of the audio output response without performing a simulation. The equalization system 230 can additionally generate equalization filters based on a combination of the model and / or the process. The equalization system 230 can be provided to the audio system 215 of the head-mounted device 210, as described in more detail below.
[0134] audio system
[0135] Figure 10 This is a block diagram of an audio system 215 according to one or more embodiments. Figure 1A or Figure 1B The audio system in the example can be an embodiment of audio system 215. In some embodiments, audio system 215 uses a personalized audio output response generated by equalization system 230 to generate and / or modify audio content for the user. Figure 2 In some embodiments, the audio system 215 includes a transducer array 1010, a sensor array 1020, and an audio controller 1030. Some embodiments of the audio system 215 have components that are different from those described herein. Similarly, in some cases, functionality may be distributed among the components in a manner different from that described herein.
[0136] Transducer array 1010 is configured to present audio content. Transducer array 1010 includes a plurality of transducers. A transducer is a device that provides audio content. A transducer can be, for example, a loudspeaker (e.g., loudspeaker 160), a tissue transducer (e.g., tissue transducer 170), some other device for providing audio content, or some combination thereof. A tissue transducer can be configured to function as a bone conduction transducer or a cartilage conduction transducer. Transducer array 1010 can present audio content via air conduction (e.g., via one or more loudspeakers), via bone conduction (via one or more bone conduction transducers), via a cartilage conduction audio system (via one or more cartilage conduction transducers), or some combination thereof. In some embodiments, transducer array 1010 may include one or more transducers to cover different portions of a frequency range. For example, a piezoelectric transducer may be used to cover a first portion of the frequency range, while a dynamic transducer may be used to cover a second portion of the frequency range.
[0137] A bone conduction transducer generates sound pressure waves by vibrating the bones / tissues of the user's head. The bone conduction transducer can be coupled to a portion of a head-mounted device and can be configured to couple to a portion of the user's skull behind the auricle. The bone conduction transducer receives vibration commands from an audio controller 1030 and vibrates a portion of the user's skull based on the received commands. The vibrations from the bone conduction transducer generate sound pressure waves that propagate through tissue, bypassing the eardrum and reaching the user's cochlea.
[0138] A cartilage transducer generates sound pressure waves by vibrating one or more portions of the auricular cartilage in a user's ear. The cartilage transducer can be coupled to a portion of a head-mounted device and can be configured to couple to one or more portions of the auricular cartilage in the ear. For example, the cartilage transducer can be coupled to the posterior part of the auricle of the user's ear. The cartilage transducer can be located anywhere along the auricular cartilage around the outer ear (e.g., the auricle, tragus, other portion of the auricular cartilage, or some combination thereof). Vibration of one or more portions of the auricular cartilage can generate: sound pressure waves propagating through the air outside the ear canal; sound pressure waves generated by tissue that cause certain portions of the ear canal to vibrate, thereby generating air-propagating sound pressure waves within the ear canal; or some combination thereof. The generated air-propagating sound pressure waves propagate along the ear canal towards the eardrum.
[0139] Transducer array 1010 generates audio content according to instructions from audio controller 1030. In some embodiments, the audio content is spatialized. Spatialized audio content is audio content that sounds to originate from a specific direction and / or target area (e.g., objects and / or virtual objects in a local area). For example, spatialized audio content could make the sound appear to be from a virtual singer across the room from the user of audio system 215. Transducer array 1010 may be coupled to a wearable device (e.g., head-mounted device 100 or head-mounted device 105). In alternative embodiments, transducer array 1010 may be multiple speakers detached from the wearable device (e.g., coupled to an external console).
[0140] In one embodiment, transducer array 1010 uses one or more personalized audio output responses generated by equalization system 230 to provide audio content to a user. Each transducer in transducer array 1010 may use the same personalized audio output response, or each transducer may correspond to a unique personalized audio output response. One or more personalized audio output responses may be received from equalization system 230 and / or sound filter module 1080.
[0141] Sensor array 1020 detects sound within a local area surrounding it. Sensor array 1020 may include multiple acoustic sensors, each detecting changes in air pressure of sound waves and converting the detected sound into an electronic format (analog or digital). The multiple acoustic sensors may be located on a head-mounted device (e.g., head-mounted device 100 and / or head-mounted device 105), a user (e.g., in the user's ear canal), a neckband, or some combination thereof. The acoustic sensors may be, for example, microphones, vibration sensors, accelerometers, or any combination thereof. In some embodiments, sensor array 1020 is configured to use at least some of the multiple acoustic sensors to monitor audio content generated by transducer array 1010. Increasing the number of sensors can improve the accuracy of information describing the sound field generated by transducer array 1010 and / or sound from a local area (e.g., directionality).
[0142] The audio controller 1030 controls the operation of the audio system 215. Figure 10 In some embodiments, the audio controller 1030 includes a data storage unit 1035, a DOA estimation module 1040, a transfer function module 1050, a tracking module 1060, a beamforming module 1070, and a sound filter module 1080. In some embodiments, the audio controller 1030 may be located inside the head-mounted device. Some embodiments of the audio controller 1030 have different components than those described herein. Similarly, functions may be distributed among the components in a manner different from that described herein. For example, some functions of the controller may be performed externally to the head-mounted device.
[0143] Data storage 1035 stores equalization filters and other data for use by audio system 215. The data in data storage 1035 may include sound, audio content, head-related transfer function (HRTF), transfer functions of one or more sensors, array transfer function (ATF) of one or more acoustic sensors, personalized audio output response, audio profile, sound source localization, virtual model of local area, direction of arrival estimation, sound filters, and other data related to use by audio system 215, or any combination thereof.
[0144] The DOA estimation module 1040 is configured to locate sound sources in a local area, in part based on information from the sensor array 1020. Localization is the process of determining the location of a sound source relative to the user of the audio system 215. The DOA estimation module 1040 performs DOA analysis to locate one or more sound sources within the local area. DOA analysis may include analyzing the intensity, spectrum, and / or time of arrival of each sound at the sensor array 1020 to determine the direction from which the sound source originates. In some cases, DOA analysis may include any suitable algorithm used to analyze the surrounding acoustic environment in which the audio system 215 is located.
[0145] For example, DOA analysis can be designed to receive an input signal from sensor array 1020 and apply digital signal processing algorithms to the input signal to estimate the direction of arrival. These algorithms can include, for example, delay and summation algorithms, where the input signal is sampled and weighted and delayed versions of the resulting sampled signals are averaged together to determine the DOA. A least mean square (LMS) algorithm can also be implemented to create an adaptive filter. This adaptive filter can then be used, for example, to identify differences in signal strength or differences in arrival time. These differences can then be used to estimate the DOA. In another embodiment, the DOA can be determined by transforming the input signal into the frequency domain and selecting specific bins in the time-frequency (TF) domain to be processed. Each selected TF bin can be processed to determine whether the bin contains a portion of the audio spectrum with a direct-path audio signal. Those bins containing a portion of the direct-path signal can then be analyzed to identify the angle at which sensor array 1020 receives the direct-path audio signal. The determined angle can then be used to identify the DOA of the received input signal. Other algorithms not listed above can also be used, either alone or in combination with the algorithms described above, to determine the DOA.
[0146] In some embodiments, the DOA estimation module 1040 can also determine the DOA relative to the absolute position of the audio system 215 within a local area. The position of the sensor array 1020 can be received from an external system (e.g., another component of the head-mounted device, an artificial reality console, a mapping server, a position sensor (e.g., position sensor 190), etc.). The external system can create a virtual model of the local area, in which the local area and the position of the audio system 215 are mapped. The received position information may include the location and / or orientation of some or all of the audio system 215 (e.g., sensor array 1020). The DOA estimation module 1040 can update the estimated DOA based on the received position information.
[0147] The transfer function module 1050 is configured to generate one or more acoustic transfer functions. Generally, a transfer function is a mathematical function that provides a corresponding output value for each possible input value. Based on the parameters of the detected sound, the transfer function module 1050 generates one or more acoustic transfer functions associated with the audio system. The acoustic transfer functions can be array transfer functions (ATFs), head-related transfer functions (HRTFs), other types of acoustic transfer functions, or some combination thereof. An ATF characterizes how a microphone receives sound from a point in space.
[0148] The ATF comprises multiple transfer functions characterizing the relationship between sound and the corresponding sound received by the acoustic sensors in sensor array 1020. Therefore, for each sound source, each acoustic sensor in sensor array 1020 has a corresponding transfer function. This set of transfer functions is collectively referred to as the ATF. Thus, for each sound source, there exists a corresponding ATF. Note that a sound source can be, for example, a person or object generating sound in a local area, a user, or one or more transducers in transducer array 1010. The ATF relative to a specific sound source location in sensor array 1020 may vary from user to user because a person's anatomy (e.g., ear shape, shoulders, etc.) affects the sound as it reaches the person's ears. Therefore, the ATF of sensor array 1020 is personalized for each user of audio system 215.
[0149] In some embodiments, the transfer function module 1050 determines one or more HRTFs for a user of the audio system 215. An HRTF characterizes how an ear receives sound from a point in space. An HRTF relative to a person's specific source location is unique to each ear (and unique to that person) because a person's anatomy (e.g., ear shape, shoulders, etc.) affects sound as it reaches the ear. In some embodiments, the transfer function module 1050 may determine the HRTF for the user using a calibration process. In some embodiments, the transfer function module 1050 may provide information about the user to a remote system. The remote system uses, for example, machine learning to determine a set of HRTFs tailored to the user and provides this tailored set of HRTFs to the audio system 215.
[0150] Tracking module 1060 is configured to track the localization of one or more sound sources. Tracking module 1060 can compare current DOA estimates and compare them to a stored history of previous DOA estimates. In some embodiments, audio system 215 can periodically recalculate DOA estimates, for example, once per second or once per millisecond. The tracking module can compare the current DOA estimate with the previous DOA estimate, and in response to a change in the DOA estimate of a sound source, tracking module 1060 can determine that the sound source has moved. In some embodiments, tracking module 1060 can detect changes in localization based on visual information received from a head-mounted device or some other external source. Tracking module 1060 can track the movement of one or more sound sources over time. Tracking module 1060 can store values regarding the number of sound sources and the localization of each sound source at each time point. In response to a change in the number of sound sources or the value of their localization, tracking module 1060 can determine that a sound source has moved. Tracking module 1060 can calculate an estimate of the localization variance. The localization variance can be used as a confidence level for each determination of a change in movement.
[0151] Beamforming module 1070 is configured to process one or more ATFs to selectively emphasize sound from a sound source within a certain area while attenuating sound from other areas. When analyzing sound detected by sensor array 1020, beamforming module 1070 can combine information from different acoustic sensors to emphasize relevant sound from a specific area of the local region while attenuating sound from outside that region. Beamforming module 1070 can isolate the audio signal associated with sound from a particular sound source from other sound sources in the local region based on, for example, different DOA estimates from DOA estimation module 1040 and tracking module 1060. Beamforming module 1070 can thus selectively analyze discrete sound sources in the local region. In some embodiments, beamforming module 1070 can enhance the signal from the sound source. For example, beamforming module 1070 can apply a sound filter that eliminates signals above, below, or in between specific frequencies. Signal enhancement is used to amplify the sound associated with a given identified sound source relative to other sounds detected by sensor array 1020.
[0152] The sound filter module 1080 determines a sound filter (e.g., an equalization filter) for the transducer array 1010. In some embodiments, the sound filter spatializes the audio content so that it sounds as if it originates from a target region. The sound filter module 1080 can generate the sound filter using HRTF and / or acoustic parameters. Acoustic parameters describe the acoustic properties of a local region. Acoustic parameters may include, for example, reverberation time, reverberation level, room impulse response, etc. In some embodiments, the sound filter module 1080 calculates one or more acoustic parameters. In some embodiments, the sound filter module 1080 requests acoustic parameters from a mapping server (e.g., as referenced below). Figure 11 (as described above). In some embodiments, the sound filter module 1080 receives one or more equalization filters or personalized equalization filters from the equalization system 230. The sound filter module 1080 provides sound filters (e.g., personalized equalization filters) to the transducer array 1010. In some embodiments, the sound filters can cause positive or negative amplification of sound according to frequency.
[0153] Figure 11 The system 1100 includes a head-mounted device 1105 according to one or more embodiments. In some embodiments, the head-mounted device 1105 may be... Figure 1A Head-mounted device 100 or Figure 1B The head-mounted device 105. The system 1100 can operate in an artificial reality environment (e.g., a virtual reality environment, an augmented reality environment, a mixed reality environment, or some combination thereof). Figure 11The system 1100 shown includes a head-mounted device 1105, an input / output (I / O) interface 1110 coupled to a console 1115, a network 1120, and a mapping server 1125. Although Figure 11 An example system 1100 is shown, including a head-mounted device 1105 and an I / O interface 1110; however, in other embodiments, the system 1100 may include any number of these components. For example, there may be multiple head-mounted devices, each with an associated I / O interface 1110, and each head-mounted device and I / O interface 1110 communicating with a console 1115. In alternative configurations, the system 1100 may include different and / or additional components. Furthermore, in some embodiments, [the following is a continuation of the previous paragraph, but the translation is incomplete]. Figure 11 The functions described by one or more components shown can be combined with Figure 11 The components are distributed in different ways. For example, some or all of the functions of the console 1115 may be provided by the head-mounted device 1105.
[0154] The head-mounted device 1105 includes a display assembly 1130, an optical block 1135, one or more position sensors 1140, and a DCA 1145. Some embodiments of the head-mounted device 1105 have a combination with... Figure 11 The components described are different components. Furthermore, in other embodiments, [they are combined with...] Figure 11 The functions provided by the various components described may be distributed differently among the components of the head-mounted device 1105, or captured in separate components remote from the head-mounted device 1105.
[0155] Display component 1130 displays content to the user based on data received from console 1115. Display component 1130 uses one or more display elements (e.g., display element 120) to display content. The display elements can be, for example, electronic displays. In various embodiments, display component 1130 includes a single display element or multiple display elements (e.g., displays for each of the user's eyes). Examples of electronic displays include: liquid crystal displays (LCDs), organic light-emitting diode (OLED) displays, active-matrix organic light-emitting diode (AMOLED) displays, waveguide displays, some other type of display, or some combination thereof. Note that in some embodiments, display element 120 may also include some or all of the functionality of optical block 1135.
[0156] Optical block 1135 can amplify image light received from an electronic display, correct optical errors associated with the image light, and present the corrected image light to one or both windows of head-mounted device 1105. In various embodiments, optical block 1135 includes one or more optical elements. Example optical elements included in optical block 1135 include: apertures, Fresnel lenses, convex lenses, concave lenses, filters, reflective surfaces, or any other suitable optical elements that affect image light. Furthermore, optical block 1135 can include combinations of different optical elements. In some embodiments, one or more optical elements in optical block 1135 may have one or more coatings, such as partially reflective or antireflective coatings.
[0157] The amplification and focusing of image light by optical block 1135 allows the electronic display to be physically smaller, lighter, and consume less power than larger displays. Additionally, the amplification increases the field of view of the content presented on the electronic display. For example, the field of view of the displayed content allows the displayed content to utilize almost the user's field of view (e.g., approximately 110 degrees diagonally), and in some cases, the entire field of view. Furthermore, in some embodiments, the amount of amplification can be adjusted by adding or removing optical elements.
[0158] In some embodiments, optical block 1135 may be designed to correct one or more types of optical errors. Examples of optical errors include barrel or pincushion distortion, longitudinal chromatic aberration, or lateral chromatic aberration. Other types of optical errors may further include spherical aberration, chromatic aberration, or errors caused by lens field curvature, astigmatism, or any other type of optical error. In some embodiments, the content provided to the electronic display for display is pre-distorted, and optical block 1135 corrects the distortion when it receives content-based image light from the electronic display.
[0159] Position sensor 1140 is an electronic device that generates data indicating the position of head-mounted device 1105. Position sensor 1140 generates one or more measurement signals in response to movement of head-mounted device 1105. Position sensor 190 is an embodiment of position sensor 1140. Examples of position sensor 1140 include: one or more IMUs, one or more accelerometers, one or more gyroscopes, one or more magnetometers, another suitable type of sensor for detecting motion, or some combination thereof. Position sensor 1140 may include multiple accelerometers measuring translational motion (forward / backward, up / down, left / right) and multiple gyroscopes measuring rotational motion (e.g., pitch, yaw, roll). In some embodiments, the IMU rapidly samples the measurement signals and calculates an estimated position of head-mounted device 1105 based on the sampled data. For example, the IMU integrates the measurement signals received from the accelerometers over time to estimate a velocity vector and integrates the velocity vector over time to determine an estimated position of a reference point on head-mounted device 1105. A reference point is a point that can be used to describe the position of the head-mounted device 1105. Although a reference point can usually be defined as a point in space, it is actually defined as a point within the head-mounted device 1105.
[0160] The DCA 1145 generates depth information for a portion of a local area. The DCA includes a DCA controller and one or more imaging devices. The DCA 1145 may also include an illuminator. The operation and structure of the DCA 1145 are described above. Figure 1A It has been described.
[0161] Audio system 1150 provides audio content to a user of head-mounted device 1105. Audio system 1150 is substantially the same as audio system 215 described above. Audio system 1150 may include one or more acoustic sensors, one or more transducers, and an audio controller. In some embodiments, audio system 1150 receives one or more equalization filters from equalization system 230 and applies the equalization filters to one or more transducers. Audio system 1150 may provide spatialized audio content to the user. In some embodiments, audio system 1150 may request acoustic parameters from mapping server 1125 via network 1120. Acoustic parameters describe one or more acoustic characteristics of a local area (e.g., room impulse response, reverberation time, reverberation level, etc.). Audio system 1150 may provide information describing at least a portion of the local area from, for example, DCA 1145 and / or positioning information of head-mounted device 1105 from position sensor 1140. The audio system 1150 can generate one or more sound filters using one or more acoustic parameters received from the mapping server 1125, and use the sound filters to provide audio content to the user.
[0162] I / O interface 1110 is a device that allows a user to send action requests and receive responses from console 1115. An action request is a request to perform a specific action. For example, an action request may be an instruction to start or stop capturing image or video data, or an instruction to perform a specific action within an application. I / O interface 1110 may include one or more input devices. Example input devices include a keyboard, mouse, game controller, or any other suitable device for receiving action requests and transmitting them to console 1115. Action requests received by I / O interface 1110 are transmitted to console 1115, which performs the action corresponding to the action request. In some embodiments, I / O interface 1110 includes an IMU that captures calibration data indicating the estimated position of I / O interface 1110 relative to its initial position. In some embodiments, I / O interface 1110 may provide haptic feedback to a user based on instructions received from console 1115. For example, haptic feedback can be provided when a motion request is received, or the console 1115 can send instructions to the I / O interface 1110, thereby causing the I / O interface 1110 to generate haptic feedback when the console 1115 performs an action.
[0163] The console 1115 provides content to the head-mounted device 1105 for processing based on information received from one or more of the following: DCA 1145, the head-mounted device 1105, and the I / O interface 1110. Figure 11 In the example shown, console 1115 includes application storage 1155, tracking module 1160, and engine 1165. Some embodiments of console 1115 have a combination with... Figure 11 The different modules or components described. Similarly, the functions further described below can be combined in different ways. Figure 11 The described manner is distributed among the components of console 1115. In some embodiments, the functions discussed herein with reference to console 1115 may be implemented in head-mounted device 1105 or a remote system.
[0164] Application storage 1155 stores one or more applications executed by console 1115. An application is a set of instructions that, when executed by a processor, generate content to be presented to a user. The content generated by the application may respond to input received from the user via movement of the head-mounted device 1105 or I / O interface 1110. Examples of applications include: game applications, conferencing applications, video playback applications, or other suitable applications.
[0165] Tracking module 1160 uses information from DCA 1145, one or more position sensors 1140, or some combination thereof, to track the movement of head-mounted device 1105 or I / O interface 1110. For example, tracking module 1160 determines the position of a reference point of head-mounted device 1105 in a mapping of a local region based on information from head-mounted device 1105. Tracking module 1160 can also determine the position of an object or virtual object. Additionally, in some embodiments, tracking module 1160 can use portions of data from position sensors 1140 indicating the position of head-mounted device 1105 and a representation of a local region from DCA 1145 to predict the future location of head-mounted device 1105. Tracking module 1160 provides engine 1165 with an estimated or predicted future position of head-mounted device 1105 or I / O interface 1110.
[0166] Engine 1165 executes the application and receives position information, acceleration information, velocity information, predicted future position, or a combination thereof from tracking module 1160 of head-mounted device 1105. Based on the received information, engine 1165 determines the content to be presented to the user on head-mounted device 1105. For example, if the received information indicates that the user has looked to the left, engine 1165 generates content for head-mounted device 1105 that reflects the user's movement in a virtual local area or in a local area enhanced with additional content. Furthermore, engine 1165 performs in-application actions on console 1115 in response to action requests received from I / O interface 1110 and provides feedback to the user that the action has been performed. The feedback provided may be visual or auditory feedback via head-mounted device 1105 or haptic feedback via I / O interface 1110.
[0167] Network 1120 couples the head-mounted device 1105 and / or console 1115 to the mapping server 1125. Network 1120 may include any combination of local area networks and / or wide area networks using wireless and / or wired communication systems. For example, network 1120 may include the Internet and mobile phone networks. In one embodiment, network 1120 uses standard communication technologies and / or protocols. Therefore, network 1120 may include links using technologies such as Ethernet, 802.11, WiMAX, 2G / 3G / 4G mobile communication protocols, Digital Subscriber Line (DSL), Asynchronous Transfer Mode (ATM), unlimited bandwidth, PCI Express Advanced Switching, etc. Similarly, network protocols used on network 1120 may include Multiprotocol Label Switching (MPLS), Transmission Control Protocol / Internet Protocol (TCP / IP), User Datagram Protocol (UDP), Hypertext Transfer Protocol (HTTP), Simple Mail Transfer Protocol (SMTP), File Transfer Protocol (FTP), etc. Data exchanged over network 1120 can be represented using technologies and / or formats including binary image data (e.g., Portable Network Graphics (PNG)), Hypertext Markup Language (HTML), Extensible Markup Language (XML), etc. Furthermore, all or some links can be encrypted using traditional encryption technologies such as Secure Sockets Layer (SSL), Transport Layer Security (TLS), Virtual Private Network (VPN), Internet Protocol Security (IPsec), etc.
[0168] Mapping server 1125 may include a database storing virtual models describing multiple spaces, where a location in the virtual model corresponds to the current configuration of a local region of head-mounted device 1105. Mapping server 1125 receives information describing at least a portion of the local region and / or location information of the local region from head-mounted device 1105 via network 1120. Based on the received information and / or location information, mapping server 1125 determines the location in the virtual model associated with the local region of head-mounted device 1105. Mapping server 1125 determines (e.g., retrieves) one or more acoustic parameters associated with the local region, in part based on the determined location in the virtual model and any acoustic parameters associated with the determined location. Mapping server 1125 may transmit the location of the local region and any acoustic parameter values associated with the local region to head-mounted device 1105.
[0169] Additional configuration information
[0170] The foregoing description of embodiments has been presented for illustrative purposes; it is not intended to be exhaustive or to limit patent rights to the precise form disclosed. Those skilled in the art will understand that many modifications and variations are possible in light of the foregoing disclosure.
[0171] Some portions of this description describe embodiments based on algorithms and symbolic representations of operations on information. Those skilled in the art of data processing commonly use these algorithmic descriptions and representations to effectively communicate the substance of their work to others skilled in the art. While these operations are described functionally, computationally, or logically, they should be understood as being implemented by computer programs or equivalent circuits, microcode, etc. Furthermore, referring to these arrangements of operations as modules has sometimes proven convenient without loss of generality. The described operations and their associated modules can be embodied in software, firmware, hardware, or any combination thereof.
[0172] Any of the steps, operations, or processes described herein may be performed or implemented using one or more hardware or software modules, individually or in combination with other devices. In one embodiment, a software module is implemented using a computer program product comprising a computer-readable medium containing computer program code that can be executed by a computer processor for performing any or all of the described steps, operations, or processes.
[0173] The embodiments may also relate to means for performing the operations herein. Such means may be specifically configured for the desired purpose, and / or may comprise a general-purpose computing device selectively activated or reconfigured by a computer program stored in a computer. Such a computer program may be stored in a non-transitory, tangible, computer-readable storage medium or any type of medium suitable for storing electronic instructions, said medium being coupled to a computer system bus. Furthermore, any computing system mentioned in the specification may include a single processor, or may be an architecture employing a multiprocessor design to enhance computing power.
[0174] The embodiments may also relate to products produced through the computational processes described herein. Such products may include information obtained from the computational processes, wherein the information is stored on a non-transitory, tangible, computer-readable storage medium, and may include any embodiment of a computer program product or other combination of data described herein.
[0175] Finally, the language used in this specification has been chosen primarily for readability and instruction purposes and may not have been chosen to describe or limit the patent rights. Therefore, the scope of the patent rights intended hereby is not limited by this detailed description, but rather by any claims published on the application based thereon. Thus, the disclosure of embodiments is intended to illustrate, rather than limit, the scope of the patent rights, which are set forth in the appended claims.
Claims
1. A method for generating an equalization filter for a user, comprising: Receive one or more images, including the user's ear; Identify one or more features of the user's ear from the one or more images; In one or more images, the user is wearing a head-mounted device, and the one or more features are identified at least in part based on the position of the head-mounted device relative to the user's ears; The head-mounted device includes an eyeglass frame with two arms, each arm coupled to the eyeglass body, and the one or more images include at least a portion of one of the two arms, the at least a portion of one of the two arms including a transducer among a plurality of transducers; The model is provided with one or more features of the user's ear, and the model is configured to predict audio output at the user's ear based on the identified one or more features; The model is configured to determine the audio output response based at least in part on the position of one of the plurality of transducers relative to the user's ear; and An equalization filter is generated based on the ratio between the predicted frequency response and the target frequency response of the audio output at the user's ear. The equalization filter is configured to adjust one or more acoustic parameters of the audio content provided to the user when applied to the audio content provided to the user.
2. The method according to claim 1, further comprising: The generated equalization filter is provided to a head-mounted device, which is configured to use the equalization filter when providing audio content to the user.
3. The method according to claim 1, further comprising: The equalization filter is provided to the online system for storage in association with the user's user profile, wherein the equalization filter can be retrieved by one or more head-mounted devices associated with the user and authorized to access the user profile for use in providing content to the user.
4. The method according to claim 1, further comprising: The model is trained using multiple labeled images, each of which identifies features of an additional user's ear, for which the audio output at the ear is known.
5. The method according to claim 1, wherein, The one or more images are depth images captured using a depth camera component.
6. The method according to claim 1, wherein, One or more features identified are anthropometric features describing the size or shape of the user's ears.
7. The method according to claim 1, further comprising: The predicted frequency response of the audio output at the user's ear is compared with the measured audio output at the user's ear; as well as The model is updated based on the comparison.
8. The method according to claim 7, wherein, The measured audio output response was measured in the following manner: Providing audio content to the user via a head-mounted device; and The audio output at the user's ear is analyzed using one or more microphones placed near the user's ear.
9. A non-transitory computer-readable storage medium having instructions stored thereon, the instructions, when executed by a processor, causing the processor to perform steps including: Receive one or more images, including the user's ear; Identify one or more features of the user's ear based on the one or more images; in, The user in the one or more images is wearing a head-mounted device, and the one or more features are identified at least in part based on the position of the head-mounted device relative to the user's ears; The head-mounted device includes an eyeglass frame with two arms, each arm coupled to the eyeglass body, and the one or more images include at least a portion of one of the two arms, the at least a portion of one of the two arms including a transducer among a plurality of transducers; The one or more features are provided to the model, which is configured to determine the audio output at the user's ear based on the identified one or more features; The model is configured to determine the audio output at the user's ear based at least in part on the position of one of the plurality of transducers relative to the user's ear; and An equalization filter is generated based on the ratio between the predicted frequency response and the target frequency response of the audio output at the user's ear. The equalization filter is configured to adjust one or more acoustic parameters of the audio content provided to the user when applied to the audio content provided to the user.
10. The non-transitory computer-readable storage medium according to claim 9, wherein, When executed by a processor, the instruction also causes the processor to perform steps including the following: The model is trained using multiple labeled images, each of which identifies features of an additional user's ear, for which the audio output response is known.
11. The non-transitory computer-readable storage medium according to claim 9, wherein, When applied to audio content provided to the user, the equalization filter adjusts one or more acoustic parameters of the audio content for the user based on the predicted audio output at the user's ear.
12. The non-transitory computer-readable storage medium according to claim 9, wherein, The one or more images are depth images captured using a depth camera component.
13. The non-transitory computer-readable storage medium according to claim 9, wherein, One or more features identified are anthropometric features describing the size or shape of the user's ears.
Citation Information
Patent Citations
An arrangement for producing head related transfer function filters
CN108885690A
Head-related transfer function (HRTF) personalization based on captured images of user
US10341803B1
Determining individualized head-related transfer functions
US20120183161A1
Binaural audio signal processing method and apparatus reflecting personal characteristics
US20170272890A1
Method for generating a customized / personalized head related transfer function
US20190014431A1