Select spatial positioning for audio personalization

By iteratively updating the HRTF set, the problem of inaccurate HRTF calculation in the prior art is solved, and more accurate sound positioning is achieved in artificial reality systems.

CN114365510BActive Publication Date: 2025-05-09CTRL-LABS CORP
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202080061706.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2019-09-05
Filing Date
2020-08-08
Publication Date
2025-05-09
Estimated Expiration
2040-08-08

AI Technical Summary

Technical Problem

The prior art is difficult to accurately calculate the user's head-dependent transfer function (HRTF), resulting in inaccurate sound positioning in artificial reality systems.

Method used

Through an iterative process, the test positioning set is generated using the initial estimated HRTF set, and the HRTF set is updated and adjusted according to the user's response to the test sound until the threshold accuracy is reached.

Benefits of technology

More accurate HRTF calculations are realized, the accuracy of sound positioning in artificial reality systems is improved, and the dependence on external audio devices is reduced.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114365510B_ABST
    Figure CN114365510B_ABST
Patent Text Reader

Abstract

The audio system generates a customized head-related transfer function (HRTF) for the user. The audio system receives an initial estimated HRTF set. The initial HRTF set can be estimated using a trained machine learning and computer vision system and a picture of the user's ear. The audio system generates a test positioning set using the initial HRTF set. The audio system presents test sounds at each test position in the initial test positioning set using the initial HRTF set. The audio system monitors the user's response to the test sounds. The audio system uses the monitored response to generate a new estimated HRTF set and a new test positioning set. The process is repeated until a threshold accuracy is reached or until a set time period expires. The audio system presents audio content to the user using the customized HRTF.
Need to check novelty before this filing date? Find Prior Art

Description

Field of the Invention

[0001] The present disclosure relates generally to artificial reality systems, and more particularly to audio systems for artificial reality systems.

[0002] background

[0003] People hear sounds in different ways. For users of an audio system, such as for users of an audio system in an artificial reality system, sounds presented by the audio system may be heard in different ways by different users. The audio system may analyze images of the users, such as images of the users' ears, to calculate head-related transfer functions and customize the sounds presented to the users.

[0004] Overview

[0005] Therefore, the present invention relates to a method and a computer program product according to the attached claims. An audio system generates or receives an initial set of head-related transfer functions (HRTFs) for a user. The initial set of HRTFs can be estimated using trained machine learning and computer vision systems and images (e.g., of a user's ears, head, etc.). The audio system generates a set of test locations using the initial set of HRTFs. The audio system presents audio content at each location in the initial set of test locations using the initial set of HRTFs. The audio system monitors the user's response to the audio content presented for each test location in the set of test locations. The audio system uses the monitored responses to generate estimates of a new set of HRTFs and a new set of test locations. The process can be repeated until a threshold accuracy is reached, until a set time period expires, until a set number of iterations is reached, and so on.

[0006] In one aspect of the present invention, a method is disclosed, the method comprising: selecting a test positioning set for a user's estimated head-related transfer function (HRTF) set; generating a customized HRTF set for the user, the generation being based in part on applying an iterative process to the estimated HRTF set, the iterative process comprising: generating test sounds for the test positioning set, wherein the generated test sounds are spatialized using the user's estimated HRTF; determining an accuracy value of the user's estimated HRTF based in part on the user's response to the generated test sounds; updating the user's estimated HRTF based in part on the accuracy value; and adjusting the test positioning set based in part on the updated estimated HRTF; repeating the iterative process until a quality metric is met; and presenting content to the user using the customized HRTF set.

[0007] In some embodiments, a method may include selecting a test positioning set for a user's estimated head-related transfer function (HRTF) set. A customized HRTF set is generated for the user, the generation being based in part on applying an iterative process to the estimated HRTF set. The iterative process may be repeated until a quality metric is met. Content is presented to the user using the customized HRTF set. For example, the iterative process may include generating test sounds for the test positioning set. The generated test sounds are spatialized using the user's estimated HRTF. The iterative process may also include determining an accuracy value for the user's estimated HRTF based in part on the user's response to the generated test sounds and updating the user's estimated HRTF based in part on the accuracy value. The iterative process may also include adjusting the test positioning set based in part on the updated estimated HRTF.

[0008] In some embodiments, a method may include selecting a first test positioning set based on a first estimated head-related transfer function (HRTF) set for a user. Generating test sounds for the first test positioning set, and calculating an accuracy value for the first estimated HRTF set for the user based on a user response to the test sounds for the first test positioning set. Calculating a second HRTF set for the user based on the accuracy value for the first estimated HRTF set. Selecting a second test positioning set based on the second HRTF set, and generating test sounds for the second test positioning set.

[0009] In an embodiment of the method of the present invention, the set of estimated HRTFs may be generated by a HRTF machine learning and computer vision module.Alternatively or additionally, the first set of estimated HRTFs is generated based on data describing physical features of the user.

[0010] In an embodiment of the method of the invention, the data describing the user may comprise an image of the user's ear.

[0011] In an embodiment of the method according to the invention, generating the test sounds for the set of test locations may comprise sequentially generating the test sounds for each of the test locations.

[0012] In an embodiment of the method according to the invention, the accuracy value of the estimated HRTF may be calculated based on the positioning difference between the test positioning for the estimated HRTF and the gaze positioning of the user in response to the test sound for the test positioning.

[0013] In an embodiment of the method according to the invention, the test locations may be selected based on the rate of change of the HRTF within the region.

[0014] In one aspect of the present invention, the present invention also discloses a second method, which includes: selecting a first test positioning set based on a first estimated head-related transfer function (HRTF) set of a user; generating a test sound for the first test positioning set; calculating an accuracy value for the first estimated HRTF set for the user based on a user response to the test sound for the first test positioning set; calculating a second HRTF set for the user based on the accuracy value of the first estimated HRTF set; selecting a second test positioning set based on the second HRTF set; and generating a test sound for the second test positioning set. In this method, the first estimated HRTF set can be generated by an HRTF machine learning and computer vision module. Additionally or alternatively, the first estimated HRTF set can be generated based on data describing a physical feature of the user. The data describing the user may include an image of the user's ear.

[0015] In an embodiment of the second method according to the present invention, generating test sounds for a set of test locations may include sequentially generating test sounds for each of the test locations. In addition, an accuracy value for the estimated HRTF may be calculated based on a positioning difference between the test locations for the estimated HRTF and the gaze positioning of the user in response to the test sounds for the test locations.

[0016] In another embodiment of the second method according to the present invention, the test locations may be selected based on the rate of change of the HRTF within the region.

[0017] In one aspect of the present invention, the present invention also discloses a computer program product, which includes a non-transitory computer-readable storage medium, which contains computer program code, for selecting a test positioning set for a user's estimated head-related transfer function (HRTF) set; generating a customized HRTF set for the user, the generation of which is based in part on applying an iterative process to the estimated HRTF set, the iterative process comprising: generating test sounds for the test positioning set, wherein the generated test sounds are spatialized using the user's estimated HRTF; determining an accuracy value of the user's estimated HRTF based in part on the user's response to the generated test sounds; updating the user's estimated HRTF based in part on the accuracy value; and adjusting the test positioning set based in part on the updated estimated HRTF; repeating the iterative process until a quality metric is met; and presenting content to the user using the customized HRTF set.

[0018] In an embodiment of the computer program product according to the present invention, the estimated HRTF set may be generated by HRTF machine learning and computer vision modules.

[0019] In another embodiment of the computer program product according to the invention, the estimated HRTF set may be generated based on data describing the user's physical characteristics. In addition, the data describing the user may include an image of the user's ear.

[0020] In another embodiment of the computer program product according to the present invention, generating test sounds for the set of test locations includes sequentially generating test sounds for each of the test locations. In addition, an accuracy value for the estimated HRTF can be calculated based on the positioning difference between the test locations for the estimated HRTF and the gaze positioning of the user in response to the test sounds for the test locations. BRIEF DESCRIPTION OF THE DRAWINGS

[0022] Figure 1A is a perspective view of a headset implemented as an eyewear device according to one or more embodiments.

[0023] Figure 1B is a perspective view of a head mounted device implemented as a head mounted display according to one or more embodiments.

[0024] Figure 2 is a block diagram of an audio system according to one or more embodiments.

[0025] Figure 3 is a schematic diagram of a head mounted device and multiple test locations according to various embodiments.

[0026] Figure 4 is a flow chart illustrating a process for generating a customized HRTF in accordance with one or more embodiments.

[0027] Figure 5 A system including a head mounted device according to one or more embodiments.

[0028] The accompanying drawings depict several embodiments for purposes of illustration only. One skilled in the art will readily recognize from the following discussion that alternative embodiments of the structures and methods shown herein may be employed without departing from the principles described herein.

[0029] Detailed Description

[0030] The head mounted device includes an audio system that presents sounds to a user using a head related transfer function (HRTF) customized for the user. The audio system uses an iterative process to refine the HRTF for the user. The iterative process may include actively or passively obtaining user feedback on the sounds presented using the HRTF.

[0031] The audio system generates or receives an initial HRTF set for the user. The initial HRTF set can be estimated using a trained machine learning and computer vision system and a description of the user, which may include a picture of the user's head, torso or ear, or an outline or measurement of the ear called anthropometric features. The audio system generates a test positioning set using the initial HRTF set. The test positioning may be selected to be positioned at a location where the HRTF varies significantly according to the position, which may indicate a relatively high level of uncertainty in the HRTF in this area, or a small inaccuracy of the HRTF may cause a significant error in the sound presented to the user. The audio system presents audio content at each test positioning of the initial test positioning set using the initial HRTF set. The audio system monitors the user's response to the audio content presented for each test positioning in the test positioning set. The response may include gaze direction, head movement, verbal response, or any other suitable detectable response from the user. The audio system may detect the response using sensors such as cameras, motion sensors, and / or microphones. The audio system uses the monitored response to generate a new estimated HRTF set and a new test positioning set. This process can be repeated until a threshold accuracy is reached or until a set time period expires.

[0032] It may be difficult to obtain accurate HRTFs for all possible sound source localizations. However, by selecting test localizations in areas where it is a priori known that the HRTFs are very sensitive, the audio system can reduce the time and computational requirements to improve HRTF estimates for comparisons to localizations that may contain inaccurately estimated HRTFs. The disclosed audio system and HRTF customization process allow the audio system to accurately calculate HRTFs for a user without the need to actively measure the HRTFs using external audio equipment. In addition, the iterative process of refining the HRTFs based on active or passive user feedback allows the audio system to obtain more accurate HRTFs compared to systems that use a static set of estimated HRTFs.

[0033] Embodiments of the present invention may include an artificial reality system or be implemented in conjunction with an artificial reality system. Artificial reality is a form of reality that has been adjusted in some way before being presented to a user, which may include, for example, virtual reality (VR), augmented reality (AR), mixed reality (MR), hybrid reality, or some combination and / or derivative thereof. Artificial reality content may include fully generated content or content generated in combination with captured (e.g., real-world) content. Artificial reality content may include video, audio, tactile feedback, or some combination thereof, any of which may be presented in a single channel or in multiple channels (e.g., stereoscopic video that produces a three-dimensional effect to the viewer). In addition, in some embodiments, artificial reality may also be associated with applications, products, accessories, services, or some combination thereof, which are used to create content in artificial reality and / or be used in other ways in artificial reality. Artificial reality systems that provide artificial reality content can be implemented on a variety of platforms, including wearable devices (e.g., head-mounted devices) connected to a host computer system, standalone wearable devices (e.g., head-mounted devices), mobile devices or computing systems, or any other hardware platform capable of providing artificial reality content to one or more viewers.

[0034] Figure 1A is a perspective view of a head mounted device 100 implemented as an eyewear device according to one or more embodiments. In some embodiments, the eyewear device is a near eye display (NED). Typically, the head mounted device 100 can be worn on a user's face such that content (e.g., media content) is presented using a display component and / or an audio system. However, the head mounted device 100 can also be used such that media content is presented to the user in a different manner. Examples of media content presented by the head mounted device 100 include one or more images, videos, audio, or some combination thereof. The head mounted device 100 includes a frame and may include, among other components, a display component including one or more display elements 120, a depth camera assembly (DCA), an audio system, and a position sensor 190. Although Figure 1A The example positioning on the head mounted device 100 shows components of the head mounted device 100, but these components may be located elsewhere on the head mounted device 100, on a peripheral device that is paired with the head mounted device 100, or some combination of these two locations. Similarly, components on the head mounted device 100 may be located on a smaller device than on the head mounted device 100. Figure 1A More or less than shown.

[0035] The frame 110 holds the other components of the head mounted device 100. The frame 110 includes a front portion that holds one or more display elements 120 and an end piece (e.g., temple) that is attached to the user's head. The front portion of the frame 110 bridges the top of the user's nose. The length of the end piece can be adjustable (e.g., adjustable temple length) to fit different users. The end piece can also include a portion that curls behind the user's ear (e.g., temple tip, ear hook).

[0036] One or more display elements 120 provide light to a user wearing the head mounted device 100. As shown, the head mounted device includes a display element 120 for each eye of the user. In some embodiments, the display element 120 generates image light provided to the eyebox of the head mounted device 100. The eyebox is the spatial location occupied by the eye when the user wears the head mounted device 100. For example, the display element 120 can be a waveguide display. The waveguide display includes a light source (e.g., a two-dimensional light source, one or more linear light sources, one or more point light sources, etc.) and one or more waveguides. The light from the light source is coupled inwardly into one or more waveguides, which output light in such a way that there is a pupil replication in the eyebox of the head mounted device 100. The inward coupling and / or outward coupling of light from one or more waveguides can be accomplished using one or more diffraction gratings. In some embodiments, the waveguide display includes a scanning element (e.g., a waveguide, a reflector, etc.) that scans the light from the light source when the light from the light source is coupled inwardly into one or more waveguides. Note that in some embodiments, one or both of the display elements 120 are opaque and do not transmit light from a local area around the head mounted device 100. The local area is the area around the head mounted device 100. For example, the local area may be a room where a user wearing the head mounted device 100 is located, or the user wearing the head mounted device 100 may be outside and the local area is an external area. In this case, the head mounted device 100 generates VR content. Alternatively, in some embodiments, one or both of the display elements 120 are at least partially transparent so that light from the local area can be combined with light from one or more display elements to produce AR and / or MR content.

[0037] In some embodiments, display elements 120 do not generate image light, but rather lenses transmit light from localized areas to the viewport. For example, one or both of display elements 120 may be non-corrective lenses (over-the-counter), or prescription lenses (e.g., single vision lenses, bifocal and trifocal lenses, or progressive lenses) to help correct for imperfections in the user's vision. In some embodiments, display elements 120 may be polarized and / or tinted to protect the user's eyes from sunlight.

[0038] It should be noted that in some embodiments, the display element 120 may include an additional optical block (not shown). The optical block may include one or more optical elements (e.g., lenses, Fresnel lenses, etc.) that direct light from the display element 120 to the viewing window. The optical block may, for example, correct aberrations in some or all image content, magnify some or all images, or some combination thereof.

[0039] The DCA determines depth information for a portion of a local area around the head mounted device 100. The DCA includes one or more imaging devices 130 and a DCA controller (not shown). Figure 1A ), and may also include an illuminator 140. In some embodiments, the illuminator 140 illuminates a portion of the local area with light. The light may be, for example, structured light (e.g., dot patterns, stripes, etc.) in infrared (IR), IR flash for time of flight, etc. In some embodiments, one or more imaging devices 130 capture an image of a portion of the local area including light from the illuminator 140. As shown in the figure, Figure 1A A single illuminator 140 and two imaging devices 130 are shown. In an alternative embodiment, there is no illuminator 140 and at least two imaging devices 130.

[0040] The DCA controller calculates depth information for a portion of the local area using the captured images and one or more depth determination techniques. The depth determination techniques may be, for example, direct time-of-flight (ToF) depth sensing, indirect ToF depth sensing, structured light, passive stereo analysis, active stereo analysis (using texture added to the scene by light from illuminator 140), some other technique for determining the depth of a scene, or some combination thereof.

[0041] The audio system provides audio content. The audio system includes a transducer array, a sensor array, and an audio controller 150. However, in other embodiments, the audio system may include different and / or additional components. Similarly, in some cases, the functions described with reference to the components of the audio system may be distributed between the components in a manner different from that described herein. For example, some or all functions of the controller may be performed by a remote server.

[0042] The transducer array presents sound to the user. The transducer array includes multiple transducers. The transducer can be a speaker 160 or a tissue transducer 170 (e.g., a bone conduction transducer or a cartilage conduction transducer). Although the speaker 160 is shown outside the frame 110, the speaker 160 can be enclosed in the frame 110. In some embodiments, instead of a separate speaker for each ear, the head mounted device 100 includes a speaker array that includes multiple speakers integrated into the frame 110 to improve the directionality of the presented audio content. The tissue transducer 170 is coupled to the user's head and directly vibrates the user's tissue (e.g., bone or cartilage) to generate sound. The number and / or positioning of the transducers can be different from the head mounted device 100. Figure 1A Different than shown.

[0043] The sensor array detects sounds in a local area of ​​the head mounted device 100. The sensor array includes a plurality of acoustic sensors 180. The acoustic sensors 180 capture sounds emitted from one or more sound sources in a local area (e.g., a room). Each acoustic sensor is configured to detect sounds and convert the detected sounds into an electronic format (analog or digital). The acoustic sensor 180 may be a sound wave sensor, a microphone, a sound transducer, or a similar sensor suitable for detecting sounds.

[0044] In some embodiments, one or more acoustic sensors 180 can be placed in the ear canal of each ear (e.g., acting as binaural microphones). In some embodiments, acoustic sensors 180 can be placed on an exterior surface of head mounted device 100, on an interior surface of head mounted device 100, separate from head mounted device 100 (e.g., as part of some other device), or some combination thereof. The number and / or positioning of acoustic sensors 180 can vary depending on the device. Figure 1A For example, the number of acoustic detection positionings can be increased to increase the amount of audio information collected and the sensitivity and / or accuracy of the information. The acoustic detection positionings can be oriented so that the microphones can detect sounds in a wide range of directions around the user wearing the head mounted device 100.

[0045] The audio controller 150 processes information from the sensor array describing the sounds detected by the sensor array. The audio controller 150 may include a processor and a computer-readable storage medium. The audio controller 150 may be configured to generate a direction of arrival (DOA) estimate, generate an acoustic transfer function (e.g., an array transfer function and / or a head-related transfer function), track the location of a sound source, form a beam in the direction of a sound source, classify a sound source, generate a sound filter for the speaker 160, or some combination thereof.

[0046] The position sensor 190 generates one or more measurement signals in response to the movement of the head mounted device 100. The position sensor 190 may be located on a portion of the frame 110 of the head mounted device 100. The position sensor 190 may include an inertial measurement unit (IMU). Examples of the position sensor 190 include: one or more accelerometers, one or more gyroscopes, one or more magnetometers, another suitable type of sensor that detects movement, a type of sensor used for error correction of the IMU, or some combination thereof. The position sensor 190 may be located outside the IMU, inside the IMU, or some combination thereof.

[0047] In some embodiments, the head mounted device 100 may provide simultaneous localization and mapping (SLAM) of the position of the head mounted device 100 and updates of local area models. For example, the head mounted device 100 may include a passive camera assembly (PCA) that generates color image data. The PCA may include one or more RGB cameras that are used to capture images of some or all local areas. In some embodiments, some or all of the imaging devices 130 of the DCA may also be used as PCAs. The images captured by the PCA and the depth information determined by the DCA may be used to determine parameters of the local area, generate a model of the local area, update the model of the local area, or some combination thereof. In addition, the position sensor 190 tracks the position (e.g., positioning and posture) of the head mounted device 100 in the room.

[0048] The head mounted device 100 includes an eye tracking unit 195. The eye tracking unit 195 may include one or more cameras that capture images of the user's eyes. The eye tracking unit 195 may also include one or more illuminators that illuminate the user's eyes. The eye tracking unit 195 estimates the angular orientation of one or both eyes of the user. In some embodiments, the eye tracking unit 195 may detect distortion in the illumination pattern projected by the illuminator to determine the angular orientation of the user's eyes. The orientation of the eyes corresponds to the direction in which the user is staring within the head mounted device 100. The orientation of the user's eyes may be in the direction of the foveal axis, which is the axis between the fovea (the area on the retina of the eye where the most photoreceptor cells are concentrated) and the center of the pupil of the eye. Typically, when the user's eyes are fixed on a point, the foveal axis of the user's eyes intersects the point. The pupil axis is another axis of the eye, which is defined as an axis perpendicular to the corneal surface that passes through the center of the pupil. Generally, the pupil axis is not directly aligned with the foveal axis. The two axes intersect at the center of the pupil, but the orientation of the foveal axis is offset from the pupil axis by approximately -1° to 8° laterally and approximately ±4° vertically. Because the foveal axis is defined based on the fovea, which is located at the back of the eye, in some eye tracking embodiments, the foveal axis may be difficult or impossible to detect directly. Therefore, in some embodiments, the orientation of the pupil axis is detected, and the foveal axis is estimated based on the detected pupil axis. However, in some embodiments, the orientation of the pupil axis can be used to estimate the angular orientation of one or both eyes of the user without adjusting for foveal axis differences.

[0049] Typically, the movement of the eye corresponds not only to the angular rotation of the eye, but also to the translation of the eye, the change in the torsion of the eye, and / or the change in the shape of the eye. The eye tracking unit 195 can also detect the translation of the eye: that is, the change in the position of the eye relative to the eye socket. In some embodiments, the translation of the eye is not directly detected, but is approximated based on a mapping from the detected angular orientation. The eye translation corresponding to the change in the position of the eye relative to the detection component of the eye tracking unit can also be detected. This type of translation may occur, for example, due to the movement of the position of the head-mounted device 100 on the user's head. The eye tracking unit 195 can also detect the torsion of the eye, that is, the rotation of the eye around the pupil axis. The eye tracking unit 195 can use the detected eye torsion to estimate the orientation of the foveal axis relative to the pupil axis. The eye tracking unit 195 can also track changes in the shape of the eye, which can be approximated as a skew or scale linear transformation or a distortion deformation (for example, due to torsion deformation). The eye tracking unit 195 can estimate the foveal axis based on some combination of the angular orientation of the pupil axis, the translation of the eye, the torsion of the eye, and the current shape of the eye.

[0050] In some embodiments, the eye tracking unit 195 may include at least one emitter that projects a structured light pattern on all or part of the eye. The pattern is then projected onto the shape of the eye, which may produce perceptible distortion in the structured light pattern when viewed from an offset angle. The eye tracking unit 195 may also include at least one camera that detects distortion, if any, of the light pattern projected onto the eye. A camera oriented on an axis different from the emitter captures the illumination pattern on the eye. This process is referred to herein as "scanning" the eye. By detecting deformations of the illumination pattern on the surface of the eye, the eye tracking unit 195 can determine the shape of the scanned portion of the eye. Thus, the captured distorted light pattern indicates the 3D shape of the illuminated portion of the eye. By deriving the 3D shape of the portion of the eye illuminated by the emitter, the orientation of the eye can be derived. The eye tracking unit can also estimate the pupil axis, the translation of the eye, the torsion of the eye, and the current shape of the eye based on the image of the illumination pattern captured by the camera.

[0051] In other embodiments, any suitable type of eye tracking system may be used. For example, the eye tracking unit 195 may capture images of the eye, capture stereo images of the eye, may utilize a ring of LEDs around the eye that illuminate in sequence and determine eye orientation based on reflections from the LEDs, may utilize time of flight measurements, etc.

[0052] Because the orientation can be determined for both eyes of the user, the eye tracking unit 195 is able to determine where the user is looking. The head mounted device 100 can use the orientation of the eyes to, for example, determine the user's interpupillary distance (IPD), determine the gaze direction, introduce depth cues (e.g., blur images outside the user's primary line of sight), collect heuristics about user interactions in VR media (e.g., time spent on any particular subject, object, or frame based on exposure to stimuli), some other functionality based in part on the orientation of at least one of the user's eyes, or some combination thereof. Determining the direction of the user's gaze may include determining a point of convergence based on the determined orientation of the user's left and right eyes. The point of convergence may be the point where the two foveal axes of the user's eyes intersect (or the closest point between the two axes). The direction of the user's gaze may be the direction of a line passing through the point of convergence and through the midpoint of the pupils of the user's eyes. The following is combined with Figure 5 Additional details regarding components of head mounted device 100 are discussed.

[0053] The audio system is calibrated to customize the HRTFs for the user. The audio system synthesizes sounds at the test location using an initial set of estimated HRTFs. The eye tracking unit 195 detects the gaze location of the user's eyes in response to the synthesized sounds. The audio system measures the accuracy of the HRTFs used to synthesize the sounds based on the user response (such as the difference between the gaze location and the test location). The audio system calculates a new set of HRTFs based on the accuracy of the HRTFs. The audio system adjusts the test location and calculates the accuracy of the new HRTFs at the adjusted test location. Reference Figure 2-Figure 4 The HRTF customization process is further described.

[0054] Figure 1B 1 is a perspective view of a head mounted device 105 implemented as an HMD according to one or more embodiments. In embodiments describing an AR system and / or an MR system, a portion of the front side of the HMD is at least partially transparent in the visible light band (approximately 380 nm to 750 nm), and a portion of the HMD between the front side of the HMD and the user's eyes is at least partially transparent (e.g., a partially transparent electronic display). The HMD includes a front rigid body 115 and a strap 175. The head mounted device 105 includes many of the same components as described above with reference to FIG. Figure 1A The same components described above are modified to integrate with the HMD form factor. For example, the HMD includes a display assembly, a DCA, an audio system, and a position sensor 190. Figure 1B Shown are an illuminator 140, a plurality of speakers 160, a plurality of imaging devices 130, a plurality of acoustic sensors 180, and a position sensor 190. The speakers 160 may be located in various locations, such as coupled to the band 175 (as shown), coupled to the front rigid body 115, or may be configured to be inserted into the ear canal of a user.

[0055] Figure 2 is a block diagram of an audio system 200 according to one or more embodiments. Figure 1A and / or Figure 1B The audio system in may be an embodiment of the audio system 200. The audio system 200 generates one or more acoustic transfer functions for a user. The audio system 200 may then use the one or more acoustic transfer functions to generate audio content for the user. Figure 2 In the embodiment of the present invention, the audio system 200 includes a transducer array 210, a sensor array 220, and an audio controller 230. Some embodiments of the audio system 200 have different components than those described here. Similarly, in some cases, functions can be distributed between components in a manner different from that described here.

[0056] The transducer array 210 is configured to present audio content. The transducer array 210 includes a plurality of transducers. A transducer is a device that provides audio content. The transducer may be, for example, a speaker (e.g., speaker 160), a tissue transducer (e.g., tissue transducer 170), some other device that provides audio content, or some combination thereof. The tissue transducer may be configured to be used as a bone conduction transducer or a cartilage conduction transducer. The transducer array 210 may present audio content via air conduction (e.g., via one or more speakers), via bone conduction (via one or more bone conduction transducers), via a cartilage conduction audio system (via one or more cartilage conduction transducers), or some combination thereof. In some embodiments, the transducer array 210 may include one or more transducers to cover different parts of a frequency range. For example, a piezoelectric transducer may be used to cover a first part of a frequency range, while a moving coil transducer may be used to cover a second part of a frequency range.

[0057] The bone conduction transducer generates sound pressure waves by vibrating the bones / tissue of the user's head. The bone conduction transducer can be coupled to a portion of the head mounted device and can be configured to couple to a portion of the user's skull behind the auricle. The bone conduction transducer receives vibration instructions from the audio controller 230 and vibrates a portion of the user's skull based on the received instructions. The vibrations from the bone conduction transducer generate tissue-propagated sound pressure waves that bypass the eardrum and propagate toward the user's cochlea.

[0058] The bone conduction transducer generates sound pressure waves by vibrating one or more portions of the ear cartilage of the user's ear. The bone conduction transducer can be coupled to a portion of the head-mounted device and can be configured to couple to one or more portions of the ear cartilage of the ear. For example, the bone conduction transducer can be coupled to the back of the auricle of the user's ear. The bone conduction transducer can be located anywhere along the ear cartilage around the outer ear (e.g., the auricle, the tragus, some other portion of the ear cartilage, or some combination thereof). Vibrating one or more portions of the ear cartilage can generate: air-borne sound pressure waves outside the ear canal; sound pressure waves generated by tissue causing certain portions of the ear canal to vibrate, thereby generating air-borne sound pressure waves within the ear canal; or some combination thereof. The generated air-borne sound pressure waves propagate along the ear canal toward the eardrum.

[0059] The transducer array 210 generates audio content according to instructions from the audio controller 230. In some embodiments, the audio content is spatialized. Spatialized audio content is audio content that sounds like it originates from a specific direction and / or target area (e.g., objects and / or virtual objects in a local area). For example, the spatialized audio content can make the sound sound like a virtual singer from across the room from the user of the audio system 200. The transducer array 210 can be coupled to a wearable device (e.g., the head mounted device 100 or the head mounted device 105). In an alternative embodiment, the transducer array 210 can be a plurality of speakers that are separate from the wearable device (e.g., coupled to an external console). The transducer array 210 generates spatialized sounds emitted from various test positions.

[0060] The sensor array 220 detects sounds in a local area around the sensor array 220. The sensor array 220 may include a plurality of acoustic sensors, each of which detects air pressure changes of sound waves and converts the detected sound into an electronic format (analog or digital). The plurality of acoustic sensors may be located on a head-mounted device (e.g., head-mounted device 100 and / or head-mounted device 105), a user (e.g., in the ear canal of the user), a neckband, or some combination thereof. The acoustic sensor may be, for example, a microphone, a vibration sensor, an accelerometer, or any combination thereof. In some embodiments, the sensor array 220 is configured to monitor the audio content generated by the transducer array 210 using at least some of the plurality of acoustic sensors. Increasing the number of sensors may improve the accuracy of information (e.g., directionality) describing the sound field generated by the transducer array 210 and / or the sound from the local area.

[0061] The audio controller 230 controls the operation of the audio system 200. Figure 2 In an embodiment of the present invention, the audio controller 230 includes a data storage 235, a DOA estimation module 240, a transfer function module 250, a tracking module 260, a beamforming module 270, a sound filter module 280, and a HRTF customization module 290. In some embodiments, the audio controller 230 can be located inside the head mounted device. Some embodiments of the audio controller 230 have different components than described herein. Similarly, the functions can be distributed among the components in a manner different from that described herein. For example, some functions of the controller can be performed outside the head mounted device.

[0062] The data storage 235 stores data for use by the audio system 200. The data in the data storage 235 may include sounds recorded in a local area of ​​the audio system 200, audio content, head-related transfer functions (HRTFs), transfer functions of one or more sensors, array transfer functions (ATFs) of one or more acoustic sensors, sound source localization, virtual models of the local area, direction of arrival estimates, sound filters, and other data related to the use of the audio system 200, or any combination thereof.

[0063] The data storage 235 includes an initial set of estimated HRTFs. The initial set of estimated HRTFs can be generated based on data describing the user. The data describing the user may include a description of the appearance features of the user's ears, which are called anthropometric features, an image of the user's head or torso, an image of the user's ears, a video of the user, etc. In some embodiments, the data describing the user may include an image of the user wearing a head-mounted device. The data describing the user can be input into an HRTF machine learning and computer vision module for calculating HRTF. For example, the data storage 235 can provide the size of the user's ears to the HRTF machine learning and computer vision module. The HRTF machine learning and computer vision module may be located on an external server, or the HRTF machine learning and computer vision module may be a component of the HRTF customization module 290. In some cases, an initial set of estimated HRTFs is generated on an external server, and subsequent iterative refinement of the HRTF is performed by the audio system 200.

[0064] The DOA estimation module 240 is configured to locate sound sources in a local area based in part on information from the sensor array 220. Localization is the process of determining the location of a sound source relative to the user of the audio system 200. The DOA estimation module 240 performs a DOA analysis to locate one or more sound sources in a local area. The DOA analysis may include analyzing the intensity, spectrum, and / or arrival time of each sound at the sensor array 220 to determine the direction from which the sound originated. In some cases, the DOA analysis may include any suitable algorithm for analyzing the ambient acoustic environment in which the audio system 200 is located.

[0065] For example, the DOA analysis may be designed to receive an input signal from the sensor array 220 and apply a digital signal processing algorithm to the input signal to estimate the direction of arrival. These algorithms may include, for example, a delay and sum algorithm in which the input signal is sampled and the final weighted and delayed versions of the sampled signals are averaged together to determine the DOA. A least mean square (LMS) algorithm may also be implemented to create an adaptive filter. The adaptive filter may then be used, for example, to identify differences in signal strength or differences in arrival time. These differences may then be used to estimate the DOA. In another embodiment, the DOA may be determined by converting the input signal into the frequency domain and selecting a specific bin in the time-frequency (TF) domain to be processed. Each selected TF bin may be processed to determine whether the bin includes a portion of the audio spectrum having a direct path audio signal. Those bins having a portion of the direct path signal may then be analyzed to identify the angle at which the sensor array 220 receives the direct path audio signal. The determined angle may then be used to identify the DOA of the received input signal. Other algorithms not listed above may also be used alone or in combination with the above algorithms to determine the DOA.

[0066] In some embodiments, the DOA estimation module 240 can also determine the DOA relative to the absolute position of the audio system 200 within the local area. The position of the sensor array 220 can be received from an external system (e.g., some other component of the head mounted device, an artificial reality console, a mapping server, a position sensor (e.g., the position sensor 190), etc.). The external system can create a virtual model of the local area, where the local area and the position of the audio system 200 are mapped. The received position information may include the location and / or orientation of some or all of the audio system 200 (e.g., the sensor array 220). The DOA estimation module 240 can update the estimated DOA based on the received position information.

[0067] The transfer function module 250 is configured to generate one or more acoustic transfer functions. In general, a transfer function is a mathematical function that gives a corresponding output value for each possible input value. Based on the parameters of the detected sound, the transfer function module 250 generates one or more acoustic transfer functions associated with the audio system. The acoustic transfer function can be an array transfer function (ATF), HRTF, other types of acoustic transfer functions, or some combination thereof. The ATF characterizes how a microphone receives sound from a point in space.

[0068] The ATF includes multiple transfer functions that characterize the relationship between a sound source and the corresponding sound received by the acoustic sensors in the sensor array 220. Therefore, for a sound source, each acoustic sensor in the sensor array 220 has a corresponding transfer function. This set of transfer functions is collectively referred to as the ATF. Therefore, for each sound source, there is a corresponding ATF. Note that the sound source can be, for example, someone or something that produces sound in a local area, a user, or one or more transducers of the transducer array 210. The ATF positioned relative to a specific sound source of the sensor array 220 may vary from user to user, because the human anatomical structure (e.g., ear shape, shoulders, etc.) affects the sound when the sound reaches the human ear. Therefore, the ATF of the sensor array 220 is personalized for each user of the audio system 200.

[0069] In some embodiments, the transfer function module 250 determines one or more HRTFs for the user of the audio system 200. HRTFs characterize how an ear receives sound from a point in space. HRTFs positioned relative to a person's specific source are unique to each ear of the person (and unique to a person) because the person's anatomical structure (e.g., ear shape, shoulders, etc.) affects the sound when it reaches the person's ears. In some embodiments, the transfer function module 250 can use a calibration process to determine the HRTF for the user. In some embodiments, the transfer function module 250 can provide information about the user to the remote system. The remote system can use, for example, machine learning and computer vision to determine an initial set of estimated HRTFs customized for the user, and provide the customized set of HRTFs to the audio system 200.

[0070] The tracking module 260 is configured to track the positioning of one or more sound sources. The tracking module 260 can compare current DOA estimates and compare them with a stored history of previous DOA estimates. In some embodiments, the audio system 200 can recalculate the DOA estimate periodically, for example once per second, or once per millisecond. The tracking module can compare the current DOA estimate with the previous DOA estimate, and in response to a change in the DOA estimate of the sound source, the tracking module 260 can determine that the sound source has moved. In some embodiments, the tracking module 260 can detect changes in positioning based on visual information received from a head-mounted device or some other external source. The tracking module 260 can track the movement of one or more sound sources over time. The tracking module 260 can store values ​​about the number of sound sources and the positioning of each sound source at each point in time. In response to changes in the number of sound sources or the values ​​of the positioning, the tracking module 260 can determine that the sound source has moved. The tracking module 260 can calculate an estimate of the positioning variance. The positioning variance can be used as a confidence level for each determination of a motion change.

[0071] The beamforming module 270 is configured to process one or more ATFs to selectively emphasize the sound from a sound source within a certain area while weakening the sound from other areas. When analyzing the sound detected by the sensor array 220, the beamforming module 270 can combine information from different acoustic sensors to emphasize the relevant sound from a specific area of ​​the local area while weakening the sound from outside the area. The beamforming module 270 can isolate the audio signal associated with the sound from a specific sound source from other sound sources in the local area based on different DOA estimates, such as from the DOA estimation module 240 and the tracking module 260. The beamforming module 270 can therefore selectively analyze discrete sound sources in the local area. In some embodiments, the beamforming module 270 can enhance the signal from the sound source. For example, the beamforming module 270 can apply a sound filter that eliminates the signal above, below, or between specific frequencies. Signal enhancement is used to enhance the sound associated with a given identified sound source relative to other sounds detected by the sensor array 220.

[0072] The sound filter module 280 determines the sound filter for the transducer array 210. In some embodiments, the sound filter causes the audio content to be spatialized so that the audio content sounds like it originates from the target area. The sound filter module 280 can generate the sound filter using HRTF and / or acoustic parameters. The acoustic parameters describe the acoustic properties of the local area. The acoustic parameters may include, for example, reverberation time, reverberation level, room impulse response, etc. In some embodiments, the sound filter module 280 calculates one or more acoustic parameters. In some embodiments, the sound filter module 280 requests the acoustic parameters from the mapping server (e.g., as described below with reference to Figure 5 described above).

[0073] The sound filter module 280 provides a sound filter to the transducer array 210. In some embodiments, the sound filter may cause positive or negative amplification of the sound depending on the frequency.

[0074] The HRTF customization module 290 generates a customized HRTF for the user. The HRTF customization module 290 selects a set of test locations for the initial estimated HRTF set. The HRTF customization module 290 can select any suitable number of test locations, such as between 25-50 test locations, or between 1-100 test locations. The test locations can be located in any direction relative to the user, such as in front of, behind, above, below, to the left, or to the right of the user. The test locations can be located at different distances from the user.

[0075] In some embodiments, the test locations may be selected based on the rate of change of the estimated HRTF according to the angle relative to the user. For example, in areas where the values ​​of the HRTFs differ greatly but are close to each other in space, the transfer function module 250 may relatively select more test locations, so that the density of the test locations is based in part on the rate of change of the values ​​of the HRTFs in a given area. The rate of change of the HRTF values ​​may be measured using a distance calculation algorithm. Such algorithms may utilize statistical learning, high-dimensional embedding, machine learning and computer vision, parametric modeling, dimensionality reduction or manifold learning techniques. For example, machine learning and computer vision or dimensionality reduction models may be trained to calculate the distance between HRTFs, or a set of prescribed rules about the structure of the HRTF signal may be obtained (enlisted), and the algorithm may use these rules to calculate the distance between two HRTFs. In some embodiments, the following methods may be used to estimate the rate of change of HRTF values: spectral difference estimation (SDE), which is the average of the spectral differences of the HRTFs for all audible frequencies; weighted SDE, in which the weights of certain frequencies are higher or lower than the average based on perceptual importance or audibility; or geometric or other distance metrics between prominent features in the HRTF. The trained model or rule set may be derived in advance in an independent study of the HRTF. The resulting distance may be a scalar or a collection of scalar profiles, and some combination or transformation of these profiles may define whether a given area would benefit from higher test location sampling.

[0076] The test positions are provided to the transducer array 210 to generate spatialized sounds emanating from the test positions. The HRTF customization module 290 instructs the transducer array 210 to generate the spatialized sounds.

[0077] The HRTF customization module 290 obtains perceptual feedback from the user for the sounds synthesized at each test location. The sounds may be synthesized sequentially such that a first sound is synthesized for a first test location, and a second sound is synthesized for a second test location after perceptual feedback is received, until responses are received for all test locations. The perceptual feedback includes a detected response from the user to the synthesized sound for each test location. The perceptual feedback may be captured by one or more sensors on the head mounted device, such as by an eye tracking module, by tactile feedback from a glove, or from a microphone. In some embodiments, the perceptual feedback may be captured by an external sensor, such as by a reference sensor. Figure 5The tracking module 560 described above can capture the perceived localization of the synthesized sound. In some embodiments, the perceived feedback can include a gaze direction of the user's eyes, indicating that the user perceives the sound emanating from the gaze direction. The perceived feedback can include a verbal response from the user, such as "front," "back," "left," or "right." The perceived feedback can include movement of the user, such as the user turning their head or pointing a finger in a direction. The perceived feedback can include selecting one or more answers from a list of options or answers or entities.

[0078] In some embodiments, sensory feedback can be obtained during an active calibration process. For example, the head mounted device can notify the user that the HRTF is being calibrated, and the head mounted device can provide the user with audio or visual instructions to look in the direction of the perceived sound source.

[0079] In some embodiments, sensory feedback can be obtained during a passive calibration process, where the user may not be aware that HRTF calibration is being performed. For example, the user may interact with the head mounted device, such as participating in a virtual reality game, and the HRTF customization module 290 may monitor the user's response to the sounds synthesized at the test location during the virtual reality game.

[0080] The HRTF customization module 290 compares the perceived feedback for each test location with the expected location of each test location, and determines the accuracy value of the HRTF at each test location. For example, the HRTF customization module 290 can assign a scalar accuracy value between 1-10, 10 indicating a highly accurate HRTF and 1 indicating a highly inaccurate HRTF. The accuracy can be determined based on the positioning difference between the test location and the perceived sound source location. For example, if the difference between the test location and the perceived sound source location is less than 1 degree from the user's perspective, the HRTF customization module 290 can assign an accuracy value of 10 to the HRTF at the test location. If the difference between the test location and the perceived sound source location is greater than 90 degrees from the user's perspective, the HRTF customization module 290 can assign an accuracy value of 1 to the HRTF at the test location. In some embodiments, the accuracy can be based on a radial difference, which is the difference between the perceived distance from the user to the test location and the expected difference between the user and the test location. In some embodiments, the accuracy can be based on a combination of radial differences and angular distances. In some embodiments, the accuracy calculation may not compare a given test response with any correct response, or the correct response may not exist. Test responses can be used to directly calculate the accuracy of the measurements without having access to any true responses.

[0081] The HRTF customization module 290 can transmit the accuracy value of the HRTF to the HRTF machine learning and computer vision module. In some embodiments, the HRTF machine learning and computer vision module can be a component of the HRTF customization module 290, and / or can be a component of the head mounted device. However, in some embodiments, the HRTF machine learning and computer vision module can be located on an external server, or can be located on a console that communicates with the head mounted device.

[0082] The HRTF machine learning and computer vision module applies machine learning and computer vision techniques to generate an HRTF model, which, when applied to the data describing the user, outputs an estimated HRTF about the positioning relative to the user. The HRTF machine learning and computer vision module can input the accuracy value into the HRTF model and update the estimated HRTF for the user. The initial positioning set can be predetermined based on ablation studies or independent studies on HRTF. This can be a 1-50 initial positioning set. The number of initial positionings can be fixed in advance. However, the specific positioning can depend on the user and can be calculated based on user data, which includes anthropometric features of the ear, one or more images or videos of the left and right ears, or the size of the head or torso. The set of initial positioning can be calculated based on an initial estimate of the HRTF from the user data or can even be fixed before acquiring the user data.

[0083] As part of generating the HRTF model, the HRTF machine learning and computer vision modules form a HRTF training set by identifying a positive training set of HRTFs that have been determined to be accurate, and in some embodiments, a negative training item set of HRTFs that have been determined to be inaccurate is formed. The training set of HRTFs can be obtained via carefully designed acoustic measurements in an anechoic chamber or a non-anechoic chamber. For each user participating in this training user study, an acoustic measurement of HRTF can be obtained by placing a microphone in the left and right ears and generating sounds at different spatial locations around the user. The signal captured by the microphone is then processed using acoustic signal processing techniques to obtain the HRTF of each participant. Such a measured HRTF set can be a training set. In some embodiments, the training set can be a simulated HRTF. For each participant in such a study, very high-resolution head, torso, and ear scans were obtained. These scans can be captured by a widely used 3d grid capture device, and the resulting scans will be processed by computer graphics and computer vision methods. The resulting head, torso, and ear processed scans can then be used to simulate HRTF using Monte Carlo methods or boundary elements or finite difference time domain or finite volume simulations.

[0084] The HRTF machine learning and computer vision module uses supervised machine learning and computer vision to train the HRTF model, taking the feature vectors of the positive training set and the negative training set as input. Different machine learning and computer vision techniques - such as linear support vector machines (linear SVM), boosting of other algorithms (e.g., AdaBoost), neural networks, logistic regression, naive Bayes, memory-based learning, random forests, bagged trees, decision trees, boosted trees, boosted stumps, nearest neighbors, k nearest neighbors, kernel machines, probability models, conditional random fields, Markov random fields, manifold learning, generalized linear models, generalized index models, kernel regression, or Bayesian regression - can be used in different embodiments. The HRTF machine learning and computer vision models, when applied to the data describing the user, output an estimated HRTF set for the user. In some embodiments, the machine learning and computer vision models, when applied to the data describing the user, output a scalar set or profile set that can be used to estimate a new HRTF.

[0085] The HRTF machine learning and computer vision module extracts feature values ​​from the HRTF of the training set, which are variables that are considered to be potentially related to whether the HRTF is accurate. Specifically, the feature values ​​extracted by the HRTF machine learning and computer vision module include sound source localization, frequency, amplitude, specific statistical irregularities of the signal defined as peaks or valleys in the signal structure, etc. The ordered list of features of the HRTF is referred to as the feature vector of the HRTF in this article. In one embodiment, the HRTF machine learning and computer vision module applies dimensionality reduction (for example, via linear discriminant analysis (LDA), principal component analysis (PCA), perceptual feature analysis, etc.) to reduce the amount of data in the feature vector of the HRTF to a smaller, more representative data set. In some embodiments, the HRTF machine learning and computer vision module uses deep representative learning to extract the necessary data for the feature vector of the HRTF.

[0086] The HRTF machine learning and computer vision module provides an updated HRTF to the audio system 200, and the audio system 200 can test the accuracy of the updated HRTF. The audio system 200 can iteratively update the HRTF for the user until the quality metric is met. The iterative process may include selecting a test location, generating a test sound at the test location, receiving feedback for the test sound, generating an updated HRTF, and selecting a new test location based on the updated HRTF. For example, the audio system 200 can update the HRTF until all test locations obtain an accuracy value of at least 9 (on a scale of 1-10), or until the average accuracy value is at least 9. In some embodiments, the audio system 200 can iteratively update the HRTF for a set number of iterations or for a set time period (such as 10 minutes), and the audio system 200 can end the calibration of the HRTF after the set number of iterations expires or after the set time period expires.

[0087] After completing the iterative HRTF customization process, the HRTF customization module 290 provides the user's customized set of HRTFs to the audio controller 230. The audio controller 230 uses the customized set of HRTFs to generate spatialized sound using the transducer array 210 for subsequent audio content provided to the user.

[0088] Figure 3 is a schematic diagram of a head mounted device 300 and multiple test positions according to one or more embodiments. Figure 1A The head mounted device 100 and Figure 1B The head mounted device 105 may be an embodiment of the head mounted device 300. The head mounted device 300 includes an audio system (such as Figure 2 In some embodiments, the initial set of estimated HRTFs may be generated by an external system and transmitted to the head mounted device 300. In other embodiments, the initial set of estimated HRTFs may be generated by a local audio system on the head mounted device 300. The initial set of estimated HRTFs may be generated based at least in part on trained machine learning and computer vision models and images of the user's ears and body.

[0089] The head mounted device 300 selects test locations 310 to test the accuracy of the initial set of estimated HRTFs. The test locations 310 may be selected based on the initial set of estimated HRTFs. For example, the test locations 310 may be selected such that the density of the test locations 310 is based on the rate of change of the estimated HRTFs as a function of distance and / or angle. In some embodiments, the test locations 310 may be separated by a minimal perceptible change in the direction of arrival, such as at least 1 degree in azimuth and 5 degrees in elevation.

[0090] In some embodiments, the positioning that captures the user's response may be unique for each user. A set of positionings may be selected based on an initial estimate of the HRTF to obtain the user's response, and as the user's responses accumulate over time, new sets of positionings to be tested may correspond to areas where the HRTF is more sensitive, noisier, or more discontinuous between the available positioning choices. Among the attributes obtained as part of the user feedback, one or more attributes that drive the selection of positioning in these subsequent iterations may also depend on the user. In some embodiments, these attributes can be personalized for the user based on some other simple questions or statistical data accumulated from the user, for example, what the user cares about in terms of sound quality, or what sounds the user may listen to more often, etc.

[0091] The head mounted device 300 synthesizes audio content for each test location 310 and presents the audio content to the user. The test sound may correspond to broadband speech or broadband noise or specific sounds commonly seen in reality. In some embodiments, the test sound may be concentrated on frequencies between 3kHz and 10kHz or higher.

[0092] The head mounted device 300 monitors the user's response to the audio content for each test positioning. In some embodiments, the initial set of HRTFs may be based on a user wearing the head mounted device, while in other embodiments, the initial set of HRTFs may be based on a user not wearing the head mounted device. For the former case, a predetermined transformation or mapping of HRTF signal changes between the head mounted device and the head mounted device is used to adjust the HRTF. These predetermined transformations may be calculated through ablation studies or other user studies. These predetermined transformations may be specific to an individual, and in some cases, they may also be calculated using machine learning and computer vision models, which again utilize user data including anthropometric features, images, or videos of the left and right ears.

[0093] For example, the head mounted device 300 may track the gaze direction of the user in response to presenting the synthesized sound for the test location 310a. The head mounted device 300 may determine, based on the gaze direction, that the synthesized sound perceived by the user originates from the perceptual location 320. The positioning difference between the test location 310a and the perceptual location 320 represents the inaccuracy of the estimated HRTF for the test location 310a. Based on the monitored response, the head mounted device 300 generates a new set of estimated HRTFs and a new set of test locations 330. For example, the accuracy value of the estimated HRTF may be input into the HRTF model, and the HRTF model may output a new set of estimated HRTFs. In some embodiments, each test location in the new set of test locations 330 may be different from each test location in the test locations 310. However, in some embodiments, at least one test location in the new set of test locations 330 may be positioned together with at least one test location in the test locations 310.

[0094] The head mounted device 300 synthesizes audio content for each test location in the new set of test locations 330 and presents the audio content to the user. The head mounted device 300 monitors the user's response to the audio content. Based on the monitored response, the head mounted device 300 generates a new set of estimated HRTFs and a new set of test locations. The head mounted device 300 may iteratively refine the estimated HRTF until a threshold accuracy is reached, until a fixed number of iterations is reached, until a time limit expires, or until a user input ends calibration.

[0095] Figure 4 is a flow chart of a method 400 of generating a customized HRTF in accordance with one or more embodiments. Figure 4 The process shown may be performed by components of an audio system (e.g., audio system 200). In other embodiments, other entities may perform Figure 4 Embodiments may include different and / or additional steps, or perform the steps in a different order.

[0096] The audio system selects 410 a set of test positions for a user's estimated HRTF set. The user's estimated HRTF set can be generated locally on the head mounted device, or received from an online system, such as an HRTF machine learning and computer vision module. The set of test positions can be selected so that a greater density of test positions is selected in areas where the estimated HRTFs differ relatively greatly as a function of angle.

[0097] The audio system generates 420 a customized set of HRTFs for a user based in part on applying an iterative process to the estimated set of HRTFs. The iterative process is described below and includes steps 430-480.

[0098] The audio system generates 430 test sounds for the set of test positionings. The audio system may instruct the transducer array to generate the test sounds. The generated test sounds are spatialized using the estimated HRTF of the user.

[0099] The audio system determines 440 an accuracy value of an estimated HRTF for the user based in part on a response of the user to the generated test sound. The user's response may be detected using sensors on the head mounted device, such as an eye tracking unit that detects a gaze location of the user. For example, if the user perceives that the generated test sound is emitted from a perceptual location that is different from the test location, the difference in location indicates an inaccuracy of the HRTF for the test location.

[0100] For active calibration, the audio system may instruct the user to perform a response, such as looking or pointing in the direction of a perceived sound. For passive calibration, the audio system may detect a user's response of looking or pointing at the location of a perceived sound without providing the user with explicit instructions for a sound response.

[0101] The audio system updates 450 the estimated HRTF of the user based in part on the accuracy value.For example, the audio system can provide the accuracy value for the estimated HRTF to a local or external HRTF model that calculates an updated HRTF.

[0102] The audio system adjusts 460 the set of test locations based in part on the updated estimated HRTF. For example, the audio system may select test locations in an area containing a relatively high rate of change of the updated estimated HRTF. The adjusted set of test locations may include a greater density of test locations in an area where the system calculates a greater inaccuracy in the estimated HRTF.

[0103] The process includes repeating 470 the iterative process until a quality metric is met. The quality metric is based in part on the accuracy value. In some embodiments, the iterative process may continue for a fixed number of iterations, a fixed amount of time, or until the audio system receives a user instruction to end the iterative process.

[0104] The process includes presenting 480 content to the user using the customized set of HRTFs. The content can be any type of audio content presented in normal use of the head mounted device. The head mounted device can restart the HRTF customization process in response to an action such as activating the head mounted device or updating some system details, or in response to a command from the user to customize the HRTFs, or at set intervals, such as once a week or more. In some embodiments, this calibration is performed only once, immediately after the user first starts the device.

[0105] Figure 55 is a system 500 including a head mounted device 505 according to one or more embodiments. In some embodiments, the head mounted device 505 may be Figure 1A The head mounted device 100 or Figure 1B The head mounted device 105. The system 500 can operate in an artificial reality environment (e.g., a virtual reality environment, an augmented reality environment, a mixed reality environment, or some combination thereof). Figure 5 The system 500 shown includes a head mounted device 505, an input / output (I / O) interface 510 coupled to a console 515, a network 520, and a mapping server 525. Figure 5 An example system 500 is shown that includes one head mounted device 505 and one I / O interface 510, but in other embodiments, any number of these components may be included in the system 500. For example, there may be multiple head mounted devices, each head mounted device having an associated I / O interface 510, each head mounted device and I / O interface 510 communicating with a console 515. In alternative configurations, different and / or additional components may be included in the system 500. Furthermore, in some embodiments, in combination with Figure 5 The functionality described by one or more of the components shown may be combined with Figure 5 The described methods are distributed between the components in different ways. For example, some or all of the functions of the console 515 can be provided by the head mounted device 505.

[0106] The head mounted device 505 includes a display assembly 530, an optical block 535, one or more position sensors 540, and a DCA 545. Some embodiments of the head mounted device 505 have a Figure 5 In addition, in other embodiments, by combining Figure 5 The functionality provided by the various components described may be distributed differently among the components of the head mounted device 505 , or may be captured in a separate component remote from the head mounted device 505 .

[0107] The display assembly 530 displays content to the user based on the data received from the console 515. The display assembly 530 displays content using one or more display elements (e.g., the display element 120). The display element can be, for example, an electronic display. In various embodiments, the display assembly 530 includes a single display element or multiple display elements (e.g., a display for each eye of the user). Examples of electronic displays include: a liquid crystal display (LCD), an organic light emitting diode (OLED) display, an active matrix organic light emitting diode display (AMOLED), a waveguide display, some other display, or some combination thereof. It should be noted that in some embodiments, the display element 120 can also include some or all of the functionality of the optical block 535.

[0108] The optical block 535 can amplify the image light received from the electronic display, correct the optical errors associated with the image light, and present the corrected image light to one or two windows of the head mounted device 505. In various embodiments, the optical block 535 includes one or more optical elements. Example optical elements included in the optical block 535 include: an aperture, a Fresnel lens, a convex lens, a concave lens, a filter, a reflective surface, or any other suitable optical element that affects the image light. In addition, the optical block 535 can include a combination of different optical elements. In some embodiments, one or more optical elements in the optical block 535 can have one or more coatings, such as a partially reflective coating or an anti-reflective coating.

[0109] The magnification and focusing of the image light by the optical block 535 allows the electronic display to be physically smaller, lighter, and consume less power than a larger display. Additionally, the magnification can increase the field of view of the content presented by the electronic display. For example, the field of view of the displayed content is such that the displayed content is presented using almost all of the user's field of view (e.g., approximately 110 degrees diagonally), and in some cases all of the field of view. Additionally, in some embodiments, the amount of magnification can be adjusted by adding or removing optical elements.

[0110] In some embodiments, the optical block 535 can be designed to correct one or more types of optical errors. Examples of optical errors include barrel or pincushion distortion, longitudinal chromatic aberration, or lateral chromatic aberration. Other types of optical errors may also include spherical aberration, chromatic aberrations, or errors due to lens field curvature, astigmatism, or any other type of optical error. In some embodiments, the content provided to the electronic display for display is pre-distorted, and when the optical block 535 receives image light generated based on the content from the electronic display, the optical block 535 corrects the distortion.

[0111] The position sensor 540 is an electronic device that generates data indicating the position of the head mounted device 505. In some embodiments, the position of the head mounted device 505 can be provided to the audio system 550 as an indication of the user's response to the test sound. The position sensor 540 generates one or more measurement signals in response to the movement of the head mounted device 505. The position sensor 190 is an embodiment of the position sensor 540. Examples of the position sensor 540 include: one or more IMUs, one or more accelerometers, one or more gyroscopes, one or more magnetometers, another suitable type of sensor that detects movement, or some combination thereof. The position sensor 540 may include multiple accelerometers that measure translational movement (forward / backward, up / down, left / right) and multiple gyroscopes that measure rotational movement (e.g., pitch, yaw, roll). In some embodiments, the IMU samples the measurement signals quickly and calculates the estimated position of the head mounted device 505 based on the sampled data. For example, the IMU integrates the measurement signals received from the accelerometer over time to estimate a velocity vector, and integrates the velocity vector over time to determine an estimated position of a reference point on the head mounted device 505. A reference point is a point that can be used to describe the position of the head mounted device 505. Although a reference point can generally be defined as a point in space, in practice, a reference point is defined as a point within the head mounted device 505.

[0112] The DCA 545 generates depth information for a portion of the local area. The DCA includes one or more imaging devices and a DCA controller. The DCA 545 may also include an illuminator. The operation and structure of the DCA 545 are described above with respect to Figure 1A Described.

[0113] The audio system 550 provides audio content to the user of the head mounted device 505. The audio system 550 is an embodiment of the audio system 200 described above. The audio system 550 may include one or more acoustic sensors, one or more transducers, and an audio controller. The audio system 550 may provide spatialized audio content to the user. In some embodiments, the audio system 550 may request acoustic parameters from the mapping server 525 through the network 520. The acoustic parameters describe one or more acoustic characteristics (e.g., room impulse response, reverberation time, reverberation level, etc.) of a local area. The audio system 550 may provide information describing at least a portion of the local area from, for example, the DCA 545 and / or positioning information of the head mounted device 505 from the position sensor 540. The audio system 550 may generate one or more sound filters using the one or more acoustic parameters received from the mapping server 525, and use the sound filters to provide audio content to the user.

[0114] The audio system 550 generates a customized HRTF for the user. In some embodiments, the audio system 550 may receive an initial set of estimated HRTFs from the HRTF machine learning and computer vision module 570. Figure 2-Figure 4 As further described, the audio system 550 performs an iterative process to customize the HRTF for the user by selecting test sound localizations, rendering the sounds using the estimated HRTFs, and detecting the user's responses to the test sounds.

[0115] The I / O interface 510 is a device that allows a user to send an action request and receive a response from the console 515. An action request is a request to perform a specific action. For example, an action request can be an instruction to start or end capturing image or video data, or an instruction to perform a specific action within an application. The I / O interface 510 may include one or more input devices. Example input devices include a keyboard, a mouse, a game controller, or any other suitable device for receiving an action request and transmitting the action request to the console 515. The action request received by the I / O interface 510 is transmitted to the console 515, and the console 515 performs the action corresponding to the action request. In some embodiments, the I / O interface 510 includes an IMU, and the IMU captures calibration data indicating the estimated position of the I / O interface 510 relative to the initial position of the I / O interface 510. In some embodiments, the I / O interface 510 can provide tactile feedback to the user according to the instruction received from the console 515. For example, tactile feedback is provided when an action request is received, or when console 515 transmits instructions to I / O interface 510 that cause I / O interface 510 to generate tactile feedback when console 515 performs the action.

[0116] The console 515 provides content to the head mounted device 505 for processing based on information received from one or more of: the DCA 545, the head mounted device 505, and the I / O interface 510. Figure 5 In the example shown, the console 515 includes an application store 555, a tracking module 560, and an engine 565. Some embodiments of the console 515 have a Figure 5 Similarly, the functions described further below may be implemented in different ways in combination with the modules or components described above. Figure 5 The described approach is distributed among the components of the console 515. In some embodiments, the functionality discussed herein with reference to the console 515 can be implemented in the head mounted device 505 or in a remote system.

[0117] The application storage 555 stores one or more applications for execution by the console 515. An application is a set of instructions that, when executed by a processor, generates content for presentation to a user. The content generated by the application may be responsive to input received from the user via movement of the head mounted device 505 or the I / O interface 510. Examples of applications include: gaming applications, conferencing applications, video playback applications, or other suitable applications.

[0118] The tracking module 560 uses information from the DCA 545, one or more position sensors 540, or some combination thereof to track the movement of the head mounted device 505 or the I / O interface 510. The tracking module 560 can detect the position or orientation of the head mounted device 505 in response to the test sound, such as by using an external camera to view the orientation of the head mounted device. The tracking module 560 can transmit the detected position of the head mounted device 505 to the audio system 550 for calculating the accuracy value of the estimated HRTF. In some embodiments, the tracking module 560 determines the position of the reference point of the head mounted device 505 in the mapping of the local area based on the information from the head mounted device 505. The tracking module 560 can also determine the position of an object or a virtual object. In addition, in some embodiments, the tracking module 560 can use the portion of the data indicating the position of the head mounted device 505 from the position sensor 540 and the representation of the local area from the DCA 545 to predict the future positioning of the head mounted device 505. The tracking module 560 provides the estimated or predicted future position of the head mounted device 505 or the I / O interface 510 to the engine 565 .

[0119] The engine 565 executes the application and receives the position information, acceleration information, velocity information, predicted future position, or some combination thereof of the head mounted device 505 from the tracking module 560. Based on the received information, the engine 565 determines the content to be provided to the head mounted device 505 for presentation to the user. For example, if the received information indicates that the user has looked to the left, the engine 565 generates content for the head mounted device 505 that reflects (mirror) the user's movement in a virtual local area or in a local area that enhances the local area with additional content. In addition, the engine 565 performs an action within an application executed on the console 515 in response to an action request received from the I / O interface 510, and provides feedback to the user that the action is performed. The feedback provided may be visual or auditory feedback via the head mounted device 505 or tactile feedback via the I / O interface 510.

[0120] The network 520 couples the head mounted device 505 and / or the console 515 to the mapping server 525. The network 520 may include any combination of local area networks and / or wide area networks using wireless and / or wired communication systems. For example, the network 520 may include the Internet and a mobile phone network. In one embodiment, the network 520 uses standard communication technologies and / or protocols. Therefore, the network 520 may include links using technologies such as Ethernet, 802.11, Worldwide Interoperability for Microwave Access (WiMAX), 2G / 3G / 5G mobile communication protocols, digital subscriber lines (DSL), asynchronous transfer mode (ATM), InfiniBand, PCI Express Advanced Switching, etc. Similarly, the network protocols used on the network 520 may include Multi-Protocol Label Switching (MPLS), Transmission Control Protocol / Internet Protocol (TCP / IP), User Datagram Protocol (UDP), Hypertext Transfer Protocol (HTTP), Simple Mail Transfer Protocol (SMTP), File Transfer Protocol (FTP), etc. Data exchanged over the network 520 may be represented using techniques and / or formats including image data in binary form (e.g., portable network graphics (PNG)), hypertext markup language (HTML), extensible markup language (XML), etc. In addition, all or part of the links may be encrypted using conventional encryption techniques, such as secure socket layer (SSL), transport layer security (TLS), virtual private network (VPN), Internet Protocol security (IPsec), etc.

[0121] The mapping server 525 may include a database storing virtual models describing a plurality of spaces, wherein a location in the virtual model corresponds to a current configuration of a local area of ​​the head mounted device 505. The mapping server 525 receives information describing at least a portion of the local area and / or location information of the local area from the head mounted device 505 via the network 520. The mapping server 525 determines a location in the virtual model associated with the local area of ​​the head mounted device 505 based on the received information and / or location information. The mapping server 525 determines (e.g., retrieves) one or more acoustic parameters associated with the local area based in part on the determined location in the virtual model and any acoustic parameters associated with the determined location. The mapping server 525 may transmit the location of the local area and any acoustic parameter values ​​associated with the local area to the head mounted device 505.

[0122] The system 500 includes a HRTF machine learning and computer vision module 570. The HRTF machine learning and computer vision module 570 applies machine learning and computer vision techniques to generate an HRTF model that, when applied to data describing a user, outputs an estimated HRTF relative to the user's location. As part of generating the HRTF model, the HRTF machine learning and computer vision module 570 forms an HRTF training set by identifying a positive training set of HRTFs that have been determined to be accurate, and in some embodiments, forms a negative training set of HRTFs that have been determined to be inaccurate. The HRTF model, when applied to data describing a user, such as a picture of a user's head, torso, or ears, or a profile or measurement of the ears known as anthropometric features, outputs a set of estimated HRTFs for the user of the head mounted device 505. The HRTF machine learning and computer vision module 570 receives feedback from the head mounted device 505 describing the accuracy of the estimated HRTFs. The HRTF machine learning and computer vision module 570 uses the feedback as an input to the HRTF model to update the HRTF for the user. Although shown as a separate component, in some embodiments, the HRTF machine learning and computer vision module 570 can be a component of the head mounted device 505 or console 515, such as part of the audio system 550.

[0123] Additional configuration information

[0124] The foregoing description of the embodiments has been presented for the purpose of illustration; it is not intended to be exhaustive or to limit the patent rights to the precise forms disclosed. Those skilled in the relevant art will appreciate that many modifications and variations are possible in light of the above disclosure.

[0125] Some parts of this specification describe embodiments from the aspect of algorithms and symbolic representations of operations on information. Those skilled in the art of data processing typically use these algorithmic descriptions and representations to effectively convey the essence of their work to other technicians in the field. Although these operations are described functionally, computationally, or logically, it should be understood that they will be implemented by computer programs or equivalent circuits, microcodes, etc. In addition, it is sometimes proven to be convenient to refer to these arrangements of operations as modules without loss of generality. The described operations and their associated modules can be embodied in software, firmware, hardware, or any combination thereof.

[0126] Any steps, operations or processes described herein may be performed or implemented using one or more hardware or software modules alone or in combination with other devices. In one embodiment, the software modules are implemented using a computer program product including a computer-readable medium containing computer program code, which may be executed by a computer processor to perform any or all of the steps, operations or processes described.

[0127] Embodiments may also relate to apparatus for performing the operations described herein. The apparatus may be specially constructed for the desired purpose, and / or it may include a general-purpose computing device selectively activated or reconfigured by a computer program stored in a computer. Such a computer program may be stored in a non-transitory, tangible computer-readable storage medium, or any type of medium suitable for storing electronic instructions, which may be coupled to a computer system bus. In addition, any computing system mentioned in this specification may include a single processor or may be an architecture that employs a multi-processor design to increase computing power.

[0128] Embodiments may also relate to products produced by the computing processes described herein. Such products may include information obtained from the computing processes, wherein the information is stored on a non-transitory, tangible computer-readable storage medium and may include any embodiment of a computer program product or other data combination described herein.

[0129] Finally, the language used in the specification is selected primarily for readability and instructional purposes, and it may not be selected to describe or limit the patent rights. Therefore, it is intended that the scope of the patent rights be limited not by this detailed description, but by any claims of an application filed based thereon. Therefore, the disclosure of the embodiments is intended to illustrate, not to limit, the scope of the patent rights set forth in the appended claims.

Claims

1. A method for presenting personalized audio content to a user, comprising: Selecting a test positioning set for a user's estimated head related transfer function HRTF set; Generating a customized set of HRTFs for the user, the generating being based in part on applying an iterative process to the estimated set of HRTFs, the iterative process comprising: generating test sounds for the set of test positionings, wherein the generated test sounds are spatialized using the estimated HRTF of the user; determining an accuracy value of an estimated HRTF for the user based in part on a response of the user to the generated test sound; updating the estimated HRTF of the user based in part on the accuracy value; and adjusting the set of test positions based in part on a rate of change of the updated estimated HRTF in the region, wherein adjusting the set of test positions comprises selecting relatively more test positions in the region where the values ​​of the HRTF vary greatly; Repeating the iterative process until the quality metric is met; and Content is presented to the user using the customized set of HRTFs.

2. The method according to claim 1, wherein: The estimated HRTF set is generated by HRTF machine learning and computer vision modules.

3. The method according to claim 1, wherein: The set of estimated HRTFs is generated based on data describing physical characteristics of the user.

4. The method according to claim 3, wherein: The data describing the user includes an image of the user's ear.

5. The method according to claim 1, wherein: Generating test sounds for the set of test locations includes sequentially generating test sounds for each of the test locations.

6. The method according to claim 5, wherein: Based on a location difference between a test location for the estimated HRTF and a gaze location of a user in response to a test sound for the test location, an accuracy value for the estimated HRTF is calculated.

7. A method for presenting personalized audio content to a user, comprising: selecting a first set of test positions based on a first estimated head related transfer function HRTF set of the user and a rate of change of the first estimated HRTF set in a region, wherein selecting the first set of test positions comprises selecting relatively more test positions in a region where the values ​​of the HRTF vary greatly; generating test sounds for the first test location set; calculating an accuracy value for the first estimated HRTF set for the user based on a user response to the test sounds of the first test positioning set; calculating a second HRTF set for the user based on the accuracy value of the first estimated HRTF set; selecting a second set of test locations based on the second set of HRTFs; and Test sounds are generated for the second set of test locations.

8. The method according to claim 7, wherein: The first estimated HRTF set is generated by HRTF machine learning and computer vision modules.

9. The method according to claim 7, wherein: The first set of estimated HRTFs is generated based on data describing a physical characteristic of the user.

10. The method according to claim 9, wherein: The data describing the user includes an image of the user's ear.

11. The method according to claim 7, wherein: Generating test sounds for the first set of test locations and generating test sounds for the second set of test locations includes sequentially generating test sounds for each of the test locations.

12. The method according to claim 11, wherein: An accuracy value for each estimated HRTF in the first set of estimated HRTFs is calculated based on a positioning difference between the test positioning for the estimated HRTF in the first set of estimated HRTFs and a gaze positioning of the user in response to the test sound for the test positioning.

13. A computer program product, comprising a non-transitory computer-readable storage medium, wherein the non-transitory computer-readable storage medium stores a computer program code, wherein when the computer program code is executed by a processor, the method according to any one of claims 1 to 6 is implemented.

Citation Information

Patent Citations

  • Method of improving localization of surround sound

    US20190246231A1