Estimating hearing loss of a user from the interaction of the user with a local environment identified from accessed audio and

By identifying the user's interaction with local areas and applying filters to generate enhanced audio, the problem of time-consuming and incomplete evaluation of hearing loss diagnosis in the prior art is solved, and hearing loss estimation and personalized compensation without testing are achieved.

CN120358983APending Publication Date: 2025-07-22CTRL-LABS CORP
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202380085796.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2023-04-24
Filing Date
2023-12-18
Publication Date
2025-07-22

AI Technical Summary

Technical Problem

In the prior art, the diagnosis of mild to moderate hearing loss requires time-consuming clinical testing and investigation by individuals, resulting in incomplete assessment of individual functional hearing, limiting the accurate estimate and compensation of hearing loss.

Method used

By using acoustic sensors and additional sensors to acquire audio and information from local areas, identify user interactions with local areas, and estimate user's hearing profile based on interactions and related attributes, filters are applied to generate enhanced audio to compensate for hearing loss.

Benefits of technology

It can estimate user hearing loss without time-consuming testing, provide more comprehensive hearing assessment and personalized hearing compensation, and improve the accuracy and compensation effect of hearing loss.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120358983A_ABST
    Figure CN120358983A_ABST
Patent Text Reader

Abstract

An audio system includes one or more acoustic sensors that acquire audio from a local area around the audio system and one or more additional sensors that acquire data describing the local area. An audio controller in the audio system identifies a user's interaction based on the collected audio and data describing the local area. The audio controller determines attributes associated with each interaction, such as metrics describing user conversations, metrics describing user requests for sharpening processing, and metrics based on themes determined from collected audio from the user. The audio controller estimates a hearing profile of the user based on the interaction and the associated attributes, the hearing profile identifying a range of frequencies of hearing loss of the user. The audio controller may utilize the hearing profile to determine one or more filters to counteract hearing loss of the user.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure generally relates to reducing hearing loss, and more particularly, to estimating a user's hearing loss and reducing the user's hearing loss. Background Art

[0002] Although mild to moderate hearing loss affects a significant number of individuals, the conventional diagnosis of such hearing loss in an individual requires the individual to visit a clinical service provider, such as an audiologist. The clinical service provider conducts a hearing test to estimate an audiogram that describes the individual's hearing for different frequencies, and may have the individual fill out a series of survey questions. The clinical service provider may also measure the individual's intelligibility and listening effort when listening to speech with different levels of background audio. The clinical service provider estimates the individual's hearing loss and the impact of the estimated hearing loss on the individual based on the individual's audiogram, the answers to the survey questions, and the results of the speech and effort measurements. However, participating in a hearing test and filling out a survey are time-consuming for the individual being evaluated. In addition, the hearing test and / or survey provide limited data about the individual's hearing, preventing either process from providing a comprehensive assessment of the individual's functional hearing. This limited assessment of the individual's functional hearing limits the ability of conventional methods to accurately estimate an individual's hearing loss and compensate for the individual's hearing loss. Summary of the Invention

[0003] According to a first aspect of the present disclosure, there is provided a method that includes: collecting audio from a local area around an audio controller using one or more acoustic sensors; collecting information describing the local area from one or more additional sensors; identifying, by the audio controller, one or more interactions of the user with the local area based on the collected audio and the information describing the local area; and estimating a sound profile of the user based on the identified one or more interactions and attributes associated with the one or more interactions, the sound profile identifying the user's hearing loss for one or more frequency ranges.

[0004] In some embodiments, the method further includes: determining, based on the sound profile, one or more filters to be applied to audio for presentation to the user; generating enhanced audio by applying the one or more filters to the audio for presentation to the user, the enhanced audio increasing the amplitude of a portion of the audio whose frequencies are within a frequency range for which the sound profile identifies a hearing loss; and presenting the enhanced audio to the user via one or more transducers.

[0005] In some embodiments, estimating a user's hearing profile based on one or more identified interactions and attributes associated with the one or more interactions includes: applying a model to the one or more identified interactions and attributes associated with the one or more interactions, the model outputting a hearing profile based on the one or more identified interactions and attributes associated with the one or more interactions.

[0006] In some embodiments, an attribute associated with an interaction is an indication of whether a sound source providing audio during the interaction is within the field of view of an imaging device, the field of view of the imaging device overlapping the field of view of the user.

[0007] In some embodiments, an audio controller is included in a device worn by the user, and an attribute associated with the interaction is the position of the sound source providing audio during the interaction relative to the audio controller.

[0008] In some embodiments, an attribute associated with an interaction includes data describing a remediation initiation made by the user during the interaction, the remediation initiation indicating a request to initiate clarification of audio from a sound source providing audio during the interaction.

[0009] In some embodiments, an attribute associated with an interaction includes one or more turn-taking metrics for the interaction, the turn-taking metrics describing the conversation activities performed by the user during the interaction.

[0010] In some embodiments, the turn-taking metrics are selected from the group consisting of: the length of time the user produces audio during the interaction; the percentage of time the user produces audio during the interaction; the amount of time between when the sound source provides audio and when the user provides audio during the interaction; and any combination thereof.

[0011] In some embodiments, an attribute associated with an interaction includes a semantic metric based on a topic determined for audio collected from the user during the interaction.

[0012] According to a second aspect of the present disclosure, there is provided a computer program product comprising a non-transitory computer-readable storage medium having instructions encoded thereon that, when executed by a processor, cause the processor to: perform the method according to the first aspect. The computer program product comprises a non-transitory computer-readable storage medium having instructions encoded thereon that, when executed by a processor, cause the processor to perform the following operations: acquire audio from a local area around an audio controller using one or more acoustic sensors; acquire information describing the local area from one or more additional sensors; identify, by the audio controller, one or more interactions of a user with the local area based on the acquired audio and the information describing the local area; and estimate a hearing profile of the user based on the identified one or more interactions and attributes associated with the one or more interactions, the hearing profile identifying the user's hearing loss for one or more frequency ranges.

[0013] In some embodiments, the non-transitory computer-readable storage medium has instructions encoded thereon that, when executed by a processor, cause the processor to perform the following operations: determine one or more filters to apply to audio for presentation to the user based on the hearing profile; generate enhanced audio by applying the one or more filters to the audio for presentation to the user, the enhanced audio increasing the amplitude of a portion of the audio whose frequency is within a frequency range for which the hearing profile identifies hearing loss; and present the enhanced audio to the user via one or more transducers.

[0014] In some embodiments, estimating a hearing profile of the user based on the identified one or more interactions and attributes associated with the one or more interactions comprises: applying a model to the identified one or more interactions and attributes associated with the one or more interactions, the model outputting a hearing profile based on the identified one or more interactions and attributes associated with the one or more interactions.

[0015] In some embodiments, an attribute associated with the interaction is an indication of whether a sound source providing audio during the interaction is within the field of view of an imaging device, the field of view of the imaging device overlapping the field of view of the user.

[0016] In some embodiments, the audio controller is included in a device worn by the user, and an attribute associated with the interaction is the position of the sound source providing audio during the interaction relative to the audio controller.

[0017] In some embodiments, an attribute associated with the interaction includes data describing a remediation initiation made by the user during the interaction, the remediation initiation indicating a request to initiate clarification of audio from a sound source providing audio during the interaction.

[0018] In some embodiments, attributes associated with an interaction include one or more turn-taking metrics for the interaction, where the turn-taking metrics describe the conversational activity of the user during the interaction.

[0019] In some embodiments, the turn-taking metrics are selected from the group consisting of: the length of time the user produces audio during the interaction; the percentage of time the user produces audio during the interaction; the amount of time between when the sound source provides audio and when the user provides audio during the interaction; and any combination thereof.

[0020] In some embodiments, attributes associated with an interaction include semantic metrics based on topics determined from audio collected from the user during the interaction.

[0021] According to a third aspect of the present disclosure, there is provided a head-mounted device including: a frame; one or more display elements coupled to the frame, each display element configured to generate image light for presentation to a user; one or more acoustic sensors configured to collect audio from a local area around the head-mounted device; one or more additional sensors configured to collect information describing the local area around the head-mounted device; and an audio controller including a processor and a non-transitory computer-readable storage medium having instructions encoded thereon that, when executed by the processor, cause the processor to: identify one or more interactions of the user with the local area based on the collected audio and the information describing the local area; and estimate a hearing profile of the user based on the identified one or more interactions and attributes associated with the one or more interactions, the hearing profile identifying the user's hearing loss for one or more frequency ranges.

[0022] In some embodiments, the head-mounted device further includes a transducer array for presenting audio to the user, and wherein the audio controller further includes instructions encoded thereon that, when executed by the processor, cause the processor to: determine one or more filters to apply to the audio for presentation to the user based on the hearing profile; generate enhanced audio by applying the one or more filters to the audio for presentation to the user, the enhanced audio increasing the amplitude of a portion of the audio whose frequencies are within a frequency range for which the hearing profile identifies hearing loss; and present the enhanced audio to the user via the one or more transducers.

[0023] To estimate a user's hearing loss, a controller can receive data collected by one or more sensors, the data including information describing the user and information describing a local area. The controller receives data collected by various sensors and analyzes the received data to generate a hearing profile of the user. In various embodiments, the controller is a component of a head-mounted device that can be used in an artificial reality environment. The controller and the user are located in a local area (e.g., a room) surrounding the controller and the user. One or more acoustic sensors are coupled to the controller and collect audio within the local area. One or more additional sensors are also coupled to the head-mounted device, where these additional sensors collect information about the user and / or information describing the local area. For example, information describing the user can include head movement, eye tracking, facial expressions, biometric data, other data characterizing the user, or some combination thereof. For example, to collect head movement, a position sensor can be implemented to track the position and / or movement of the head-mounted device, and / or an imaging device can collect image data of the local area to locate the head-mounted device relative to the local area. To collect eye tracking information and / or facial expressions, an imaging device facing inward toward the user's face can be provided on the head-mounted device to collect image data of the user's face. The controller can analyze the image data to determine eye tracking information, facial expressions, or some combination thereof. Other example biometric sensors include a heart sensor, a blood pressure sensor, a blood oxygen level sensor, an electroencephalogram sensor, a blood glucose monitor, other sensors capable of measuring biometric data of the user. Information describing the local area can include image data, audio data, other data characterizing the local area. For example, the additional sensor is an imaging device that collects video or images of the local area. As another example, the additional sensor is a position sensor that collects the position or movement of the head-mounted device or the user within the local area.

[0024] In one or more embodiments, the controller may utilize the collected audio data of the user to determine the user's hearing profile. In such an embodiment, the controller may be a component of an audio system, such as an audio controller. The controller may also utilize other information describing the user, information describing the local area, or some combination thereof to determine the hearing profile. The controller identifies one or more interactions of the user with the local area from the collected audio and any other information that can be combined. The user's interaction is the user's response to a stimulus in the local area. For example, the interaction is the user's movement (e.g., the user's posture) or the user generating audio. Additionally, the controller determines one or more attributes associated with the identified interaction. These attributes may describe the user's speech or movement during the interaction (e.g., asking for repetition; moving closer to the sound source), or may identify the conditions in the local area during the interaction (e.g., background noise, the position of the sound source relative to the audio controller). The controller estimates the user's hearing profile based on the identified one or more interactions and the attributes associated with the identified one or more interactions. The hearing profile describes the degree of the user's hearing loss and the configuration of the user's hearing loss, such as identifying the frequency range in which the user's hearing is impaired. In various embodiments, the controller applies a trained model to the identified one or more interactions and their attributes to estimate the user's hearing profile.

[0025] In one or more embodiments, the controller may utilize the user's eye-tracking information to determine the user's hearing profile. In such an embodiment, the controller may be a component of an eye-tracking system including the controller and one or more imaging devices. The controller may also utilize other information describing the user, information describing the local area, or some combination thereof to determine the hearing profile.

[0026] In one or more embodiments, the controller may utilize the user's facial expression information to determine the user's hearing profile. In such an embodiment, the controller may be a component of a face-tracking system including the controller and one or more imaging devices that are set inward (where the user's face is within the field of view). The controller may also utilize other information describing the user, information describing the local area, or some combination thereof to determine the hearing profile.

[0027] In some embodiments, the controller compensates for the hearing loss of the user identified by the hearing profile. The controller determines one or more filters based on the hearing profile, and the one or more filters are configured to counteract one or more effects of the hearing loss. For example, one determined filter increases the amplitude of an audio signal whose frequency is within the frequency range where the hearing profile indicates that the user has hearing loss. The controller applies the one or more filters to the audio to generate enhanced audio. Continuing with the above example, in the enhanced audio, the audio within the frequency range where the hearing profile indicates the user's hearing loss is amplified relative to other frequencies. The enhanced audio is presented to the user through one or more transducers (e.g., speakers), increasing the likelihood that the user understands the enhanced audio compared to the audio without the application of the one or more filters.

[0028] In various embodiments, one or more acoustic sensors collect audio from a local area around the audio controller. One or more additional sensors collect information describing the local area. Examples of information describing the local area include video of the local area, the movement of the user in the local area, the background noise level in the local area, or other data describing the conditions in the local area. The audio controller identifies one or more interactions of the user with the local area based on the collected audio, video, head tracking, eye tracking, and information describing the local area, and estimates the user's hearing profile based on the identified one or more interactions and the attributes associated with the one or more interactions. The hearing profile identifies the user's hearing loss for one or more frequency ranges.

[0029] In some embodiments, a computer program product includes a non-transitory computer-readable storage medium having instructions encoded thereon that, when executed by a processor, cause the processor to: collect audio from a local area around the audio controller using one or more acoustic sensors and collect information describing the local area from one or more additional sensors. Additionally, the instructions, when executed, cause the processor to: identify, by the audio controller, one or more interactions of the user with the local area based on the collected audio and information describing the local area. Execution of the instructions causes the processor to: estimate the user's hearing profile based on the identified one or more interactions and the attributes associated with the one or more interactions, where the hearing profile identifies the user's hearing loss for one or more frequency ranges.

[0030] In some embodiments, a head-mounted device includes one or more display elements coupled to a frame, each display element configured to generate image light for presentation to a user. The head-mounted device further includes one or more acoustic sensors and one or more additional sensors, the one or more acoustic sensors configured to collect audio from a local area around the head-mounted device, and the one or more additional sensors configured to collect information describing the local area around the head-mounted device. The head-mounted device further includes an audio controller having a processor and a non-transitory computer-readable storage medium having instructions encoded thereon that, when executed by the processor, cause the processor to: identify, by the audio controller, one or more interactions of the user with the local area based on the collected audio and the information describing the local area, and estimate a hearing profile of the user based on the identified one or more interactions and attributes associated with the one or more interactions, wherein the hearing profile identifies the user's hearing loss for one or more frequency ranges.

[0031] It will be appreciated that any features described herein as being suitable for incorporation into one or more aspects or embodiments of the present disclosure are intended to be generalizable across any and all aspects and embodiments of the present disclosure. Those skilled in the art can understand other aspects of the present disclosure based on the specification, claims, and drawings of the present disclosure. The foregoing general description and the following detailed description are merely exemplary and explanatory and are not restrictive of the claims. BRIEF DESCRIPTION OF THE DRAWINGS

[0032] Figure 1A is a perspective view of a head-mounted device implemented as a glasses device according to one or more embodiments.

[0033] Figure 1B is a perspective view of a head-mounted device implemented as a head-mounted display according to one or more embodiments.

[0034] Figure 2 is a block diagram of an audio system according to one or more embodiments.

[0035] Figure 3 is a flowchart showing a method for estimating a hearing loss of a user of a head-mounted device according to one or more embodiments.

[0036] Figure 4 is a conceptual diagram of a head-mounted device determining a hearing profile of a user based on collected audio from a local area and information describing the local area according to one or more embodiments.

[0037] Figure 5 is a system including a head-mounted device according to one or more embodiments.

[0038] These figures depict various embodiments for illustrative purposes only. As discussed below, those skilled in the art will readily recognize that alternative embodiments of the structures and methods shown herein can be employed without departing from the principles described herein. Detailed Description

[0039] The controller receives data collected by various sensors and analyzes the received data to estimate the user's hearing loss. In some embodiments, the controller also compensates for the estimated hearing loss of the user. In various embodiments, the controller is coupled to one or more acoustic sensors that collect audio from a local area around the controller. Additionally, the controller is coupled to one or more additional sensors that collect other information, which includes information describing the user, information describing the local area around the controller, or some combination thereof. Examples of additional sensors include imaging devices and position sensors. In various embodiments, the controller is included in a head-mounted device, such as a virtual reality (VR) head-mounted device or an artificial reality (AR) head-mounted device.

[0040] The controller identifies the user's interaction with the local area based on information describing the user and information describing the local area from various sensors. The user's interaction with the local area is the user's response to a stimulus (e.g., audio) in the local area. The controller also determines the attributes of the identified user interaction, where these attributes describe the user's actions or conditions in the local area associated with the identified interaction. The controller estimates the user's hearing profile based on the identified one or more interactions and associated attributes. This profile identifies the user's hearing loss for different frequencies of audio. For example, the identified one or more interactions and associated attributes are input into a trained model, which outputs the user's hearing profile. In some embodiments, the hearing profile is an audiogram. As another example, the controller maintains a correlation between the combination of interactions and attributes with the local area and the audiogram, and selects the audiogram based on the collected audio and information describing the local area. In various embodiments, the user wears a head-mounted device that includes the controller.

[0041] In various embodiments, the controller determines one or more filters to be applied to subsequent audio based on a hearing profile determined according to the identified user's interaction with the local area. The filters determined by the controller mitigate the hearing loss of the user identified by the hearing profile estimated for the user. Thus, applying the determined one or more filters to the subsequent audio modifies the subsequent audio to cancel out the estimated hearing loss of the user within one or more frequency ranges identified by the hearing profile estimated for the user.

[0042] Embodiments of the present disclosure may include or be implemented in conjunction with an artificial reality system. Artificial reality is a form of reality that has been adjusted in some way before being presented to a user, and artificial reality may include, for example, virtual reality (VR), augmented reality (AR), mixed reality (MR), hybrid reality, or some combination and / or derivative thereof. Artificial reality content may include fully generated content or generated content combined with captured (e.g., real-world) content. Artificial reality content may include video, audio, tactile feedback, or some combination thereof, any of which may be presented in a single channel or multiple channels (e.g., stereoscopic video that gives a viewer a three-dimensional effect). Additionally, in some embodiments, artificial reality may also be associated with an application, product, accessory, service, or some combination thereof, which are used to create content in artificial reality and / or otherwise used in artificial reality. An artificial reality system that provides artificial reality content may be implemented on various platforms, including wearable devices (e.g., head-mounted devices) connected to a host computer system, standalone wearable devices (e.g., head-mounted devices), mobile devices or computing systems, or any other hardware platform capable of providing artificial reality content to one or more viewers.

[0043] Figure 1Ais a perspective view of a head-mounted device 100 implemented as a glasses device according to one or more embodiments. In some embodiments, the glasses device is a near eye display (NED). Generally, the head-mounted device 100 can be worn on a user's face such that a display component and / or an audio system are used to present content (e.g., media content). However, the head-mounted device 100 can also be used to present media content to the user in different ways. Examples of media content presented by the head-mounted device 100 include one or more images, videos, audio, or some combination thereof. The head-mounted device 100 includes a frame and can include, among other components: a display component (which includes one or more display elements 120), a depth camera assembly (DCA), an audio system, a controller 150, and a position sensor 190. Although Figure 1A illustrates example locations of the components of the head-mounted device 100 on the head-mounted device 100, these components can be: at other locations on the head-mounted device 100; on a peripheral device paired with the head-mounted device 100; or some combination thereof. Similarly, there can be more components or fewer components on the head-mounted device 100 than Figure 1A shown.

[0044] The frame 110 holds the other components of the head-mounted device 100. The frame 110 includes a front component that holds one or more display elements 120 and end components (e.g., temples) that attach to the user's head. The front component of the frame 110 spans the top of the user's nose. The length of the end components can be adjustable (e.g., adjustable temple length) to fit different users. The end components can also include portions that curve behind the user's ears (e.g., temple tips, ear pieces).

[0045] One or more display elements 120 provide light to a user wearing the head-mounted device 100. As shown, the head-mounted device includes a display element 120 for each eye of the user. In some embodiments, the display element 120 generates image light that is provided to the eyebox of the head-mounted device 100. The eyebox is the location in the space occupied by the user's eyes when wearing the head-mounted device 100. For example, the display element 120 may be a waveguide display. The waveguide display includes a light source (e.g., a two-dimensional source, one or more line sources, one or more point sources, etc.) and one or more waveguides. Light from the light source is coupled into the one or more waveguides, and the one or more waveguides output light in such a way that pupil replication exists in the eyebox of the head-mounted device 100. Coupling light into and / or out of the one or more waveguides may be accomplished using one or more diffraction gratings. In some embodiments, the waveguide display includes a scanning element (e.g., a waveguide, a mirror, etc.) that scans the light when the light from the light source is coupled into the one or more waveguides. Note that in some embodiments, one or both of the two display elements 120 are opaque and do not transmit light from a local area around the head-mounted device 100. The local area is the area around the head-mounted device 100. For example, the local area may be the room in which the user wearing the head-mounted device 100 is located, or the user wearing the head-mounted device 100 may be outdoors and the local area is the outdoor area. In this context, the head-mounted device 100 generates VR content. Alternatively, in some embodiments, one or both of the respective display elements 120 are at least partially transparent, such that light from the local area can be combined with light from the one or more display elements to produce AR content and / or MR content.

[0046] In some embodiments, the display element 120 does not generate image light, but rather acts as a lens that transmits light from the local area to the eyebox. For example, one or both of the respective display elements 120 may be an uncorrected (non-prescription) lens or a prescription lens (e.g., a single vision lens, a bifocal lens, and a trifocal lens, or a progressive lens) that helps correct the user's vision defect. In some embodiments, the display element 120 may be polarized and / or colored to protect the user's eyes from the sun.

[0047] In some embodiments, the display element 120 may include another optical block (not shown). The optical block may include one or more optical elements (e.g., lenses, Fresnel lenses, etc.) that direct light from the display element 120 to the eyebox. The optical block may, for example, correct aberrations in some or all of the image content, magnify some or all of the image, or some combination thereof.

[0048] The DCA determines depth information for a portion of a local area around the head-mounted device 100. The DCA includes one or more imaging devices 130 and a DCA controller ( Figure 1A not shown in the figure), and may also include an illuminator 140. In some embodiments, the illuminator 140 uses light to illuminate a portion of the local area. The light can be, for example, structured light in the infrared (IR) (such as a dot pattern, bar, etc.), an IR flash for time-of-flight, etc. In some embodiments, one or more imaging devices 130 acquire an image of a portion of the local area that includes light from the illuminator 140. As shown in the figure, Figure 1A a single illuminator 140 and two imaging devices 130 are shown. In alternative embodiments, there is no illuminator 140 and at least two imaging devices 130 are present.

[0049] The DCA controller uses the acquired images and one or more depth determination techniques to calculate the depth information for this portion of the local area. Depth determination techniques can be, for example, direct time-of-flight (ToF) depth sensing, indirect ToF depth sensing, structured light, passive stereoscopic analysis, active stereoscopic analysis (using textures added to the scene by light from the illuminator 140), some other techniques for determining the depth of a scene, or some combination thereof.

[0050] The audio system provides audio content. The audio system includes a transducer array, a sensor array, and an audio controller. However, in other embodiments, the audio system may include different components and / or additional components. Similarly, in some cases, the functions described with reference to the components of the audio system may be distributed among the multiple components in a manner different from that described herein. For example, some or all of the functions of the audio controller may be performed by a remote server.

[0051] The transducer array presents sound to the user. The transducer array includes a plurality of transducers. The transducers can be speakers 160 or tissue transducers 170 (such as bone conduction transducers or cartilage conduction transducers). Although the speaker 160 is shown outside the frame 110, the speaker 160 can be enclosed within the frame 110. In some embodiments, the head-mounted device 100 includes a speaker array instead of a separate speaker for each ear, the speaker array including a plurality of speakers integrated into the frame 110 to improve the directivity of the presented audio content. The tissue transducer 170 is coupled to the user's head and directly vibrates the user's tissue (such as bone or cartilage) to generate sound. The number and / or position of the transducers can be different from Figure 1A the number and / or position shown in the figure.

[0052] The sensor array detects sound within a local area of the head-mounted device 100. The sensor array includes a plurality of acoustic sensors 180. The acoustic sensors 180 collect sound emitted from one or more sound sources in the local area (e.g., a room). Each acoustic sensor is configured to detect sound and convert the detected sound into an electronic format (analog format or digital format). The acoustic sensors 180 can be acoustic wave sensors, microphones, sound transducers, or similar sensors suitable for detecting sound.

[0053] In some embodiments, one or more of the acoustic sensors 180 can be placed in the ear canals of each ear (e.g., acting as a stereo microphone). In some embodiments, these acoustic sensors 180 can be placed on the outer surface of the head-mounted device 100, on the inner surface of the head-mounted device 100, separate from the head-mounted device 100 (e.g., as part of some other device), or some combination of the above positions. The number and / or position of the acoustic sensors 180 can be different from Figure 1A the number and / or position shown. For example, the number of acoustic detection positions can be increased to increase the amount of audio information collected and to improve the sensitivity and / or accuracy of that information. The acoustic detection positions can be oriented such that the microphones can detect sound in a wide range of directions around the user wearing the head-mounted device 100.

[0054] The audio controller processes information from the sensor array that describes the sound detected by the sensor array. The audio controller can include a processor and a computer-readable storage medium. The audio controller can be configured to generate a direction of arrival (DOA) estimation result, generate an acoustic transfer function (e.g., an array transfer function and / or a head-related transfer function), track the position of a sound source, form a beam in the direction of the sound source, classify the sound source, generate a sound filter for the speaker 160, or some combination thereof.

[0055] The controller 150 controls the operation of the head-mounted device 100. The controller 150 processes the data collected by various components of the head-mounted device 100 to determine the user's hearing profile. The data used to determine the hearing profile may include information describing the user, information describing the local area, or some combination thereof. The information describing the user includes audio detected by a sensor array (e.g., one or more acoustic sensors 180). The information describing the local area including the head-mounted device 100 may include ambient sound, image data of the local area, the position of the head-mounted device 100 in the local area, other data characterizing the local area, or some combination thereof. Additional sensors may be included in the head-mounted device 100, such as an imaging device 130 or a position sensor 190. Additionally or alternatively, one or more additional sensors are located outside the head-mounted device 100. For example, the additional sensor is a position sensor included in another wearable device worn by the user together with the head-mounted device 100, or another type of sensor included in the other wearable device.

[0056] The controller 150 may also process the data collected by various components to determine information describing the user, information describing the local area, or some combination thereof. For example, the controller 150 may perform eye-tracking analysis on the image data collected by the imaging device 130 positioned according to the user's eyes. As another example, the controller 150 may perform facial expression detection on the image data collected by the imaging device 130 positioned according to the user's face. In yet another example, the controller 150 may determine the position and / or movement of the head-mounted device 100 based on the image data collected by the imaging device 130 positioned according to the local area, the position data collected by the position sensor 190, or some combination thereof.

[0057] In one or more embodiments, the controller 150 identifies the interaction of the user of the head-mounted device 100 within the local area from the audio collected by one or more acoustic sensors 180 and other data. As used herein, the user's interaction is an action taken by the user in response to a stimulus within the local area. For example, if the stimulus within the local area is a sound from a sound source, the user's interaction is one or more actions taken by the user in response to that sound. Example interactions of the user include generating audio (e.g., talking), performing body postures, other movements of the user, or the user not moving. As described below in connection with Figures 2 to 4Further described, the controller 150 also identifies attributes associated with each interaction. Attributes associated with an interaction describe characteristics of a local area for the interaction and may describe actions made by the user during the interaction. For example, the controller 150 identifies the user's reaction to a sound and determines whether the source of the sound is within the field of view of the head-mounted device 100. As another example, the controller 150 determines the user's reaction to a sound and the position of the sound relative to the head-mounted device 100. In other embodiments, the controller 150 analyzes the audio generated by the user in response to the sound to describe the audio generated by the user. Other example attributes of an interaction include: the amplitude of the sound that includes a stimulus to the user, the frequency range of the sound that includes a stimulus to the user, the position of the sound that includes a stimulus to the user relative to the head-mounted device 100, or other information that describes the sound that includes a stimulus to the user. Thus, the controller 150 determines a combination of: the user's interactions identified by the controller 150; and the attributes associated with the identified interactions.

[0058] The controller 150 estimates the user's hearing profile based on the identified interactions and associated attributes. In some embodiments, the controller 150 applies a model to the combination of the identified interactions and associated attributes, where the model outputs the user's hearing profile based on the interactions and associated attributes. The model may also input stimuli in the local area. In other embodiments, the controller 150 maintains a set of rules: the set of rules maps a particular combination of one or more interactions and associated attributes to different hearing profiles or multiple parts of a hearing profile. A hearing profile describes the degree of the user's hearing loss and the configuration of the user's hearing loss (e.g., the frequency bands in which the user has hearing impairment). For example, the degree of hearing loss may include 0 (indicating no hearing loss), 1 (indicating mild hearing loss), 2 (indicating moderate hearing loss), or 3 (indicating severe hearing loss). Other gradings may be used, such as a continuous range. The hearing profile may also grade the degree of hearing loss for each ear in different frequency ranges, different environments, or some combination thereof. In some embodiments, the hearing profile also identifies one or more functional impacts of the hearing loss on the user. Example functional impacts of hearing loss include speech intelligibility deficits, or increased listening effort, self-perceived hearing loss. In some embodiments, the hearing profile is an audiogram that describes the user's hearing according to frequency, thereby allowing identification of the user's hearing loss for different frequencies. In other embodiments, the hearing profile is a response to a hearing survey, a measure of speech intelligibility, a measure of listening effort, or a combination thereof.

[0059] In some embodiments, controller 150 determines one or more filters to be applied to the audio presented to the user based on the hearing profile. In some embodiments, the audio controller can be part of controller 150. For example, the audio controller can determine one or more of the plurality of filters to be applied to the presented audio. The one or more filters enhance the audio, for example, to compensate for deficiencies in the user's hearing identified through the hearing profile. For example, if the hearing profile identifies a particular frequency range that the user has difficulty hearing, the filter amplifies the audio in that particular frequency range relative to other frequencies. By applying the determined one or more filters to the audio, the audio system generates enhanced audio for presentation to the user. The enhanced audio is provided to one or more speakers 160 or to one or more tissue transducers 170 for presentation to the user.

[0060] Position sensor 190 generates one or more measurement signals in response to the movement of the head-mounted device 100. Position sensor 190 can be located on a portion of the frame 110 of the head-mounted device 100. Position sensor 190 can include an inertial measurement unit (IMU). Examples of position sensor 190 include: one or more accelerometers, one or more gyroscopes, one or more magnetometers, another suitable type of sensor for detecting movement, a type of sensor for error correction of the IMU, or some combination thereof. Position sensor 190 can be located outside the IMU, inside the IMU, or some combination thereof.

[0061] In some embodiments, the head-mounted device 100 can provide simultaneous localization and mapping (SLAM) for the position of the head-mounted device 100 and the update of the model of the local area. For example, the head-mounted device 100 can include a passive camera assembly (PCA) that generates color image data. The PCA can include one or more red, green, and blue (RGB) cameras that capture images of some or all of the local area. In some embodiments, some or all of the imaging devices 130 of the DCA can also be used as the PCA. The images captured by the PCA and the depth information determined by the DCA can be used to determine the parameters of the local area, generate a model of the local area, update the model of the local area, or some combination thereof. In addition, the position sensor 190 tracks the position (e.g., location and orientation) of the head-mounted device 100 within the room. Additional details regarding the various components of the head-mounted device 100 are discussed below in conjunction with Figure 5 Additional details regarding the various components of the head-mounted device 100 are discussed.

[0062] Figure 1B is a perspective view of a head-mounted device 105 implemented as an HMD according to one or more embodiments. In embodiments describing an AR system and / or an MR system, a portion of the front of the HMD is at least partially transparent in the visible band (from approximately 380 nanometers (nm) to 750 nm), and a portion of the HMD between the front of the HMD and the user's eyes is at least partially transparent (e.g., a partially transparent electronic display). The HMD includes a front rigid body 115 and a strap 175. The head-mounted device 105 includes many of the same components as described above with reference to Figure 1A many of the same components of the plurality of identical components described, but these components have been modified to be integrated with the HMD form factor. For example, the HMD includes a display assembly, a DCA, an audio system, a controller 150, and a position sensor 190. Figure 1B shows an illuminator 140, a plurality of speakers 160, a plurality of imaging devices 130, a plurality of acoustic sensors 180, and a position sensor 190. These speakers 160 may be located in various positions, such as being coupled to the strap 175 (as shown), being coupled to the front rigid body 115, or may be configured to be inserted into the user's ear canal.

[0063] Figure 2 is a block diagram of an audio system 200 according to one or more embodiments. Figure 1A The audio system in Figure 1B or the audio system in Figure 2 may be an embodiment of the audio system 200. The audio system 200 generates one or more acoustic transfer functions for the user. Then, the audio system 200 may use the one or more acoustic transfer functions to generate audio content for the user. In embodiments of

[0064] The transducer array 210 is configured to present audio content. The transducer array 210 includes a plurality of transducers. A transducer is a device that provides audio content. The transducer can be, for example, a speaker (e.g., speaker 160), a tissue transducer (e.g., tissue transducer 170), some other device that provides audio content, or some combination thereof. The tissue transducer can be configured to function as a bone conduction transducer or a cartilage conduction transducer. The transducer array 210 can present audio content by: air conduction (e.g., via one or more speakers), bone conduction (via one or more bone conduction transducers), a cartilage conduction audio system (via one or more cartilage conduction transducers), or some combination thereof. In some embodiments, the transducer array 210 can include one or more transducers to cover different portions of the frequency range. For example, piezoelectric transducers can be used to cover a first portion of the frequency range, and dynamic coil transducers can be used to cover a second portion of the frequency range.

[0065] A bone conduction transducer generates a sound pressure wave by vibrating the bone / tissue of the user's head. The bone conduction transducer can be coupled to a portion of the head-mounted device and can be configured to be behind the auricle coupled to a portion of the user's skull. The bone conduction transducer receives vibration instructions from the audio controller 230 and vibrates a portion of the user's skull based on the received instructions. The vibration from the bone conduction transducer generates a tissue-borne sound pressure wave that travels around the eardrum towards the user's cochlea.

[0066] A cartilage conduction transducer generates a sound pressure wave by vibrating one or more portions of the ear cartilage of the user's both ears. The cartilage conduction transducer can be coupled to a portion of the head-mounted device and can be configured to be coupled to one or more portions of the ear cartilage of the ear. For example, the cartilage conduction transducer can be coupled to the rear portion of the auricle of the user's ear. The cartilage conduction transducer can be located anywhere along the ear cartilage around the outer ear (e.g., auricle, tragus, some other portion of the ear cartilage, or some combination thereof). Vibrating one or more portions of the ear cartilage can generate: an air-borne sound pressure wave outside the ear canal; a tissue-borne sound pressure wave that causes certain portions of the ear canal to vibrate and thereby generate an air-borne sound pressure wave within the ear canal; or some combination thereof. The generated air-borne sound pressure wave travels down the ear canal towards the tympanic membrane.

[0067] The transducer array 210 generates audio content according to instructions from the audio controller 230. In some embodiments, the audio content is spatialized. The spatialized audio content is audio content that appears to originate from a specific direction and / or target area (e.g., an object and / or virtual object in a local area). For example, the spatialized audio content can make the sound appear to originate from a virtual singer located at the other end of the user's room in the audio system 200. The transducer array 210 can be coupled to a wearable device (e.g., the head-mounted device 100 or the head-mounted device 105). In alternative embodiments, the transducer array 210 can be multiple speakers separate from the wearable device (e.g., coupled to an external console).

[0068] The sensor array 220 detects sounds within a local area around the sensor array 220. The sensor array 220 can include multiple acoustic sensors, each of which detects the air pressure change of sound waves and converts the detected sound into an electronic format (analog format or digital format). The multiple acoustic sensors can be positioned on a head-mounted device (e.g., the head-mounted device 100 and / or the head-mounted device 105), on the user (e.g., in the user's ear canal), on a neckband, or some combination thereof. The acoustic sensors can be, for example, microphones, vibration sensors, accelerometers, or any combination thereof. In some embodiments, the sensor array 220 is configured to monitor the audio content generated by the transducer array 210 using at least some of the multiple acoustic sensors. Increasing the number of sensors can improve the accuracy of information (e.g., directivity) describing the sound field generated by the transducer array 210 and / or the sounds from the local area.

[0069] The audio controller 230 controls the operation of the audio system 200. In Figure 2 embodiments, the audio controller 230 includes a data repository 235, a DOA estimation module 240, a transfer function module 250, a tracking module 260, a beamforming module 270, a sound filter module 280, and a hearing loss estimation module 290. In one or more embodiments, the audio controller 230 can be a controller of a head-mounted device (e.g., Figure 1A the head-mounted device 100 in Figure 1Ba part of the controller 150 of the head-mounted device 105 in []. In some embodiments, the audio controller 230 may be located inside the head-mounted device. Some embodiments of the audio controller 230 have components different from those described herein. Similarly, the various functions may be distributed among the components in a manner different from that described herein. For example, some functions of the controller may be performed outside the head-mounted device. The user may opt in to allow the audio controller 230 to transmit data collected by the head-mounted device to a system outside the head-mounted device, and the user may select privacy settings to control access to any such data.

[0070] The data repository 235 stores data for use by the audio system 200. The data in the data repository 235 may include: sounds recorded in a local area of the audio system 200; audio content; head-related transfer function (HRTF); transfer functions for one or more sensors; array transfer function (ATF) for one or more of the acoustic sensors; sound source location; virtual model of the local area; direction-of-arrival estimation results; sound filters; and other data related to the use of the audio system 200; or any combination thereof.

[0071] In various embodiments, the data repository 235 includes a trained model that outputs a user's hearing profile based on one or more interactions and attributes associated with each of the one or more interactions. The training of the model will be further described below. Alternatively or additionally, the data repository 235 includes a set of rules that associate interactions and attributes of the interactions with a hearing profile or a part of a hearing profile. The rules include the association between one or more interactions and attributes associated with the one or more interactions and a hearing profile (or a part of a hearing profile). In addition, the data repository 235 includes the association between the identifier of the user and the estimated hearing profile of the user, as further described below.

[0072] The user may opt in to allow the data repository 235 to record data collected by the audio system 200. In some embodiments, the audio system 200 may employ always on recording, where the audio system 200 records all sounds collected by the audio system 200 to improve the user experience. The user may opt in or opt out to allow or prevent the audio system 200 from recording, storing, or transmitting the recorded data to other entities.

[0073] The DOA estimation module 240 is configured to localize sound sources in a local area based in part on information from the sensor array 220. Localization is the process of determining where a sound source is located relative to a user of the audio system 200. The DOA estimation module 240 performs DOA analysis to localize one or more sound sources within the local area. The DOA analysis may include: analyzing the intensity, spectrum, and / or arrival time of each sound at the sensor array 220 to determine from which direction the sound originated. In some cases, the DOA analysis may include any suitable algorithm for analyzing the surrounding acoustic environment in which the audio system 200 is located.

[0074] For example, the DOA analysis may be designed to receive input signals from the sensor array 220 and apply digital signal processing algorithms to the input signals to estimate the direction of arrival. These algorithms may include, for example, the delay and sum algorithm, in which the input signals are sampled and the final weighted and delayed versions of the sampled signals are averaged together to determine the DOA. The least mean squared (LMS) algorithm may also be implemented to create an adaptive filter. Then, the adaptive filter may be used to identify, for example, differences in signal strength or differences in arrival time. Then, these differences may be used to estimate the DOA. In another embodiment, the DOA may be determined by transforming the input signals into the frequency domain and selecting specific bins within the time-frequency (TF) domain to be processed. Each selected TF bin may be processed to determine whether the bin includes a portion of the audio spectrum with a direct path audio signal. Then, those bins with a portion of the direct path signal may be analyzed to identify the angle at which the sensor array 220 received the direct path audio signal. Then, the determined angles may be used to identify the DOA of the received input signals. Other algorithms not listed above may also be used alone or in combination with the above algorithms to determine the DOA.

[0075] In some embodiments, the DOA estimation module 240 may also determine the DOA relative to the absolute position within the local area of the audio system 200. The position of the sensor array 220 may be received from an external system (e.g., certain other components of a head-mounted device, an artificial reality console, a map building server, a position sensor (e.g., position sensor 190), etc.). The external system may create a virtual model of the local area in which the local area and the position of the audio system 200 are drawn. The received position information may include the location and / or orientation of some or all of the audio system 200 (e.g., the sensor array 220). The DOA estimation module 240 may update the estimated DOA based on the received position information.

[0076] The transfer function module 250 is configured to generate one or more acoustic transfer functions. Generally, a transfer function is a mathematical function that gives a corresponding output value for each possible input value. The transfer function module 250 generates one or more acoustic transfer functions associated with the audio system based on the parameters of the detected sound. The acoustic transfer function can be an array transfer function (ATF), a head-related transfer function (HRTF), other types of acoustic transfer functions, or some combination thereof. The ATF characterizes how a microphone receives sound from a point in space.

[0077] The ATF includes a number of transfer functions that characterize the relationship between a sound source and the corresponding sound received by each acoustic sensor in the sensor array 220. Thus, for a sound source, there is a corresponding transfer function for each acoustic sensor in each acoustic sensor of the sensor array 220. And this set of transfer functions is collectively referred to as the ATF. Thus, for each sound source, there is a corresponding ATF. Note that the sound source can be, for example, a person or something that produces sound in a local area, a user, or one or more transducers in the transducer array 210. Since the human anatomical structure (e.g., ear shape, shoulders, etc.) affects the sound as it travels towards a person's ear, the ATF for a particular sound source position relative to the sensor array 220 may vary from user to user. Thus, these ATFs of the sensor array 220 are personalized for each user of the audio system 200.

[0078] In some embodiments, the transfer function module 250 determines one or more HRTFs for a user of the audio system 200. The HRTF characterizes how the ear receives sound from a point in space. Since the human anatomical structure (e.g., ear shape, shoulders, etc.) affects the sound as it travels towards a person's ear, the HRTF for a particular source position relative to a person is unique for each ear of that person (and thus unique for that person). In some embodiments, the transfer function module 250 can use a calibration process to determine the HRTF for the user. In some embodiments, the transfer function module 250 can provide information about the user to a remote system. The user can adjust the privacy settings to allow or prevent the transfer function module 250 from providing information about the user to any remote system. The remote system uses, for example, machine learning to determine a customized set of HRTFs for the user and provides the customized set of HRTFs to the audio system 200.

[0079] The tracking module 260 is configured to track the location of one or more sound sources. The tracking module 260 can compare multiple current DOA estimation results and compare these current DOA estimation results with the stored history of previous DOA estimation results. In some embodiments, the audio system 200 can recalculate the DOA estimation results according to a periodic schedule (e.g., once per second or once per millisecond). The tracking module can compare the current DOA estimation results with the previous DOA estimation results, and the tracking module 260 can determine that the sound source has moved in response to a change in the DOA estimation result of the sound source. In some embodiments, the tracking module 260 can detect a change in location based on received visual information from a head-mounted device or some other external source. The tracking module 260 can track the movement of one or more sound sources over time. The tracking module 260 can store the numerical value of the sound source and the location of each sound source at each time point. The tracking module 260 can determine that the sound source has moved in response to a change in the numerical value or location of the sound source. The tracking module 260 can calculate an estimate of the localization variance. The localization variance can be used as the confidence level for each determination of a movement change.

[0080] The beamforming module 270 is configured to process one or more ATFs to selectively highlight the sound from sound sources within a certain area while weakening the sound from other areas. When analyzing the sound detected by the sensor array 220, the beamforming module 270 can combine the information from different acoustic sensors to highlight the sound associated with a specific area of the local area while weakening the sound from outside that area. The beamforming module 270 can, for example, separate the audio signal associated with the sound from a specific sound source from other sound sources in the local area based on different DOA estimation results from the DOA estimation module 240 and the tracking module 260. Thus, the beamforming module 270 can selectively analyze discrete sound sources in the local area. In some embodiments, the beamforming module 270 can enhance the signal from a sound source. For example, the beamforming module 270 can apply sound filters that eliminate signals above certain frequencies, below certain frequencies, or between certain frequencies. The signal enhancement is used to enhance the sound associated with a given identified sound source relative to other sounds detected by the sensor array 220.

[0081] The sound filter module 280 determines a sound filter for the transducer array 210. In some embodiments, the sound filter spatializes the audio content such that the audio content appears to originate from a target region. The sound filter module 280 may use HRTF and / or acoustic parameters to generate the sound filter. These acoustic parameters describe the acoustic characteristics of the local region. The acoustic parameters may include, for example, reverberation time, reverberation level, room impulse response, etc. In some embodiments, the sound filter module 280 calculates one or more of the plurality of acoustic parameters. In some embodiments, the sound filter module 280 requests acoustic parameters from a map building server (e.g., as described below with respect to Figure 5 ). As further described below in connection with Figure 3 and Figure 4 , in various embodiments, the sound filter module 280 determines one or more filters to be applied to the audio based on the hearing profile determined for the user by the hearing loss estimation module 290.

[0082] The sound filter module 280 provides the sound filter to the transducer array 210. In some embodiments, the sound filter may amplify the sound positively or negatively according to frequency.

[0083] The hearing loss estimation module 290 receives the audio collected by the sensor array 220 and may also receive other data from additional sensors, where the other data includes other information describing the user, information describing the local region around the audio system 200, or some combination thereof. In some embodiments, the additional sensors may be included in the sensor array 220. One or more of the plurality of additional sensors may be separate from the audio system 200, for example, included in one or more separate devices, as further described above in connection with Figure 1A . The additional sensors may be an imaging device 130, a position sensor 190, or other sensors that collect information describing the user or the local region. The hearing loss estimation module 290 identifies the interaction of the user of the audio system 200 (e.g., the user wearing the head-mounted device 100 including the audio system 200) within the local region, where the user's interaction is an action taken by the user in response to a stimulus within the local region. Examples of user interactions are further described below in connection with Figure 3 . The hearing loss estimation module 290 identifies the attributes of each interaction, which may take into account the conditions of the local region during the interaction or the actions taken by the user, as further described below in connection with Figure 3 .

[0084] The hearing loss estimation module 290 estimates the user's hearing profile based on the identified interactions and the attributes associated with these interactions. In some embodiments, the audio controller 230 applies a model to a combination of the identified interactions and the associated attributes, where the model outputs the user's hearing profile based on these interactions and the associated attributes. In various embodiments, the model can be a trained machine learning model, where the machine learning model receives one or more combinations of the identified user interactions and the corresponding attributes and outputs the user's hearing profile. In some embodiments, the model determines an embedding for each combination of the identified interactions and attributes and estimates the hearing profile based on the determined embeddings. For example, the model is a set of weights that includes parameters for determining the hearing profile corresponding to one or more combinations of the interactions and the associated attributes, so that these weights transform the input data received by the model into output data. The weights can be generated through a training process in which the model is trained based on a set of training examples and the labels associated with these training examples. The training examples include combinations of interactions and associated attributes and the following labels: the label is the hearing profile associated with the combination of the interaction and the associated attributes. The training process of the model can include: applying the model to the training examples, generating a score of the model by comparing the output of the model with the label of the training example, and updating the weights associated with the model based on the score through a backpropagation process.

[0085] In other embodiments, the hearing loss estimation module 290 or the data repository 235 maintains a set of rules that map a specific combination of an interaction and the associated attributes to a hearing profile or multiple parts of a hearing profile. The hearing profile describes the degree of the user's hearing loss and the configuration of the user's hearing loss (e.g., the frequency bands in which the user has hearing impairment). In some embodiments, the hearing profile is an audiogram that describes the user's hearing according to frequency, thus allowing identification of the user's hearing loss for different frequencies. In other embodiments, the hearing profile is an answer to a hearing survey, a measure of speech intelligibility, a measure of listening effort, or a combination thereof.

[0086] In some embodiments, the hearing loss estimation module 290 may store the hearing profile generated for the user in the data repository 235. As described above, the hearing profile may include the degree of hearing loss of the user. In a later time period, the hearing loss estimation module 290 may use one or more of the above techniques (e.g., using a trained machine learning model, using one or more rule sets, etc.) to update the user's hearing profile based on subsequent interactions and their attributes. The hearing loss estimation module 290 may also implement one or more feedback loops based on the enhanced audio provided to the user. For example, the sound filter module 280 may apply a sound filter generated based on the estimated hearing profile to present enhanced audio content to the user. The user may interact in such a way as to notify adjustments to the hearing profile and / or the sound filter. For example, if the user is still turning up the volume, the hearing loss estimation module 290 may infer that further adjustment of the hearing profile is needed.

[0087] In some embodiments, the hearing loss estimation module 290 may be part of a more general controller. The controller may receive data including the following information through the hearing loss estimation module 290 to determine the user's hearing profile: information describing the user; information describing the local area; or some combination thereof. In a first example, the hearing loss estimation module 290 may determine the hearing profile based at least on the user's acoustic data, and may also determine the hearing profile based on other information describing the user, information describing the local area, or some combination thereof. In a second example, the hearing loss estimation module 290 may determine the hearing profile based at least on the user's eye tracking data, and may also determine the hearing profile based on other information describing the user, information describing the local area, or some combination thereof. In a third example, the hearing loss estimation module 290 may determine the hearing profile based at least on the user's facial expression data, and may also determine the hearing profile based on other information describing the user, information describing the local area, or some combination thereof.

[0088] In various embodiments, the hearing loss estimation module 290 and / or the sound filter module 280 determine one or more filters to be applied to the audio presented to the user based on the hearing profile. In some embodiments, the hearing loss estimation module 290 sends the determined hearing profile of the user to the sound filter module 280, and the sound filter module 280 determines or generates one or more filters based on the hearing profile. The filters determined based on the hearing profile compensate for the deficiencies in the user's hearing identified by the hearing profile, as described below in connection with Figure 3This is further described. Thus, the hearing loss estimation module 290 or the sound filter module 280 generates enhanced audio for presentation to the user by applying one or more filters determined for the user to the audio before presenting the audio to the user. Applying the determined one or more filters enhances the audio by amplifying portions of the audio whose frequencies are within the frequency range in which the hearing profile has identified the user's hearing loss. This allows the estimated hearing profile of the user to compensate for the user's hearing loss when presenting audio to the user subsequently.

[0089] Figure 3 is a flowchart of a method for estimating the hearing loss of a user of the head-mounted device 100 according to one or more embodiments. Figure 3 The process shown in can be performed by a head-mounted device (e.g., Figure 1A the head-mounted device 100 in and / or Figure 1B the head-mounted device 105 in), or can be performed by components of an audio system (e.g., the audio system 200). In other embodiments, other entities can perform Figure 3 some or all of the multiple steps in. Various embodiments can include different and / or additional steps, or perform these steps in a different order.

[0090] In various embodiments, the head-mounted device 100 (e.g., as described above in connection with Figure 1A and Figure 1B includes a controller and one or more sensors for collecting data. According to method 300, the one or more sensors collect 310 information describing the user, information describing the local area, or some combination thereof. In one or more embodiments, the one or more sensors include the acoustic sensor 180, which collects audio from the user and / or the local area.

[0091] In addition, one or more additional sensors collect information describing a local area. For example, the additional sensor is an imaging device 130, and the imaging device 130 collects an image or video of a local area within the field of view of the imaging device 130. As another example, the additional sensor is a position sensor 190, and the position sensor 190 determines the movement of a part of the user within the local area. In some embodiments, the position sensor 190 is in the same device as the controller 150. For example, the controller 150 and the position sensor 190 are included in the head-mounted device 100. Alternatively, the position sensor 190 is located in a device different from the controller 150. For example, the position sensor 190 is located in a wearable device (such as a smartwatch, etc.) worn by the user, while the controller 150 is located in the head-mounted device 100 worn by the user. Different position sensors 190 may be included in different devices at different positions on the user's body, thereby allowing collection of information describing the movement or position of different parts of the user's body. One or more additional sensors may be included in the device including the controller 150, and other additional sensors are included in one or more different devices. In various embodiments, different additional sensors collect different types of information describing the local area, different types of information describing the user, or some combination thereof. The controller 150 receives various information.

[0092] The controller 150 identifies 315 one or more interactions of the user with the local area from the collected audio and the information describing the local area. The user's interaction is an action taken by the user in response to a stimulus within the local area. Example interactions of the user include generating audio (such as talking), making a body gesture (such as moving or repositioning one or more parts of the user's body). In various embodiments, the stimulus within the local area is a sound within the local area, and the interaction is the action of the user adjusting the volume of the source of the sound. For example, the controller 150 identifies 315 the user's interaction in response to the collected audio including a word or phrase from the user requesting an increase in the sound volume, or in response to detecting an increase in the sound amplitude after the collected audio includes audio from the user. As another example, the controller 150 identifies 315 the user's interaction in response to the collected audio from the user including a word or phrase associated with a request for clarification of the content. In another example, the controller 150 identifies 315 the interaction in response to the audio collected within at least a threshold amount of time after the collected audio includes a sound from the local area not including audio from the user (indicating the user's lack of response to the sound).

[0093] In other embodiments, the controller 150 identifies 315 interactions based on indications of the user's social activities. For example, the controller 150 identifies 315 the user's interaction in response to information describing the local area including one or more specific postures or movements of the user. A specific posture may be a movement of a part of the user's body towards a sound source in the local area, a specific movement of a specific part of the user's body (e.g., the user pinching the user's ear with a hand).

[0094] One or more interactions identified 315 by the controller 150 are between the user and one or more other users in the local area. For example, the controller 150 identifies that the user is talking to the additional user based on the collected audio including audio from the user and from the additional user. For an interaction with an additional user, the controller 150 identifies the duration of the conversation between the user and the additional user and the attributes of the audio exchanged between the user and the additional user, as further described below. In some embodiments, the controller 150 determines the number of times the user talks to an additional user in the local area.

[0095] When the controller 150 identifies 315 an interaction between the user and a sound source (e.g., an additional user), the controller 150 identifies the attributes of the interaction from the collected audio and the information describing the local area. Example attributes of an interaction with an additional user include the depth level of the conversation (which depth level is based on a topic extracted from the collected audio and the duration of the collected audio (including audio from the user and the additional user)), the category of the additional user (e.g., the user's friend, the user's family member, the user's spouse, a stranger, etc.), the identity of the additional user (determined based on the collected audio from the additional user, from video including the additional user, from the identifier of the additional user's head-mounted device 100, etc.), or other information describing the additional user involved in the interaction. As further described below, the controller 150 may process the collected audio from the user and the additional user (or other sources) to identify one or more attributes associated with the interaction.

[0096] In some embodiments, the properties of the interaction between a user and a sound source in a local area (e.g., an additional user) include one or more turn-taking metrics of the interaction. Turn-taking metrics describe the user's conversation activities. For example, a turn-taking metric identifies the length of time that the user produces audio during the interaction or the length of time that the sound source (e.g., an additional user) produces audio during the interaction. As another example, a turn-taking metric identifies the percentage of time that the user produces audio during the interaction or the percentage of time that the sound source (e.g., an additional user) produces audio during the interaction. In another example, a turn-taking metric identifies the amount of time (e.g., average amount of time, median amount of time) between the sound source (e.g., an additional user) producing audio during the interaction and the user producing audio during that interaction (or vice versa). Additional turn-taking metrics can be determined based on the audio produced by the user and the sound source (e.g., an additional user) during the interaction.

[0097] In various embodiments, the properties of the interaction between a user and an additional user (or other sound source) include data describing repair initiation. Repair initiation is determined based on words or phrases identified in the collected audio in which the user requests repair of the interaction, where the repair indicates that the user requests clarification of the audio from the sound source (e.g., an additional user). For example, the properties of the interaction specify the number of repair initiations made by the user during the interaction or the frequency with which repair initiations occur by the user during the interaction. In various embodiments, the controller 150 identifies a repair initiation for an interaction by: identifying the audio from the user collected during the interaction; extracting words or phrases from the collected audio from the user; and comparing the extracted words or phrases with the stored data corresponding to repair initiation (e.g., phrases, words, grammar). The controller 150 identifies a repair initiation when the extracted words or phrases match the stored data and uses the identified one or more repair initiations to determine one or more properties of the interaction. The controller 150 can also specify whether the repair initiation is open-ended or closed-ended. Closed-ended repair initiations include one or more indications of the topic that the user is tracking in the conversation and / or the latest statement of another user. Open-ended repair initiations lack any indication of the topic that the user is tracking and / or the latest statement of another user.

[0098] As another example, the attributes of the interaction between the user and the sound source include one or more semantic metrics. In various embodiments, the semantic metrics are based on the topics determined for the audio collected from the user. In some embodiments, the semantic metrics also take into account the topics of the audio collected from the sound source. In various embodiments, the controller 150 determines the semantic metrics of the interaction by applying a natural language model to the audio collected from the user, where the natural language model outputs the topic of the audio collected from the user. Additionally, the controller 150 applies the natural language module to the audio collected from the sound source (e.g., an additional user) involved in the interaction to determine the topic of the audio collected from that sound source. In various embodiments, the semantic metrics are determined based on the difference between the topic of the audio collected from the user and the topic of the audio collected from the sound source. For example, the semantic metric is a measure of the similarity between the topic of the audio collected from the user and the topic of the audio collected from the sound source. Thus, in various embodiments, the semantic metric represents a measure of the similarity between the topic of the audio collected from the user and the topic of the audio collected from the sound source involved in the interaction, thereby allowing the semantic metric to indicate whether the topic of the user's audio deviates from the topic of the sound source's audio. In various embodiments, the controller 150 determines the semantic metrics for different combinations of the user and different sound sources involved in the identified interaction. In other embodiments, the controller 150 determines the semantic metrics between the topic determined for the audio collected from the user and a reference value or reference topic.

[0099] The controller 150 can also learn the characteristics of the user's speech. The controller 150 can apply a natural language processing model to extract the characteristics of the user's speech. Example characteristics can include word choice, grammar tendencies, pitch variations in the user's speech, the user's accent, etc. Such characteristics can be attributes of the user interaction (i.e., the user's speech) that can be used when determining the user's hearing profile.

[0100] Additionally, when identifying the 315 interaction, the controller 150 determines the attributes of the local area based on the information describing the local area and the audio collected. For example, when identifying the 315 interaction, the attribute of the local area is the amount of ambient noise or background noise in the local area. When identifying the 315 interaction based on the video of the local area or the location information from the location sensor 190, the controller 150 can identify the type or category of the local area. Additionally or alternatively, other characteristics of the local area when identifying the 315 interaction can be determined based on the information describing the local area.

[0101] When identifying the interaction of the 315 user with the sound source, the controller 150 determines an attribute indicating whether the sound source is within the user's field of view. For example, one or more imaging devices 130 are positioned on the head-mounted device 100 (or another device) to have a field of view that at least partially overlaps with the user's field of view. The controller 130 determines whether the position of the sound source in the interaction is within the field of view of the imaging device 130, the field of view of which overlaps with the user's field of view (as an attribute associated with the interaction). This allows the audio controller 130 to determine whether the sound source involved in the interaction is visible to the user during the interaction. In various embodiments, the controller 150 stores a visibility indication associated with the identified interaction, where the visibility indication has a first value in response to the sound source being within the field of view of the imaging device 130 that overlaps with the user's field of view, and has a second value in response to the sound source not being within the field of view of the imaging device 130 that overlaps with the user's field of view.

[0102] In some embodiments, the controller 150 determines, based on the direction of arrival of the collected audio, the attribute associated with the interaction as the position of the sound source involved in the interaction relative to the controller 150, as described above in connection with Figure 2 Further described. The position of the sound source relative to the controller 150 is used as a proxy for the position of the sound source relative to the user, thereby allowing the controller 150 to identify the position of the sound source relative to the user for the identified interaction. The controller 150 stores the position of the source relative to the controller 150 in association with the identified interaction, thereby allowing the identified interaction to consider the position of the sound source relative to the controller 150 in the identified interaction (representing the position of the sound source relative to the user).

[0103] The controller 150 estimates 320 the user's hearing profile based on the identified interaction and the attribute associated with the identified interaction. In some embodiments, the controller 150 applies a model to the combination of the identified interaction and the associated attribute. As described above in connection with Figure 2 Further described, the model outputs the user's hearing profile based on the identified interaction and its associated attribute. In various embodiments, the model is a trained machine learning model, where the machine learning model receives one or more combinations of the identified interaction and the associated attribute as input and outputs the user's hearing profile. The model is a set of weights that includes parameters for determining the hearing profile corresponding to one or more combinations of the interaction and the associated attribute, where these weights are generated through a training process in which the model is applied to a set of training examples and the labels associated with these training examples. The training examples include combinations of the interaction and the associated attribute and the following label: the label is the hearing profile associated with the combination of the interaction and the associated attribute.

[0104] In various embodiments, the training process of the model includes applying the model to each training example in a set of training examples. Applying the model to a training example outputs a hearing profile based on the one or more interactions and associated attributes. The controller 150 scores the model based on the hearing profile for the training example output by the model and the label of the training example. The score is based on the difference between the hearing profile output by the model and the label applied to the training example. In various embodiments, the controller 150 generates a score for the model based on a loss function. The loss function is a function that generates a score for the model, so that when the model performs poorly, the score is high, and when the model performs well, the score is low. In various embodiments, the loss function is based on the difference between the hearing profile output by the model and the label applied to the training example. Example loss functions include mean squared error function, mean absolute error, hinge loss function, and cross-entropy loss function.

[0105] The controller 150 updates a set of parameters of the model based on the score generated by the loss function. For example, the controller 150 updates one or more parameters (e.g., weights) that make up the model by backpropagation based on the score from applying the model to the training example. The controller 150 can update one or more parameters that make up the model until one or more criteria are met. For example, the controller 150 modifies the parameters until the value of the loss function is less than a threshold. After training the model, the model is applied to the identified interactions and associated attributes to estimate the hearing profile of the user 320.

[0106] In other embodiments, the controller 150 maintains a set of rules, where each rule maps a combination of one or more identified interactions and associated attributes to multiple portions of the hearing profile. Different rules map different combinations of the identified interactions and associated attributes to different portions of the hearing profile or different hearing profiles. For example, one rule maps the following combination to a hearing profile indicating hearing loss for audio having a frequency greater than a threshold frequency: the combination of an interaction with an attribute indicating that the user has difficulty perceiving audio from a sound source and indicating that the background noise is at least a threshold level during that interaction; and an additional interaction with an attribute indicating that the user has no difficulty perceiving audio from a sound source and indicating that the background noise is less than the threshold level during that additional interaction. In another example, one rule maps the following combination to a hearing profile indicating hearing loss for audio having a frequency greater than an alternative threshold frequency (the alternative threshold frequency being lower than the threshold frequency in the previous example): the combination of an interaction with an attribute indicating that the user has difficulty perceiving audio from a sound source and indicating that the background noise is at least a threshold level during that interaction; and an additional interaction with an attribute indicating that the user has difficulty perceiving audio from a sound source and indicating that the background noise is less than the threshold level during that additional interaction. As another example, one rule maps the combination of an interaction with associated attributes indicating that the user has difficulty perceiving audio from a sound source and that the background noise is less than the threshold level during that interaction to a hearing profile indicating hearing loss for audio having a frequency range (the frequency range having an upper threshold frequency and a lower threshold frequency). In an additional example, one rule maps an interaction associated with attributes indicating that the user has difficulty perceiving audio from a sound source and indicating that the sound source is not within the user's field of view to a hearing profile indicating hearing loss for audio having a frequency range (the frequency range having an upper threshold frequency and a lower threshold frequency). In another example, one rule maps an interaction associated with the following attributes to a hearing profile indicating hearing loss for audio within a wider frequency range between a minimum frequency and a maximum frequency: the attributes indicating that the user does not respond to audio from a sound source and indicating that the sound source is not within the user's field of view.

[0107] The hearing profile describes the degree of the user's hearing loss and the user's hearing loss configuration (e.g., the frequency band or range in which the user has hearing impairment). In some embodiments, the hearing profile also identifies one or more functional impacts of the hearing loss on the user. Example functional impacts of hearing loss include speech intelligibility deficits, increased listening effort, self-perceived hearing loss. In some embodiments, the hearing profile is an audiogram that describes the user's hearing according to frequency, thereby allowing identification of the user's hearing loss for different frequencies. In other embodiments, the hearing profile is the response to a hearing survey, a measure of speech intelligibility, a measurement result of listening effort, or a combination thereof. Thus, estimating the hearing profile 320 enables the controller 150 to estimate the user's hearing loss and the user's hearing loss for different frequencies based on the identified user interactions and their associated attributes (rather than based on the user's response to specific questions or measuring the user's hearing response to different sounds presented to the user). This simplifies the estimation of the user's hearing loss while providing more information about the user's hearing loss in a wider range of scenarios than those covered by conventional hearing tests.

[0108] The controller 150 may store the estimated hearing profile 320 of the user in association with the user identifier of the user. In some embodiments, the controller 150 stores the hearing profile in association with the user's identifier in a data repository (e.g., the data repository 235 of the audio system 200). Additionally or alternatively, the controller 150 transmits the user's hearing profile and identifier to an external device (e.g., the map building server 525 described further in conjunction with Figure 5 the map building server 525) or other servers.

[0109] In various embodiments, in addition to estimating the hearing profile 320 of the user that identifies the user's hearing loss, the controller 150 modifies the subsequent audio presented to the user to mitigate the hearing loss identified by the hearing profile. To compensate for the hearing loss identified by the hearing profile, the controller 150 determines 325 one or more filters to be applied to the subsequent audio presented to the user. In some embodiments, the controller 150 identifies from the hearing profile the frequency range in which the user has hearing loss and determines 325 a filter that amplifies the identified frequency range relative to other frequencies. Thus, the determined filter amplifies the audio within the frequency range of the determined user's hearing loss. In various embodiments, the different filters determined 325 by the controller 150 correspond to different frequency ranges, so the controller 150 determines 325 a filter corresponding to each frequency range in which the user has hearing loss as identified by the hearing profile. In some embodiments, Figure 2The audio controller 230 determines the one or more filters based on the hearing profile. In other embodiments, the controller 150 includes an audio controller 230 that can utilize the sound filter module 280 to obtain the user's hearing profile and determine 325 the one or more filters based on the hearing profile. In some embodiments, the controller 150 stores the determined one or more filters in association with the user's user identifier, thereby accelerating subsequent retrieval of the filters determined for the user.

[0110] After determining 325 the one or more filters, when audio is to be presented to the user, the controller 150 generates 330 enhanced audio by applying the one or more filters to the audio to be presented to the user. Applying the one or more filters amplifies the following portions of the audio: the frequencies of these portions of the audio are within the range where the hearing profile indicates that the user has hearing loss relative to other frequencies. Thus, the enhanced audio increases the amplitude of the frequencies of the audio where the hearing profile indicates that the user experiences hearing loss.

[0111] The controller 150 presents 335 the enhanced audio to the user. For example, the controller 150 plays the enhanced audio through the transducer array 205. As an example, the speaker 160 or the transducer 170 receives the enhanced audio from the controller 150 and plays the enhanced audio to the user. When the enhanced audio amplifies the audio having frequencies where the hearing profile indicates that the user has hearing loss, the enhanced audio compensates for the user's hearing loss. This allows the user to more clearly hear the audio enhanced by the controller 150 using the hearing profile estimated 320 based on the identified user's interaction in the local area.

[0112] Figure 4 is a conceptual diagram of the head-mounted device 100 determining the hearing profile 435 of the user 405 based on information 420 describing the local area, information 422 describing the user, or some combination thereof. The information 422 describing the user can include the user's captured audio and other matters described throughout this disclosure. The information 420 describing the local area can include the captured audio from the local area and information describing the local area, other matters described throughout this disclosure. As Figure 4 shown, the user 405 wears the head-mounted device 100 as further described above in connection with Figure 1A The head-mounted device 100 includes a controller 150. In other embodiments, the controller 150 is included in a different device.

[0113] The head-mounted device 100 includes one or more acoustic sensors 180 configured to collect audio within a local area 400 as information 420 describing the local area. Alternatively or additionally, one or more acoustic sensors are separate from the head-mounted device 100 and collect audio in the local area 400 that is transmitted to the controller 150. In Figure 4 the example of, the local area 400 includes a sound source 410 that emits audio 415 into the local area 400. For example, the sound source 410 is a person, such as another user of another head-mounted device 100. As another example, the sound source 410 is a device such as a television, a mobile device, or other device capable of emitting audio. One or more acoustic sensors 180 collect the audio 415 from the local area 400.

[0114] In addition, the head-mounted device 100 includes one or more additional sensors, such as an imaging device 130 or a position sensor 190. In some embodiments, one or more additional sensors are separate from the head-mounted device 100 and communicatively coupled to the head-mounted device 100. For example, a position sensor 190 in a wearable device on a part of the user 405's body is communicatively coupled to the head-mounted device 100. The one or more additional sensors collect information 420 describing the local area, information 422 describing the user, or some combination thereof. As described above in connection with Figure 3 further described, examples of information 420 describing the local area may include video of the local area 400, the position of the head-mounted device 100 in the local area 400, the position of a part of the user's body in the local area 400, or other data describing movement of the local area 400 or the user in the local area 400. As described elsewhere in this disclosure, information 422 describing the user may include collected audio of the user 405, eye-tracking information, facial expression information, or movement of the user 405's head (or another body part).

[0115] As described above in connection with Figure 3 further described, the controller 150 of the head-mounted device 100 identifies one or more interactions of the user 405 with the local area 400 and attributes 430 associated with each identified interaction based on information 420 describing the local area, information 422 describing the user, or some combination thereof. In Figure 4 the example of, the controller 150 identifies an interaction 425 from the audio 415 and information 420 describing the local area 400, and identifies an attribute 430 corresponding to the interaction 425. As described above in connection with Figure 3 further described, the interaction 425 represents an action taken by the user 405 in response to a stimulus within the local area 400. In Figure 4In the example, interaction 425 represents an action taken by user 405 in response to audio 415 from sound source 410. As described above in connection with Figure 3 As further described, controller 150 identifies interaction 425 based on the following: movement of user 405 in the local area, audio from the user captured after audio 415 from source 410 has been captured, other data from one or more acoustic sensors 180 or one or more additional sensors, or any combination thereof. In various embodiments, controller 150 may identify any number of interactions of user 405 with the local area.

[0116] In addition, controller 150 determines an attribute 430 associated with interaction 425, allowing controller 150 to consider conditions within local area 400 or actions of the user when identifying interaction 425. As described above in connection with Figure 3 As further described, attribute 430 describes local area 400, information about user 405, information about actions made by the user during interaction 425, or a combination thereof. Example attributes 430 include the background noise level in the local area, the position of source 410 relative to head-mounted device 100, an indication of whether source 410 is within the field of view of head-mounted device 100, or other information describing conditions within local area 400. Other example attributes 430 include the duration of interaction of user 405 with sound source 410, the number of times user 405 has spoken with sound source 410, an indication of whether user 405 has understood audio 415 from source 410, or other information describing the user's reaction to audio 415 from sound source 410.

[0117] Controller 150 estimates a hearing profile 435 of user 405 based on interaction 425 and associated attribute 430. As described above in connection with Figure 3 As further described, controller 150 applies a trained model to interaction 425 and associated attribute 430, where the model outputs hearing profile 435 based on interaction 425 and associated attribute 430. In other embodiments, controller 150 estimates a hearing profile or one or more portions of hearing profile 435 based on a set of rules, where each rule in the set associates a combination of one or more interactions 425 and their associated attributes 430 with one or more portions of hearing profile 435. This allows controller 150 to utilize the identified interactions 425 of user 405 with local area 400 and associated attributes 430 to estimate hearing profile 435, which describes the degree of hearing loss of user 400 and the hearing loss configuration of user 405 (e.g., the frequency bands in which the user has hearing impairment). For example, hearing profile 435 identifies the degree of hearing loss of user 405 for different frequencies, or identifies one or more frequency ranges of audio for which user 405 has at least a threshold amount of hearing loss.

[0118] As described above in connection with Figure 3 and further described, in various embodiments, the controller 150 utilizes the hearing profile 435 to determine one or more filters to apply to audio that is subsequently presented to the user 405. The one or more filters increase the amplitude of the audio in a frequency range that is identified by the hearing profile 435 as the range in which the user 405 experiences hearing loss. Applying the one or more filters to the audio generates enhanced audio that increases the amplitude of portions of the audio (portions of the audio having frequencies at which the user 405 has hearing loss as indicated by the hearing profile 435), thereby allowing the enhanced audio to mitigate the user's hearing loss and increase the likelihood that the user understands the enhanced audio relative to the audio to which the one or more filters were not applied.

[0119] Figure 5 is a system 500 including a head-mounted device 505 according to one or more embodiments. In some embodiments, the head-mounted device 505 can be Figure 1A the head-mounted device 100 or Figure 1B the head-mounted device 105. The system 500 can operate in an artificial reality environment (e.g., a virtual reality environment, an augmented reality environment, a mixed reality environment, or some combination thereof). Figure 5 The system 500 shown includes a head-mounted device 505, an input / output (I / O) interface 510 coupled to a console 515, a network 520, and a map building server 525. Although Figure 5 an example system 500 is shown including one head-mounted device 505 and one I / O interface 510, in other embodiments, the system 500 can include any number of these components. For example, there can be multiple head-mounted devices, each having an associated I / O interface 510, where each head-mounted device and I / O interface 510 communicate with the console 515. In an alternative configuration, the system 500 can include different and / or additional components. Additionally, in some embodiments, the functions described in connection with Figure 5 one or more of the components shown in Figure 5 can be distributed among the components in a manner different from the manner described in connection with

[0120] The head-mounted device 505 includes a display component 530, an optical block 535, one or more position sensors 540, a DCA 545, an audio system 550, and a controller 555. Some embodiments of the head-mounted device 505 have components different from those described in connection with Figure 5 In addition, in other embodiments, by combining Figure 5The functions provided by the various components described can be distributed differently among the components of the head-mounted device 505 or embodied in separate components remote from the head-mounted device 505.

[0121] The display component 530 displays content to the user based on data received from the console 515. The display component 530 uses one or more display elements (e.g., display element 120) to display the content. The display element can be, for example, an electronic display. In various embodiments, the display component 530 includes a single display element or multiple display elements (e.g., one display for each eye of the user). Examples of electronic displays include: liquid crystal display (LCD), organic light emitting diode (OLED) display, active-matrix organic light-emitting diode display (AMOLED), waveguide display, some other display, or some combination thereof. Note that in some embodiments, the display element 120 can also include some or all of the functions of the optical block 535.

[0122] The optical block 535 can magnify the image light received from the electronic display, correct the optical errors associated with the image light, and present the corrected image light to one or both eyes of the head-mounted device 505. In various embodiments, the optical block 535 includes one or more optical elements. Example optical elements included in the optical block 535 include: aperture, Fresnel lens, convex lens, concave lens, filter, reflective surface, or any other suitable optical element that affects the image light. Additionally, the optical block 535 can include a combination of different optical elements. In some embodiments, one or more of the multiple optical elements in the optical block 535 can have one or more coatings, such as a partial reflection coating or an anti-reflection coating.

[0123] The magnification and focusing of the image light by the optical block 535 allows the electronic display to be physically smaller, lighter in weight, and lower in power consumption than a larger display. Additionally, the magnification can increase the field of view of the content presented by the electronic display. For example, the field of view of the content displayed is such that almost all of the user's field of view (e.g., approximately 110 degrees diagonal) is used to present the content displayed, and in some cases, all of the user's field of view is used to present the content displayed. Additionally, in some embodiments, the amount of magnification can be adjusted by adding or removing optical elements.

[0124] In some embodiments, the optical block 535 can be designed to correct for one or more types of optical errors. Examples of optical errors include barrel distortion or pincushion distortion, longitudinal chromatic aberration or lateral chromatic aberration. Other types of optical errors can also include spherical aberration; chromatic aberration; or errors due to lens field curvature, astigmatism; or any other type of optical error. In some embodiments, the content provided for display to the electronic display is pre-distorted, and the optical block 535 corrects the distortion when it receives the image light generated from the electronic display based on the content.

[0125] The position sensor 540 is an electronic device that generates data indicating the position of the head-mounted device 505. The position sensor 540 generates one or more measurement signals in response to the movement of the head-mounted device 505. The position sensor 190 is an embodiment of the position sensor 540. Examples of the position sensor 540 include: one or more IMUs, one or more accelerometers, one or more gyroscopes, one or more magnetometers, another suitable type of sensor for detecting movement, or some combination thereof. The position sensor 540 can include multiple accelerometers for measuring translational movement (forward / backward, up / down, left / right), and multiple gyroscopes for measuring rotational movement (e.g., pitch, yaw, roll). In some embodiments, the IMU samples the measurement signals quickly and calculates the estimated position of the head-mounted device 505 based on the sampled data. For example, the IMU integrates the measurement signals received from the accelerometer over time to estimate the velocity vector, and integrates the velocity vector over time to determine the estimated position of a reference point on the head-mounted device 505. The reference point is a point that can be used to describe the position of the head-mounted device 505. Although the reference point can generally be defined as a point in space, in fact, the reference point is defined as a point within the head-mounted device 505.

[0126] The DCA 545 generates depth information for a portion of a local area. The DCA includes one or more imaging devices, as well as a DCA controller. The DCA 545 can also include an illuminator. The above description Figure 1A has been made regarding the operation and structure of the DCA 545.

[0127] The audio system 550 provides audio content to a user of the head-mounted device 505. The audio system 550 is an embodiment of the audio system 200 described above. The audio system 550 may include one or more acoustic sensors, one or more transducers, and an audio controller. The audio system 550 may provide spatialized audio content to the user. In some embodiments, the audio system 550 may request acoustic parameters from the map building server 525 via the network 520. The acoustic parameters describe one or more acoustic characteristics of a local area (e.g., room impulse response, reverberation time, reverberation level, etc.). The audio system 550 may provide information describing at least a portion of the local area from, for example, the DCA 545, and / or location information of the head-mounted device 505 from the location sensor 540. The audio system 550 may use one or more of the received acoustic parameters from the map building server 525 to generate one or more sound filters and use these sound filters to provide audio content to the user.

[0128] The controller 555 controls the operation of various components of the head-mounted device 505. In one or more embodiments, the controller 555 performs a hearing loss estimate. The hearing loss estimate requires estimating a hearing profile that describes the degree of hearing loss of the user based on data including information describing the user, information describing the local area, or some combination thereof. The controller 555 may estimate the hearing profile by determining one or more interactions of the user and the attributes of these interactions based on the data. These interactions and attributes may be input into a trained model or otherwise applied to a set of rules to estimate the hearing profile. The controller 555 (or the audio system 550) may generate one or more sound filters based on the hearing profile to enhance the audio content to be presented to the user. In some embodiments, the one or more filters include the determined filters that increase the amplitude of the audio in the frequency range where the hearing profile indicates that the user has experienced hearing loss.

[0129] The I / O interface 510 is a device that allows a user to send action requests and receive responses from the console 515. An action request is a request to perform a specific action. For example, an action request can be an instruction to start or end the acquisition of image data or video data, or an instruction to perform a specific action within an application. The I / O interface 510 may include one or more input devices. Example input devices include: a keyboard, a mouse, a game controller, or any other suitable device for receiving action requests and transmitting these action requests to the console 515. The action requests received by the I / O interface 510 are transmitted to the console 515, which performs the action corresponding to the action request. In some embodiments, the I / O interface 510 includes an IMU that acquires calibration data indicating an estimated position of the I / O interface 510 relative to an initial position of the I / O interface 510. In some embodiments, the I / O interface 510 may provide haptic feedback to the user according to instructions received from the console 515. For example, haptic feedback is provided when an action request is received, or when the console 515 performs an action, the console 515 transmits an instruction to the I / O interface 510, causing the I / O interface 510 to generate haptic feedback.

[0130] The console 515 provides content for processing to the head-mounted device 505 based on information received from one or more of the following: the DCA 545; the head-mounted device 505; and the I / O interface 510. In Figure 5 the example shown, the console 515 includes an application repository 560, a tracking module 565, and an engine 570. Some embodiments of the console 515 have modules or components different from those described in connection with Figure 5 the modules or components described. Similarly, the functions further described below may be distributed among the components of the console 515 in a manner different from the manner described in connection with Figure 5 the description. In some embodiments, the functions described herein with respect to the console 515 may be implemented in the head-mounted device 505 or a remote system.

[0131] The application repository 560 stores one or more applications for the console 515 to execute. An application is a set of instructions that, when executed by a processor, generates content for presentation to a user. The content generated by an application may be in response to input received from a user via movement of the head-mounted device 505 or the I / O interface 510. Examples of applications include: a game application, a conferencing application, a video playback application, or other suitable applications.

[0132] Tracking module 565 uses information from DCA 545, one or more position sensors 540, or some combination thereof to track the movement of the head-mounted device 505 or the I / O interface 510. For example, tracking module 565 determines the position of a reference point of the head-mounted device 505 within a drawn local area based on information from the head-mounted device 505. Tracking module 565 may also determine the position of an object or virtual object. Additionally, in some embodiments, tracking module 565 may use a portion of the data from position sensor 540 indicating the position of the head-mounted device 505 and a representation of the local area from DCA 545 to predict a future position of the head-mounted device 505. Tracking module 565 provides an estimated or predicted future position of the head-mounted device 505 or the I / O interface 510 to engine 570.

[0133] Engine 570 executes an application and receives position information, acceleration information, velocity information, a predicted future position, or some combination thereof of the head-mounted device 505 from tracking module 565. Engine 570 determines the content to be provided to the head-mounted device 505 for presentation to the user based on the received information. For example, if the received information indicates that the user has looked to the left, engine 570 generates the following content for the head-mounted device 505: the content reflects the movement of the user within a virtual local area or within a local area that has been enhanced with additional content. Additionally, engine 570 performs an action within the application executed on console 515 in response to an action request received from the I / O interface 510 and provides feedback to the user that the action has been performed. The feedback provided may be visual feedback or auditory feedback via the head-mounted device 505, or tactile feedback via the I / O interface 510.

[0134] Network 520 couples the head-mounted device 505 and / or the console 515 to the map construction server 525. Network 520 can include any combination of local area networks and / or wide area networks that use both wireless communication systems and / or wired communication systems. For example, network 520 can include the Internet and mobile phone networks. In one embodiment, network 520 uses standard communication technologies and / or protocols. Thus, network 520 can include links that use technologies such as Ethernet, 802.11, Worldwide Interoperability for Microwave Access (WiMAX), 2G / 3G / 4G mobile communication protocols, Digital Subscriber Line (DSL), Asynchronous Transfer Mode (ATM), InfiniBand, PCI Express Advanced Switching, etc. Similarly, the networking protocols used on network 520 can include Multiprotocol Label Switching (MPLS), Transmission Control Protocol / Internet Protocol (TCP / IP), User Datagram Protocol (UDP), Hypertext Transfer Protocol (HTTP), Simple Mail Transfer Protocol (SMTP), File Transfer Protocol (FTP), etc. Data exchanged over network 520 can be represented using technologies and / or formats such as image data in binary form (e.g., Portable Network Graphics (PNG)), Hypertext Markup Language (HTML), Extensible Markup Language (XML), etc.In addition, conventional encryption techniques can be used to encrypt all or some of the links. These conventional encryption techniques are, for example, Secure Sockets Layer (SSL), Transport Layer Security (TLS), Virtual Private Network (VPN), Internet Protocol Security (IPsec), etc.

[0135] The map building server 525 can include a database storing a virtual model describing multiple spaces, where a location in the virtual model corresponds to the current configuration of a local area of the head-mounted device 505. The map building server 525 receives, via the network 520, information describing at least a part of the local area and / or location information of the local area from the head-mounted device 505. The user can adjust the privacy settings to allow or prevent the head-mounted device 505 from sending information to the map building server 525. The map building server 525 determines the location in the virtual model associated with the local area of the head-mounted device 505 based on the received information and / or location information. The map building server 525 determines (e.g., retrieves) one or more acoustic parameters associated with the local area, at least in part based on the determined location in the virtual model and any acoustic parameters associated with the determined location. The map building server 525 can send the location of the local area and the values of any acoustic parameters associated with the local area to the head-mounted device 505.

[0136] One or more components of the system 500 can include a privacy module that stores one or more privacy settings for user data elements. The user data elements describe the user or the head-mounted device 505. For example, the user data elements can describe the physical characteristics of the user, the actions made by the user, the location of the user of the head-mounted device 505, the location of the head-mounted device 505, the HRTF for the user, etc. The privacy settings (or "access settings") for the user data elements can be stored in any suitable manner, such as storing them associated with the user data elements, storing them in an index on an authorization server, storing them in another suitable manner, or any suitable combination thereof.

[0137] Privacy settings for user data elements specify how the user data elements (or specific information associated with the user data elements) can be accessed, stored, or otherwise used (e.g., viewed, shared, modified, copied, executed, displayed, or identified). In some embodiments, the privacy settings for user data elements can specify a “blacklist” of entities that may not have access to certain information associated with the user data elements. The privacy settings associated with user data elements can specify any suitable granularity of access allowed or denied. For example, some entities may have the right to ascertain the existence of a particular user data element, some entities may have the right to view the content of a particular user data element, and some entities may have the right to modify a particular user data element. The privacy settings can permit the user to allow other entities to access or store the user data elements for a limited period of time.

[0138] The privacy settings can allow the user to specify one or more geographical locations from which the user data elements can be accessed. Access to or denial of access to the user data elements can depend on the geographical location of the entity attempting to access the user data elements. For example, the user can allow access to the user data elements and specify that the user data elements are only accessible to an entity when the user is in a particular location. If the user leaves that particular location, the user data elements may no longer be accessible to that entity. As another example, the user can specify that the user data elements are only accessible to entities within a threshold distance of the user (e.g., another user of a head-mounted device within the same local area as the user). If the user then changes location, the entities that had access to the user data elements may lose access, and a new set of entities can gain access when they come within the user's threshold distance.

[0139] System 500 can include one or more authorization / privacy servers for implementing the privacy settings. A request from an entity for a particular user data element can identify the entity associated with the request, and the user data element can be sent to the entity only if the authorization server determines, based on the privacy settings associated with the user data element, that the entity is authorized to access the user data element. If the requesting entity is not authorized to access the user data element, the authorization server can prevent the requested user data element from being retrieved or can prevent the requested user data element from being sent to the entity. Although this disclosure describes implementing the privacy settings in a particular manner, this disclosure contemplates implementing the privacy settings in any suitable manner.

[0140] Additional configuration information

[0141] The above description of the embodiments has been presented for illustrative purposes and is not intended to be exhaustive or to limit the patent rights to the exact forms disclosed. Those skilled in the relevant art will appreciate that many modifications and variations are possible in light of the above disclosure.

[0142] Some portions of this specification describe multiple embodiments in terms of algorithms and symbolic representations of operations on information. These algorithmic descriptions and representations are commonly used by those skilled in the data processing arts to effectively convey the substance of their work to others skilled in the art. Although these operations are described functionally, computationally, or logically, they are to be understood as being implemented by a computer program or equivalent circuitry, microcode, etc. Additionally, it has sometimes proven convenient to refer to the arrangement of these operations as modules, without loss of generality. The described operations and their associated modules can be embodied in software, firmware, hardware, or any combination thereof.

[0143] Any steps, operations, or processes described herein can be executed or implemented singly or in combination with other devices by one or more hardware or software modules. In one embodiment, a software module is implemented using a computer program product that includes a computer-readable medium containing computer program code that can be executed by a computer processor to perform any or all of the described steps, operations, or processes.

[0144] The embodiments can also relate to an apparatus for performing the operations herein. Such an apparatus can be specially constructed for the required purposes, and / or such an apparatus can include a general-purpose computing device selectively activated or reconfigured by a computer program stored in the computer. Such a computer program can be stored in a non-transitory, tangible computer-readable storage medium or in any type of medium suitable for storing electronic instructions that can be coupled to a computer system bus. Additionally, any computing system mentioned in this specification can include a single processor or can be an architecture employing a multi-processor design for increased computing power.

[0145] The embodiments can also relate to a product produced by the computing processes described herein. Such a product can include information produced by the computing process, where the information is stored on a non-transitory, tangible computer-readable storage medium, and such a product can include any embodiment of the computer program product or other data combinations described herein.

[0146] Finally, the language used in this specification has been principally selected for readability and guidance, and it may not have been selected to circumscribe or delimit the patent rights. Accordingly, it is intended that the scope of the patent rights not be limited by the specific embodiments described, but rather by any claims that are published on the basis thereof in this application. Therefore, the disclosure of the embodiments is intended to illustrate rather than limit the scope of the patent rights claimed in the appended claims.

Claims

1. A method, comprising: Collecting audio from a local area around an audio controller using one or more acoustic sensors; Collecting information describing the local area from one or more additional sensors; Identifying, by the audio controller, one or more interactions of a user with the local area based on the collected audio and the information describing the local area; And Estimating a hearing profile of the user based on the identified one or more interactions and attributes associated with the one or more interactions, the hearing profile identifying hearing loss of the user for one or more frequency ranges.

2. The method according to claim 1, further comprising: Determining one or more filters to be applied to audio for presentation to the user based on the hearing profile; Generating enhanced audio by applying the one or more filters to the audio for presentation to the user, the enhanced audio increasing the amplitude of a portion of the audio whose frequencies are within a frequency range for which the hearing profile identifies hearing loss; And Presenting the enhanced audio to the user via one or more transducers.

3. The method according to claim 1 or 2, wherein Estimating the hearing profile of the user based on the identified one or more interactions and attributes associated with the one or more interactions comprises: Applying a model to the identified one or more interactions and attributes associated with the one or more interactions, the model outputting the hearing profile based on the identified one or more interactions and attributes associated with the one or more interactions.

4. The method according to any one of the preceding claims, wherein, An attribute associated with an interaction is an indication of whether a sound source providing audio during the interaction is within the field of view of an imaging device, the field of view of the imaging device overlapping the field of view of the user.

5. The method according to any one of the preceding claims, wherein, The audio controller is included in a device worn by the user, and an attribute associated with an interaction is the position of a sound source providing audio during the interaction relative to the audio controller.

6. The method according to any one of the preceding claims, wherein, An attribute associated with an interaction includes data describing a repair initiation made by the user during the interaction, the repair initiation indicating a request to initiate clarification of audio from a sound source providing audio during the interaction.

7. The method according to any one of the preceding claims, wherein, An attribute associated with an interaction includes one or more turn-taking metrics for the interaction, the turn-taking metrics describing conversation activities of the user during the interaction.

8. The method according to claim 7, wherein The turn-taking metric is selected from the group consisting of: the length of time the user produces audio during the interaction; the percentage of time the user produces audio during the interaction; the amount of time between when the sound source provides audio and when the user provides audio during the interaction; And any combination thereof.

9. The method according to any one of the preceding claims, wherein, An attribute associated with an interaction includes a semantic metric based on a topic determined for audio collected from the user during the interaction.

10. A computer program product, comprising a non-transitory computer-readable storage medium having instructions encoded thereon that, when executed by a processor, cause the processor to: Execute the method according to any one of the preceding claims.

11. A head-mounted device, comprising: A frame; One or more display elements coupled to the frame, each display element configured to generate image light for presentation to a user; One or more acoustic sensors configured to collect audio from a local area around the head-mounted device; One or more additional sensors configured to collect information describing the local area around the head-mounted device; And An audio controller, the audio controller including a processor and a non-transitory computer-readable storage medium having instructions encoded thereon that, when executed by the processor, cause the processor to perform the following operations: Identify, by the audio controller, one or more interactions of the user with the local area based on the collected audio and the information describing the local area; And Estimate a hearing profile of the user based on the identified one or more interactions and attributes associated with the one or more interactions, the hearing profile identifying hearing loss of the user for one or more frequency ranges.

12. The head-mounted device according to claim 11, further comprising a transducer array for presenting audio to a user, and wherein, The audio controller further includes instructions encoded thereon that, when executed by the processor, cause the processor to perform the following operations: Determine one or more filters to be applied to the audio for presentation to the user based on the hearing profile; Generate enhanced audio by applying the one or more filters to the audio for presentation to the user, the enhanced audio increasing the amplitude of a portion of the frequencies of the audio that are in a frequency range for which the hearing profile identifies hearing loss; And Present the enhanced audio to the user via one or more transducers.